The assessment arrives addressed to the leadership team. It gets completed by someone three levels down.
That was the version I saw most often in fourteen years of walking institutions through these things. Someone capable, usually in IT or the strategy office, was handed it and told to sort it out. They chased colleagues for answers, filled the gaps with what seemed reasonable, and sent it back up. The leadership saw the score, not the questions.
When the team does sit down together, what happens is this. The question is about oversight of AI systems. Five levels, from no arrangements at all through to continuous practice. Someone senior reads them and picks the fourth. Governed.
Underneath sits a list of practices. A register of deployed systems. A named owner for each one. A documented authority to pause a system. A record of what the systems decided.
Nobody ticks any of them.
Both answers were honest. They were given thirty seconds apart by the same people. And in almost every assessment in circulation, the second answer changes nothing. The report records Governed.
The question behind the question
This is not a story about institutions overstating themselves. I went looking for evidence that they systematically do, and did not find it.
It is a story about what the instrument asks.
A maturity level is a summary judgment. Asked whether oversight of AI is governed in their institution, a leadership team answers the question it hears, which is whether the institution takes this seriously, whether there is a committee, whether someone is thinking about it. On that question, Governed is frequently the honest answer.
The indicator list asks something else entirely. Not whether the institution takes this seriously, but whether a specific artefact exists. Not whether oversight is intended, but whether a named person can pause a system this afternoon.
Two different questions produce two different answers, and both are truthful. The problem is that most instruments collect the second and score the first.
Where this came from
I should be plain about my position, because this is an argument about maturity assessments made by someone who sells one.
Threshold's diagnostic was built the way I have just criticised. An institution declared a level, the indicators sat underneath, and the score followed the declaration.
Before launch it went to chief AI officers, chief information officers and practitioners inside the institutions it was built for. The objection came back from more than one of them: nothing stops a leadership team selecting the level it would like to have. They were right, and the instrument changed before it was used with anyone.
The change has a cost worth stating. Institutions score lower. Some score considerably lower than they expected to, in front of colleagues, which is not a comfortable way to begin.
What changes when the evidence decides
An institution's declared level is a claim. The practices beneath it are the evidence for that claim. If the instrument records the lower of the two, three things happen immediately.
The score becomes defensible. A leadership team that records Developing because it confirmed the practices of Developing has a number it can put in front of a board without qualification.
The instrument stops being flattering. Nobody arrives at a good result by reading the level descriptions carefully and choosing well.
And the distance between the claim and the evidence becomes visible, which is the part that matters. That distance is not a scoring artefact. It is a finding, and it is frequently the most useful thing the assessment produces.
An institution that declares Governed and confirms nothing has told you something precise: the intent exists and the apparatus does not. That is a different problem from an institution that declares Exposed, and it needs a different first move. The first has agreement at the top and nothing underneath. The second has neither, and knows it.
Only one of those two institutions is at risk of being surprised.
Why this matters more for AI than for anything else
Governance frameworks have always been self-assessed. It worked because most of what they cover fails slowly. An institution that overstates its records management can be wrong for years before anything tests it.
AI does not offer that grace period. The question arrives attached to a specific decision and a specific date. Who approved this system. Who answers for what it decided. Regulators in this region have begun writing that requirement into procedure, with response windows measured in days rather than quarters. I set out how that works in Saudi Arabia, Bahrain and the DIFC in The Vendor Is Not Accountable.
An institution that has rated itself Governed and cannot produce a name does not discover the gap gradually. It discovers it on the day it is asked, with a clock running.
That is the hazard of a self-assessment nobody can fail. It does not just produce a wrong number. It produces confidence, and confidence stops the work.
The distance is the debt
The distance between what an institution has deployed and what it can answer for is accountability debt. An instrument that lets an institution declare its own maturity without evidence does not measure that debt. It hides it, in the institution's own words.
Measuring it is not complicated. It requires only that the instrument ask both questions and let the second one decide.
Do this with whatever you already use
You do not need a new tool to test this. Take the last governance assessment your institution completed, on any framework, and run it a second time under one rule.
For every level you claimed, name the artefact you would hand to a regulator to support it. Not a plan, not a committee minute recording an intention, not a policy that describes what should happen. A record of what does.
Where you can name it, the claim stands.
Where you cannot, you have not found a failure. You have found the difference between what your institution believes about itself and what it can currently demonstrate. Better that you calculate it than that someone else does.