Governance is easy to buy and hard to justify a year later, because the thing it prevents does not happen and the thing it produces is paperwork. This page is about what to measure so that the question can be answered with evidence rather than conviction.
It measures method rather than outcomes, and that distinction is deliberate. LABUSA does not publish a return figure for governance work, because we have no customer result we are able to publish, and a number invented to fill that gap would be worse than the gap.
This is not the same question as what it cost
The first mistake is to treat governance as an investment with a return, and then to look for the return in savings. Governance is not primarily a cost-reduction activity. It occasionally reduces cost, usually by stopping duplicate tool purchases, but if that is the case being made then a procurement review would have been cheaper.
What governance actually buys is the ability to answer questions and to make decisions faster than you could before. Both are measurable, and neither shows up as a saving.
There is a third framing that is worth naming and setting aside: insurance. Governance is sometimes sold as protection against a future penalty, and the pitch is effective because the penalties are real and the numbers are large. It is a poor basis for measurement, though, because it produces a programme optimised for the argument you would make after something went wrong rather than for the decisions you make every week. Those decisions are the thing that has value, and they are the thing that can be counted.
There is a second version of the same mistake, which is measuring the absence of incidents. An organization with no AI incidents may have excellent governance, or it may have very little AI, or it may have incidents it has not detected because nobody is looking. Absence is not evidence until you know what detection looks like.
What to measure
Five families of measure hold up, and each answers a question somebody actually asks.
Coverage. What proportion of the AI systems in use are on the inventory, classified, and have a named owner. This is the foundational measure because every other one depends on it, and it is the one most likely to reveal that the programme is smaller than believed. Coverage falls between reviews without anybody doing anything wrong; that is what makes it worth tracking rather than establishing once.
Time to decision. How long it takes to answer may we use this tool. Measure it from the request rather than from when the committee first saw it, because the queue in front of the committee is usually the larger part. This is the measure staff feel, and it is the one that determines whether people route around the process.
Time to answer. How long it takes to complete a customer security questionnaire, an insurer form or a procurement due-diligence pack. Organizations with a maintained deliverable set answer these from documents; organizations without one convene meetings. The difference is large and easy to record.
Exception rate. How often a system is in use without having gone through the route, discovered afterwards. A rising exception rate means the process is too slow or too obscure, not that staff are careless.
Review currency. What proportion of classifications, policies and approvals are within their stated review interval. This is the measure that catches a programme quietly stopping. Coverage can look perfect for a year after everyone has lost interest.
Two of the five need care in how they are counted. Coverage has a denominator problem: the proportion of systems on the inventory is only meaningful if you have some independent way of knowing how many exist, which usually means a periodic sweep of expense claims, browser telemetry or procurement records rather than asking teams to self-report. And exception rate is a rate, so it needs the same denominator; a raw count of exceptions falls when adoption falls, which reads as an improvement and is not.
The NIST AI Risk Management Framework is organized around four functions, GOVERN, MAP, MEASURE and MANAGE, so measurement sits in the structure rather than being reporting bolted on afterwards. MAP precedes MEASURE, which is the practical point here: establish what you have, then measure it. Coverage first, everything else after.
What good looks like
Absolute targets are not transferable, because they depend on how much AI an organization uses and how much risk it carries. Direction and consistency are transferable, and they are what to look for.
Coverage should be high and stable rather than high once. A programme that reached ninety per cent at the end of an engagement and is at sixty a year later has not degraded slightly; it has stopped.
Time to decision should be predictable more than it should be short. Staff can plan around three weeks. They cannot plan around a process that takes three days sometimes and three months at others, and unpredictability is what drives people to adopt tools quietly.
Exception rate should fall and then flatten, not reach zero. A zero exception rate over a long period usually means nobody is checking, and a small steady rate that gets caught is a healthier signal than silence.
Review currency should be near total, because it is entirely within your control. Unlike coverage, nothing external moves it. A low figure here is a statement about attention rather than about difficulty.
How long it takes
The measures mature at different rates, and expecting them all at once produces a disappointing first review.
Coverage and review currency can be reported from the moment the inventory exists, which is usually within the first engagement. Time to answer improves as soon as the documents are maintained and findable, so it moves early too.
Time to decision needs enough requests to be meaningful, and a small organization may take two or three quarters to accumulate them. Exception rate needs longer still, because it depends on detection improving, and detection improves gradually as people learn what to look for.
Anyone promising a measurable return in the first quarter is describing coverage and calling it something else.
A reporting cadence that matches those maturities is more useful than a single annual review. Coverage and review currency quarterly, because they move and because a fall is worth catching early. The rest annually, with the first year treated openly as baseline-setting rather than as performance. Reporting a measure before it can move is how a genuinely working programme acquires a reputation for producing flat numbers.
The measures that mislead
Four are common and each is worse than reporting nothing.
Policy count. The number of AI policies published measures writing, not governance. Twelve policies nobody reads is a worse position than two that staff can recite, and the count moves in the wrong direction.
Training completion. A useful hygiene measure and a poor governance measure. It records that people clicked through a module, which is not evidence that a decision was made differently.
Incidents avoided. Unknowable. Any figure attached to it is constructed, and constructing it damages the credibility of the measures that are real. If a board asks for it, the honest answer is that it cannot be measured and here are five things that can.
Maturity scores. Useful for orienting, misleading as a target. A score is a compression of many judgements into one number, and once it becomes the goal, the judgements are made to move the number. Report the underlying measures and let the score be a summary rather than an objective.
A fifth trap is subtler than the four above because the measure itself is sound: reporting a measure without its denominator. Coverage of ninety per cent means one thing across two hundred systems and another across nine, and a board given the percentage alone cannot tell which it is looking at. Report both numbers, always.
The GAO AI Accountability Framework puts continuous monitoring alongside governance, data and performance as a distinct principle, which is a useful corrective: monitoring is a practice with its own evidence, not a number produced at the end of the others.
Writing it into the purchase
Decide the measures before the engagement rather than after, for a reason that is not about accountability. Several of them require a baseline, and a baseline can only be taken before the work changes anything. Time to answer in particular is unrecoverable once the documents exist, because nobody remembers how long the last questionnaire took.
Three things worth agreeing in the scope of work: which measures will be reported, who produces the figure, and what the first baseline was. The third is the one usually skipped, and it is the reason so many governance programmes cannot demonstrate improvement they genuinely achieved.
Record what you decided not to measure as well, and why. Federal reviewing bodies have found repeatedly that organizations fail to capture lessons in a form later work can use, most recently in GAO-26-107859, and a measurement decision with no recorded rationale is re-argued at every review.
Where to go next
To establish a baseline before commissioning anything, the AI governance checklist gives a defensible starting position in an afternoon. The artifacts these measures are taken from are described in what you own at the end of an AI governance engagement, and the domains they sit within are in the AI governance framework.
If you want to agree the measures and the baseline before any work begins, that is exactly the right order and it is how LABUSA scopes governance work. Talk to us about what you would want to be able to show a year from now.