Twelve months after an assessment, somebody who did not authorise it will ask what it achieved. It is a fair question and it is usually asked at budget time, which is the worst moment to start looking for an answer.
The answer is easier to give if the measures were chosen before the work began. That is the whole method, and most of what follows is about choosing them well rather than about arithmetic.
This is not the same question as what it cost
Cost is a separate matter, and this cluster deliberately publishes no price for the assessment. What follows is about whether the work changed anything, which is answerable whatever it cost. If you are still deciding whether to commission one at all, what an AI readiness assessment is covers scope, and what happens after an AI readiness assessment covers the weeks that follow delivery.
What to measure
Four families of measure between them cover almost everything worth knowing, and they answer different questions for different audiences.
Did the stated problem move. An assessment is commissioned because something specific was unclear: whether the data would support a use case, whether governance existed, whether the infrastructure could carry the load. Write that question down at the start in the form of a sentence somebody could later mark true or false. The most common failure here is not measuring badly. It is never having written the question down, so that any outcome can be described afterwards as the one intended.
What it cost you internally. Not the invoice, but the hours your own people spent in interviews, gathering records and reviewing drafts. This number is almost never captured and it is a real cost. It is also the number that makes a second engagement easier to plan honestly.
Did capability stay in the organization. If the only person who understands the findings is the supplier, the engagement bought a document rather than a capability. A practical test: can somebody internal explain the sequence of work and why it is in that order, without opening the report.
Did the next decision get easier. This is the use most assessments are actually put to. A requirement written from evidence, a supplier conversation that started from your conditions rather than their product, a board paper that took an afternoon instead of a fortnight. What the deliverable set contains and how it gets reused is covered in what you own at the end of an AI readiness assessment.
There is a useful distinction in the UK Government's AI Playbook between metrics that describe how well a technology is performing and metrics that describe whether users' needs and business goals are being met. The Playbook notes the two can diverge considerably. The same split applies to an advisory engagement: a report can be accurate, thorough and on time while nothing downstream of it moves.
Re-scoring against the same dimensions
There is one measure available here that is not available for most consulting work, and it is the most direct evidence you will get.
Because the assessment scores the organization against a stable set of dimensions, described in the AI readiness framework, the same instrument can be run again later and the two results compared. Improvement in the dimensions you actually invested in is meaningful evidence. Improvement in dimensions you did nothing about is a prompt to ask what else changed.
Two cautions, both of which matter more than the comparison itself. The score is a decision-support indicator rather than a measurement of a physical quantity, a point made plainly on the methodology page for an AI readiness assessment from LABUSA, so a move from 54 to 61 is a direction rather than a quantity. And the comparison only works if the method was recorded the first time. A re-score carried out differently is a new assessment, not a measurement of progress. You can see the instrument yourself through the self-assessment.
What good looks like
A good measure has four properties, and they are worth writing into the engagement rather than discovering later.
It is checkable. Somebody can look at it and say yes or no without interpreting. "Governance improved" is not checkable. "An AI use policy exists, is approved, and names an owner" is.
It has a named confirmer who is neither the supplier nor the sponsor. The supplier has an obvious interest. The sponsor, less obviously, has one too, having asked for the budget.
It has a date. A measure with no date is never assessed, because there is never a moment at which it is due.
It was written before the work started. A measure chosen afterwards will, with the best intentions in the world, be one the work happened to satisfy.
Two or three such statements are enough. A list of fifteen is a list nobody will ever revisit.
How long it takes
Different kinds of value arrive on different schedules, and checking too early is the commoner error.
The advisory value is available immediately. Within days of delivery you either can or cannot write a better requirement, brief a board, or explain the sequence to a successor.
The implementation value arrives when the thing you were warned about would next have happened. If an assessment finds that a data quality problem would undermine a planned use case, the value shows up at the point that use case would have failed. That may be two months away or three quarters.
The capability value shows on the second occurrence. The first time a policy is applied it is being tested. The second time it is being used.
Re-scoring is worth doing at six or twelve months. Sooner than six and you are largely measuring enthusiasm.
The measures that mislead
Some numbers are easy to produce, tend to arrive unprompted, and say very little about whether the work mattered.
Hours delivered. This records that the engagement happened, and under a time based contract it is also the invoice, which makes it the least independent evidence available.
Volume of activity. Interviews conducted, systems reviewed, pages produced. Activity rises when work is going well and also when it is going badly.
Satisfaction. Worth collecting and easy to over read. It usually reflects how an engagement felt to the people in it, which is genuine and is not the same as whether anything changed.
On time and on budget. Measured against an estimate your own organization approved, this mostly measures the estimate.
Anything that moved for another reason. Systems get replaced, staff arrive and leave, funding cycles turn. Write down at the start what else was expected to change in the same period. It costs one sentence and it is the difference between evidence and coincidence.
Be most careful with a calculated saving. A figure produced by multiplying an assumed time saving by an assumed hourly rate will be asked about eventually, and it will not survive being asked about. This is not a small organization failing. In April 2026 the U.S. Government Accountability Office reported that agencies acquiring AI were not consistently collecting and applying lessons learned, and that some had difficulty understanding what their AI efforts actually cost. If federal agencies with dedicated oversight staff find attribution hard, a confident return figure attached to an advisory engagement deserves the same skepticism.
LABUSA does not publish a return figure for this service, and would not be able to support one if it did. What it can support is the method above.
Writing it into the purchase
All of this is cheap to arrange in advance and awkward to arrange afterwards. Three or four lines in the order form are enough: the question the assessment is meant to answer, the two or three checkable statements that would show it was answered, who confirms them and by when, and whether a re-score is included or is a separate piece of work.
The habit itself is not unusual in public sector practice. The accountability framework published by the U.S. Government Accountability Office organizes oversight of an AI programme around governance, data, performance and monitoring, and treats performance and monitoring as continuing obligations rather than one time checks. The measure functions of the NIST AI Risk Management Framework take the same view. For federal agencies specifically, and for them only, OMB memorandum M-25-21 requires that an expected benefit be supported by specific metrics or qualitative analysis compared against existing processes. State and local bodies are not bound by it, and it is a reasonable model to borrow.
If you would like help deciding which two or three statements are the right ones for your organization, that conversation is part of how we scope a readiness engagement.
Sources and further reading
- U.S. Government Accountability Office, Artificial Intelligence: An Accountability Framework for Federal Agencies and Other Entities, GAO-21-519SP, June 2021.
- U.S. Government Accountability Office, AI Acquisitions: Agencies Should Collect and Apply Lessons Learned, GAO-26-107859, April 2026.
- National Institute of Standards and Technology, AI Risk Management Framework.
- UK Government Digital Service, Artificial Intelligence Playbook for the UK Government, February 2025.
- Executive Office of the President, Office of Management and Budget, M-25-21, Accelerating Federal Use of AI, April 2025. Binds federal agencies only.
Every source above was opened and read on 19 August 2026. To talk through how you would measure a readiness engagement in your own organization, get in touch.