Resources 8 min read

Measuring the Value of an AI Readiness Assessment

How to tell whether an AI readiness assessment was worth commissioning: what to measure, what good looks like, when to check, and the measures that quietly mislead.

Woman in glasses reviews data dashboards on dual monitors at an enterprise office desk; taking notes with a pen.

Twelve months after an assessment, somebody who did not authorise it will ask what it achieved. It is a fair question and it is usually asked at budget time, which is the worst moment to start looking for an answer.

The answer is easier to give if the measures were chosen before the work began. That is the whole method, and most of what follows is about choosing them well rather than about arithmetic.

This is not the same question as what it cost

Cost is a separate matter, and this cluster deliberately publishes no price for the assessment. What follows is about whether the work changed anything, which is answerable whatever it cost. If you are still deciding whether to commission one at all, what an AI readiness assessment is covers scope, and what happens after an AI readiness assessment covers the weeks that follow delivery.

What to measure

Four families of measure between them cover almost everything worth knowing, and they answer different questions for different audiences.

Did the stated problem move. An assessment is commissioned because something specific was unclear: whether the data would support a use case, whether governance existed, whether the infrastructure could carry the load. Write that question down at the start in the form of a sentence somebody could later mark true or false. The most common failure here is not measuring badly. It is never having written the question down, so that any outcome can be described afterwards as the one intended.

What it cost you internally. Not the invoice, but the hours your own people spent in interviews, gathering records and reviewing drafts. This number is almost never captured and it is a real cost. It is also the number that makes a second engagement easier to plan honestly.

Did capability stay in the organization. If the only person who understands the findings is the supplier, the engagement bought a document rather than a capability. A practical test: can somebody internal explain the sequence of work and why it is in that order, without opening the report.

Did the next decision get easier. This is the use most assessments are actually put to. A requirement written from evidence, a supplier conversation that started from your conditions rather than their product, a board paper that took an afternoon instead of a fortnight. What the deliverable set contains and how it gets reused is covered in what you own at the end of an AI readiness assessment.

There is a useful distinction in the UK Government's AI Playbook between metrics that describe how well a technology is performing and metrics that describe whether users' needs and business goals are being met. The Playbook notes the two can diverge considerably. The same split applies to an advisory engagement: a report can be accurate, thorough and on time while nothing downstream of it moves.

Re-scoring against the same dimensions

There is one measure available here that is not available for most consulting work, and it is the most direct evidence you will get.

Because the assessment scores the organization against a stable set of dimensions, described in the AI readiness framework, the same instrument can be run again later and the two results compared. Improvement in the dimensions you actually invested in is meaningful evidence. Improvement in dimensions you did nothing about is a prompt to ask what else changed.

Two cautions, both of which matter more than the comparison itself. The score is a decision-support indicator rather than a measurement of a physical quantity, a point made plainly on the methodology page for an AI readiness assessment from LABUSA, so a move from 54 to 61 is a direction rather than a quantity. And the comparison only works if the method was recorded the first time. A re-score carried out differently is a new assessment, not a measurement of progress. You can see the instrument yourself through the self-assessment.

What good looks like

A good measure has four properties, and they are worth writing into the engagement rather than discovering later.

It is checkable. Somebody can look at it and say yes or no without interpreting. "Governance improved" is not checkable. "An AI use policy exists, is approved, and names an owner" is.

It has a named confirmer who is neither the supplier nor the sponsor. The supplier has an obvious interest. The sponsor, less obviously, has one too, having asked for the budget.

It has a date. A measure with no date is never assessed, because there is never a moment at which it is due.

It was written before the work started. A measure chosen afterwards will, with the best intentions in the world, be one the work happened to satisfy.

Two or three such statements are enough. A list of fifteen is a list nobody will ever revisit.

How long it takes

Different kinds of value arrive on different schedules, and checking too early is the commoner error.

The advisory value is available immediately. Within days of delivery you either can or cannot write a better requirement, brief a board, or explain the sequence to a successor.

The implementation value arrives when the thing you were warned about would next have happened. If an assessment finds that a data quality problem would undermine a planned use case, the value shows up at the point that use case would have failed. That may be two months away or three quarters.

The capability value shows on the second occurrence. The first time a policy is applied it is being tested. The second time it is being used.

Re-scoring is worth doing at six or twelve months. Sooner than six and you are largely measuring enthusiasm.

The measures that mislead

Some numbers are easy to produce, tend to arrive unprompted, and say very little about whether the work mattered.

Hours delivered. This records that the engagement happened, and under a time based contract it is also the invoice, which makes it the least independent evidence available.

Volume of activity. Interviews conducted, systems reviewed, pages produced. Activity rises when work is going well and also when it is going badly.

Satisfaction. Worth collecting and easy to over read. It usually reflects how an engagement felt to the people in it, which is genuine and is not the same as whether anything changed.

On time and on budget. Measured against an estimate your own organization approved, this mostly measures the estimate.

Anything that moved for another reason. Systems get replaced, staff arrive and leave, funding cycles turn. Write down at the start what else was expected to change in the same period. It costs one sentence and it is the difference between evidence and coincidence.

Be most careful with a calculated saving. A figure produced by multiplying an assumed time saving by an assumed hourly rate will be asked about eventually, and it will not survive being asked about. This is not a small organization failing. In April 2026 the U.S. Government Accountability Office reported that agencies acquiring AI were not consistently collecting and applying lessons learned, and that some had difficulty understanding what their AI efforts actually cost. If federal agencies with dedicated oversight staff find attribution hard, a confident return figure attached to an advisory engagement deserves the same skepticism.

LABUSA does not publish a return figure for this service, and would not be able to support one if it did. What it can support is the method above.

Writing it into the purchase

All of this is cheap to arrange in advance and awkward to arrange afterwards. Three or four lines in the order form are enough: the question the assessment is meant to answer, the two or three checkable statements that would show it was answered, who confirms them and by when, and whether a re-score is included or is a separate piece of work.

The habit itself is not unusual in public sector practice. The accountability framework published by the U.S. Government Accountability Office organizes oversight of an AI programme around governance, data, performance and monitoring, and treats performance and monitoring as continuing obligations rather than one time checks. The measure functions of the NIST AI Risk Management Framework take the same view. For federal agencies specifically, and for them only, OMB memorandum M-25-21 requires that an expected benefit be supported by specific metrics or qualitative analysis compared against existing processes. State and local bodies are not bound by it, and it is a reasonable model to borrow.

If you would like help deciding which two or three statements are the right ones for your organization, that conversation is part of how we scope a readiness engagement.

Sources and further reading

Every source above was opened and read on 19 August 2026. To talk through how you would measure a readiness engagement in your own organization, get in touch.

About LABUSA

LAB Information Technology Incorporated (LABUSA) is a trusted provider of managed IT solutions, empowering organizations with secure, efficient, and scalable technologies. With expertise spanning cybersecurity, cloud services, enterprise software, and data management, LABUSA helps clients modernize operations, strengthen compliance, and optimize performance. Our customer-focused approach ensures tailored solutions that align with organizational goals while maintaining the highest standards of reliability and security. Headquartered in Houston, Texas, LABUSA serves government agencies, corporations, and nonprofits across the United States and internationally.