The question arrives eventually, usually from whoever approved the budget: is this working? It is a fair question and most organizations cannot answer it — not because the platform failed, but because nothing was measured before it arrived, so there is nothing to compare against.
That is the whole discipline in one sentence. The expensive mistake is not choosing the wrong metric; it is starting without a baseline, which converts every subsequent claim into an assertion.
Baseline first, and it is nearly free
Almost everything worth measuring can be captured before a single change is made, from systems you already run. Do it during discovery, while it is still cheap and while nobody has an interest in the answer.
What to capture:
- Site-search behaviour. Query volume, the proportion returning nothing, the queries repeatedly refined, and how often someone searches and then leaves. Your search log is the single richest source here and is almost always already recording.
- Support and inbound volume by topic. Which questions arrive that your content was supposed to answer.
- Editorial cycle time. How long from a change being requested to it being live, measured over a handful of recent examples rather than estimated.
- Content health. How many items exist, how many have a named owner, how many are inside their review period, how many duplicate something else.
- Task completion on the journeys that matter — whether someone who arrives looking for a specific thing reaches it.
The last two are the ones organizations skip and later wish they had. Content health in particular tends to be the measure that improves earliest and most convincingly, because the preparatory work moves it directly.
What is worth measuring, and what each one tells you
Discovery
Zero-result rate is the most useful single number in this whole area: it is unambiguous, it is easy to capture, and it maps directly onto a fixable cause. A falling zero-result rate means people are finding things that exist; a stubborn one usually means the content does not exist rather than that search is poor, which is a different and more valuable finding.
Search refinement — how often someone reformulates rather than clicking — measures whether results are relevant, not merely present. Click depth on results tells you whether the right answer is at the top or merely somewhere in the list.
Self-service
Deflected contact is what most business cases are really about, and it is harder to measure honestly than it looks. The defensible version is a fall in inbound volume on specific topics that correspond to content you improved, compared against topics you did not touch. That comparison controls for seasonality and general traffic changes, which a raw before-and-after does not.
Editorial
Cycle time from request to publication, and rework rate — how often published content is corrected shortly afterwards. Rework is the more interesting of the two: if cycle time falls while rework rises, the process got faster and worse.
Content health
Share of content with a named owner, share inside its review period, and duplicate count. These are leading indicators — they move before the reader-facing measures do, and a decline predicts trouble months ahead of anyone noticing degraded answers.
Where an assistant is in use
Grounded-answer rate (answers with a traceable source), appropriate-refusal rate (declining when it should), and reported wrong answers. A very low refusal rate is a warning, not a success — it usually means the system is improvising rather than that it knows everything.
Metrics that mislead
Several measures are easy to collect, look impressive on a slide, and tell you nothing about whether the platform is working.
- Volume of AI-generated content. Measures activity, not value. Publishing more is not the goal and is frequently the opposite of it.
- Total searches. Rising search volume is ambiguous — more people finding the platform useful, or more people failing to find things and trying again. Only paired with success rate does it mean anything.
- Suggestion acceptance rate. Reads as a quality measure and behaves as a fatigue measure. A very high acceptance rate is as likely to indicate rubber-stamping as accuracy — see human-in-the-loop content management.
- Time saved, self-reported. People are poor estimators of their own time, and the estimate is influenced by whether they want the tool to succeed.
- Pageviews. A visitor who finds the answer on the first page generates fewer pageviews than one who does not. Improving the platform can push this down.
- Model-level benchmarks. Published scores measure the model on someone else's task. They predict nothing about your content.
The common thread: each measures effort or activity rather than whether anyone got what they came for.
Be careful with financial claims
The pressure to express this as a return figure is real, and it is where measurement most often becomes fiction. A number built by multiplying a self-reported time saving by a blended hourly rate is arithmetic, not evidence, and it tends to collapse under the first serious question.
A more defensible approach is to report the observable measures honestly and let the reader attach value to them. "Contact on these six topics fell while contact on untouched topics did not" is a claim you can stand behind. "This delivered a 340% return" is a claim someone will eventually ask you to substantiate.
Deciding what to measure? Schedule an AI CMS consultation — agreeing the measures before implementation is what makes the result arguable afterwards.
Measure the process, not only the output
Once volume passes what anyone can read, the platform's quality becomes a property of the process producing content rather than something inspectable item by item. A fixed evaluation set — real questions with known-correct sources, re-run after any change to prompts, models, or content — is what turns that into an observable trend. The method belongs to quality assurance; what matters here is that the results are tracked over time rather than checked once at launch.
This is also the mechanism that detects drift, which is the failure mode that produces no error and no complaint until confidence has already gone.
Report on a cadence, and expect a shape
Benefits do not arrive together, and a report that expects them to will read as failure part-way through. Discovery measures move first, because better retrieval can often be applied to existing content early. Content-health measures move next, as the modelling and ownership work lands. Editorial and self-service measures move last, because both depend on habit rather than configuration.
Saying that in advance is worth more than any single metric. A business case built entirely on the late-arriving measures will be under pressure long before those measures move.
How LABUSA approaches measurement
We agree the measures and capture the baseline during discovery, before anything changes, because that is the only moment it can be done cheaply and without an interested party. In practice that means a short list — usually zero-result rate, content ownership coverage, editorial cycle time, and topic-level inbound volume — instrumented properly rather than a dashboard of everything available.
An AI content management capability that cannot be evaluated is one nobody will renew, so we treat instrumentation as scope rather than as reporting. Where a measure would require assumptions we cannot support, we will say so rather than produce a number.
Frequently asked questions
We have already launched. Is it too late to baseline?
For the reader-facing measures, largely — though historical search logs and support records often reach back far enough to reconstruct one. For content health you can baseline today, since it changes slowly. Start now rather than not at all.
What is the single best metric?
Zero-result search rate, if you must pick one. It is unambiguous, cheap to capture, and points at a specific fixable cause.
How often should this be reviewed?
Frequently enough to act and rarely enough for signal to exceed noise — monthly for operational measures, quarterly for the trend. Re-measure immediately after any change to prompts, models, or major content sets.
Should we expect improvement in the first month?
In discovery measures, plausibly. In editorial and self-service measures, no — those depend on adoption, and expecting them early is a reliable way to declare failure prematurely.
What if the numbers do not improve?
Then you have learned something for the price of instrumentation, which is the cheapest part of the programme. Most often it points at content that does not exist rather than at a platform that does not work — and that is an actionable finding.
Related reading
- Benefits of AI Content Management — what each benefit depends on, qualitatively.
- AI Content Quality Assurance — measuring whether the content itself is good.
- AI CMS Cost and Budget Considerations — the other side of the equation.
- Enterprise AI Content Strategy — where success measures should first be agreed.