Resources 11 min read

Public-Sector AI Case Studies: Lessons for Organizations Beginning with AI

Six lessons from documented government AI programmes, drawn from GAO evaluations and the UK Government AI Playbook, with what each one suggests for a public organization starting out.

American flag flying above a brick government building with white columns; surrounded by mature trees and manicured lawn.

Public organizations have been deploying AI long enough that there is now something better than speculation to learn from. Federal agencies publish inventories of what they run, auditors examine those programmes against a published accountability framework, and one national government has collected its departments' own write-ups of what worked and what did not.

What follows is drawn entirely from that documentation. None of these are LABUSA projects and none of the organizations named is a client. The difficulties are consistent: not the models, but the governance, data, workforce and monitoring around them.

Each lesson separates what the source found from what we think an organization can take from it. The gap between those two is where most vendor case studies quietly do their work.

Lesson 1: establish governance before scaling

What the source found

In December 2023 the U.S. Government Accountability Office reported on how federal agencies were implementing AI requirements, including the requirement to maintain an inventory of their AI use cases. Examining the inventories of 23 civilian agencies, GAO found five had provided comprehensive information for each reported use case and the other 15 had instances of incomplete and inaccurate data. Some omitted required fields such as the system's lifecycle stage. Two included entries the agencies themselves later determined were not AI. GAO made 35 recommendations to 19 agencies.

GAO's own summary of the consequence: without accurate inventories, the government's management of its use of AI will be hindered by incomplete and inaccurate data.

What LABUSA takes from it

These are large, well-resourced organizations working to an explicit federal requirement, with a deadline and an auditor, and they still could not produce an accurate list of what they ran. An organization with no requirement and no auditor should assume its picture is less complete, not more.

An inventory is therefore not overhead to be done later. It is what makes every other decision possible, because a governance position cannot be applied to systems nobody has written down, and the first pass is a conversation rather than a project.

What this means for TIPS members

A district, city or nonprofit will usually find more AI in use than it expected, because staff adopt useful free tools without a purchase order and so without a procurement event to notice. Ask without consequences attached and the answer is honest. Ask as an investigation and the number comes back smaller and wrong.

Relevant LABUSA readiness dimensions: governance and compliance; cybersecurity and privacy; strategy and leadership.

Lesson 2: data reliability decides what is possible

What the source found

In February 2024 GAO examined the Department of Homeland Security's use of AI for cybersecurity, assessing an automated system for detecting personally identifiable information against 11 practices from GAO's own AI Accountability Framework. Four were fully implemented, five partially, and two not at all.

The two that were not implemented were both about data: documenting the sources and origins of the data used to develop the capability, and assessing that data's reliability. GAO made eight recommendations, DHS concurred with all of them, and as of January 2025 all eight were recorded as implemented.

What LABUSA takes from it

The instructive detail is which practices were missed. This is a capable security organization running a working system, and the gaps were not technological. They were in knowing where the data came from and whether it could be relied on, which is invisible while a system works and decisive the moment somebody asks how it reached a conclusion.

The story also ends well: the gaps were found against a published framework, the recommendations accepted, the work done. That is a reasonable model for an organization that would rather find its own gaps than have somebody else find them.

What this means for TIPS members

Before scoping a candidate, establish where the information it depends on lives, who owns it, who may grant access, and whether anyone has checked it for accuracy. In most organizations the answer to the last is no, which is a finding rather than a failure.

Relevant LABUSA readiness dimensions: data readiness; governance and compliance; technology and infrastructure.

Lesson 3: keep a person between the output and the consequence

What the source found

The UK Government Digital Service published, in the appendix of its AI Playbook, a case study of GOV.UK Chat, an experimental assistant for finding answers across more than 700,000 pages of government content. The team used retrieval rather than fine-tuning, because the content changes constantly, and hundreds of users tested it in a controlled environment.

Roughly 70 percent of users found the responses useful and about 65 percent were satisfied, preferring it to conventional search for complex questions. The same write-up reports that the system occasionally produced answers that were not true. Rather than launching, the team concentrated on accuracy and made a wider pilot conditional on an accuracy threshold.

What LABUSA takes from it

A satisfaction rate near 70 percent is a good result, and the team still did not ship. User satisfaction and factual reliability are different measurements, and a system can score well on the first while failing the second in ways users never notice.

NIST has a precise term for this. Its Generative AI Profile calls it confabulation, meaning a system generating and confidently presenting content that is erroneous or false. The confidence is the problem, because it is what stops a reader checking.

What this means for TIPS members

Decide in advance which outputs a person reviews before they take effect, and write that down rather than assuming it. Then confirm the reviewer can actually judge the output in their own subject matter. Review by somebody who cannot tell whether an answer is right is not a control, it is a signature.

Relevant LABUSA readiness dimensions: governance and compliance; workforce and adoption; AI use cases.

Lesson 4: capability is the constraint, not licences

What the source found

In July 2025 GAO reported on generative AI use across 12 federal agencies. The growth was substantial: reported generative AI use cases across the selected agencies rose from 32 to 282 between 2023 and 2024. Officials at ten of the twelve told GAO that existing federal policy, data privacy policy among it, could present obstacles to adoption. Four said the pace at which the technology was changing made it difficult to establish governance and usage guidance.

Agencies also reported difficulty attracting and developing staff with generative AI expertise, competing with the private sector for the same people, and sustaining training that dates as fast as the technology.

What LABUSA takes from it

Federal agencies have budget, hiring authority and scale. If capability binds them, it will bind a smaller organization too, and the version that matters is usually not a shortage of specialists. It is whether the ordinary staff who rely on AI output day to day can recognize when it is wrong.

That is a training question with a specific shape: not general AI awareness, which is quickly forgotten, but role-specific competence in judging output.

What this means for TIPS members

Plan training by role rather than one session for everyone, and talk to affected staff about how their work would change before a tool arrives. Most resistance to AI in public organizations is a reasonable response to having been told nothing.

Relevant LABUSA readiness dimensions: workforce and adoption; strategy and leadership.

Lesson 5: the work does not end at go-live

What the source found

GAO's AI Accountability Framework devotes one of its four principles entirely to monitoring, with key practices covering plans for continuous monitoring, the range of data and model drift that is acceptable, documenting monitoring results and any corrective action, assessing whether a system is still relevant to its current context, and identifying the conditions under which it may be scaled.

The UK AI Playbook makes the same point in one of its ten principles: manage the full AI lifecycle, from selecting and implementing a system through maintaining it to decommissioning it securely, monitoring for drift and bias throughout.

What LABUSA takes from it

Both treat the deployed system, not the project, as the thing governed. An AI service is not a website that keeps working: the data underneath changes, the supplier changes the model without telling anyone, and performance degrades quietly rather than failing visibly.

This is the part organizations are least ready for, and it is why the readiness questions ask whether deployed tools are reviewed on a schedule, whether one could be switched off without disrupting dependent work, and whether an incident caused by an AI system's output would be recognized.

What this means for TIPS members

Before a pilot begins, name who owns the tool after the project team disbands, when it is next reviewed, and what evidence would justify turning it off. A pilot with no defined end is how an unevaluated system becomes permanent.

Relevant LABUSA readiness dimensions: governance and compliance; technology and infrastructure; cybersecurity and privacy.

Lesson 6: start with the problem, not the product

What the source found

In April 2026 GAO examined 13 AI acquisitions across four federal agencies and found that none of them was systematically collecting lessons learned, despite guidance directing agencies to share knowledge through a common repository. Officials at all four agencies said they had no policy requiring it, which left them unable to share good practice or to avoid repeating mistakes. Agencies also reported difficulty accessing AI technical expertise and understanding AI-related costs. GAO made four recommendations, one to each agency, and all four concurred.

The contrast in the UK Playbook's appendix is instructive. NHS England's automated moderation of NHS.UK reviews began from a defined operational problem, moderating hundreds of thousands of reviews against published policies, and was evaluated with confusion matrices, clerical review of false positives and negatives, and latency testing. The Crown Commercial Service measured its recommendation system against a control group that received none.

What LABUSA takes from it

The projects with a clearly stated problem produced measurable results, because a defined problem is what makes measurement possible. Difficulty understanding costs is a symptom of the opposite pattern: an organization that has not defined the problem cannot size the solution, so cannot forecast what running it will cost. No governance framework asks that; the evidence that it matters comes from procurement.

What this means for TIPS members

Name the process, the team and the rough volume before evaluating any product, and write down what a good outcome looks like and the number that would move. Then keep a short record of what happened, including what did not. Almost no public organization does this, and it is close to free.

Relevant LABUSA readiness dimensions: AI use cases; strategy and leadership; technology and infrastructure.

What should your organization do first

The lessons converge on an order. It is deliberately unexciting, and the first two steps cost nothing.

  1. Find out what AI is already in use, without attaching consequences to the answer.
  2. Name one executive accountable for AI decisions, with authority to approve and to stop work.
  3. Identify the mission problems worth solving, before looking at any product.
  4. Establish where the relevant information lives, who owns it and whether software can reach it.
  5. Assess the privacy and cybersecurity obligations attaching to the records in scope.
  6. Write a short governance position: who approves a tool, what staff may put into one, and which outputs a person reviews.
  7. Prioritize candidates on value and feasibility, scored separately.
  8. Pilot one contained use case where a knowledgeable person checks the output.
  9. Measure it against the outcome defined at step three.
  10. Scale only what the measurement supports, and set the next review date.

For a structured version of steps one to seven, work through the public-sector AI readiness checklist, take the AI Readiness Self-Assessment for a scored benchmark in about eight minutes, or read how the LABUSA assessment works.

A note on what these examples are

Every organization named above is somebody else's. The federal findings are published work of the U.S. Government Accountability Office, and the UK examples are departments' own write-ups collected in the Government Digital Service's AI Playbook, which notes that they were submitted in spring 2024, are not exhaustive and should not be treated as formal advice. None of these organizations is a LABUSA client, none has endorsed LABUSA, and nothing here describes work LABUSA performed.

They are here because documented public-sector experience is more useful than a vendor's account of its own successes, and the difficulties in them are the ones a structured readiness review exists to find first.

Sources and further reading

If you would like to talk through where your organization sits against any of this, get in touch.

About LABUSA

LAB Information Technology Incorporated (LABUSA) is a trusted provider of managed IT solutions, empowering organizations with secure, efficient, and scalable technologies. With expertise spanning cybersecurity, cloud services, enterprise software, and data management, LABUSA helps clients modernize operations, strengthen compliance, and optimize performance. Our customer-focused approach ensures tailored solutions that align with organizational goals while maintaining the highest standards of reliability and security. Headquartered in Houston, Texas, LABUSA serves government agencies, corporations, and nonprofits across the United States and internationally.