Resources 8 min read

AI Risk Assessment

How to run an AI risk assessment: assess use cases rather than systems, ten dimensions, a three-tier model, and who is entitled to accept what remains.

A hand holding a pen over a clipboard while taking notes during an assessment.

An AI risk assessment answers a narrow question: for this use of this system, what could go wrong, how badly, who would be affected, and what are we going to do about it.

This page is about the method. What the risks actually are is covered in more depth in the AI cybersecurity risks organizations should assess before deployment, and repeating that list here would help nobody.

Assess use cases, not systems

The most common structural error is assessing a product. One assistant may be entirely appropriate for drafting internal documents, unacceptable for anything touching personal data, and a genuinely open question for customer correspondence. A single verdict on the product cannot express that, so it either blocks something useful or approves something it should not.

Assess the use case: this system, doing this job, on this data, for these users, with this exposure. That is also the unit governance decisions attach to, which is why the inventory records use cases rather than only systems.

What to establish before assessing anything

Four facts, and an assessment that starts without them will produce a confident answer to the wrong question.

What the system is actually doing. Not what it is marketed as doing. Ask whoever uses it to describe a real task end to end.

What data reaches it. Including data the user does not think of as data: the document pasted for context, the ticket history retrieved automatically, the file attached for reference.

Where the output goes. Internal only, to a customer, into a record, into a decision, or published.

Who is affected if it is wrong. This is the question that determines how much of the rest of the assessment is worth doing.

The dimensions worth assessing

Ten, and most use cases will be trivial on seven of them. The value of the list is that it makes the three that matter visible rather than assumed.

Accuracy and confabulation. NIST uses the term confabulation rather than hallucination for output that is fluent, plausible and wrong. The assessment question is not whether it happens, which it does, but whether anything downstream would catch it.

Consequence to a person. Does the output bear on eligibility, employment, discipline, grading, safety, credit or a benefit. If yes, most of the remaining dimensions become mandatory rather than optional.

Privacy and data exposure. What personal or confidential information enters the system, where it is stored, how long, and what the vendor may do with it.

Security. How the system is authenticated, what it can reach, and what an attacker could do with it. Access inheritance is the recurring surprise: a retrieval system answers using everything the requesting account can open.

Bias and differential performance. Whether the system performs differently for different groups in a way that matters given what it is used for.

Transparency. Whether the people affected know AI was involved, and whether anyone can explain a specific output well enough to defend it.

Reliability and dependency. What happens when it is unavailable, and whether a process has quietly come to depend on it.

Vendor and model dependency. What happens when the vendor changes the model, the terms or the price, or withdraws the capability.

Legal and contractual exposure. Whether the use touches an obligation the organization already has, including sector rules and customer contracts. This is a question for counsel, not for an assessment template.

Human oversight. Whether a person reviews before the output takes effect, whether they have the information to judge it, and whether they can in practice overturn it.

A tiering model that survives contact with real requests

Sort by consequence, not by technology. Three tiers cover most organizations, and a coarse scheme that gets applied beats a fine one that gets argued.

Tier 1, contained. Output affects only the person using it. Drafting, summarising for personal use, exploratory work. Light-touch: the acceptable-use position and the data rule are sufficient, and no per-use assessment is needed.

Tier 2, external or operational. Output reaches someone outside the immediate user, or enters a record or a business process. Customer correspondence, published content, internal knowledge systems. Assessment required, human review defined, review date set.

Tier 3, consequential to a person. Output bears on a decision about an individual. Full assessment, mandatory human review with genuine authority to overturn, documented rationale, shorter review cycle, and an explicit decision recorded by an accountable owner.

Two rules keep the scheme honest. Tier is a property of the use case rather than the system, and a system can hold use cases in different tiers. And anything unclassified is treated as tier 2 until somebody classifies it, because the alternative is a default of no oversight.

Running the assessment

For a tier 2 or tier 3 use case, the assessment is a conversation with a record attached, not a form-filling exercise. It works best with the system owner, someone from security or privacy, and whoever will actually use it in the room together, because the four establishing facts above are usually spread across those three people.

Work through the dimensions, note which are material, and for each material one record the risk, what control addresses it, who owns that control, and what residual risk remains after it. The residual line is the part organizations skip and the part a governing body will ask about.

Reach an explicit decision: approved, approved with conditions, declined, or deferred pending something specific. A deferral without a named blocker becomes a permanent no by attrition.

Residual risk, and who is entitled to accept it

No control eliminates a risk, and an assessment that concludes every risk is mitigated has not been done properly. What remains after controls is residual risk, and somebody has to accept it by name.

That person should be the system owner for tier 2 and an executive for tier 3. Recording the acceptance is not bureaucracy: it is the difference between a decision the organization made and a thing that happened to it, and it is what allows a later reviewer to see that the risk was known rather than missed.

Two failure modes in how assessments are written up

The assessment that assesses the vendor. A long document describing the supplier's certifications, data centers and encryption, and nothing about what your staff will do with the product. Supplier posture is one input to one dimension. It is not the assessment, and a strong vendor does not make a weak use case safe.

The assessment nobody can act on. A risk register that lists concerns without naming an owner, a control or a decision. If a reader cannot tell what changed as a result of the assessment, nothing did. Every material risk should end in one of three states: a control with an owner, an accepted residual risk with a name against it, or a condition on approval.

A useful discipline is to write the decision first and the analysis underneath it. It forces the assessment to reach a conclusion, and it puts the part a reviewer needs at the top.

Where the recognized frameworks fit

The NIST AI Risk Management Framework covers this territory in its MAP and MEASURE functions, and its Generative AI Profile enumerates risks specific to generative systems, including confabulation, information integrity and data privacy. Both are voluntary guidance rather than requirements.

ISO/IEC 42005:2025 addresses AI system impact assessment directly and is the published standard closest to the method on this page. ISO/IEC 23894:2023 covers AI risk management guidance more broadly. Both are purchasable standards; neither is reproduced here.

The GAO AI Accountability Framework is organized around Governance, Data, Performance and Monitoring and is addressed to federal agencies and other entities. It is useful as a comparison point for a governing body rather than as an obligation on a private organization.

Reassessment is the part that gets dropped

An assessment describes a system on a date. Vendors change models, enable capabilities by default and revise terms, and use cases drift from what was approved.

Give every tier 2 and tier 3 use case a review date and a trigger list. The triggers that matter are a model or terms change, a new default-on capability, a change in the data reaching the system, and an incident. None of those arrives with a notification addressed to you, which is why the trigger list has to be somebody's job.

Where to go next

Assessment takes the inventory as input and produces the tiering that the rest of the governance framework depends on. If a supplier is the subject, AI vendor risk assessment covers the questions and contract terms specifically. To work through your own position quickly, use the AI governance checklist.

LABUSA runs AI risk assessments as part of an governance and responsible AI services, including the tiering scheme and the reassessment cadence. If you have a specific system in front of you and need a view on it, start there.

About LABUSA

LAB Information Technology Incorporated (LABUSA) is a trusted provider of managed IT solutions, empowering organizations with secure, efficient, and scalable technologies. With expertise spanning cybersecurity, cloud services, enterprise software, and data management, LABUSA helps clients modernize operations, strengthen compliance, and optimize performance. Our customer-focused approach ensures tailored solutions that align with organizational goals while maintaining the highest standards of reliability and security. Headquartered in Houston, Texas, LABUSA serves government agencies, corporations, and nonprofits across the United States and internationally.