Every AI governance program eventually hits the same wall. The inventory comes back with 60, 100, 200 systems, and the review process was designed as if there would be 12. Committees back up. Business teams wait weeks for approvals on tools that draft meeting notes. Meanwhile the one system that can move money sits in the same queue as everything else, waiting its turn.
The wall is not a resourcing problem. It is a design problem. A program that treats every AI system the same has quietly decided that a grammar assistant and an autonomous refund agent deserve equal scrutiny. They do not, and pretending otherwise costs you twice: heavy process where it is not needed, and diluted attention where it is.
Risk tiering is the fix. Score each system when it enters the program, and let the score decide how much governance it gets. Done well, most of your portfolio moves fast through a light path while your review capacity concentrates on the small set of systems that can actually hurt you.
Start with impact, then stop pretending it is enough
Traditional risk tiering scores business impact. What breaks if this system is wrong? How much money, how many customers, how much regulatory exposure? That question still matters and should anchor your tiers. A reasonable starting scale runs from minimal impact, through moderate and significant, up to critical, where an error creates serious financial, legal, or safety harm.
For classic software, impact plus data sensitivity told you most of what you needed. AI breaks that sufficiency, for one reason: impact scoring assumes you know who is making the decisions. With AI, that is now a variable.
Two systems can score identical on impact while being wildly different in practice. One drafts a response for a human to review and send. The other sends it on its own, and can trigger a payment while it is at it. Same data, same business process, same impact score. Completely different exposure, because one has a person between the decision and the action and the other does not.
Impact tells you how bad an error could be. It says nothing about how fast an error can happen, how many can happen before a person notices, or whether anyone can stop it.
The dimensions impact misses
To tier AI honestly, score a few more things alongside impact. Each one answers a question your impact score cannot.
Autonomy
How independently does the system act? A scale helps here, from a system that only responds when asked, through one that suggests actions for humans to take, up to one that plans and executes multi-step work on its own. Be skeptical of self-reported answers. Teams consistently describe their systems as more supervised than they are, not from dishonesty but from optimism.
Authority
What is it allowed to affect? Reading data is one level. Writing to internal records is another. Communicating with customers, moving money, or changing entitlements is another again. Authority is the blast radius. A highly autonomous system with read-only access is a research assistant. The same autonomy with payment authority is something else entirely.
Human oversight
Who checks the output, and when? Before each action, in real time with power to intervene, after the fact on a sample, or never. This dimension moves risk more than any other, because oversight is the compensating control for everything else. A system nobody reviews earns a higher tier than its impact alone suggests.
Criticality
How dependent is the business on it? A system can be low impact when it errs and still be tier-worthy because everything stops when it goes down. Dependence is its own kind of risk, and it grows quietly as adoption spreads.
Let the worst dimension win
With multiple dimensions in play, you need a combination rule, and the safe one is simple. The most restrictive dimension sets the tier. A system that scores low on impact but acts fully autonomously with broad authority gets governed like the risk it actually is, not the average of its scores.
Averaging feels fairer and is quietly dangerous. It lets one alarming dimension hide behind several benign ones. The system that averages to "moderate" while holding unsupervised payment authority is exactly the system that ends up in an incident report. Most restrictive wins, automatically, with no meeting required to confirm it.
Make the tier mean something
A tier is only useful if consequences attach to it mechanically. Decide once what each tier requires, write it down, and let the classification trigger it. A workable pattern looks like this.
- The lowest tier gets a lightweight intake record and an annual glance. Approval in days, not weeks.
- The middle tiers get a named owner, a documented review, defined monitoring, and a periodic re-check.
- The top tier gets committee sign-off before launch, a written decision record, continuous monitoring, a tested way to shut it down fast, and a short review clock.
The point of the pattern is not the specific artifacts. It is that nobody negotiates governance system by system. The tier decides, the requirements follow, and your committee spends its time on the 10 systems that deserve it instead of the 100 that do not.
Re-tier on change, not just on schedule
An AI system's tier is a snapshot. Vendors ship new capabilities into products you already approved. A team widens an agent's authority to close a workflow gap. Usage spreads until a convenience becomes a dependency. Any of those can move a system across a tier boundary without anyone deciding it should.
So build 2 triggers into the program. A scheduled review, more frequent for higher tiers. And an event trigger: re-score whenever autonomy, authority, oversight, or data scope changes. The second trigger matters more than the first. Systems rarely drift into higher risk on your review calendar.
Where to start
If you are staring at a long inventory and a short runway, do this. Score everything quickly on the dimensions above, even roughly. Sort by tier. Take the top tier, usually a single-digit percentage of the list, and govern it properly this quarter. Put the lowest tier on the fast path immediately so the business feels the program helping, not just checking. The middle can move in waves.
Rough tiers beat perfect paralysis. You can refine scores forever. What you cannot do is concentrate scrutiny where it matters until you have decided, explicitly, that not everything matters the same amount.
Tiering, already built
The AI Governance Accelerator classifies every AI system across five dimensions, applies the most-restrictive rule automatically, and derives the required controls, reviews, and cadence for each tier. Ten deliverables and a working Excel engine.
Get the Toolkit →