One governance standard applied uniformly to every agent is a common way these programs stall. A read-only summarizer waits behind the same review board as a payment agent, so teams route around the board. Tier the autonomy, then tier the individual action, and let regulatory classification and data sensitivity cap the result rather than dissolve into it.
The model drafts, suggests or summarizes. Nothing leaves the session and nothing is written anywhere. A person types the final action themselves.
The agent calls read APIs and search, chains several steps and returns an answer with citations. It cannot change any system state.
The agent prepares a concrete change: a ticket, a draft email, a config diff, a code pull request. A named person approves before anything lands.
The agent writes to systems without per-action approval, but only within a declared envelope: named tools, value limits, tenant scope, time window and step budget. Anything outside the fence escalates.
The agent owns an outcome end to end, coordinates other agents and adapts its own plan. Reserve this for processes where the failure cost is understood, bounded and insurable.
A single customer-service agent can do four things of very different consequence in the same conversation. Registering one tier against the agent and stopping there hides that. Score the use case, then score each action class inside it.
Look up the policy terms, summarize the claim history, answer a coverage question with citations. Nothing changes.
Ceiling up to R3 · reversible, no state change, entitlement-filtered reads
Compose the customer reply, prepare the case note, assemble the settlement recommendation. A person still sends or commits it.
Ceiling up to R3 · reversible before commit, review is the control
Update an address, reschedule an appointment, reopen a case, apply a small goodwill credit. Real change, contained and undoable.
Ceiling R2, or R3 with a proven envelope · needs compensation and rollback
Issue a large payment, terminate a service, decline a claim, send an external communication. Money, entitlements or reputation.
Ceiling R2, or R3 only after independent review · approval and dual control
These do not average into the score. They sit on top of it, and any one of them can lower the ceiling the autonomy tier would otherwise allow.
If the process falls under a high-risk regime, the obligations attach regardless of how reversible the action looks. Determine whether you act as provider or deployer before you scope controls, because the duties differ.
Special-category, payment and regulated data raise the floor on isolation, retention and evidence even for a read-only action. An A1 lookup over sensitive records is still an A1 action with A3 data handling.
Where an outcome changes someone's money, access, employment or care, add outcome monitoring across relevant segments and a route to challenge, whatever the technical reversibility says.
An error nobody would notice without an audit deserves more control than an equally costly error that is obvious within the hour. Silent failure modes are systematically under-governed.
If you are standing up an agent platform, tightening the controls on one you already have, or preparing for an audit that now includes agents, I am happy to look at it with you.
Tell me where you are with agents and what you are trying to make safe. I reply to every enquiry within two business days.