Deciding which agent actions need a person

Sort the actions by what happens if the agent is wrong and how hard the mistake is to undo. Confidence scores are the wrong gate.

Sort every action the agent can take by two questions. What happens if it is wrong, and how hard is the mistake to undo. Then require a person on the actions where the cost of being wrong exceeds the value of not waiting. That is the whole method, and it deliberately does not involve the model’s confidence score.

Confidence is the popular answer and it is a poor one. Models are systematically miscalibrated: a high stated certainty is not reliable evidence of being right, and worse, the errors that survive to production are often the confident ones. A confident mistake on a refund is still a refund. Confidence scores are useful for deciding which items to sample when you are reviewing quality. They are not a control.

Four bands, in order of consequence

Reading. The agent looks something up. Wrong answers are free, provided nothing acts on them automatically. No approval. Log it and move on.

Reversible internal writes. Setting a routing tag, adding an internal note, populating a field nobody outside the company sees. If the team can put it back in seconds and nobody outside notices, this runs unsupervised. Most of the volume in a well-scoped first workflow lives here, which is why the first release can feel underwhelming and still remove most of the work.

Anything that leaves the building. An email to a customer, a message in a shared channel with a supplier, a document sent out. These are not reversible in any meaningful sense, because the recipient has already read it. Our default is that the agent drafts and a person sends, and that default does not change on the strength of good pilot results.

Anything that changes money, a legal position, or a permission. Refunds, credits, discounts, payment releases, contract changes, bank detail updates, access grants. A person signs, every time, with their name against it.

The line that matters is between the second and third band. Almost every argument about scope is really an argument about where a specific action sits, and having the four bands written down turns that into a ten-minute conversation with the process owner instead of a standoff two days before go-live.

Blast radius is a separate axis

The same action is a different risk at a different scale. Updating one record is not the equivalent of updating four hundred in a batch, even though the action class is identical.

So cap the batch size on any action that touches more than one record, and treat crossing the cap as its own approval event. This catches a specific and expensive failure: a rule that is slightly wrong applied uniformly at speed. One wrong record is a correction. Four hundred wrong records is a data recovery exercise and an awkward conversation about who authorised it.

Rate is worth capping too. An agent that suddenly starts acting far more often than its normal pattern is usually reacting to something upstream that has changed, and stopping it to ask is cheaper than reading the aftermath.

What the approval actually looks like

An approval request that says “the agent wants to update this record, approve or reject” is not an approval, it is a coin toss with a signature attached. The reviewer needs enough to decide without going and looking.

Ours carry the intent in plain words, the exact values that will be written, the evidence the agent used, a risk tag, and a link to the full trace for anyone who wants it. The reviewer can approve, reject, edit and approve, or ask for more information. Edit and approve matters more than it looks: it is where most of your improvement signal comes from, because a reviewer changing one field repeatedly is telling you precisely which rule is wrong.

Two implementation details that are easy to get wrong and painful to discover later.

First, bind the approved values on resume. If the agent re-reasons after approval and produces slightly different values, you approved one thing and executed another. The approved payload is what runs, full stop.

Second, put an idempotency key on every outbound action. Approvals get clicked twice, queues get replayed, processes get restarted. An approved refund must not be capable of firing twice.

The run itself has to checkpoint durably and release its compute while it waits. An approval can sit overnight, or over a weekend, or until someone comes back from leave, and nothing should be holding a process open in the meantime.

Timeouts, and what happens when nobody answers

Every approval needs an expiry and an explicit behaviour on expiry. The default should be to deny and notify, never to approve. An automation that quietly proceeds because nobody looked is the worst of both designs: you carry the risk and you have a record showing a person was supposed to have checked.

Alongside that, decide the escalation route before go-live. Who is the backup reviewer. What happens when the named reviewer is on leave. Whether the clock respects business hours, which it usually should, because a refund approval waiting at three in the morning is not an incident.

Approval fatigue is the real failure mode

The design that fails most often is not too little supervision. It is too much, in the wrong places.

If a reviewer is clearing approvals faster than they could have read them, they are not reviewing. They are clicking, and you have built a rubber stamp that produces an audit trail implying scrutiny that did not happen. That is worse than no approval step, because it is misleading.

So measure the review. Specifically, measure the rate at which reviewers change or reject anything. If nothing is ever changed on a class of action, you have a decision to make: widen the automation for that class, or narrow the review to a sample. Leaving it as a mandatory approval that always says yes is the option that quietly costs you a person’s attention forever.

The corollary is that supervision has a real price, and it belongs in the business case. If the review load on a workflow is larger than the work the agent removed, the scope is wrong. That is a legitimate reason to walk away from an automation, and we would rather find it on a call than in month three.

Widen slowly, and by class

Start narrow. Give the agent read access and reversible internal writes. Watch it on real items. Then widen one action class at a time, with the reroute rate or the correction rate as the evidence, not the general feeling that it has been going well.

Good pilot results do not buy an agent the right to send external messages or move money. Those stay behind a person because of what they are, not because of how the agent has been performing.

A note on obligations, carefully

For organisations in Europe running systems that fall into the high-risk category under the EU AI Act, the deployer duties in Article 26 point in the same direction as the design above: human oversight assigned to named people who have the competence and the actual authority to intervene, and retention of automatically generated logs under your control for at least six months.

Three caveats, because this area is routinely oversold. Most support, sales operations and document workflows are unlikely to be high-risk under the Act’s categories, so the duty often will not apply. The application timetable has been amended during 2026 and the dates should be checked rather than assumed. And no vendor can take these duties off you by contract, so treat anyone claiming to make you compliant with suspicion.

The reason to design this way is that it is the right way to run an agent against systems that matter. The evidence it produces is a by-product, and a useful one if a regulator, an auditor or an insurer ever asks.

Where we draw the line

A person still signs anything that can change money or a customer. The agent reads the item, checks what it is allowed to check, takes the approved action and writes back. That boundary is the same on all three workflows we start with, and it is described on the customer support and document processing pages in the terms each team uses.

The commercial shape, one workflow live and then a monthly run with a named owner, is on the pricing page. If you want to argue about where a specific action sits, that is a good use of a free thirty-minute assessment.