The happy path is what gets demoed. The exception path is what you are still paying for in month four.
A predictable share of items cannot be finished by the agent. They are not equally interesting, and they do not belong to the same person. An invoice over the purchase order is procurement. A quantity that does not match the goods receipt is receiving. A ticket asking for money back is policy. Dump them into one list called Exceptions and you have built a second manual process, staffed by whoever feels guiltiest.
We design the exception path with the same care as the write path. That is a large part of what $2,900 a month is for.
Separate by cause, not by severity
Severity feels like management. Cause is what you can fix.
For documents, typical causes are: the join key is missing (no purchase order number), the join key is present and the values disagree, the document is the wrong type for this queue, the extract is unreadable, and the vendor is a duplicate of a bill you already paid. Those five want five owners, or at least five filters the same owner can work in order.
For a helpdesk, the useful split is usually: the agent was not sure, the agent was sure and a human later moved the ticket, the record was missing a field the rule needs, and the customer asked for something the agent is forbidden to do. The last one should never have been in the automated slice.
Give each type a name your team already uses. If they say “price variance” and you name the queue “Type B”, nobody will open it.
Name an owner, and make age visible
An exception type without a named human is a folder. Write the name on the type, with a deputy for leave. Put the age of each item on the list, in working hours, not in a hidden timestamp.
Agree how old is too old before go-live. A price variance that sits for a week is a different operational failure from a mis-tagged ticket that sits for an hour. Do not use one SLA for every type.
The reviewer should be deciding, not investigating. Attach the source (PDF, ticket thread, CRM snapshot), the values the agent read, and the action it proposed. If they have to go hunting in the ERP, the packet is wrong.
Feed decisions back into the rules
The same two or three types will dominate the volume. Two of them are usually fixable upstream: a supplier who never puts the PO on the invoice, a form that writes the company name into notes, a macro that still fights the agent on priority.
Once a week, look at the types that moved. Change the rule, the extract, or the upstream filter. Then shrink the slice the agent is allowed to see, or widen it, based on that, not based on a confidence score.
Confidence scores are a sampling tool. They are a poor control. An item can be confidently wrong. Consequence decides whether a person is in the way, which is how we already gate writes.
The failure mode, named
One queue. No owner. No age. No evidence. A coordinator copies from it the way they used to copy from the inbox. Leadership thinks the agent is live. The team knows they now have two inboxes.
If that is what your last pilot produced, the next conversation is about exception design, not about a better model. Bring a sample of the current pile to the assessment. If the pile has no rule, we will say the workflow is not ready, which is the same finding as the document piles that are not ready.