A pilot that impressed the room and then quietly died usually failed on four things, and the model is not one of them. The data it ran on was a copy. The credentials it ran under belonged to a person. Nobody owned the items it could not finish. And nobody had written down what it was allowed to change.
Each of those is fixable. All four are cheaper to fix before the build than after, which is the practical reason we spend the first half hour on the workflow rather than on the technology.
The demo data was clean because somebody cleaned it
Every pilot dataset is a selection. Somebody exported a few hundred tickets, or a folder of invoices, or a list of leads, and they picked the ones that made the point. That is not dishonest. It is how you get a demo built in a fortnight.
Production is the rest of the distribution. The ticket that arrived through a channel nobody remembered was connected, so the subject line is empty. The invoice from the supplier who bills three purchase orders on one PDF. The lead where the company name is in the notes field because the form that captured it was retired two years ago and nobody migrated the mapping.
The pilot never met those, so nobody costed them. When they turn up in week one of production they present as a broken agent, and the project acquires a reputation problem it does not deserve.
What helps is unglamorous. Before we build, we pull a real sample from the live system, including the ugly tail, and count the shapes. Not to fix them. To decide which shapes the agent handles, which it refuses and routes, and which we agree it will never see because a person filters them upstream. That decision belongs to the client’s process owner, and it has to be made with the real distribution in front of them.
It ran as a person, and production will not allow that
The most common technical reason a pilot cannot be promoted is that it was built under a human login, usually an administrator’s.
In Salesforce, this shows up immediately. A query written while the builder was an administrator returns everything. The same query, scoped properly for a service context, enforces object permissions and field-level security, and it either returns less than the agent expected or fails outright. Salesforce’s own guidance for agent actions now points at running queries and data changes in user mode for exactly this reason, and at scoping the integration user or connected app to the minimum each action needs. That guidance is right. It is also the thing that breaks demos, because the demo was never scoped.
The same pattern shows up elsewhere with different vocabulary. A helpdesk integration built against an admin token will happily update any ticket field, merge tickets, and read every organisation record. A production service account should not be able to do most of that, and the moment you narrow it, you discover which reads the agent had been quietly relying on.
We treat the permission set as a deliverable rather than a configuration step. It gets drafted early, reviewed by whoever owns access in the client’s business, and cut down until it is narrow. What the agent cannot do is a more useful document than what it can.
Nothing owned the cases it could not finish
In a demo, every item resolves. That is the point of a demo.
In production, a predictable share of items cannot be finished by the agent, and they are not evenly interesting. An invoice where the price is above the purchase order is a procurement question. An invoice where the quantity does not match the goods receipt is a receiving question. A ticket where the customer is asking for money back is a policy question. Those go to three different people, and they have three different acceptable waiting times.
The failure mode is a single queue called Exceptions that nobody owns. It fills up. Within a month it is a second manual process, staffed by whoever feels guiltiest, and the original problem is now worse because there is a new place for work to hide.
So the exception path gets designed with the same care as the happy path. Separate the exception types by cause, not by severity. Name an owner for each type. Make the age of each item visible. Attach the source document and the action the agent proposed, so the reviewer is deciding rather than investigating. And let the reviewer’s decisions feed back into the rules, because the same three exception types will account for most of the volume and two of them are usually fixable upstream.
This is the part of the work that decides whether the thing is still running in month four. It is also the least demonstrable part, which is why pilots skip it.
Nobody had agreed what it was allowed to change
The last one is a governance gap dressed as a technical gap.
Ask three people in the room which writes the agent may make unsupervised and you will get three answers. The Head of Support thinks a refund under a certain value is fine. Finance does not. Nobody has asked Legal. So the question stays open, and it surfaces two days before go-live as an objection that stops everything.
The way through is to settle it early and narrowly. We write down every action the agent can take, then sort them by what happens if the agent is wrong and how hard the mistake is to reverse. Reading a record is free. Setting a routing tag is reversible in seconds. Sending an email to a customer is not reversible at all. Changing a payment fact is reversible and expensive and someone will ask who authorised it.
A lot of published advice suggests gating this on a model confidence score. We do not. Those scores are poorly calibrated, and in any case an agent’s certainty says nothing about the size of the mistake it is about to make. The consequence is what decides. So anything that can change money or a customer’s position waits for a person, regardless of how certain the agent claims to be. The agent drafts, routes, and stops.
When we tell a client not to do it
Some workflows should not get an agent yet, and the assessment is where that gets said rather than discovered later.
If nobody in the business can state the decision rule, an agent cannot execute it. Not because the technology is weak, but because there is nothing to implement. That situation needs a process decision first, and we will say so.
If the system holding the work is being replaced inside the next few months, build after the migration. Wiring an agent into a tenant that is about to be decommissioned is spending the go-live twice.
If the volume is small enough that one person absorbs it between other tasks, a person is cheaper and the monthly retainer will not pay for itself. We would rather say that on a free call than invoice for it.
What we actually do about it
We put one supervised agent on one existing workflow, inside the systems the client already runs, and a person still signs anything that can change money or a customer. The sequence is a look at the real process, a map of inputs and systems and approvals, a build wired into the client’s own tools, and a monthly run with a named owner from our employed team.
The four gaps above are the reason the first step is a free thirty-minute assessment rather than a proposal. Half an hour on a real queue tells us whether the data, the permissions, the exception owner and the write scope can be settled. If they cannot, that is a finding, and the client keeps the workflow map either way.
If you want the sequence in more detail before booking, it is set out on how we work. If you already know which queue is the problem, the three we start from are customer support, sales operations and document processing.