Checking an agent on real items before it is allowed to write

Read-only first, then shadow mode against decisions a person already made, then a graduated write scope. A sandbox hides the parts that break you.

Do not give an agent write access because the demo classified fifty clean tickets. Give it write access after it has seen the live tail, under the permission set production will actually grant, and after a person has looked at the places it disagrees with history.

We check on real items before go-live. That line on how we work is the whole method below, not a slogan.

Read-only on the live queue

The first run should change nothing. Same integration user you will use later, same objects, same field-level security. If a query that worked for an administrator now returns fewer rows, that is the finding. Fix the permission set or drop the action. Do not widen the user to make the demo green again.

Pull a sample that includes the ugly tail: empty subjects, three purchase orders on one PDF, the company name sitting in a notes field. Count the shapes. Decide which the agent will handle, which it will refuse, and which a person will filter upstream. That decision belongs to the process owner, with the real distribution in front of them.

Shadow mode

Once reads work, let the agent propose on items a person has already finished. Do not tell it the answer. Compare.

A disagreement is not automatically an error. The person may have been routing by habit. The agent may be reading a field the person ignores. Split the disagreements by cause before you keep score: missing data, a rule nobody wrote down, a restricted picklist the agent guessed, a historic decision that was itself wrong.

What you want from shadow mode is a list of exception types and a list of writes you are willing to allow. Accuracy percentages from a vendor deck do not survive contact with that list, so we do not collect them as a go-live gate.

Graduated writes

Routing tags, an internal note, a proposal field: reversible in seconds. Do those first, on a slice of the queue, with a human still watching the first week.

Public comments, emails, stage changes, payment facts: later, and only if the reversible path has become boring. Anything that can change money or a customer waits for a person in the first release. That is the package, not a temporary precaution.

Duplicate webhook delivery belongs in this phase too. The same ticket-created event will arrive twice. Key the run on the record and the event, and make the write safe to repeat, or you will get two internal notes and a confused team.

What a sandbox hides

A sandbox is a copy from a date. Permissions have usually been loosened so the project could move. The ugly tail was often not refreshed. Automations that exist only in production are missing, and automations that exist only in the sandbox will surprise you the other way.

Use the sandbox to develop. Sign off on a sample from production, under production permissions, even if those writes are still mocked. The assessment is where we look at whether that sample can be obtained. If it cannot, we should not be talking about a go-live date.