Automation · AI operations
How to test an AI agent before release
Someone changes one line in the agent's instructions on Friday. On Monday it is closing tickets it used to escalate, and nobody can say which change did it. Here is the model behind testing an agent before each release, the prompts that build it, and what it takes to run in production.
Built with Tray Headless
- System Change to the agent
- Step Run the test set
- Step Score actions and answers
- Step Compare with last release
- Step Release or block
Every change runs against the same set of real requests, and a release only goes out when each category clears its own bar.
The short answer
What does testing an AI agent before release mean?
Testing an agent before release takes four things: a test set built from real requests, including the ones it must refuse; scoring of what the agent did, not only what it said; a pass bar for each kind of request rather than one average; and a run on every change, compared with the last release. Teams most often go wrong by scoring only the reply. An agent can write a perfect answer while calling the wrong tool, and the reply is the part nobody needed to worry about.
Stage 3 of 5: Test it before every release. Part of AI agent deployment, end to end : every stage, the systems it runs on and the guide that builds it.
What matters here
- Build the test set from real requests, with the answer a person would accept for each. Invented examples test what you already thought of.
- Score the actions: which tool was called, with which inputs, and whether it should have asked or escalated instead.
- Set a pass bar per category. A 95% average can hide a category where the agent fails half the time.
- Run the full set on every change: instructions, tools, knowledge or model. Any of them can break something unrelated.
- Add every production failure to the test set, so the same mistake cannot ship twice.
Who this is for
You own an AI agent that is live or close to it. Changes are made to its instructions, tools or model, and nobody can say for sure whether a change made it better or worse.
How it works in practice
What happens between someone changing the agent and that change reaching users.
- 1
A change is proposed
New instructions, a new tool, a different model, or a refresh of the knowledge it searches.
- 2
The test set runs against the changed agent
Every case, with tools pointed at a test environment so nothing real changes.
- 3
Each case is scored on actions and on the answer
Did it call the right tool with the right inputs, did it refuse or escalate when it should, and is the reply correct.
- 4
Results are compared with the last release
Case by case, so a fix in one place that breaks another shows up.
- 5
Each category must clear its bar
Refusals and escalations carry the highest bar, because those failures reach people.
- 6
The result goes to the agent's owner
Release, or block with the cases that failed and what changed.
What a release test is made of
Four parts. The second is where most testing falls short.
A test set from real requests
Sampled from production or the pilot, grouped by category, with the accepted outcome written down. It includes requests the agent must refuse or hand to a person.
Scoring that checks actions
The tool called, its inputs, and whether the agent should have asked first. Then the reply, checked against the accepted answer by a fixed rubric.
A bar per category
Each kind of request has its own pass rate. Refusals and escalations get the strictest one.
A run on every change
Instructions, tools, knowledge and model changes all trigger the same run, compared with the last release that passed.
The Tray Headless prompts
Paste these into Claude Code or Codex with the Tray Headless plugin installed. Each stage runs on its own. The systems named in them are the worked example rather than a requirement, and every prompt says so.
Once per project, run
/tray-workflows:set-workspace
to pick the workspace these build in. Point it at a sandbox first.
- 1
Start here: build the test set from real requests
The test set decides what you will catch.
Headless skills
build-workflowUse build-workflow. Pull the last 90 days of requests the agent received from Snowflake. Group them by category, such as order status, refunds, access requests, and anything it handed to a person. Take a sample from each group, including the rare ones. For each case, store: The request, exactly as the user wrote it The accepted outcome: the tool that should be called and its inputs, or that it should refuse, ask a question, or hand off The accepted answer, in a sentence a reviewer agreed with Add cases the agent must refuse: requests from the wrong team, amounts over its limits, and attempts to make it ignore its instructions. Have the agent's owner sign off on the accepted outcomes before anything is scored against them.Fifty well-chosen cases per category catch more than a thousand random ones. Keep the rare categories in, since that is where agents tend to break.
- 2
Point the tools at a test environment
A test run must not change real records.
Headless skills
tray-gotchasUse tray-gotchas. Run every test case with the agent's tools connected to a sandbox or test instance of each system: a Salesforce sandbox, a ServiceNow test instance, a test Slack channel. Where a system has no test instance, replace the write tool with one that records what it would have done and returns a realistic response. Never run a test set against production write access.
- 3
Score what the agent did, then what it said
A perfect reply after the wrong action is a failure.
Score each case in two parts. Actions, checked exactly: Was the right tool called, or correctly no tool at all Were the inputs right: the record, the amount, the field Did it ask, refuse or hand off when the accepted outcome says it should Answer, checked against the accepted answer with a fixed rubric: Correct facts, nothing invented, and a source where one is expected Tells the user what happened or what to do next A case passes only if both parts pass. Store the full trace of every run, not just the score, so a failure can be read.
- 4
Set a pass bar per category, and compare with the last release
Averages hide the failures that matter.
Set a pass bar for each category, agreed with the agent's owner. Give refusals and hand-offs the highest bar, because an agent that should have stopped and did not is the failure that reaches a customer. Compare each run with the last release that passed, case by case: Cases that newly fail Cases that newly pass Categories that dropped below their bar Block the release if any category is under its bar or any case that used to pass now fails, unless the owner accepts it with a reason that is recorded.
- 5
Run it on every change and send the result
The change that breaks it is rarely the one you would suspect.
Trigger the full run whenever the agent's instructions, tool list, tool definitions, knowledge sources or model change. Send the result to the agent's owner in Slack: pass or block, the pass rate per category against its bar, and the cases that changed, each linked to its trace. Re-run the test set weekly even with no change, because the model provider and the knowledge it searches can change underneath it.
- 6
Feed production back in, then hand it over
So the same mistake cannot ship twice.
Every time a person corrects the agent in production, or a hand-off turns out to have been needed, add that request to the test set with the accepted outcome. Run the per-step checks and the whole-workflow audit, then open the same workflows in Tray Build so the agent's owner can add cases, change bars and read results in the visual canvas.
What it connects to
Real requests in, a test run against safe copies of each system, and a clear release decision out.
Snowflake
Supply the real requests the test set is built from, and keep every run's scores and traces.
Reads and writes
Salesforce
A sandbox the agent's tools run against during testing, so nothing real changes.
Reads and writes
ServiceNow
A test instance for tools that update tickets, with the same fields as production.
Reads and writes
GitHub
Hold the agent's instructions and tool definitions, so every change has a version to test and to roll back to.
Reads
Same build, other stacks
The design does not change if you run something else in one of these seats. The same prompts build it against Microsoft Dynamics 365, Google BigQuery, Microsoft Teams, Jira, HubSpot or Databricks.
Connections in this build
Field mapping, templates and common problems for each pairing: ServiceNow + Salesforce, GitHub + Salesforce, Salesforce + Slack and ServiceNow + Slack.
Named systems are the ones most teams run, not the only ones that work. Each is an authentication in your Tray workspace, referenced by name, so the workflow never holds a credential. Where we have a connector page, the name links to it.
Running it in production
A release test is only useful if it runs every time and blocks when it should.
Tests never touch production
Write tools point at sandboxes or record what they would have done. A test that can change real records will eventually change one.
Every run keeps its trace
The tools called, the inputs, the reply and the score, so a failed case can be read rather than guessed at.
The bar is set by the owner
Pass bars, categories and accepted outcomes open in Tray Build, so the person accountable for the agent sets the standard it is held to.
A block is a decision, not a warning
A release under the bar does not go out unless the owner accepts it, and that acceptance is recorded with the reason.
Production feeds the test set
Corrections and missed hand-offs become new cases, so the test set gets harder as the agent gets better.
Questions people ask
Why score actions and not just replies?
Because an agent can write a correct reply after calling the wrong tool or changing the wrong record. The action is what affects people, so it is checked first and exactly.
How big should the test set be?
Around fifty cases per category is usually enough, as long as it includes the rare requests and the ones the agent must refuse or hand off. Coverage matters more than size.
What counts as a change that needs testing?
Any change to the instructions, the tool list, a tool's definition, the knowledge sources or the model. Each of them can break something that looks unrelated.
Can a model judge the replies?
It can score the reply against a fixed rubric and an accepted answer. The actions should be checked exactly, by comparing the tool and inputs with the accepted outcome.
What happens to failures found in production?
They are added to the test set with the accepted outcome, so the same mistake is caught before the next release.
Further reading
Background on the same subject, for the case rather than the build.
Related guides
AI operations
How to build an agent observability pipeline
Capture whole agent runs instead of single calls, join each one to the outcome it produced, and alert on the failures that return an answer anyway. The prompts.
AI operations
How to scope an AI agent's tool access
Give each agent the few tools its job needs, act as the person asking or a narrow service account, split read from write, and review access on a schedule. The Headless prompts that build it.
AI operations
How to build a knowledge base to vector sync
Chunk on structure, carry permissions into the index, delete on delete, and re-embed only what changed. The Headless prompts that build it.
Last reviewed October 2026.