Automation · Platform engineering
How to handle exceptions and retries in automated workflows
A workflow retries a failed step three times and creates three invoices, or it fails once and nobody finds out for a month. Here is how exception handling that does neither actually works, the prompts that build it, and what running it demands.
Built with Tray Headless
- Step Workflow step
- Step Sort the failure
- Step Retry with a wait
- Step Check before writing
- System NetSuite
Each failure is sorted first: a passing problem is retried safely, a real one is parked with its context and sent to the person who owns it.
The short answer
What is exception handling in an automated workflow?
Exception handling in an automated workflow sorts every failure into one of two kinds, a passing problem worth retrying and a real one a person has to fix, makes each retry safe to repeat, parks anything that still fails with the record and the reason attached, and alerts the person who owns that kind of problem. Most workflows get one of two things wrong. Retrying everything turns a bad record into duplicates, and retrying nothing turns a two-second outage into a missing order nobody notices.
Stage 6 of 7: Handle exceptions. Part of Process automation, end to end : every stage, the systems it runs on and the guide that builds it.
What matters here
- Sort the failure before reacting to it. A timeout and a rejected record need opposite responses.
- Retry passing problems with a growing wait, and a limit. Hammering a busy system makes it busier.
- Make every write safe to repeat: check whether it already happened before doing it again.
- Park what still fails, with the record and the reason, where someone can fix it and run it again.
- Alert the person who owns the problem, not a shared channel everyone has muted.
- Count parked items by reason. A reason that keeps coming back is a fix, not a chore.
Who this is for
You build or own automated workflows across business systems. Most runs succeed, the failures land in an email nobody reads, and the cost shows up later as a duplicate or a gap.
How it works in practice
What has to happen between a step failing and the work being done anyway.
- 1
The failure is sorted
A timeout, a rate limit or a brief outage is retried; a rejected record, a missing field or a permission error is not.
- 2
Passing problems are retried with a growing wait
A short wait, then a longer one, up to a limit, so a busy system has time to recover.
- 3
Each retry checks before it writes
If the earlier attempt actually succeeded, the retry finds the record and moves on instead of creating a second one.
- 4
What still fails is parked with its context
The record, the step, the error and what was already done, kept where it can be fixed.
- 5
The owner is alerted
The person who owns that kind of problem, with a link to the parked item and the reason in plain words.
- 6
Fixed items run again from where they stopped
Without repeating the steps that already succeeded.
What exception handling is made of
Four pieces, and the third is the one that makes retries safe.
A sort for every failure
Rules that say which errors are worth retrying and which need a person. Held in one place and shared by every workflow, so each one does not invent its own.
Retries with a growing wait
A limit on attempts and a wait that grows between them. A system that is down for a minute recovers; one that is down for an hour gets parked.
Writes that are safe to repeat
Before creating anything, look for it by the source record's ID. If it exists, update or skip. A retry should never be able to create a second invoice.
A holding queue with owners
Parked items with their record, step, error and history, each routed to the team that owns that kind of problem, with a way to run it again.
The Tray Headless prompts
Paste these into Claude Code or Codex with the Tray Headless plugin installed. Each stage runs on its own. The systems named in them are the worked example rather than a requirement, and every prompt says so.
Once per project, run
/tray-workflows:set-workspace
to pick the workspace these build in. Point it at a sandbox first.
- 1
Sort failures before handling them
A timeout and a rejected record need opposite responses.
Headless skills
build-workflowtray-patternsUse build-workflow and tray-patterns. The systems in play are Salesforce, NetSuite, Slack, Jira and Snowflake, or whatever we run in those seats. Set every step that calls another system to manual error handling, and send its errors to one shared sorting step. Sort each error into: Retry: timeouts, rate limits, brief outages, locked records Fix: rejected or invalid data, missing required fields, a record that does not exist, a permission error Keep the sorting rules in one place that every workflow uses, so a new error type is sorted once for all of them.
Treat an error you have not seen before as one to fix, not one to retry. An unknown error retried five times is five chances to make it worse.
- 2
Retry passing problems with a growing wait
Hammering a busy system makes it busier.
For errors sorted as retry, wait and try again: a few seconds, then longer each time, up to five attempts or the limit the target system asks for. If the system says how long to wait, use that. Stop retrying when the limit is reached and park the item. Never retry in a loop with no limit; a system that is down for an hour will otherwise receive thousands of calls when it comes back.
- 3
Make every write safe to repeat
A retry should never create a second invoice.
Headless skills
tray-gotchasUse tray-gotchas. Before any step that creates a record, look for it first by the source record's ID (store that ID on the record you create). If it exists, update it or skip; only create it if it does not. Do this for every create, not only the ones that have failed before. A call that timed out may still have succeeded on the other side, and the retry is the moment you find out. Record on each run which steps completed, so a run that is resumed skips them.
- 4
Park what still fails, with its context
A failure nobody can see is a gap nobody fixes.
When an item is sorted as fix, or runs out of retries, park it in a holding table in Snowflake with: the source record, the workflow and step, the error in plain words, every step that already completed, and a link to the run. Create a Jira ticket for the team that owns that kind of problem (finance for an invoice rejected by NetSuite, sales operations for a Salesforce record missing a required field), and post the ticket in that team's Slack channel with the reason. Group repeats. Fifty records failing for the same reason is one ticket with fifty items, not fifty tickets.
- 5
Run fixed items again from where they stopped
Fixing the data should be the only manual step.
Add a way to run a parked item again once someone has fixed the cause: a button on the Jira ticket or a Slack action. The run starts at the step that failed and skips the ones that completed. Close the ticket automatically when the item succeeds, and report by reason each week: how many items were parked, how long they waited, and which reasons keep coming back. A reason that recurs every week is a fix to the workflow or the data, not a chore for the owning team.
- 6
Validate it, then hand the sorting rules to the owners
Because new errors appear as systems change.
Run the per-step checks and the whole-workflow audit before this goes live. Test a timeout, a rate limit, a rejected record, a permission error, an error you have never seen, and a call that timed out but actually succeeded. Then open the same workflow in Tray Build so the owning teams can change the sorting rules, the retry limits and who gets alerted in the visual canvas without a deployment.
What it connects to
Failures are sorted inside the workflow, parked where they can be seen, and sent to the people who can fix them.
Salesforce
Be checked before writing, so a retried step finds the record an earlier attempt created.
Reads and writes
NetSuite
The same check before every create, so a retry never books a second order or invoice.
Reads and writes
Snowflake
Hold every parked item with its context and history, which is where recurring reasons show up.
Writes
Same build, other stacks
The design does not change if you run something else in one of these seats. The same prompts build it against Microsoft Dynamics 365, SAP S/4HANA, Google BigQuery, Microsoft Teams, ServiceNow or HubSpot.
Connections in this build
Field mapping, templates and common problems for each pairing: NetSuite + Salesforce, Jira + Salesforce, Salesforce + Slack and Jira + Slack.
Named systems are the ones most teams run, not the only ones that work. Each is an authentication in your Tray workspace, referenced by name, so the workflow never holds a credential. Where we have a connector page, the name links to it.
Running it in production
Exception handling is what decides whether an automation can be trusted with real work.
It runs on the platform, not on your laptop
Retries and alerts have to fire at three in the morning, when the outage that caused them actually happens.
One set of rules for every workflow
Sorting, retry limits and owners in one shared place, so a new workflow inherits them instead of reinventing them.
Every run is logged
Each step, input and error, so the question of what happened to a record has an answer.
Credentials are managed, never in code
The lookups before each write use the same managed credentials as the writes, with no extra access.
Owners see their own problems
Each team gets the failures it can fix, in its own channel, and the sorting rules are open in Tray Build.
Questions people ask
Which errors should be retried?
Passing ones: timeouts, rate limits, brief outages and locked records. Rejected data, missing fields and permission errors will fail the same way every time, so they go to a person.
How do retries create duplicates?
A call can time out after it has already succeeded on the other side. Retrying it without checking creates the record a second time. Looking for the record first, by the source record's ID, prevents that.
Where should failed items go?
To a holding queue with the record, the step, the error and what already completed, routed to the team that owns the problem, with a way to run it again once fixed.
How is this different from data quality monitoring?
That guide checks data that has already landed. This handles a workflow step that failed while it was running, so the work still gets done.
Which metric matters most?
Parked items by reason, week over week. A reason that keeps coming back is the next fix to make.
Related guides
Finance
How to automate approval routing
Hold who approves what as data, send each approver the facts they need where they work, set a deadline that escalates, and keep the record. The prompts.
Data operations
How to build data quality monitoring
Test what breaks decisions, alert the owner rather than a channel, and say what depends on a failure. The Headless prompts that build it.
Platform engineering
How to build webhook fan-out
Acknowledge before you fan out, verify the signature first, retry per consumer, and make replay possible. The Headless prompts that build it.
Last reviewed October 2026.