Automation · Platform engineering
How to build error alerts for customer integrations
A customer's sync breaks on Tuesday, and they find out on the following Monday when a report is wrong. How to catch every failed run across every customer, the prompts that sort and route them, and how to tell the customer before they notice.
Built with Tray Headless
- System Customer instance
- Step Catch the failure
- Step Who can fix it?
- Step Group repeats
- System Zendesk
Every failed run is sorted by who can fix it. The customer hears about what they can fix, your team about the rest, and repeats are grouped.
The short answer
How do you alert on errors in customer integrations?
Error alerts for customer integrations come down to four things: one place that catches a failed run from any customer's instance, a sort that decides who can fix it (the customer, for a broken connection or a missing field; your team, for a bug or an outage at the other app), grouping so fifty customers hit by one outage open one incident rather than fifty tickets, and a message to the customer in your product before they notice. With Tray Embedded the catch is a solution alert workflow that receives errors from every instance. Most teams route everything to support. Then support relays the customer's own fix back to them, a day late.
Stage 5 of 7: Errors reach whoever can fix them. Part of Embedded integration, end to end : every stage, the systems it runs on and the guide that builds it.
What matters here
- Catch failures in one place, for every customer. Checking instances one at a time doesn't scale past ten.
- Sort by who can fix it. A revoked connection is the customer's to fix, a bug is yours.
- Group repeats. One outage at the other app should open one incident, not fifty tickets.
- Tell the customer in your product, with the fix. They should hear it from you before their data is wrong.
- Keep the run that failed. Support can't answer "what happened to my record" without it.
Who this is for
You run the integrations your product offers its customers. Failures reach you as support tickets, usually after the customer has found the gap in their own data.
How it works in practice
From a customer's run failing to the right person knowing what to do about it.
- 1
A run fails in a customer's instance
A step error, a refused call, a timeout, a record the other app won't accept.
- 2
The failure reaches one alert workflow
With the customer, the instance, the step and the error, from any instance of any integration.
- 3
It is sorted by who can fix it
Customer: connection, permission, missing field, bad value. Your team: a bug, a change at the other app, an outage.
- 4
Repeats are grouped
The same error from the same customer once an hour is one alert. The same error across many customers is one incident.
- 5
The customer hears first, with the fix
A banner in your product and an email, saying what failed and the one thing to do.
- 6
Your team gets what only it can fix
A ticket or incident with the run attached, sent to whoever owns that integration.
What integration error alerts are made of
Four parts. The second decides whether support is a relay or not.
One catch for every instance
A single alert workflow that receives failures from every customer's instance of every integration, with the context to act.
A who-can-fix-it sort
Rules over the error and the step that put each failure with the customer or with your team, and a short list for anything unclear.
Grouping
Repeats from one customer folded into one alert, and the same failure across customers raised as one incident.
Messages with the fix
To the customer in your product, in plain words, with the button that fixes it: reconnect, open setup, retry.
The Tray Headless prompts
Paste these into Claude Code or Codex with the Tray Headless plugin installed. Each stage runs on its own. The systems named in them are the worked example rather than a requirement, and every prompt says so.
Once per project, run
/tray-workflows:set-workspace
to pick the workspace these build in. Point it at a sandbox first.
- 1
Catch every failure in one place
Checking instances one at a time doesn't scale.
Headless skills
build-workflowUse build-workflow. Build one solution alert workflow in Tray Embedded that receives failures from every instance of every integration we offer. The systems in play are Zendesk, Slack, Datadog and SendGrid, or whatever we run in those seats. For each failure, gather: the customer, the instance, the integration, the workflow and step, the error text, the record involved if any, and a link to the run. If we already send logs to Datadog, send these there as well, so the integration failures sit next to the rest of our product's errors.
- 2
Sort by who can fix it
So support stops relaying the customer's own fix.
Headless skills
build-workflowtray-gotchasUse build-workflow and tray-gotchas. Sort each failure: Customer can fix: expired or revoked connection, missing permission, a mapped field that was deleted, a value their app refuses, a limit on their plan We must fix: an error in our own workflow, a change at the other app that breaks every customer, a timeout or outage Keep the rules in a table we can edit. Anything that matches no rule goes to our team, and the table gets a new row once someone decides.
- 3
Group repeats into one alert or incident
One outage should open one incident.
Fold the same error from the same instance into one alert per day, with a count, rather than one per run. If the same error appears across more than a handful of customers of one integration within an hour, open one incident instead: in Slack, to the owner of that integration, with the customers affected and the first failing run. While an incident is open, tell affected customers once that we know and are fixing it, and don't send them the per-run alerts.
- 4
Tell the customer first, with the fix
They should hear it from you.
For failures the customer can fix, show a banner on the integration in our product and email their admin: what stopped, which records are waiting, and one button that fixes it (reconnect, open setup, or retry). For failures we must fix, open a Zendesk ticket against the customer's organisation with the run attached, so support starts from the run rather than from a question. When the next run succeeds, clear the banner and close the alert.
Customers trust an integration that tells them it broke far more than one that seems never to break and is quietly wrong.
What it connects to
The failure comes from the instance. Who hears about it depends on who can fix it.
Zendesk
Open a ticket against the customer's organisation for failures your team must fix, with the run attached.
Writes
Slack
Raise one incident to the integration's owner when many customers hit the same failure.
Writes
Same build, other stacks
The design does not change if you run something else in one of these seats. The same prompts build it against Microsoft Teams, Jira, Google Chat, ServiceNow, Jira Service Desk or Freshservice.
Connections in this build
Field mapping, templates and common problems for each pairing: Slack + Zendesk, Datadog + Slack and SendGrid + Zendesk.
Named systems are the ones most teams run, not the only ones that work. Each is an authentication in your Tray workspace, referenced by name, so the workflow never holds a credential. Where we have a connector page, the name links to it.
Running it in production
Customers judge the integration by what happens when it breaks.
Every failure is caught
One alert workflow for every instance, so nothing depends on a customer noticing.
The customer hears first
A banner and an email with the fix, before the gap shows up in their own reporting.
One outage is one incident
Grouping keeps your team on the fix and the customer informed once, rather than buried in alerts.
The run is always attached
Support starts from what happened, step by step, rather than from asking the customer to describe it.
Questions people ask
Should customers see integration errors?
The ones they can fix, yes, in plain words with the button that fixes it. The ones only you can fix, tell them once that you know and are on it.
How do you stop alert floods during an outage?
Group the same failure across customers into one incident, tell affected customers once, and hold the per-run alerts until it is resolved.
Can integration errors go to Datadog?
Yes. Send them from the alert workflow, or stream logs there, so integration failures sit with the rest of your product's errors.
Who should own each integration's alerts?
A named person or team per integration. An incident that goes to a shared channel waits until somebody decides it is theirs.
Related guides
Platform engineering
How to let customers connect their own accounts
Let customers authorise their own apps inside your product, under your brand, with expired connections caught and fixed before a sync stops. The prompts.
Platform engineering
How to roll out integration changes to every customer
Ship a new version of a customer-facing integration to every customer: test it as a customer first, release in waves, and ask for setup again only when it must. The prompts.
AI operations
How to build an agent observability pipeline
Capture whole agent runs instead of single calls, join each one to the outcome it produced, and alert on the failures that return an answer anyway. The prompts.
Last reviewed October 2026.