Skip to content

Automation  ·  Platform engineering

How to build error alerts for customer integrations

A customer's sync breaks on Tuesday, and they find out on the following Monday when a report is wrong. How to catch every failed run across every customer, the prompts that sort and route them, and how to tell the customer before they notice.

Built with Tray Headless

  1. System Customer instance
  2. Step Catch the failure
  3. Step Who can fix it?
  4. Step Group repeats
  5. System Zendesk
Also Tell the customer

Every failed run is sorted by who can fix it. The customer hears about what they can fix, your team about the rest, and repeats are grouped.

The short answer

How do you alert on errors in customer integrations?

Error alerts for customer integrations come down to four things: one place that catches a failed run from any customer's instance, a sort that decides who can fix it (the customer, for a broken connection or a missing field; your team, for a bug or an outage at the other app), grouping so fifty customers hit by one outage open one incident rather than fifty tickets, and a message to the customer in your product before they notice. With Tray Embedded the catch is a solution alert workflow that receives errors from every instance. Most teams route everything to support. Then support relays the customer's own fix back to them, a day late.

Stage 5 of 7: Errors reach whoever can fix them. Part of Embedded integration, end to end : every stage, the systems it runs on and the guide that builds it.

What matters here

  • Catch failures in one place, for every customer. Checking instances one at a time doesn't scale past ten.
  • Sort by who can fix it. A revoked connection is the customer's to fix, a bug is yours.
  • Group repeats. One outage at the other app should open one incident, not fifty tickets.
  • Tell the customer in your product, with the fix. They should hear it from you before their data is wrong.
  • Keep the run that failed. Support can't answer "what happened to my record" without it.

Who this is for

You run the integrations your product offers its customers. Failures reach you as support tickets, usually after the customer has found the gap in their own data.

How it works in practice

From a customer's run failing to the right person knowing what to do about it.

  1. 1

    A run fails in a customer's instance

    A step error, a refused call, a timeout, a record the other app won't accept.

  2. 2

    The failure reaches one alert workflow

    With the customer, the instance, the step and the error, from any instance of any integration.

  3. 3

    It is sorted by who can fix it

    Customer: connection, permission, missing field, bad value. Your team: a bug, a change at the other app, an outage.

  4. 4

    Repeats are grouped

    The same error from the same customer once an hour is one alert. The same error across many customers is one incident.

  5. 5

    The customer hears first, with the fix

    A banner in your product and an email, saying what failed and the one thing to do.

  6. 6

    Your team gets what only it can fix

    A ticket or incident with the run attached, sent to whoever owns that integration.

What integration error alerts are made of

Four parts. The second decides whether support is a relay or not.

One catch for every instance

A single alert workflow that receives failures from every customer's instance of every integration, with the context to act.

A who-can-fix-it sort

Rules over the error and the step that put each failure with the customer or with your team, and a short list for anything unclear.

Grouping

Repeats from one customer folded into one alert, and the same failure across customers raised as one incident.

Messages with the fix

To the customer in your product, in plain words, with the button that fixes it: reconnect, open setup, retry.

The Tray Headless prompts

Paste these into Claude Code or Codex with the Tray Headless plugin installed. Each stage runs on its own. The systems named in them are the worked example rather than a requirement, and every prompt says so.

Once per project, run /tray-workflows:set-workspace to pick the workspace these build in. Point it at a sandbox first.

  1. 1

    Catch every failure in one place

    Checking instances one at a time doesn't scale.

    Headless skills build-workflow

    Use build-workflow. Build one solution alert workflow in Tray Embedded
    that receives failures from every instance of every integration we
    offer. The systems in play are Zendesk, Slack, Datadog and SendGrid,
    or whatever we run in those seats.
    
    For each failure, gather: the customer, the instance, the integration,
    the workflow and step, the error text, the record involved if any, and
    a link to the run.
    
    If we already send logs to Datadog, send these there as well, so the
    integration failures sit next to the rest of our product's errors.
  2. 2

    Sort by who can fix it

    So support stops relaying the customer's own fix.

    Headless skills build-workflow tray-gotchas

    Use build-workflow and tray-gotchas. Sort each failure:
    
      Customer can fix: expired or revoked connection, missing
      permission, a mapped field that was deleted, a value their app
      refuses, a limit on their plan
      We must fix: an error in our own workflow, a change at the other
      app that breaks every customer, a timeout or outage
    
    Keep the rules in a table we can edit. Anything that matches no rule
    goes to our team, and the table gets a new row once someone decides.
  3. 3

    Group repeats into one alert or incident

    One outage should open one incident.

    Fold the same error from the same instance into one alert per day,
    with a count, rather than one per run.
    
    If the same error appears across more than a handful of customers of
    one integration within an hour, open one incident instead: in Slack,
    to the owner of that integration, with the customers affected and the
    first failing run.
    
    While an incident is open, tell affected customers once that we know
    and are fixing it, and don't send them the per-run alerts.
  4. 4

    Tell the customer first, with the fix

    They should hear it from you.

    For failures the customer can fix, show a banner on the integration in
    our product and email their admin: what stopped, which records are
    waiting, and one button that fixes it (reconnect, open setup, or retry).
    
    For failures we must fix, open a Zendesk ticket against the customer's
    organisation with the run attached, so support starts from the run
    rather than from a question.
    
    When the next run succeeds, clear the banner and close the alert.

    Customers trust an integration that tells them it broke far more than one that seems never to break and is quietly wrong.

What it connects to

The failure comes from the instance. Who hears about it depends on who can fix it.

Zendesk

Open a ticket against the customer's organisation for failures your team must fix, with the run attached.

Writes

Slack

Raise one incident to the integration's owner when many customers hit the same failure.

Writes

Datadog

Receive every integration failure alongside the rest of your product's errors.

Writes

SendGrid

Email the customer's admin about a failure they can fix, with the fix in one click.

Writes

Same build, other stacks

The design does not change if you run something else in one of these seats. The same prompts build it against Microsoft Teams, Jira, Google Chat, ServiceNow, Jira Service Desk or Freshservice.

Connections in this build

Field mapping, templates and common problems for each pairing: Slack + Zendesk, Datadog + Slack and SendGrid + Zendesk.

Named systems are the ones most teams run, not the only ones that work. Each is an authentication in your Tray workspace, referenced by name, so the workflow never holds a credential. Where we have a connector page, the name links to it.

Running it in production

Customers judge the integration by what happens when it breaks.

Every failure is caught

One alert workflow for every instance, so nothing depends on a customer noticing.

The customer hears first

A banner and an email with the fix, before the gap shows up in their own reporting.

One outage is one incident

Grouping keeps your team on the fix and the customer informed once, rather than buried in alerts.

The run is always attached

Support starts from what happened, step by step, rather than from asking the customer to describe it.

Questions people ask

Should customers see integration errors?

The ones they can fix, yes, in plain words with the button that fixes it. The ones only you can fix, tell them once that you know and are on it.

How do you stop alert floods during an outage?

Group the same failure across customers into one incident, tell affected customers once, and hold the per-run alerts until it is resolved.

Can integration errors go to Datadog?

Yes. Send them from the alert workflow, or stream logs there, so integration failures sit with the rest of your product's errors.

Who should own each integration's alerts?

A named person or team per integration. An incident that goes to a shared channel waits until somebody decides it is theirs.

Last reviewed October 2026.