Skip to content

Integration  ·  AI operations

How to build an MCP audit trail

Security asks what an agent changed last Tuesday, and the only record says a service account called the API. Here is what an MCP audit trail has to capture, the prompts that build it, and what it takes to keep it useful.

Built with Tray Headless

  1. System MCP tool call
  2. Step Attach person and assistant
  3. Step Mask sensitive values
  4. Step Write the record
  5. System Splunk
Also Alert on unusual calls

Every call is written once, with the person and the assistant attached, then sent to the tools security already searches.

The short answer

What is an MCP audit trail?

An MCP audit trail has four parts: one record per tool call naming the person, the AI assistant, the server and the tool, the inputs and the result stored with sensitive values masked, the records sent to the tools security already searches, and alerts on calls that look wrong. The usual gap is the person. A log that says a service account called Salesforce at 3pm cannot answer the only question anyone asks, which is who asked the agent to do it.

Stage 7 of 8: Audit every call. Part of AI and MCP governance, end to end : every stage, the systems it runs on and the guide that builds it.

What matters here

  • Write one record per call, with the person who asked. A log that names a service account names nobody.
  • Store what went in and what came back, not just that a tool ran. Rebuilding what an agent did needs both.
  • Mask sensitive values before they are stored. An audit log full of customer data is its own risk.
  • Send the records to the tools security already uses. A log nobody searches is not an audit trail.
  • Decide how long to keep records before the first audit asks. Retention set after the fact is retention nobody can prove.

Who this is for

You run security, compliance or AI operations. Agents call tools on company systems every day, and you need to be able to say what they did and for whom.

How it works in practice

What has to happen between an agent calling a tool and that call being something an auditor can find.

  1. 1

    Every call produces one record

    Server, tool, time, how long it took and whether it worked, written whether the call succeeded or not.

  2. 2

    The person and the assistant are attached

    From the signed-in session, never from what the model sends, so the record cannot be talked into naming someone else.

  3. 3

    Inputs and results are stored, with masking

    Card numbers, personal details and secrets are replaced before the record is written.

  4. 4

    Records go to the security team's tools

    Streamed to the SIEM and kept in the warehouse, so investigations use the search people already know.

  5. 5

    Unusual calls raise an alert

    A burst of changes, a tool used at 3am, or a refused call retried twenty times.

  6. 6

    Retention follows a written rule

    Kept for as long as policy says, then removed on schedule, with the removal itself logged.

What an MCP audit trail is made of

Four parts. Without the first, the other three describe calls nobody can attribute.

A named person on every call

Taken from the signed-in session. The record says Maya Chen asked Claude to move a deal, not that an integration user did.

Inputs and results, masked

Enough to rebuild what happened, with sensitive values removed before storage rather than after.

Delivery to existing tools

Splunk, Datadog or the warehouse, so a question about an agent is answered with the same search as any other incident.

Alerts and retention

Rules that flag calls worth a look, and a keep-and-delete schedule written down before the first audit.

The Tray Headless prompts

Paste these into Claude Code or Codex with the Tray Headless plugin installed. Each stage runs on its own. The systems named in them are the worked example rather than a requirement, and every prompt says so.

Once per project, run /tray-workflows:set-workspace to pick the workspace these build in. Point it at a sandbox first.

  1. 1

    Set up and write one record per call

    The record is the unit everything else depends on.

    Headless skills build-workflow

    Use build-workflow. The systems in play are Splunk, Datadog, Snowflake
    and Slack, or whatever we run in those seats. Before you plan
    anything, tell me which of them are already authenticated in the
    workspace, because I do not want a connector stubbed that I have not
    authenticated.
    
    For every MCP tool call on our managed servers, write one record with:
    the time, the server, the tool, the person who asked, the AI assistant
    that made the call, how long it took, whether it worked, and the error
    if it did not.
    
    Write the record for refused and failed calls too. A refused call is
    often the most interesting line in the log.
  2. 2

    Attach the person from the session

    A record the model can edit is not a record.

    Headless skills build-workflow

    Use build-workflow. Take the person and the assistant from the signed-in
    session for the call, never from the tool's inputs, because the model
    writes the inputs and could put any name there.
    
    If a call arrives with no person attached, write the record anyway,
    mark it as unattributed, and send it to security. Unattributed calls
    should be rare, and each one is a gap worth closing.
  3. 3

    Store inputs and results, masked

    Enough to rebuild what happened, without storing what should not be kept.

    Headless skills tray-patterns

    Store the inputs the tool received and the result it returned, so
    anyone reading the record can see what the agent asked for and what
    changed.
    
    Before writing, mask card numbers, bank details, government ID numbers,
    passwords and tokens, and anything our data policy lists as personal.
    Keep the shape of the value so a reviewer can tell a field was there.
    
    Cap the size of stored results. A tool that returns ten thousand rows
    should be recorded as a count and a sample, not ten thousand rows.

    Mask before the record is written, not in a later clean-up job. Anything stored unmasked, even for a minute, is already in the backups.

  4. 4

    Send records to security's tools

    A log nobody searches is not an audit trail.

    Headless skills build-workflow

    Use build-workflow. Stream every record to Splunk, so the security team
    searches agent activity next to everything else, and send call counts,
    errors and timings to Datadog for the people who run the servers.
    
    Load the same records into Snowflake each hour for reporting and for
    anything that needs a longer history than the SIEM keeps.
    
    If Splunk is down, hold the records and send them when it is back.
    Never drop a record because the place it was going was busy.
  5. 5

    Alert on calls that look wrong

    Most records are read only when something has already gone wrong. Alerts are how you hear first.

    Headless skills tray-gotchas

    Use tray-gotchas, then add alert rules that post to the security
    channel in Slack:
    
      More than 50 changes by one person in ten minutes
      A tool that changes records used outside working hours
      A refused call retried more than five times
      Any unattributed call
      A tool called for the first time by someone outside its usual group
    
    Each alert links to the records behind it. An alert that says
    "something happened" and makes people search for it gets muted.
  6. 6

    Set retention, then hand the rules to security

    Because what counts as unusual, and how long to keep it, is security's call.

    Run the per-step schema checks and the whole-workflow audit before this
    touches production. Make a test call that includes a fake card number
    and confirm it is masked in every place the record lands.
    
    Keep records for the period our policy sets, remove them on schedule,
    and log each removal. Then open the same workflow in Tray Build so
    security can change masking rules, alert thresholds and retention in
    the visual canvas.

What it connects to

Records start at the tool call and end wherever security and compliance already look.

Splunk

Receive every record as it is written, so agent activity is searched like any other security event.

Writes

Datadog

Take call counts, errors and timings per server and tool, for the people who keep the servers healthy.

Writes

Snowflake

Keep the full history for reporting, access reviews and anything older than the SIEM holds.

Writes

Okta

Look up the person's team and manager from the signed-in identity, so reviews can group calls by team.

Reads

Slack

Post alerts to the security channel with links to the records behind them.

Writes

Same build, other stacks

The design does not change if you run something else in one of these seats. The same prompts build it against Google BigQuery, Microsoft Teams, Azure Active Directory, Databricks, Google Chat or AWS Redshift.

Connections in this build

Field mapping, templates and common problems for each pairing: Splunk HTTP Event Collector + Slack, Datadog + Slack and Okta + Slack.

Named systems are the ones most teams run, not the only ones that work. Each is an authentication in your Tray workspace, referenced by name, so the workflow never holds a credential. Where we have a connector page, the name links to it.

Running it in production

This is the record of what agents did with company systems. If it is wrong or missing, nothing else on this list can be proved.

It runs on the platform, not on your laptop

Calls happen at any hour, and every record is written on the same engine with retries, so a busy SIEM delays a record instead of losing it.

Masking happens before storage

Sensitive values never reach the log, the warehouse or the backups, and the masking rules are tested with fake data before launch.

Credentials live in the workspace, never in the repo

The SIEM, monitoring and warehouse connections are separate authentications that can only write, so the audit trail cannot be used to read anything back.

Security owns the rules

Masking, alert thresholds and retention open in Tray Build, so the people accountable for the log decide what goes in it.

Test with the question you will be asked

Before production, pick a test call and answer who asked, through which assistant, what changed and when, from the records alone.

Questions people ask

Why does the record need the person and not the account?

Because the question is always who asked the agent to do it. A record that names a shared account cannot answer that, and neither can anyone reading it.

Should inputs and results be stored?

Yes, with sensitive values masked first. Without them you know a tool ran but not what it did, and rebuilding an agent's actions needs both.

Where should the records go?

To the tools security already uses, such as Splunk or Datadog, plus the warehouse for long-term history. A separate log nobody opens does not count.

What should raise an alert?

Bursts of changes by one person, tools that change records used out of hours, refused calls retried many times, and any call with no person attached.

How long should records be kept?

As long as your policy and your auditors require, written down before launch. Removal should run on schedule and be logged itself.

Further reading

Background on the same subject, for the case rather than the build.

Last reviewed October 2026.