Integration · AI operations
How to build an MCP audit trail
Security asks what an agent changed last Tuesday, and the only record says a service account called the API. Here is what an MCP audit trail has to capture, the prompts that build it, and what it takes to keep it useful.
Built with Tray Headless
- System MCP tool call
- Step Attach person and assistant
- Step Mask sensitive values
- Step Write the record
- System Splunk
Every call is written once, with the person and the assistant attached, then sent to the tools security already searches.
The short answer
What is an MCP audit trail?
An MCP audit trail has four parts: one record per tool call naming the person, the AI assistant, the server and the tool, the inputs and the result stored with sensitive values masked, the records sent to the tools security already searches, and alerts on calls that look wrong. The usual gap is the person. A log that says a service account called Salesforce at 3pm cannot answer the only question anyone asks, which is who asked the agent to do it.
Stage 7 of 8: Audit every call. Part of AI and MCP governance, end to end : every stage, the systems it runs on and the guide that builds it.
What matters here
- Write one record per call, with the person who asked. A log that names a service account names nobody.
- Store what went in and what came back, not just that a tool ran. Rebuilding what an agent did needs both.
- Mask sensitive values before they are stored. An audit log full of customer data is its own risk.
- Send the records to the tools security already uses. A log nobody searches is not an audit trail.
- Decide how long to keep records before the first audit asks. Retention set after the fact is retention nobody can prove.
Who this is for
You run security, compliance or AI operations. Agents call tools on company systems every day, and you need to be able to say what they did and for whom.
How it works in practice
What has to happen between an agent calling a tool and that call being something an auditor can find.
- 1
Every call produces one record
Server, tool, time, how long it took and whether it worked, written whether the call succeeded or not.
- 2
The person and the assistant are attached
From the signed-in session, never from what the model sends, so the record cannot be talked into naming someone else.
- 3
Inputs and results are stored, with masking
Card numbers, personal details and secrets are replaced before the record is written.
- 4
Records go to the security team's tools
Streamed to the SIEM and kept in the warehouse, so investigations use the search people already know.
- 5
Unusual calls raise an alert
A burst of changes, a tool used at 3am, or a refused call retried twenty times.
- 6
Retention follows a written rule
Kept for as long as policy says, then removed on schedule, with the removal itself logged.
What an MCP audit trail is made of
Four parts. Without the first, the other three describe calls nobody can attribute.
A named person on every call
Taken from the signed-in session. The record says Maya Chen asked Claude to move a deal, not that an integration user did.
Inputs and results, masked
Enough to rebuild what happened, with sensitive values removed before storage rather than after.
Delivery to existing tools
Splunk, Datadog or the warehouse, so a question about an agent is answered with the same search as any other incident.
Alerts and retention
Rules that flag calls worth a look, and a keep-and-delete schedule written down before the first audit.
The Tray Headless prompts
Paste these into Claude Code or Codex with the Tray Headless plugin installed. Each stage runs on its own. The systems named in them are the worked example rather than a requirement, and every prompt says so.
Once per project, run
/tray-workflows:set-workspace
to pick the workspace these build in. Point it at a sandbox first.
- 1
Set up and write one record per call
The record is the unit everything else depends on.
Headless skills
build-workflowUse build-workflow. The systems in play are Splunk, Datadog, Snowflake and Slack, or whatever we run in those seats. Before you plan anything, tell me which of them are already authenticated in the workspace, because I do not want a connector stubbed that I have not authenticated. For every MCP tool call on our managed servers, write one record with: the time, the server, the tool, the person who asked, the AI assistant that made the call, how long it took, whether it worked, and the error if it did not. Write the record for refused and failed calls too. A refused call is often the most interesting line in the log.
- 2
Attach the person from the session
A record the model can edit is not a record.
Headless skills
build-workflowUse build-workflow. Take the person and the assistant from the signed-in session for the call, never from the tool's inputs, because the model writes the inputs and could put any name there. If a call arrives with no person attached, write the record anyway, mark it as unattributed, and send it to security. Unattributed calls should be rare, and each one is a gap worth closing.
- 3
Store inputs and results, masked
Enough to rebuild what happened, without storing what should not be kept.
Headless skills
tray-patternsStore the inputs the tool received and the result it returned, so anyone reading the record can see what the agent asked for and what changed. Before writing, mask card numbers, bank details, government ID numbers, passwords and tokens, and anything our data policy lists as personal. Keep the shape of the value so a reviewer can tell a field was there. Cap the size of stored results. A tool that returns ten thousand rows should be recorded as a count and a sample, not ten thousand rows.
Mask before the record is written, not in a later clean-up job. Anything stored unmasked, even for a minute, is already in the backups.
- 4
Send records to security's tools
A log nobody searches is not an audit trail.
Headless skills
build-workflowUse build-workflow. Stream every record to Splunk, so the security team searches agent activity next to everything else, and send call counts, errors and timings to Datadog for the people who run the servers. Load the same records into Snowflake each hour for reporting and for anything that needs a longer history than the SIEM keeps. If Splunk is down, hold the records and send them when it is back. Never drop a record because the place it was going was busy.
- 5
Alert on calls that look wrong
Most records are read only when something has already gone wrong. Alerts are how you hear first.
Headless skills
tray-gotchasUse tray-gotchas, then add alert rules that post to the security channel in Slack: More than 50 changes by one person in ten minutes A tool that changes records used outside working hours A refused call retried more than five times Any unattributed call A tool called for the first time by someone outside its usual group Each alert links to the records behind it. An alert that says "something happened" and makes people search for it gets muted.
- 6
Set retention, then hand the rules to security
Because what counts as unusual, and how long to keep it, is security's call.
Run the per-step schema checks and the whole-workflow audit before this touches production. Make a test call that includes a fake card number and confirm it is masked in every place the record lands. Keep records for the period our policy sets, remove them on schedule, and log each removal. Then open the same workflow in Tray Build so security can change masking rules, alert thresholds and retention in the visual canvas.
What it connects to
Records start at the tool call and end wherever security and compliance already look.
Splunk
Receive every record as it is written, so agent activity is searched like any other security event.
Writes
Datadog
Take call counts, errors and timings per server and tool, for the people who keep the servers healthy.
Writes
Snowflake
Keep the full history for reporting, access reviews and anything older than the SIEM holds.
Writes
Okta
Look up the person's team and manager from the signed-in identity, so reviews can group calls by team.
Reads
Same build, other stacks
The design does not change if you run something else in one of these seats. The same prompts build it against Google BigQuery, Microsoft Teams, Azure Active Directory, Databricks, Google Chat or AWS Redshift.
Connections in this build
Field mapping, templates and common problems for each pairing: Splunk HTTP Event Collector + Slack, Datadog + Slack and Okta + Slack.
Named systems are the ones most teams run, not the only ones that work. Each is an authentication in your Tray workspace, referenced by name, so the workflow never holds a credential. Where we have a connector page, the name links to it.
Running it in production
This is the record of what agents did with company systems. If it is wrong or missing, nothing else on this list can be proved.
It runs on the platform, not on your laptop
Calls happen at any hour, and every record is written on the same engine with retries, so a busy SIEM delays a record instead of losing it.
Masking happens before storage
Sensitive values never reach the log, the warehouse or the backups, and the masking rules are tested with fake data before launch.
Credentials live in the workspace, never in the repo
The SIEM, monitoring and warehouse connections are separate authentications that can only write, so the audit trail cannot be used to read anything back.
Security owns the rules
Masking, alert thresholds and retention open in Tray Build, so the people accountable for the log decide what goes in it.
Test with the question you will be asked
Before production, pick a test call and answer who asked, through which assistant, what changed and when, from the records alone.
Questions people ask
Why does the record need the person and not the account?
Because the question is always who asked the agent to do it. A record that names a shared account cannot answer that, and neither can anyone reading it.
Should inputs and results be stored?
Yes, with sensitive values masked first. Without them you know a tool ran but not what it did, and rebuilding an agent's actions needs both.
Where should the records go?
To the tools security already uses, such as Splunk or Datadog, plus the warehouse for long-term history. A separate log nobody opens does not count.
What should raise an alert?
Bursts of changes by one person, tools that change records used out of hours, refused calls retried many times, and any call with no person attached.
How long should records be kept?
As long as your policy and your auditors require, written down before launch. Removal should run on schedule and be logged itself.
Further reading
Background on the same subject, for the case rather than the build.
Related guides
AI operations
How to build MCP server access control
Open each MCP server to identity provider groups, make every caller sign in as themselves, cap calls, and close access the day someone leaves. The Headless prompts that build it.
AI operations
How to expose an internal system as an MCP tool
Scope the tool to a job rather than an API, resolve identity per call, make writes idempotent, and log every invocation. The Headless prompts that build it.
AI operations
How to build an agent observability pipeline
Capture whole agent runs instead of single calls, join each one to the outcome it produced, and alert on the failures that return an answer anyway. The prompts.
Last reviewed October 2026.