Automation · AI operations
How to manage the MCP server lifecycle
A tool changes its inputs on a Friday, every agent that used it starts failing, and nobody can say what the old version looked like. Here is how MCP servers should move from draft to live to retired, the prompts that build it, and what it takes to run.
Built with Tray Headless
- System GitHub
- Step Test the change
- Step Owner approves
- Step Publish new version
- System MCP server
Every change to a server goes through the same path, and every live version can be put back.
The short answer
What is the MCP server lifecycle?
Managing the MCP server lifecycle comes down to four things: every change goes through a test and an owner's approval before it is live, every published version is kept so the last good one can be put back, the agents and people who depend on a tool hear about a change before it happens, and tools nobody uses are retired on a date. The common failure is editing a live tool in place. Agents learn a tool's inputs, and changing them without warning breaks every agent that relied on the old ones.
Stage 4 of 8: Publish and retire servers. Part of AI and MCP governance, end to end : every stage, the systems it runs on and the guide that builds it.
What matters here
- Never edit a live tool in place. Publish a new version, and keep the old one until nothing uses it.
- Keep every published version. The fastest fix for a bad change is putting the last good version back.
- Tell the people and agents that use a tool before it changes. Usage records show exactly who they are.
- Retire tools on a date, with a warning first. A tool nobody calls still holds a live credential.
- Every server has an owner who approves its changes. A server anyone can change is a server nobody controls.
Who this is for
You run platform engineering or AI operations. MCP servers are on a managed path, and changes to them now affect agents across the company.
How it works in practice
What happens between someone wanting to change a tool and the change being safe for every agent that uses it.
- 1
The change is proposed against a version
In source control, so the difference from what is live is visible line by line.
- 2
It is tested against real calls
Recent calls from the audit trail are replayed against the new version to see what would break.
- 3
The owner approves
With the test results and the list of agents and people who use the tool attached.
- 4
Users hear before it goes live
Anyone whose calls would change gets a message with the date and what is different.
- 5
The new version is published, the old one kept
So putting it back is one step, not a rebuild.
- 6
Unused tools are retired on a date
Warned, switched off, and their credentials removed.
What a managed lifecycle is made of
Four parts. The second is the one teams wish they had the first time something breaks.
A path for every change
Propose, test, approve, publish. The same path for a new tool, a changed input and a new server.
A kept version history
Every published version stored, with who approved it and when, so the last good one can go back in a single step.
Warning before change
Usage records name the people and agents that call a tool, and they hear about a change before it lands.
Retirement on a schedule
Tools with no calls for a set period are warned, switched off and stripped of their credentials.
The Tray Headless prompts
Paste these into Claude Code or Codex with the Tray Headless plugin installed. Each stage runs on its own. The systems named in them are the worked example rather than a requirement, and every prompt says so.
Once per project, run
/tray-workflows:set-workspace
to pick the workspace these build in. Point it at a sandbox first.
- 1
Set up and put every server under version control
You cannot put back a version you never kept.
Headless skills
build-workflowUse build-workflow. The systems in play are GitHub, Jira, Slack and Snowflake, or whatever we run in those seats. Before you plan anything, tell me which of them are already authenticated in the workspace, because I do not want a connector stubbed that I have not authenticated. Keep the definition of every managed MCP server and its tools in GitHub: tool names, inputs, descriptions, which workflow each one runs, the owner and the access groups. Every published version gets a tag. A change to a live server starts as a pull request against the current tag, so the difference from what is live is visible before anything happens.
- 2
Test changes against real calls
The calls agents made last week are the best test there is.
Headless skills
tray-patternsWhen a change is proposed, take a sample of last week's calls to that tool from the audit records in Snowflake and run them against the new version in a test workspace. Report which calls would now fail, which would return something different, and which inputs no longer exist. Attach the report to the change. A change that would break real calls needs a new version and a warning, not an edit.
- 3
Get the owner's approval, with the facts attached
The owner decides, and should not have to go looking for anything.
Headless skills
build-workflowUse build-workflow. Open a Jira ticket for every change with the difference, the test report, and the list of people and agents that called the tool in the last 30 days. Send the owner a Slack message with approve and reject buttons. Only publish after the owner approves. If the owner has not answered in two working days, remind them once, then escalate to their manager instead of publishing anyway.
- 4
Warn users, publish, keep the old version
Agents learn a tool's inputs. Changing them without warning breaks every one of them.
Headless skills
tray-gotchasUse tray-gotchas. For a change that alters inputs or results, message everyone who called the tool in the last 30 days with what changes and when, at least a week ahead. Publish the change as a new version and keep the old one live alongside it until its calls drop to zero or the date passes. Never change the inputs of the live version in place. Add a one-step way back: putting the previous version live again, with the owner told and the reason logged.
- 5
Retire what nobody uses
A tool nobody calls still holds a live credential.
Headless skills
build-workflowUse build-workflow. Each week, find tools with no calls in 90 days and servers whose tools are all unused. Message the owner with the list and a retirement date 30 days out. On the date, switch the tool off, remove its credentials from the workspace, and record the retirement with the last version kept in GitHub, so it can be brought back on purpose if someone needs it.
- 6
Check it end to end, then hand the path to platform
Because the release rules will change, and owners should be able to change them.
Run the per-step schema checks and the whole-workflow audit before this touches production. Push a change that removes an input and confirm the test report flags it, the owner is asked, users are warned, and the old version stays live. Then open the same workflow in Tray Build so platform engineering can change the warning period, the retirement threshold and the approval route in the visual canvas.
What it connects to
Changes start in source control, decisions happen where owners already work, and usage comes from the audit records.
GitHub
Hold every server and tool definition, with a tag for each published version and a pull request for each change.
Reads and writes
Snowflake
Supply recent calls for testing and the list of who uses each tool, from the audit records.
Reads
Same build, other stacks
The design does not change if you run something else in one of these seats. The same prompts build it against Google BigQuery, Microsoft Teams, ServiceNow, Databricks, Google Chat or Jira Service Desk.
Connections in this build
Field mapping, templates and common problems for each pairing: GitHub + Jira, GitHub + Slack and Jira + Slack.
Named systems are the ones most teams run, not the only ones that work. Each is an authentication in your Tray workspace, referenced by name, so the workflow never holds a credential. Where we have a connector page, the name links to it.
Running it in production
This decides what every agent in the company can call tomorrow. It has to be boring and predictable.
It runs on the platform, not on your laptop
Tests, approvals, warnings and retirements run on the same engine as the servers they manage, with a full history of each step.
Every version can go back
The last good version is kept live until nothing needs it, and putting it back is one step with the reason logged.
Credentials live in the workspace, never in the repo
Server definitions in GitHub name a workspace authentication and never hold a secret, so a retired tool's credentials can be removed without touching the history.
Owners approve their servers
Each change waits for the server's owner. The approval route and the warning periods open in Tray Build.
Test with a breaking change
Before production, remove an input on a test tool and confirm every safeguard fires before anyone's agent is affected.
Questions people ask
Why not edit a live tool in place?
Agents learn a tool's inputs. Changing them in place breaks every agent that relied on the old ones, with no warning and no earlier version to go back to.
How do you know who a change will affect?
From the audit records. They list every person and agent that called the tool, so the warning goes to exactly the right people.
How should a bad change be undone?
By putting the previous version back live, which is one step because it was kept. The owner is told and the reason is logged.
When should a tool be retired?
When it has had no calls for a set period, often 90 days. The owner gets a warning and a date, and on that date the tool is switched off and its credentials removed.
Who approves changes to an MCP server?
The server's named owner, with the test results and the list of users attached. A server with no owner should not be live.
Further reading
Background on the same subject, for the case rather than the build.
Related guides
AI operations
How to expose an internal system as an MCP tool
Scope the tool to a job rather than an API, resolve identity per call, make writes idempotent, and log every invocation. The Headless prompts that build it.
AI operations
How to build MCP server access control
Open each MCP server to identity provider groups, make every caller sign in as themselves, cap calls, and close access the day someone leaves. The Headless prompts that build it.
AI operations
How to build an MCP audit trail
Log every MCP tool call with the person, the assistant, the tool, what went in and what came back, mask what should not be stored, and send it to security's tools. The Headless prompts that build it.
Last reviewed October 2026.