Skip to content

Automation  ·  AI operations

How to manage the MCP server lifecycle

A tool changes its inputs on a Friday, every agent that used it starts failing, and nobody can say what the old version looked like. Here is how MCP servers should move from draft to live to retired, the prompts that build it, and what it takes to run.

Built with Tray Headless

  1. System GitHub
  2. Step Test the change
  3. Step Owner approves
  4. Step Publish new version
  5. System MCP server
Also Put back last version

Every change to a server goes through the same path, and every live version can be put back.

The short answer

What is the MCP server lifecycle?

Managing the MCP server lifecycle comes down to four things: every change goes through a test and an owner's approval before it is live, every published version is kept so the last good one can be put back, the agents and people who depend on a tool hear about a change before it happens, and tools nobody uses are retired on a date. The common failure is editing a live tool in place. Agents learn a tool's inputs, and changing them without warning breaks every agent that relied on the old ones.

Stage 4 of 8: Publish and retire servers. Part of AI and MCP governance, end to end : every stage, the systems it runs on and the guide that builds it.

What matters here

  • Never edit a live tool in place. Publish a new version, and keep the old one until nothing uses it.
  • Keep every published version. The fastest fix for a bad change is putting the last good version back.
  • Tell the people and agents that use a tool before it changes. Usage records show exactly who they are.
  • Retire tools on a date, with a warning first. A tool nobody calls still holds a live credential.
  • Every server has an owner who approves its changes. A server anyone can change is a server nobody controls.

Who this is for

You run platform engineering or AI operations. MCP servers are on a managed path, and changes to them now affect agents across the company.

How it works in practice

What happens between someone wanting to change a tool and the change being safe for every agent that uses it.

  1. 1

    The change is proposed against a version

    In source control, so the difference from what is live is visible line by line.

  2. 2

    It is tested against real calls

    Recent calls from the audit trail are replayed against the new version to see what would break.

  3. 3

    The owner approves

    With the test results and the list of agents and people who use the tool attached.

  4. 4

    Users hear before it goes live

    Anyone whose calls would change gets a message with the date and what is different.

  5. 5

    The new version is published, the old one kept

    So putting it back is one step, not a rebuild.

  6. 6

    Unused tools are retired on a date

    Warned, switched off, and their credentials removed.

What a managed lifecycle is made of

Four parts. The second is the one teams wish they had the first time something breaks.

A path for every change

Propose, test, approve, publish. The same path for a new tool, a changed input and a new server.

A kept version history

Every published version stored, with who approved it and when, so the last good one can go back in a single step.

Warning before change

Usage records name the people and agents that call a tool, and they hear about a change before it lands.

Retirement on a schedule

Tools with no calls for a set period are warned, switched off and stripped of their credentials.

The Tray Headless prompts

Paste these into Claude Code or Codex with the Tray Headless plugin installed. Each stage runs on its own. The systems named in them are the worked example rather than a requirement, and every prompt says so.

Once per project, run /tray-workflows:set-workspace to pick the workspace these build in. Point it at a sandbox first.

  1. 1

    Set up and put every server under version control

    You cannot put back a version you never kept.

    Headless skills build-workflow

    Use build-workflow. The systems in play are GitHub, Jira, Slack and
    Snowflake, or whatever we run in those seats. Before you plan
    anything, tell me which of them are already authenticated in the
    workspace, because I do not want a connector stubbed that I have not
    authenticated.
    
    Keep the definition of every managed MCP server and its tools in
    GitHub: tool names, inputs, descriptions, which workflow each one runs,
    the owner and the access groups. Every published version gets a tag.
    
    A change to a live server starts as a pull request against the
    current tag, so the difference from what is live is visible before
    anything happens.
  2. 2

    Test changes against real calls

    The calls agents made last week are the best test there is.

    Headless skills tray-patterns

    When a change is proposed, take a sample of last week's calls to that
    tool from the audit records in Snowflake and run them against the new
    version in a test workspace.
    
    Report which calls would now fail, which would return something
    different, and which inputs no longer exist. Attach the report to the
    change. A change that would break real calls needs a new version and a
    warning, not an edit.
  3. 3

    Get the owner's approval, with the facts attached

    The owner decides, and should not have to go looking for anything.

    Headless skills build-workflow

    Use build-workflow. Open a Jira ticket for every change with the
    difference, the test report, and the list of people and agents that
    called the tool in the last 30 days. Send the owner a Slack message
    with approve and reject buttons.
    
    Only publish after the owner approves. If the owner has not answered
    in two working days, remind them once, then escalate to their manager
    instead of publishing anyway.
  4. 4

    Warn users, publish, keep the old version

    Agents learn a tool's inputs. Changing them without warning breaks every one of them.

    Headless skills tray-gotchas

    Use tray-gotchas. For a change that alters inputs or results, message
    everyone who called the tool in the last 30 days with what changes and
    when, at least a week ahead.
    
    Publish the change as a new version and keep the old one live alongside
    it until its calls drop to zero or the date passes. Never change the
    inputs of the live version in place.
    
    Add a one-step way back: putting the previous version live again, with
    the owner told and the reason logged.
  5. 5

    Retire what nobody uses

    A tool nobody calls still holds a live credential.

    Headless skills build-workflow

    Use build-workflow. Each week, find tools with no calls in 90 days and
    servers whose tools are all unused. Message the owner with the list and
    a retirement date 30 days out.
    
    On the date, switch the tool off, remove its credentials from the
    workspace, and record the retirement with the last version kept in
    GitHub, so it can be brought back on purpose if someone needs it.
  6. 6

    Check it end to end, then hand the path to platform

    Because the release rules will change, and owners should be able to change them.

    Run the per-step schema checks and the whole-workflow audit before this
    touches production. Push a change that removes an input and confirm the
    test report flags it, the owner is asked, users are warned, and the old
    version stays live.
    
    Then open the same workflow in Tray Build so platform engineering can
    change the warning period, the retirement threshold and the approval
    route in the visual canvas.

What it connects to

Changes start in source control, decisions happen where owners already work, and usage comes from the audit records.

GitHub

Hold every server and tool definition, with a tag for each published version and a pull request for each change.

Reads and writes

Jira

Track each change with the difference, the test report and the owner's decision.

Writes

Slack

Ask owners to approve, warn users before a change, and announce retirements.

Writes

Snowflake

Supply recent calls for testing and the list of who uses each tool, from the audit records.

Reads

Same build, other stacks

The design does not change if you run something else in one of these seats. The same prompts build it against Google BigQuery, Microsoft Teams, ServiceNow, Databricks, Google Chat or Jira Service Desk.

Connections in this build

Field mapping, templates and common problems for each pairing: GitHub + Jira, GitHub + Slack and Jira + Slack.

Named systems are the ones most teams run, not the only ones that work. Each is an authentication in your Tray workspace, referenced by name, so the workflow never holds a credential. Where we have a connector page, the name links to it.

Running it in production

This decides what every agent in the company can call tomorrow. It has to be boring and predictable.

It runs on the platform, not on your laptop

Tests, approvals, warnings and retirements run on the same engine as the servers they manage, with a full history of each step.

Every version can go back

The last good version is kept live until nothing needs it, and putting it back is one step with the reason logged.

Credentials live in the workspace, never in the repo

Server definitions in GitHub name a workspace authentication and never hold a secret, so a retired tool's credentials can be removed without touching the history.

Owners approve their servers

Each change waits for the server's owner. The approval route and the warning periods open in Tray Build.

Test with a breaking change

Before production, remove an input on a test tool and confirm every safeguard fires before anyone's agent is affected.

Questions people ask

Why not edit a live tool in place?

Agents learn a tool's inputs. Changing them in place breaks every agent that relied on the old ones, with no warning and no earlier version to go back to.

How do you know who a change will affect?

From the audit records. They list every person and agent that called the tool, so the warning goes to exactly the right people.

How should a bad change be undone?

By putting the previous version back live, which is one step because it was kept. The owner is told and the reason is logged.

When should a tool be retired?

When it has had no calls for a set period, often 90 days. The owner gets a warning and a date, and on that date the tool is switched off and its credentials removed.

Who approves changes to an MCP server?

The server's named owner, with the test results and the list of users attached. A server with no owner should not be live.

Last reviewed October 2026.