Here is a pair of numbers I keep coming back to. In New Relic’s 2026 State of AI Coding survey, as reported by IT Brief, 94% of technology leaders rated AI-generated code higher quality than human-written code at review. 78% saw more incidents once that code went live.
The code looks fine right up until it ships, and teams ship far more of it now. With Claude Code, Codex, or Cursor, a working version of an app takes an afternoon, and the next five versions take the rest of the week. Every one of those versions is a deploy.
Faster shipping moves the risk to the release
Harness asked 700 engineering practitioners this year and found that “For very frequent AI coding tool users, 22% of code deployments result in a rollback, hotfix, or customer-impacting incident.” Ship ten times a week at that rate and you are cleaning up after two bad releases a week. Google’s DORA research saw the same pattern in 2025: AI adoption raised delivery throughput, and it was still tied to lower delivery stability.
Most controls try to stop a bad release before it ships. They should, and some bad releases will get through anyway, because a change that passes every check can still fail in front of real users. Gartner’s research on the AI risk gap tells leaders to “shift your strategy away from pure prevention to prioritize damage control and fast recovery.” The most mature operators already work this way. After two global outages in late 2025, Cloudflare now gives configuration changes “progressive rollout, real-time health monitoring, and automated rollback to configuration deployments by default.”
So assume a bad release will reach production, and make going back to the last version that worked an ordinary thing to do.
Deployment management is a governance control now
Governance conversations about AI-built apps usually start with who can build and what an app can reach. I think deployment gets far less attention than it deserves, because it is the moment a change reaches the people who depend on the app. Gartner’s guidance on governing vibe coding is direct about it: “Provide standardized deployment pipelines with built-in security checks, environment separation, and policy enforcement.” Its guidance on scaling vibe coding adds that “No vibe-coded asset should enter production without named ownership.”
Gartner makes the same point about the tools themselves. Its market guide for enterprise vibe coding platforms finds that “All platforms produce a functional prototype from a prompt. What separates them is everything after generation: deployment pipelines, audit trails, compliance controls, back-end integrations, and team collaboration.” Governance applied at deploy is simply part of going live. Governance applied afterwards is a cleanup project that never finishes.
For IT, that comes down to three questions about every AI-built app in production. Does it reach production the same way every other app does? Who owns it, by name? The person who prompted the app is rarely the person who gets the call when it breaks at 3am. And when a release goes wrong, how fast can you get back to the version that worked? If the honest answer to the last one is “someone rebuilds it by hand from their laptop,” the app has no way back.
What going back looks like
Say a finance team runs an expense-approvals app that someone built in Claude Code. On Friday afternoon the builder ships a change to the approval rules, and by Friday evening nobody can submit an expense.
On Tray Helix, the app’s owner opens the project’s Deployments tab and chooses Restore this version on Thursday’s deployment. Helix rebuilds Thursday’s source through the normal build pipeline and ships it as a new deployment. Friday’s version keeps serving while that happens, and a few minutes later submissions work again. Nobody gets paged, and nobody rebuilds anything by hand.
- Replaced Wednesday An earlier version.
- Worked Thursday The last version that worked.
- Broken Friday Breaks submissions. Keeps serving while the restore builds.
- Live Friday, restored Rebuilt from Thursday’s source and shipped as a new deployment.
How restore works on Tray Helix
We built restore as a deploy on purpose, and most of how it behaves follows from that. Going back needs the same access as going forward, so it is governed the same way. It takes minutes, like any deploy, and the version that was live keeps serving until the restore succeeds, so trying never takes the app offline.

Restore is part of Deployment Management in Helix Enterprise Edition.
What restore does not bring back
Restore brings back code. Data stays where the newer version left it: the app’s key-value store, and anything the app wrote to a connected third-party service, are not reverted. If Friday’s release had also written bad approvals into the finance system, restoring Thursday’s code would stop new ones, and someone would still need to correct the ones already written.
Plan for that, and decide who owns the data fix before you need one.
Start with the app you would miss most
Pick the AI-built app your business would miss most and ask those three questions about it. It is often one that started small. Apps pick up users and reach more important systems over time, and nobody re-reviews them as they do. If nobody can say how fast it gets back to the version that worked, that is where I would start, because that answer sets how long its next bad release lasts.
Then time it. Restore that app through the platform, and time the builder’s own fix from their laptop. If the laptop wins, that is the path people will take under pressure, and your way back only exists on paper.
What’s new in Helix
Restore a previous deployment
Step by step in the Helix dashboard: how long it keeps each deployment, what it brings back, and what it leaves alone.
Read how restore works →
Sources +
- Harness, The State of DevOps Modernization Report 2026, 700 engineering practitioners and managers, fielded February 2026.
- Google Cloud, Announcing the 2025 DORA Report, 23 September 2025.
- New Relic, 2026 State of AI Coding, 200 US technology decision-makers, as reported by IT Brief, 11 June 2026.
- Cloudflare, Code Orange: Fail Small is complete, 1 May 2026.
- Gartner, Market Guide for Enterprise Vibe Coding Platforms, C.A. Swan, Manjunath Bhat, Bill Blosen, 28 April 2026, G00844703.
- Gartner, How to Scale Vibe Coding Using Low-Code Engineering Principles, Mukul Saha, 11 August 2026, G00857758.
- Gartner, Govern Vibe Coding for Citizen Developers With Self-Service Platforms, Nitish Tyagi, 29 June 2026, G00858202.
- Gartner, The AI Value Gap Is Really an AI Risk Gap, Rita Sallam, Mary Mesaglio, Pete Shoard, 1 July 2026, G00858675.
GARTNER is a registered trademark and service mark of Gartner, Inc. and/or its affiliates in the U.S. and internationally and is used herein with permission. All rights reserved. Gartner does not endorse any vendor, product or service depicted in its research publications and does not advise technology users to select only those vendors with the highest ratings or other designation. Gartner research publications consist of the opinions of Gartner’s research organization and should not be construed as statements of fact. Gartner disclaims all warranties, expressed or implied, with respect to this research, including any warranties of merchantability or fitness for a particular purpose.
Paul Turner is VP of Product Marketing and Market Strategy at Tray.ai.