Writing
Engineering
AI Agent Rollbacks in 2026: Why the Most Governed Teams Pull the Most Agents
74% of enterprises have pulled a live AI agent. Among the most mature teams, it’s 81%. That inversion is the most useful number in enterprise AI right now, and almost everyone is reading it backwards.
In May, Sinch published survey results from 2,527 senior decision-makers across ten countries. One number got picked up everywhere: 74% of enterprises had rolled back or shut down a live AI customer communications agent after it went to production.
The number underneath it got almost no attention. Among organizations that described their guardrails as fully mature - the most governed, most monitored AI programs in the study - the rollback rate was 81%.
More governance. More monitoring. More rollbacks.
That should stop you for a second, because it breaks the assumption the entire enterprise AI governance market is built on. Every framework, every maturity model, every vendor deck sells the same promise: govern properly, and you won’t have to pull things out. The data says the opposite happened.
What “rollback” actually means here
It’s worth being precise, because this gets conflated with the other AI failure story constantly.
Pilot purgatory is the well-covered one. Deloitte’s 2026 Tech Trends puts the agent production failure rate at 89%. Teradata found 78% of enterprises running at least one pilot and only 14% scaling one across the organization. Those are projects that died in the lab. Sad, cheap, forgettable.
A rollback is a different animal. Sinch defined it as shutting down or significantly reversing an agent that had already gone live to real customers. It shipped. It worked well enough to launch. It touched people who were paying you money. And then someone made the call to take it offline.
That is a far more expensive failure and a much more politically painful one. Somebody had to walk into a room and say the thing we announced last quarter is coming down.
62% of the enterprises surveyed already had agents live in production. So this isn’t a story about companies stuck at the starting line. It’s a story about what happens after you cross it, which is a question the market mostly stopped asking around the time it decided “shipped” was the finish line.
Why the mature teams pull more
The obvious reading is that mature governance programs are somehow worse. That’s not it.
Daniel Morris, Sinch’s CPO, put it plainly: the most advanced organizations aren’t failing less; they’re seeing failures sooner. A team with real instrumentation, defined ownership, and explicit rollback criteria catches a bad outcome in hours instead of quarters, and has a pre-agreed path to act on it. A team without those things doesn’t catch it at all. It just accumulates.
Governance doesn’t prevent production failures. It detects them and permits you to respond. Those are completely different products, and the industry has been selling the first while shipping the second.
Once you accept that, the 81% figure stops being alarming and starts being a control signal. Those companies are running fire drills. They know what’s on fire.
The number that should actually worry you
Here’s the part I haven’t seen anyone say out loud.
If rollback rates rise with monitoring maturity, then the organizations reporting no rollbacks are not the benchmark. Some of them are genuinely running clean. The rest are simply blind.
There’s a quarter of the market claiming an unblemished record with live agents in production. Given what the maturity curve does to the numbers, I’d want to know how many of those companies could detect a serious failure if they had one. Not “do you have a governance policy” - everyone has a governance policy. Can you tell me, right now, how your agent performed in the last 24 hours, and would you know within an hour if it started leaking customer data?
If the answer is no, a zero rollback rate isn’t a clean record. It’s an unread inbox.
What’s breaking isn’t the model
The triggers Sinch identified are mundane and specific: PII exposure, hallucinated responses, and missing audit trails. Each one needs a different fix, which is part of why blanket governance policies underperform. You cannot solve an audit trail problem with a stricter approval workflow.
None of these are model quality problems. The models are fine. The models have been fine for a while.
The report’s own conclusion is that infrastructure was the strongest predictor of successful deployment - stronger than governance maturity, stronger than budget. They put the correlation at 0.52, the highest across thousands of variable pairs they tested. The most common complaints about existing providers were reliability at scale, weak multi-channel support, and missing platform integrations. Plumbing, in other words.
I’d add a caveat the coverage mostly skipped: Sinch sells communications infrastructure. A report from an infrastructure vendor concluding that infrastructure is the decisive variable deserves a raised eyebrow. Take that finding as directional.
The governance inversion is the more trustworthy result, precisely because it’s inconvenient - it undercuts the tidy “buy governance, get outcomes” story that would have been easier to sell. Findings that cut against the sponsor’s interest are usually the ones worth keeping.
There’s a corroborating number that isn’t vendor-sourced, though. 84% of AI engineering teams report spending at least half their time on guardrails rather than on building capability. Whatever you think of the infrastructure conclusion, that is what your engineers are actually doing this year.
This is about to get formalized
Gartner’s forecast from the same month sharpens it: by 2027, 40% of enterprises will demote or decommission autonomous AI agents because of governance gaps that only became visible after a production incident.
Read the second half of that sentence again. Not gaps identified in review. Not gaps flagged in design. Gaps that surfaced after something went wrong in front of customers.
That’s the same phenomenon Sinch measured, projected forward two years and given a bigger denominator. The rollback wave isn’t an anomaly in the adoption curve. It’s a stage of it.
What to change on Monday
Stop treating “agents in production” as the success metric. It’s a vanity number, and it’s actively misleading, because it counts launches and ignores survivals.
Three things are worth more:
Time to detect. How long between an agent doing something wrong and a human knowing about it? If you can’t measure this, you don’t have monitoring; you have logging.
Pre-defined rollback triggers. Decide what conditions take the agent offline before you launch it, not during the incident call. Teams that write these down roll back faster and argue less.
A named owner per agent. One human, accountable when it does something expensive. Diffused ownership is the single most consistent predictor of failure across every study I looked at, and it’s free to fix.
The uncomfortable framing
98% of the enterprises in the Sinch study said they’re increasing AI investment this year. They pulled agents and doubled down at the same time. Which is the right response, and worth noting for anyone reading the rollback numbers as a sign the whole thing is collapsing. It isn’t. Investment is going up. Trust, security, and compliance is now the single largest spend category inside AI programs, ahead of AI development itself.
What’s changing is the definition of a competent AI program. Twelve months ago it meant shipping fast. It now means knowing, quickly and specifically, when to stop.
The best-run AI programs in the world turned something off this year. If yours didn’t, the interesting question isn’t whether you got lucky.
It’s whether you’d know.
Explore More:
Website: kashishmahant.com
Podcast: Spotify | Amazon | Apple Podcasts
Instagram: Instagram
Email: k@kashishmahant.com
LinkedIn: LinkedIn
Keywords
- AI Agents
- Artificial Intelligence
- Enterprise Technology
- Leadership
- Technology