Why Only 11% of AI Agents Reach Production

Tokenomics & AI Economics · Day-2 AI The bill arrives after the demo does. 75% of companies plan to invest in agentic AI. Only 11% actually get one running in production. Here's what the 11% did differently. The problem

The demo always works. Someone shows the room an agent that drafts the reply, closes the ticket, updates the record, and the room says yes. Two minutes, and the project gets a green light.

Then it actually ships. The bill, the behavior, the first incident, none of it looks like what got approved in that room. That's not the exception. That's what usually happens.

Why most agentic AI projects actually fail
40%+of agentic AI projects cancelled by 2027 · Gartner, via Forbes
95%of GenAI deployments return no measurable value · MIT
11%of companies have an agent actually in production, of the 75% investing · Deloitte

Put those three numbers next to each other and you get the same answer every time: this isn't about the model. Forbes dug into the Gartner forecast and found the projects headed for cancellation aren't dying because the model wasn't good enough. They're dying from bad governance, data nobody can actually reach, no clear owner, and an ROI case nobody wrote down before launch. Two numbers make that case hard to argue with: out of the thousands of companies now claiming they've "gone agentic," Gartner found only about 130 were building anything that deserved the name. Everyone else was running a chatbot or an automation script with a new label on it. Forrester found something similar, roughly three out of four enterprises say they've adopted agentic AI, but only a small slice actually have it running in production.

The Deloitte 11% number is the same story, just earlier in the timeline. Most companies never even get to the point where cancellation is a risk, because their agent never made it to production to begin with. That gap, between "we're investing in this" and "this is actually live," is where most of the money disappears. And it's usually for the same reason every time: nobody defined success, nobody sorted out data access, nobody's name is on it when something breaks.

Five things to watch before you scale an agent

These sit between a working demo and a system that actually holds up in production. None of them show up on the model bill, which is exactly why they catch people off guard.

1 Tokens: the unit price drops, the unit count explodes Token prices keep dropping, so it's tempting to think the cost problem is solving itself. It's not. An agent doesn't make one call, it plans, calls tools, retries, hands off, and resends its whole history at every step. The price per token goes down. The number of tokens goes up faster. Your bill still climbs. Cost isn't what you pay per token, it's how much useful work you get for a dollar, and you only know that if someone's actually watching it happen.
2 Data: the agent reasons on whatever you hand it Here's a real way this goes wrong: a support agent nails the demo, then starts misrouting tickets the moment it's live, because "customer" means one thing in the CRM and something else in billing, and nobody ever told it that. The agent isn't confused. It just never had the context it needed, and by the time anyone notices, it's already made a few hundred bad calls.
3 Assessment: the check you skip is the failure you ship Checking whether an agent actually finished the job, used the right tool, or made something up gets expensive at scale, so most teams just run it less. That's backwards. If you only check a sample, you only catch a sample of what's wrong.
4 Guardrails: you can't bolt safety on after the fact Any agent that can act on its own makes your exposure bigger, not smaller. A person can't review decisions as fast as an agent makes them, and a guardrail that only gets checked once a quarter isn't really a guardrail. It's paperwork for an incident that already happened.
5 Ownership: nobody agreed on who's watching Forbes calls this the actual driver behind the coming wave of cancellations. Before you give an agent real authority, somebody needs to answer three plain questions: what does success look like, does it really have the access it needs, and whose job is it when it gets something wrong? If nobody can answer those today, the other four problems are just a matter of time.
How OnStak works

All five of these come back to the same root cause: you can't control a cost you never see, and you can't govern a decision you never watched happen. That's the layer we build from, not something we bolt on top afterward.

Worth being precise here: your agentic AI isn't part of AI Correlation Fabric, and AI Correlation Fabric isn't a module bolted inside your agent. It's the other way around. AI Correlation Fabric is the layer underneath, and every agent you run sits on top of it. The agent still plans, calls tools, and makes decisions exactly as it would on its own, Correlation Fabric just means someone can actually see that happening in real time instead of finding out from a bill or an incident report.

OnStak AI Correlation Fabric covers the first three. It connects what's happening across your infrastructure, your data, and every tool call your agents make, so you can actually see token usage, data quality, and assessment coverage in one place, before any of it turns into a surprise on an invoice.

OnStak AI Assurance, the Yes Layer, covers the last two. It runs guardrails continuously and keeps an audit trail automatically while the agent's running, so nobody has to reconstruct what happened after the fact. Accountability just comes built in. This is what our AI and Agent Services actually do, and it's the same gap our AI implementation services are built to close before an agent ever reaches production, not after.

Why the 11% actually succeeded

Flip that Deloitte number around and you get a better question: what did that 11% actually do differently? Neither Deloitte nor Forbes says it came down to a better model. Read them side by side and the answer is the fifth item on the list above. The agents that made it to production belonged to teams who already had answers before launch, not scrambling for them afterward.

That's the part worth paying attention to. It didn't come down to which model they picked, how big the budget was, or how good the demo looked. It came down to whether ownership, data access, and a real success metric were nailed down on day one, the same basics that decide whether a project gets killed two years later. That 11% didn't get lucky. They just did the boring parts first.

That's not something we noticed after the fact, it's why our own delivery model starts with Discover and Prove before it ever gets to Build and Operate. The boring parts, deciding what success actually looks like, checking the agent can reach the data it needs, agreeing on who's on the hook, happen before an agent goes live. Not after it's already somebody's problem.

The fix

You don't need to wire up your whole portfolio to find out where you stand. Just take the one agent closest to production and get three straight answers. A CIO needs the first one before defending the budget, a CISO needs the second before signing off on access: what does it actually cost per completed task, not per model call. What can it reach today, and has anyone actually reviewed that. And if it gets something wrong, whose name is on it, and can they actually shut it down.

If you can't answer those three yet, that's the real finding, not a side note. A proper AI readiness assessment is the fastest way to get all five costs on the table before they're locked into production, and before someone has to explain them after the fact.

Conclusion

None of this happens because the model was bad. Production was never the finish line, that's Day 2: the point where these five costs either stay in view, or quietly pile up until somebody has to explain them. Whether you're still trying to get an agent live, or you're already worried yours might get cut, the gaps are the same five. Close them before launch, and you're headed toward the 11% who made it, not stuck with everyone else who didn't. That's exactly what our AI agent services are built for, closing these five gaps as part of AI implementation, before an agent ever reaches production, not after.

Frequently Asked Questions
Do AI implementation services help with cost, or just deployment?+
Both, when they're done right. OnStak's AI implementation services fold cost visibility, data readiness, and assessment coverage into the deployment itself, so an agent doesn't reach production without the instrumentation needed to catch these five costs early.
Why do only 11% of agentic AI projects reach production?+
Per Deloitte's 2026 State of AI in the Enterprise, the gap has little to do with model capability. It tracks with unresolved data access, undefined success metrics, and unclear ownership, the same issues Forbes ties to the coming cancellation wave.
Why do agent costs rise even as token prices fall?+
Because the price per token and the number of tokens per task move in opposite directions. Agents loop, call tools, retry, and resend context at every step, so the token count per task climbs faster than the per-token price drops.
What's the fastest way to find out if this is happening already?+
Pick your closest-to-production agent and ask what it costs per completed task, not per model call, and who owns the outcome when it's wrong. If those answers aren't readily available, the gap already exists.
How is OnStak's approach different from adding more monitoring tools?+
AI Correlation Fabric and AI Assurance sit underneath your agentic AI, not inside it. They correlate and generate evidence from the infrastructure and data layer up, so your agent runs on top of that visibility from day one, rather than a dashboard getting bolted on top of an agent that's already in production.
Part of the OnStak enterprise AI consulting series. Explore more on the OnStak Debrief. Keep Reading
Know what one decision costs.
Before the invoice tells you.
See how OnStak AI Correlation Fabric and AI Assurance make agent cost, data quality, and guardrail behavior visible from day one, not after the bill arrives. Let's Talk →

Coffee? No Slides

No pitch deck. No 47-page proposal. Just a straight talk about what’s broken and what to fix first.

  • Solutions
  • Debrief
  • About Us