Ask around. Everyone has an agent demo. Almost nobody has an agent in production.
A weekly roundup I read puts numbers on the gap: about 85% of large companies experimenting with agents, about 5% with anything in production, 11 to 14% of pilots ever scaling. Take those as directional, not gospel. They come from secondary analysis, not audited filings. But the shape matches everything I have seen. The funnel is a cliff.
Gartner said the same thing with a steeper warning. Gartner predicts over 40% of agentic AI projects will be canceled by the end of 2027, blaming escalating costs, unclear business value, and inadequate risk controls. One of Gartner’s analysts put it blunter: most projects right now are early experiments driven by hype. Gartner also estimates only about 130 of the thousands of agentic AI vendors are real. The rest is what Gartner calls agent washing. Old products with a new label.
I think the diagnosis is simpler than any of those numbers. The models were good enough months ago. What kills a pilot is the boring list. Which systems the agent may call. What data it may see. Who approves a write. What gets logged. What gets rolled back when it does something stupid at 2am. Nobody demos those. Every production deployment is those.
That is why the most interesting agent announcement this month came from an ERP vendor. On September 29, Oracle announced Fusion Claw, a governed execution runtime inside its Fusion Applications. The agents are not the interesting part. The envelope is.
Oracle’s word for it is the Enterprise Operating Envelope, and it reads like a production checklist: objectives, standard procedures, policies, permissions, risk thresholds, decision rights, approval requirements, escalation boundaries. All defined inside the ERP, enforced on every run. When a run finishes, the customer gets what Oracle calls an Outcome Receipt. Authority applied, evidence used, actions executed, result. That is an audit trail with a marketing name, and I mean that as praise.
The architecture splits the work the same way I would. A frontier model reasons and plans. Deterministic enterprise computation executes at volume. Intelligence where judgment is needed, cheap deterministic code for everything else. Oracle says this plainly: apply AI reasoning only where intelligence is needed. That separation is the economics of the whole thing.
There is a pattern here worth naming. Agents ship when orchestration sits inside a plane that already has identity, permissions, and audit. The ERP already knows who can touch the ledger. The agent inherits that instead of reinventing it in a sidecar gateway nobody in compliance has approved.
And agents stay shipped when they are treated as a service with an SLO. Success rate, escalation rate, cost per outcome, time to rollback. The same boring gauges as everything else in production. A demo answers whether the model can do the task. An SLO answers whether the business can depend on it. Those are different questions, and only the second one ships.
So my advice to anyone with a pilot stuck in month six: stop tuning the prompt. Write down who the agent is allowed to be, what it may touch, who approves, and what the receipt looks like. Do that inside a system your auditors already trust. The model was never your problem. The plumbing was.
