In 2019 the State of DevOps report killed the change advisory board. It surveyed about a thousand practitioners and found that teams needing approval from an external body for significant changes were 2.6 times more likely to be low performers. It also found no evidence that those approvals lowered change failure rates at all. The control cost throughput and bought nothing.
I deleted some of those boards. So did most platform teams I worked with between 2019 and 2023. It was the right call and I would make it again.
We are about to put all of it back. Not because 2019 was wrong. Because the reason it was right has stopped holding.
The thesis
The approval layer was safe to delete because accountability lived inside the actor. A named engineer with a manager and a career pushed the change. An agent has no career. Take the approver out of an agent workflow and nothing is holding the risk.
So every company scaling agents is about to rebuild the approval middleware it spent a decade deleting. It will cost real money. Almost nobody has a budget line for it, because the same business case that created the need deleted the line.
Two findings about the same budget
McKinsey published its 2026 State of AI on 25 August, from 1,719 respondents across 97 nations surveyed between 4 May and 8 June. Forty percent of organisations above a billion dollars in revenue now report scaling AI agents, up from 27 percent a year earlier. Thirty-one percent of large enterprises are scaling coding agents specifically.
Then the line that should stop a CFO. Thirty-two percent of organisations report deciding against buying one or more software products or features, because they could build it in house with coding agents instead.
Hold that against Gartner. On 26 May 2026 Gartner predicted that by 2027, 40 percent of enterprises will demote or decommission autonomous AI agents because of governance gaps found only after a production incident. Shiva Varma’s framing is four autonomy tiers: Observe, Advise, Act with Approval, Act Autonomously.
Look at what three of those four tiers actually require. Observe needs logging and retention. Advise needs a queue and a named reviewer. Act with Approval needs policy, routing, pending state, evidence and reversal. None of that is a model capability. It is middleware, and it is the same five things a 2008 BPM suite sold.
The software you stopped buying was the workflow layer: approvals, routing, exception handling, audit trail. You cancelled it because the tool that creates the need made building look cheap.
The two denominators
TheAgentCompany benchmark, built at Carnegie Mellon and published in the NeurIPS 2025 datasets track, stood up a simulated software firm on GitLab, OwnCloud, Plane and RocketChat and gave agents 175 real tasks. The best model tested, Gemini 2.5 Pro, finished 30.3 percent of them outright and scored 39.3 percent with partial credit. Claude 3.7 Sonnet finished 26.3 percent. GPT-4o finished 8.6 percent.
Read that as an operating cost rather than a capability score.
Verification cost scales with attempted work. Value scales with completed work. At roughly 30 percent completion the checking layer handles about 3.3 items for every one that ships. Every agent business case I have been shown sizes human review against the completions. Not one has sized it against the attempts. That is a factor of three error, and it widens as you deploy more agents, because you deploy them into the categories where they fail most.
Where the failures cluster
TheAgentCompany found agents strongest on software engineering and weakest on administrative, finance and communication work. RocketChat tasks, which means talking to a colleague, scored worst across every model tested. OwnCloud tasks, which means office documents, were close behind.
Now put Gartner’s CIO poll from 19 May 2025 next to it. Of 125 respondents, 52 percent said their agent deployments target internal administration, and they named IT, HR and accounting. 23 percent named customer-facing work.
The enterprise is aiming agents at the exact category the benchmark says they are worst at. It is also the category that already carries statutory controls: segregation of duties, approval thresholds, a signed audit trail.
Engineering is where agents perform best and where verification is cheapest, because code review already exists as a funded institution with tooling, a norm and a named owner per repository. Accounting has controls too. They are procedural, held in people and spreadsheets, and not expressed as an engineering artefact an agent can be routed through.
That mismatch explains Gartner’s demote-or-decommission number better than “governance gaps” does. The gap is not a missing policy document. It is a missing product.
There is a third turn. Build versus buy inverts here. You build what differentiates and you buy what regulates. A verification layer earns its keep by being boring, complete and auditable, not by fitting your business. Coding agents made everything look buildable, which pushes teams to build the one category where buildability was never the right test. Twelve months later you own a half finished approvals queue with no single sign-on, no retention policy and no evidence export, maintained by whoever drew the short straw.
Where this breaks
Four places, and two of them are aimed at my own side.
The strongest objection is that the layer arrives as a product before your budget cycle does. Gartner expects guardian agent technologies to take 10 to 15 percent of agentic AI markets by 2030, and 70 percent of AI applications to be multi-agent by 2028. If verification ships as a feature of tools you already license, the budget problem dissolves and this argument expires quietly. I think it is half right. The plumbing will commoditise. The policy about who may approve a $40,000 credit note is yours and nobody can ship that for you. But I would not bet a build against the plumbing.
Second, the benchmark is a simulation. TheAgentCompany is a fictional firm with synthetic colleagues, and 30.3 percent on 175 tasks is not 30.3 percent inside your finance function. I am using the ratio as a shape, not a forecast.
Third, and this is the one that cuts hardest, Gartner’s warning about uniform governance runs directly against me. Putting an approval gate on every agent action is the failure mode on the other side. An agent that turns a meeting into a draft nobody sends does not need a control, and if you route it through one you have rebuilt the change advisory board and earned the 2.6x. Proportionality is the entire skill and I have no clean rule for it.
Fourth, scope. If your agents only touch reversible, low value, internally visible actions, none of this applies. Build nothing. The argument starts where an action has a counterparty.
What I would do on Monday
Count attempts, not completions. Take one agent already in production. Pull the number of outputs it produced last month against the number a human accepted unchanged. That ratio is the real throughput requirement for your verification layer. If nobody can produce the number, that is the finding.
Classify every agent you run into the four tiers, on one page, this week. It takes two hours and it almost always demotes something. Anything sitting in Act Autonomously that touches money, a customer or a record of account moves to Act with Approval until a person signs for it.
Name an owner for the approve step, with authority to say no and time in their week to use it. In most plans this seat does not exist, because the headcount it would have come from was removed in the same business case that bought the agent.
Get a price on buying the layer before you build it. If your workflow or service management vendor already sells approvals, routing and an evidence trail, the honest comparison is their licence against your twelve months plus your maintenance.
Write the reversal procedure before the first autonomous action, not after the first incident. What undoes it, who can trigger the undo, and how long you have to do it.
Close
The 2019 finding was never that approval is waste. It was that approval by someone with no stake and no context is waste. We collapsed that into “approvals are slow” because the short version was easier to sell, and for ten years the shortcut cost nothing, because the actor pushing the change had a name, a manager and a career.
The agent has none of those. Somebody still has to hold the risk, and it will not be the model.
The approve step is the part of an agent program nobody volunteers to own, which is why it usually gets discovered after the incident instead of before it. If you want the verification layer in your agent plan counted and staffed before deployment rather than after, my calendar is here.

