I have sat through a lot of agent demos this year. They end the same way every time. The agent reads the ticket, checks the system, drafts the reply, files the update, and the room makes a small approving noise. Then the demo stops. Nobody runs it a second time in front of you. The second run is where the product actually lives, and the second run is not what is being sold.
Most companies should not build agents this year. Not because the models cannot do the work. Because the number that decides an agent programme is the cost of finding out when it was wrong, and almost nobody prices that before they start. Get it wrong and the failure is not a pilot that dies quietly. It is a pilot that works, ships, and puts a permanent checking tax on every team downstream.
The number the demo does not show you
TheAgentCompany is the least forgiving public benchmark I know of. It drops a model inside a simulated firm with its own code host, its own chat and its own colleagues to ask, then gives it 175 tasks across software engineering, project management, data science, administration, human resources and finance. Real work, in sequence, with the context scattered the way it is in your company.
Gemini 2.5 Pro, the best of them, scored 30.3 percent full autonomous completion. With partial credit for getting most of the way there, the same model reaches 39.3 percent. It spends an average of 27.2 model calls and $4.20 per task attempted.
Now do the division nobody in the vendor deck does. $4.20 is the price of an attempt, not the price of an outcome. At 30.3 percent, a completed task costs $13.86 and burns roughly 90 model calls. And that figure is still wrong, because it counts none of the human time spent reading the other two attempts out of three to work out which bucket they landed in.
That is the argument in one line. The agent is cheap. Knowing whether it worked is not.
Reliability is not on the same curve as capability
A Princeton group published a preprint on 24 August 2026, Towards a Science of AI Agent Reliability. Rabanser, Kapoor, Kirgis, Liu, Utpala and Narayanan ran 14 models across two benchmarks and made one distinction that should reorganise how you buy this.
Pass at k is the probability the agent succeeds at least once in k attempts. That is the demo. Pass to the power of k is the probability it succeeds every time. That is your Tuesday. Benchmarks report the first, which is best case capability with a friendly operator in the chair.
Their headline finding is the one for your board. Across 18 months of model releases, accuracy climbed steadily while consistency, robustness and calibration did not keep pace. Capability and reliability are two curves and the industry reports one.
This matters for the wait or build decision, and most people have it backwards. Wait a year and what improves is capability. What you needed was calibration: knowing which runs to check. That is not obviously arriving on schedule.
The arithmetic that decides it
Take a step the agent gets right 95 percent of the time. Excellent by current standards. Chain five of them and the workflow completes 77 percent of the time. Chain ten and you are at 60. Chain twenty and you are at 36. To reach 95 percent end to end across twenty steps, every individual step needs to be right 99.74 percent of the time. Nothing in production is at 99.74 percent.
So the question is never whether the model can do the step. It is how many steps run before a human looks. That is a process design question, answered by how your work is arranged rather than by which model you licensed, which is why it belongs to whoever owns the operating model and not to the platform team that is excited about it.
Thirty years in, the step count is the first number I ask for. Before the model, before the vendor, before the budget.
Who is already doing it, and why
About 31 percent of enterprises now run at least one agent in production, on mid 2026 numbers from S&P Global Market Intelligence and McKinsey. Banking and insurance sit near 47 percent, well ahead of everyone else.
The usual reading is that regulated industries have more money and better engineers. Wrong. They already had what an agent programme requires, built for regulators decades before anyone said the word agent. A pending transaction that settles later. Four eyes above a threshold. An audit trail, so a wrong action can be traced and undone. Checkpoints and reversibility, bought in advance for a different reason.
Everyone else is sold the agent first and asked to build the checkpoints afterwards. Wrong order, and the order almost every programme is running in.
Gartner expects over 40 percent of agentic AI projects cancelled by the end of 2027, and names three causes: escalating costs, unclear business value, inadequate risk controls. All three are verification problems wearing different hats. Escalating cost is cost per completed run discovered late. Unclear value is nobody having defined correct. Inadequate risk control is an irreversible action surface.
Gartner also estimates that of the thousands of vendors marketing agentic products, roughly 130 are the real thing. Read that as a buyer. The base rate of any given vendor being genuine is under five percent, so honest diligence on a shortlist costs more than the pilot it was meant to de risk. In that market the correct default is not to buy.
The supporting numbers rhyme. IBM’s 2025 CEO study found 25 percent of AI initiatives delivered the expected return. Median time to value runs about 5.1 months on BCG and Forrester figures. Data quality is named the biggest blocker by 52 percent, which is another way of saying nobody wrote down what correct looks like.
The four questions
Here is the test. Four questions, in order. A no stops the programme there, not six months and two hundred thousand dollars later.
One. Can you write down what correct looks like, in a file, for thirty real past cases, before anyone builds? Not a description of correct. The thirty actual inputs and their actual right answers. If your team cannot produce that file in a week, you have an aspiration rather than a specification, and no model fixes that.
Two. Is a wrong action reversible? If the agent can send the money, email the customer or close the account and there is no undo, it proposes and a person commits. That is not timidity. It is what banking did.
Three. How many steps run between human checkpoints? Count them on a whiteboard. Over ten and the arithmetic says you are buying a system that finishes 60 percent of the time at best, whoever’s model is inside it.
Four. Do you know your cost per completed run? Not cost per call. Cost per call divided by success rate, plus the loaded cost of the minutes a person spends checking.
Clear all four and build it this quarter. That set is much smaller than the market implies and much larger than zero.
Where this breaks
Four honest problems with what I just argued, two of them aimed at me.
The benchmark is a benchmark. 175 tasks in a simulated company is not your company, and narrowly scoped deployments routinely beat general benchmark numbers by a wide margin. If your agent does one classification on one document type, 30.3 percent is not your number and I should not pretend it is.
Worse for my case: that benchmark’s frontier is moving, and quoting any fixed completion rate to argue against building is the same stale evidence move I complain about when vendors do it in reverse. The honest version of my position is that the ratio between capability and verification cost has not changed much, not that the capability number is fixed.
Gartner’s 40 percent is a forecast, not a measurement. Forecasts from firms that sell advice on avoiding the outcome they forecast deserve a discount.
And the strongest objection is one I cannot fully answer. The organisations at 47 percent got there by starting before the economics were legible. Learning by doing is real, capability compounds inside teams, and a rule that says wait until you can price it would have told them to sit out. Part of the cost of my position is a year of institutional learning the people who ignored me will have and you will not. I think the trade is right for most companies. I do not think it is free.
Where the thesis does not apply at all: high volume, low stakes, fully reversible internal work. Ticket triage, drafting, classification, retrieval, summarising a call. Build those today.
What I would do on Monday
Pick the workflow everyone keeps pointing at. Count the steps between human checkpoints and write the number on the wall. If it is over ten, this quarter’s project is shortening the workflow, not buying an agent.
Ask the team for the file: thirty historical cases with their correct outputs. Give them a week. What comes back tells you more about your readiness than any vendor evaluation.
Recalculate the business case as cost per completed run. Vendor per call price, divided by their own quoted success rate, plus the checking minutes at a loaded rate. Most cases do not survive this, which is the point of doing it first.
List every action the agent could take and mark each reversible or not. Everything in the not column becomes propose and commit before a line is written.
Ask any vendor for pass to the power of five rather than pass at one. Their reaction is worth more than the number.
Close
The demo always ends after the first run. That is the tell. It is not dishonest, it is just how demos work: everyone shows their best attempt.
Your company does not get the best attempt. It gets the average attempt, every day, forever, with someone downstream quietly checking the output because they do not trust it yet and nobody has told them when they can stop.
Ask to see the second run.
The expensive part of software has never been the build. If you want someone to look at your agent roadmap and tell you which two of the twelve are worth doing, my calendar is here.

