In July 2026 the DORA team at Google Cloud published a financial model for a 500 person engineering organisation adopting AI assisted development. The headline input is a 12.5% net time gain per developer, about an hour a day. Run that out. It is 62 engineers of freed capacity sitting inside a 500 person org. The same report tells you not to cut staff. Gartner asked 350 executives at companies above a billion dollars in revenue whether they had, and roughly 80% said yes, some by as much as a fifth.
Those two numbers describe the same 62 engineers. The capacity can only be spent once.
The thesis
Headcount is not a result. It is an input you happen to control, which is exactly why it gets managed as if it were an outcome. AI frees one pool of capacity, and you convert that pool into a smaller payroll or into more change shipped. Both are defensible. Doing both, in the same year, on the same pool, is a double count, and it is the most common accounting error in enterprise AI right now.
The only figure that exposes it is cost per shipped unit of change. Almost nobody computes it.
What the metric actually is
Take everything it costs to move change into production over a quarter. Fully loaded engineering payroll, contractors, tooling, cloud, and the AI line: seats, tokens, and the API keys somebody expensed. That is the numerator. Then count the units of change that reached production and were still there thirty days later. That is the denominator. Divide.
The numerator is the easy half and most finance teams already have it, scattered. The AI component is now large enough to matter. Gartner reported in June 2026 that nearly a quarter of technology leaders spend between $200 and $500 per developer per month on AI coding tokens alone, and about 6% spend more than $2,000. Jellyfish published a worked example of a 40 developer organisation: $1,560 a month in licences, $2,100 in overages, $480 in extra tools, $1,400 on raw API keys and $300 in personal subscriptions that went through on cards. Total $5,840 a month, or $146 a developer. Inside that same company the platform team ran at $556 per developer and the frontend team at $61. That is a nine times spread, and nobody who has not built a denominator can tell you which of those two numbers is the good one.
The denominator is where the work is, because you have to decide what a unit is and then defend it. A merged pull request is too small and too gameable. A deployment is closer. A change that a customer or an internal user could describe is better.
Look at what happens when the unit is chosen badly. DORA’s own model books its value as time saved, 12.5% per developer, then reports the delivery outcome as deployments rising from 50 to 56 a year against a first year investment of $8.4 million. That is $1.4 million per additional deployment. The model is not wrong. It is measuring the return in one unit and the outcome in another, and the arithmetic only looks absurd because I forced the two units into the same fraction. That is the whole point. When the unit you book value in differs from the unit you measure delivery in, a double count can sit between them for a year and never show up on a slide.
Run the numbers on your own team
DX studied more than 400 organisations and found a median throughput gain of 7.76% from AI assisted development, a mean of 13.1%, and 43.9% at the ninetieth percentile. Take the median and a hundred engineers. Assume a fully loaded cost of $200,000 each, and adjust that to your own market. The AI spend at $350 a developer a month is $420,000 a year. The capacity gain is 7.76 engineers, worth $1.55 million. Close to four to one.
That ratio holds if you take the capacity as output. It also holds if you take it as payroll: cut 7.76 roles, save $1.55 million, ship the same volume, and the cost per shipped change still falls. Both paths work. The failure mode is the third one, where the payroll saving is banked in the January budget and the output increase is promised in the same board pack. Two quarters later the volume is down, because the 62 engineers you modelled were the same 62 you let go, and the ratio is worse than before you started.
Gartner’s finding is the tell. Workforce reduction rates were close to identical between the organisations reporting strong returns and the ones reporting nothing or worse. Helen Poitevin put it plainly: workforce reductions may create budget room, but they do not create return. If cutting correlated with returns you would see it in 350 companies with a billion dollars of revenue each. You do not.
What I count
I run this test on my own ventures before I take it into anyone else’s company.
D30 pulls a three statement fact file out of an ASX annual report and runs eight articulation gates over it. The unit is one completed fact file that passes all eight. I can tell you what a fact file costs to produce, including model spend, and I can tell you how that number has moved. I have never managed the business to analyst hours saved, because hours saved is a number you can claim twice and a fact file is a number you can count.
D23 sells managed Apache Superset. The unit is a dashboard running in production that somebody opened this month. Not a deploy. A deploy that nobody looks at is cost with no denominator.
SearchFIT counts one tracked prompt scanned across the answer engines. When a model price drops, that unit cost drops, and I can see it in a week rather than infer it from a vendor’s slide.
None of those units are obvious. Each took an argument to settle. That is the actual work, and it is why the metric is rare.
Where this breaks
Any denominator you publish becomes a target. Count merged pull requests and you will get more of them, smaller, split across branches, and your ratio will improve while nothing ships. The thirty day survival test slows that down. It does not stop it. If you put this number in front of a board you have to expect the number to be gamed, and the honest position is that this metric is a diagnostic for the people running delivery, not a compensation input.
It also does not travel outside software delivery. If your AI value is sitting in a claims queue or a contact centre, the unit is a resolved case, not a merged change, and forcing a delivery frame onto it will mislead you.
The hardest objection is timing. Cost per shipped change is noisy for two or three quarters, and a CFO under pressure this quarter cannot wait that long. Headcount is measurable today, it is unambiguous, and it moves the P&L in the period. That is not stupidity. That is a real constraint, and it is exactly why the wrong metric keeps winning.
What I would do on Monday
Write the unit down in one sentence and get the CTO and the CFO to sign the same sentence. If they cannot agree what a shipped change is, stop there, because that disagreement is the finding.
Total the AI line properly. Seats, token overages, API keys on personal cards, and the two tools nobody put through procurement. Split it by team. Expect a spread like the nine times gap in the Jellyfish example.
Instrument the denominator with a thirty day survival window. Count changes that reached production and stayed. Reversals do not count.
Compute the ratio for the four quarters before AI landed and every quarter since. You need the baseline more than you need the current number.
Put a freeze on any headcount decision justified by AI capacity until you have two quarters of the ratio. If the capacity is real it will still be there in six months.
Sixty two engineers is a real number inside a real model published in July. It is either coming off your payroll or going into your output. Decide which. Then check, in two quarters, whether the ratio agrees with you.
Brightlume does this work with enterprise teams. If the gap between the AI pilot and the P&L is the problem you have, talk to me.

