AI Makes Code Cheap. Your Organization Is Still Expensive.

Crowds of fun-seekers exploring a city on foot, "

Arjan Franzen

29 September 2026

Illustratie: een koerier op een step levert een pakketje af bij een rij van vijf bureaus met stempels en kalenders, met een dichte deur aan het eind

In early 2025, METR ran a randomised controlled trial on sixteen experienced open-source developers, working on repositories they had contributed to for an average of five years. Tasks were randomly assigned to allow or forbid AI tools, mostly Cursor Pro with Claude 3.5 and 3.7 Sonnet. Before starting, the developers expected AI to cut their completion time by 24 percent. Afterwards, having done the work, they estimated it had cut it by 20 percent.

Measured, they were 19 percent slower.

Sixteen developers is a small sample. The tooling was early-2025. And these were people with deep knowledge of mature codebases, which is close to the least favourable case for AI assistance. The authors are careful about all of this, and so should you be. What the study establishes is narrower and more useful than "AI doesn't work": the people best placed to judge were wrong about their own working week, and wrong in the flattering direction.

I bring it up because almost every conversation I have about AI and delivery runs on self-report. Someone feels faster. The team feels faster. Nobody has looked.

The sketch everyone draws

When you do look, and the delivery numbers have not moved the way the feeling suggested, most engineering leaders reach for the same picture:

Idea → AI implementation: 20 minutes → organization: 12 days → production.

Those numbers are illustrative. I chose them to have something to point at, and nobody should repeat them as data. The shape, though, is one most people recognise on sight: the writing collapsed and everything around the writing did not. Refinement. The architecture call. The security review. The change board. The team that owns the other service. The reviewer who is in workshops until Thursday.

It is a satisfying picture, and it lets everyone off the hook. It is also not quite what the evidence says happened.

The organizations did get faster

DORA's 2025 report, State of AI-assisted Software Development, published on 23 September 2025, reversed one of its own earlier findings. The 2024 edition had associated AI adoption with decreased delivery throughput. The 2025 edition states: "Unlike last year, we observe a positive relationship between AI adoption on both software delivery throughput and product performance."

So the twelve days compressed. With 90 percent of respondents using AI at work, output went up and delivery throughput went up with it. If your mental model is that organizations are simply immovable, that is the number that should bother you.

The same report, in the next breath: "However, AI adoption does continue to have a negative relationship with software delivery stability."

What they paid for it

Read those two sentences together and the story stops being about speed. Organizations absorbed the extra output and shipped more of it. What they gave up was confidence that what they shipped worked.

DORA's explanation is a mechanism rather than a metaphor: "AI accelerates software development, but that acceleration can expose weaknesses downstream. Without robust control systems, like strong automated testing, mature version control practices, and fast feedback loops, an increase in change volume leads to instability."

The framing the report settles on is that AI acts as an amplifier. It does not fix an organization and it does not break one. It takes what the organization already does and does more of it, sooner. A team with fast tests, small changes and short feedback loops gets a faster version of itself. A team whose testing was already thin gets to discover that at four times the volume.

That is a better argument than the sketch, and a less comfortable one. Your organization is not expensive because it is slow. It is expensive because it is what gets multiplied.

The minutes are not where the days are

We have written about the tactical layer of this more than once. Waiting costs real money and almost nobody measures it. Review capacity becomes the binding constraint. Your CI/CD pipeline is probably the slowest thing you own. All of that is true and all of it is worth fixing.

But notice the unit. Pipelines and reviews are measured in minutes and hours. The twelve days in that sketch are not made of minutes. They are made of decisions: refinement that has not happened, an architectural question waiting on the one person who can answer it, a dependency on a team with its own quarter, an approval that exists because of an incident in 2019 nobody has revisited since.

Decision latency does not respond to throughput. It is set by calendars and by who is allowed to say yes. You can double your engineers' output without moving it at all, because the constraint was never capacity. AI does not amplify decisions either. It amplifies everything queued behind them, which is why the queue is now visibly longer.

What to measure before you invest

Three things, none of which need a programme:

  • Split lead time into phases. Writing, review, testing, waiting for a decision, deploying. A single commit-to-production number hides exactly what you need to see. Once the phases are separate, the largest one announces itself, and it is rarely the one people expected.
  • Watch change failure rate next to throughput. This is the DORA finding turned into a local instrument. Throughput rising on its own means little; throughput rising while failure rate rises means you are moving faster into the same wall.
  • Count the humans between knowing and shipping. From "someone on the team knows what to do" to "it is running in production", how many people have to agree? That number is your decision latency, and it is almost never on a dashboard.

We wrote separately about what happens to the DORA metrics themselves once AI is writing a serious share of your commits: the four numbers stay, their meanings shift.

What we do about this

ZEN's position on the first half of this argument is already on the record: code is getting cheap and engineering isn't. This is the second half. Cheap implementation is not a productivity result, it is an input. Whether it turns into a result depends on whether the organization around it can absorb change safely, and that is an engineering problem before it is a tooling one.

That is the work: AI-native software engineering, architecture, testing and delivery that can take the extra volume, and modernising what you already have so that changing it quickly is safe.

If your AI numbers look good and your delivery numbers don't, that gap is the interesting part. Talk to us, and bring the delivery numbers.

no image placeholder

The Agent Writes It. Who Reviews It?

How we can help

AI engineering

Use AI where it genuinely helps, with accountability staying with people.

See AI engineering →