Arrio

Measuring software development

The AI productivity paradox

More is being produced than at any point in the history of software, and leadership can account for less of it than ever. The tools feel transformative. The delivery numbers, measured honestly, often do not move. This is what the evidence shows, and what to do about it.

The paradox, defined

The felt gain and the measured gain have come apart.

AI coding tools make individual developers feel dramatically faster, and they do produce far more output. But productivity is not output, it is output that reaches an outcome. When the extra code creates larger changes, longer reviews, more defects and more rework, the felt speed is spent before it arrives.

The result is a paradox leadership can see in the mood of a team but not in its results: everyone is moving faster, and nothing is arriving sooner.

A quarter of production code is now AI-authored, and productivity moved about 10%. (DX, 2026)
26.9% / 10%
A quarter of production code is now AI-authored, and productivity moved about 10%. · DX, 2026
More incidents per pull request, even as task throughput rose 33.7% with AI. (Faros AI, 2026)
+242.7%
More incidents per pull request, even as task throughput rose 33.7% with AI. · Faros AI, 2026
Slower in measured time when experienced developers used AI on familiar code, though they felt 24% faster. (METR, 2025)
19%
Slower in measured time when experienced developers used AI on familiar code, though they felt 24% faster. · METR, 2025

Usage up, trust down

Adoption is near total, and confidence is falling.

84% of developers now use AI coding tools, and 42% of committed code is AI-generated or assisted (Stack Overflow, 2025; SonarSource, 2026). Adoption is no longer the question.

Trust is. Developer trust in AI accuracy fell from 43% to 29% in a single year (Stack Overflow, 2025), 96% do not fully trust AI output, and two-thirds say they spend more time fixing almost-right code than they would have spent writing it. The people closest to the work are the least convinced by the output.

The gain is also unevenly distributed. A randomised trial across roughly 5,000 developers found an average 26% productivity gain, but juniors gained 35 to 39% while seniors gained only 8 to 16%, and around a third of engineers did not adopt the tools at all (Microsoft and Accenture, 2024).

An organisation reporting high AI adoption can still be seeing almost none of the value, and have no way to tell which.

More code, more rework

The volume arrives. The quality bill arrives just after.

Faster generation does not remove the work of getting software right, it moves it downstream and enlarges it. Code churn rose 861% after AI adoption, defects rose even in healthy codebases, and a large share of AI-generated code shipped with a security vulnerability. The saving at the keyboard is repaid, with interest, in review, in debugging and in the debt left behind.

Task throughput rose by a third. Incidents per pull request rose by nearly two and a half times.

Faros AI Engineering Report 202622,000 developers, 4,000 teams, two years of telemetry
Rise in code churn after AI adoption: written fast, then rewritten. (Faros AI, 2026)
+861%
Rise in code churn after AI adoption: written fast, then rewritten. · Faros AI, 2026
More AI-induced defects, even in codebases rated healthy. (CodeScene, 2026)
+60%
More AI-induced defects, even in codebases rated healthy. · CodeScene, 2026
Of AI-generated code introduced a security vulnerability. (Veracode, 2025)
45%
Of AI-generated code introduced a security vulnerability. · Veracode, 2025

Why it happens

Three mechanisms turn speed into motion.

Effort stopped tracking output

Every legacy measure, story points, tickets, hours, assumed that more effort meant more result. AI severed that link. The instruments still read effort, so they now measure the wrong thing.

The verification tax

Code that is almost right is expensive. Someone has to read it, test it and trust it. Two-thirds of developers say fixing almost-right AI code costs them more than writing it would have, and that cost is invisible to output counts.

The review bottleneck

More and larger changes pile up at the one stage that did not get faster: human review. Larger pull requests and longer reviews convert individual speed into organisational queueing.


What to do about it

Measure the value the AI produces, not the volume.

The paradox is not an argument against AI. It is an argument against flying blind. The organisations that win the AI shift will be the ones that can see, independently, which of their teams and vendors turn AI into delivered value and which turn it into rework.

That requires measuring the AI-built share of the work against what it delivers, read from the work itself, at the level of teams and vendors, and tracked over time. Not licences bought, not lines generated, not a survey of how it feels.

That is the read Arrio provides. Independent, from the code, in business terms. See how measurement works, or how it applies to software development ROI.


Questions

The questions worth asking

Related reading: measuring software developmentin full, and the evidence behind these figures. Defined plainly:AI productivity paradox.

What is the AI productivity paradox?

It is the gap between how productive AI coding tools feel and what they measurably deliver. Developers report moving faster and produce far more output, yet at the level of the whole organisation, delivery, quality and business outcomes often do not improve to match. More is produced than ever, and less of it can be accounted for.

Is AI-assisted coding actually making developers faster?

It depends what you measure. In a controlled study, experienced developers working on familiar codebases were 19% slower with AI while believing they were 24% faster (METR, 2025). A large randomised trial found an average 26% gain, but concentrated in juniors, with seniors seeing 8 to 16% and a third of engineers not adopting the tools at all (Microsoft and Accenture, 2024). Speed varies by context, and the felt gain consistently overstates the measured one.

If more code ships, why does delivery not improve?

Because volume is not value. Across 22,000 developers and 4,000 teams, task throughput rose 33.7% with AI while incidents per pull request rose 242.7% and review times rose 441% (Faros AI, 2026). A quarter of production code is now AI-authored, and productivity moved about 10% (DX, 2026). The extra output creates larger changes, longer reviews and more rework, which absorb the gains before they reach the outcome.

Is AI making code quality worse?

The evidence points to more rework and more risk. Code churn rose 861% after AI adoption (Faros AI, 2026), AI-induced defects rose 60% even in healthy codebases (CodeScene, 2026), and 45% of AI-generated code introduced a security vulnerability (Veracode, 2025). Trust reflects it: developer trust in AI accuracy fell from 43% to 29% (Stack Overflow, 2025), and 96% do not fully trust AI output (SonarSource, 2026).

How do we know whether our own AI investment is paying off?

By measuring the AI-built share of your work against what it actually delivers, over time, rather than by counting adoption or output. That means an independent read of output, quality and cost at the level of teams and vendors, so you can spread what works and stop what does not. Counting licences or pull requests will not tell you.

Sources

  1. METR, 2025 Experienced developers were 19% slower with AI on familiar codebases, while feeling 24% faster.
  2. Faros AI, 2026 Engineering Report 2026 (Acceleration Whiplash), 22,000 developers and 4,000+ teams over two years: task throughput up 33.7% and epics per developer up 66%, against code churn up 861%, incidents per pull request up 242.7%, review time up 441% and bugs per developer up 54%.
  3. DX, 2026 Survey of 121,000 developers across 450+ companies: 26.9% of production code is AI-authored, up from 22% the previous quarter, while productivity gains plateaued at about 10% despite 93% adoption.
  4. Microsoft and Accenture RCT, 2024 26% average productivity gain; juniors 35 to 39%, seniors 8 to 16%; around a third did not adopt.
  5. CodeScene, 2026 60% increase in AI-induced defects even in codebases rated healthy.
  6. Veracode, 2025 45% of AI-generated code introduced a security vulnerability.
  7. Stack Overflow Developer Survey, 2025 84% use AI coding tools; trust in accuracy fell from 43% to 29%; 66% spend more time fixing almost-right code.
  8. SonarSource State of Code, 2026 42% of committed code is AI-generated or assisted; 96% do not fully trust AI output.

Find out whether your AI investment is actually paying off.