The evidence, 2026
The state of software development in the AI era
Spending is rising faster than the ability to prove its return. This is the public evidence, gathered in one place and cited: the scale of the investment, the ROI gap, the productivity paradox, the quality bill, and the stakes of getting it wrong.
About this page
Every figure below comes from a named third party and is cited as such. Arrio gathered and organised this evidence; we did not produce it. Figures run from 2024 to 2026, weighted to the most recent, and where a newer figure supersedes an older one the newer is used. The full source list is at the foot of the page.
Arrio does produce its own measurement, continuously, inside client engagements. None of it appears here. What a client's software produces belongs to that client, and we do not publish it. This page is the public evidence, kept deliberately separate.
01 / The scale of the investment
More money than ever is going into software and AI.
- Forecast worldwide IT spending in 2026, up 14.2% on 2025. (Gartner, 2026)
- $6.37tn
- Forecast worldwide IT spending in 2026, up 14.2% on 2025. · Gartner, 2026
- Forecast worldwide AI spending in 2026, up 47% year on year. (Gartner, 2026)
- $2.59tn
- Forecast worldwide AI spending in 2026, up 47% year on year. · Gartner, 2026
- Of enterprises expect their AI budget to rise in 2026. (NVIDIA, 2026)
- 86%
- Of enterprises expect their AI budget to rise in 2026. · NVIDIA, 2026
02 / The ROI gap
The return is rising more slowly than the spend, and almost no one can prove it.
- Of enterprise AI pilots fail to deliver measurable financial return. (MIT, 2025)
- 95%
- Of enterprise AI pilots fail to deliver measurable financial return. · MIT, 2025
- Of enterprises achieve substantial ROI from their AI investment. (Master of Code, 2026)
- 5%
- Of enterprises achieve substantial ROI from their AI investment. · Master of Code, 2026
- Of CIOs name the assessment of technology ROI as a point of contention with their CFO. (KPMG, 2025)
- 49%
- Of CIOs name the assessment of technology ROI as a point of contention with their CFO. · KPMG, 2025
- Of leaders who see AI productivity gains can link them to a business outcome. (Deloitte, 2025-2026)
- <33%
- Of leaders who see AI productivity gains can link them to a business outcome. · Deloitte, 2025-2026
- Of CFOs are satisfied with their visibility into software spend. (KPMG, 2025)
- 20%
- Of CFOs are satisfied with their visibility into software spend. · KPMG, 2025
- Of AI initiatives deliver the returns organisations expected. (IBM, 2025)
- 25%
- Of AI initiatives deliver the returns organisations expected. · IBM, 2025
- Of CEOs say AI delivered both revenue growth and cost reduction. 56% saw neither. (PwC, 2026)
- 12%
- Of CEOs say AI delivered both revenue growth and cost reduction. 56% saw neither. · PwC, 2026
03 / The AI productivity paradox
Adoption is near total. The measured gain is not.
- Slower, in measured time, when experienced developers used AI on familiar code, though they felt 24% faster. (METR, 2025)
- 19%
- Slower, in measured time, when experienced developers used AI on familiar code, though they felt 24% faster. · METR, 2025
- More incidents per pull request, even as task throughput rose 33.7% with AI. (Faros AI, 2026)
- +242.7%
- More incidents per pull request, even as task throughput rose 33.7% with AI. · Faros AI, 2026
- A quarter of production code is now AI-authored, and productivity moved about 10%. (DX, 2026)
- 26.9% / 10%
- A quarter of production code is now AI-authored, and productivity moved about 10%. · DX, 2026
- Of developers now use AI coding tools. (Stack Overflow, 2026)
- 84%
- Of developers now use AI coding tools. · Stack Overflow, 2026
- Of committed code is now AI-generated or assisted. (SonarSource, 2026)
- 42%
- Of committed code is now AI-generated or assisted. · SonarSource, 2026
- Heavy AI users outproduce non-users by 4 to 10x. Against their own past output, the gain is 25%. (GitClear, 2026)
- 4-10x / 25%
- Heavy AI users outproduce non-users by 4 to 10x. Against their own past output, the gain is 25%. · GitClear, 2026
The fuller picture is in the AI productivity paradox.
04 / The trust gap
Eighty-four percent use it. Three percent trust it.
The clearest signal in the 2026 surveys is not adoption, which is close to universal. It is the distance between how much AI is used and how much it is believed. Teams are shipping work they do not fully trust, and checking it consumes the time the tools were meant to save.
- Of developers highly trust AI-generated code, though 84% use it. (Stack Overflow, 2026)
- 3%
- Of developers highly trust AI-generated code, though 84% use it. · Stack Overflow, 2026
- Actively distrust the accuracy of AI output. (Stack Overflow, 2026)
- 46%
- Actively distrust the accuracy of AI output. · Stack Overflow, 2026
- Name "almost right, but not quite" as their single biggest frustration. (Stack Overflow, 2026)
- 66%
- Name "almost right, but not quite" as their single biggest frustration. · Stack Overflow, 2026
- Of tasks are fully delegated to AI, even where engineers use it in 60% of their work. (Anthropic, 2026)
- 0-20%
- Of tasks are fully delegated to AI, even where engineers use it in 60% of their work. · Anthropic, 2026
- Daily AI users merge 60% more pull requests. 29% trust the output. (DX and Atlassian, 2026)
- 60% / 29%
- Daily AI users merge 60% more pull requests. 29% trust the output. · DX and Atlassian, 2026
05 / The quality bill
The volume arrives fast. The quality bill arrives just after.
- Rise in code churn after AI adoption: written fast, then rewritten. (Faros AI, 2026)
- +861%
- Rise in code churn after AI adoption: written fast, then rewritten. · Faros AI, 2026
- More AI-induced defects, even in codebases rated healthy. (CodeScene, 2026)
- +60%
- More AI-induced defects, even in codebases rated healthy. · CodeScene, 2026
- Of AI-generated code introduced a security vulnerability. (Veracode, 2025)
- 45%
- Of AI-generated code introduced a security vulnerability. · Veracode, 2025
- Of new code is now rewritten within two weeks, up from 3.3% before AI. (GitClear, 2026)
- 7.1%
- Of new code is now rewritten within two weeks, up from 3.3% before AI. · GitClear, 2026
- Increase in duplicated code blocks. (GitClear, 2026)
- +81%
- Increase in duplicated code blocks. · GitClear, 2026
- Of developers do not fully trust AI output. (SonarSource, 2026)
- 96%
- Of developers do not fully trust AI output. · SonarSource, 2026
06 / Why leadership does not see it
AI improves exactly the things most organisations already measure.
This is the part that makes it a measurement problem rather than a tooling one. AI raises speed and volume, which is what existing dashboards count, while the costs it creates land downstream in review, rework, incidents and comprehension, which they do not count. The dashboards improve and the outcomes do not, and nothing in the reporting shows the difference (EAB Global, 2026). Thoughtworks has a name for the newest version of this: cognitive debt, the widening gap between what a system does and what the team still understands about how it works (Technology Radar, 2026). DORA puts it more bluntly still: AI is a mirror and a multiplier, amplifying whatever an organisation already is.
- Theoretical AI capability against observed real-world use in technical roles. The gap is adoption, not capability. (Anthropic, 2026)
- 94% / 33%
- Theoretical AI capability against observed real-world use in technical roles. The gap is adoption, not capability. · Anthropic, 2026
- Of developers will spend more time orchestrating and architecting than writing code by the end of 2026. (Gartner, 2026)
- 75%
- Of developers will spend more time orchestrating and architecting than writing code by the end of 2026. · Gartner, 2026
- Of leaders reporting AI productivity gains can link them to a business outcome. (Deloitte, 2025-2026)
- <33%
- Of leaders reporting AI productivity gains can link them to a business outcome. · Deloitte, 2025-2026
07 / The stakes
Getting software wrong is expensive, and increasingly measurable.
- Of deal value is what failed technical due diligence costs acquirers, on average. (Deloitte and Patsnap, 2026)
- 25%
- Of deal value is what failed technical due diligence costs acquirers, on average. · Deloitte and Patsnap, 2026
- Estimated annual waste from engineers contributing little or no meaningful work. (Stanford, 2024)
- $90bn
- Estimated annual waste from engineers contributing little or no meaningful work. · Stanford, 2024
- The estimated cost of poor software quality in the US. (CISQ, 2022)
- $2.41tn
- The estimated cost of poor software quality in the US. · CISQ, 2022
Is this Arrio's own data?
No, and deliberately so. Every figure on this page comes from a named third party, such as Gartner, MIT, PwC, Deloitte, METR, GitClear, Faros AI and Stack Overflow, and is cited as such. Arrio gathered and organised the evidence; we did not produce it.
Does Arrio have data of its own?
Yes. Arrio produces measurement continuously inside client engagements, and it is the most detailed view of what software development produces that we have seen. None of it appears here. What a client's software produces belongs to that client, and publishing it is not something we do. Where we publish original research, it is drawn from open-source repositories and is labelled as such.
What is the single clearest finding?
That spending on software and AI is rising fast while the ability to prove its return is not. 95% of AI pilots fail to deliver measurable return, only 5% of enterprises see substantial AI ROI, and fewer than a third of leaders who report productivity gains can connect them to a business outcome. The gap is measurement.
Does the evidence say AI is not worth it?
No. It says the value is real but uneven, and mostly unmeasured. Adoption is near total and output is up sharply, yet delivery and quality often do not follow, and almost no one can see which of their teams turn AI into value. The conclusion is not to stop, it is to measure.
How current are these figures?
They are drawn from 2024 to 2026 sources, weighted towards the most recent. Each entry names its source and year so you can weigh and check it. Where a newer figure supersedes an older one, we use the newer.
Sources
- Gartner, 2026 Worldwide IT spending $6.15 trillion; worldwide AI spending forecast $2.52 trillion, up 44%.
- MIT, 2025 95% of enterprise AI pilots fail to deliver measurable ROI.
- Master of Code, 2026 Only 5% of enterprises achieve substantial ROI from AI.
- IBM, 2025 Only 25% of AI initiatives delivered the expected ROI (2025 CEO Study, 2,000 CEOs).
- KPMG, 2025 49% of CIOs (against 39% of CFOs) name the assessment of technology ROI as a point of contention; 20% of CFOs satisfied with visibility.
- Deloitte, 2025-2026 79% report productivity gains; fewer than 33% can link them to business outcomes.
- METR, 2025 Experienced developers 19% slower with AI on familiar codebases, while feeling 24% faster.
- Faros AI, 2026 Engineering Report 2026 (Acceleration Whiplash), 22,000 developers and 4,000+ teams over two years: task throughput up 33.7% and epics per developer up 66%, against code churn up 861%, incidents per pull request up 242.7%, review time up 441% and bugs per developer up 54%.
- DX, 2026 Survey of 121,000 developers across 450+ companies: 26.9% of production code is AI-authored, up from 22% the previous quarter, while productivity gains plateaued at about 10% despite 93% adoption.
- Stack Overflow Developer Survey, 2026 49,000 respondents across 177 countries: 84% use AI coding tools, only 3% highly trust AI-generated code and 46% actively distrust it; 66% name "almost right, but not quite" as their biggest frustration.
- PwC Global CEO Survey, 2026 4,454 CEOs across 95 countries: 56% see neither revenue gains nor cost reductions from AI; only 12% report both.
- Anthropic, 2026 Agentic Coding Trends: engineers use AI in 60% of their work but fully delegate only 0 to 20% of tasks. Economic Index (July 2026): theoretical capability in computer and maths roles is 94% against 33% observed real-world exposure.
- DX and Atlassian, 2026 State of AI Impact Q2 2026, 135,000 developers across 450+ companies: daily AI users merge 60% more pull requests and save 3.6 hours a week, while only 29% trust the output.
- Gartner, 2026 Planning Guide for Software Engineering: by the end of 2026, 75% of developers will spend more time orchestrating and architecting than writing code directly.
- EAB Global, 2026 The AI Productivity Paradox: AI failures are hard to spot because AI improves what organisations already measure, speed and volume, while the costs land downstream.
- Thoughtworks Technology Radar, 2026 Volume 34: "cognitive debt", the growing gap between a system's implementation and the team's understanding of how and why it works.
- DORA, 2025 AI is "a mirror and a multiplier" of existing strengths and weaknesses; it does not automatically improve delivery.
- SonarSource State of Code, 2026 42% of committed code is AI-generated or assisted; 96% do not fully trust AI output.
- Stanford, 2024 9.5% of engineers contribute minimal work; an estimated $90 billion wasted annually (Yegor Denisov-Blanch).
- CodeScene, 2026 60% increase in AI-induced defects even in codebases rated healthy.
- Veracode, 2025 45% of AI-generated code introduced a security vulnerability.
- GitClear, 2026 Two studies. AI adoption research (2,172 developer-weeks): heavy AI users outproduce non-users by 4 to 10x, but against their own past output the velocity gain is 25%, so most of the gap pre-dated AI. Maintainability Gap study, 211M lines of code: the share of new code rewritten within two weeks rose from 3.3% to 7.1%; duplicated code blocks up 81%; cross-file reuse down 35%; refactoring moves down 70%.
- Deloitte and Patsnap, 2026 Failed technical due diligence costs acquirers around 25% of deal value.
- CISQ, 2022 Cost of poor software quality in the US at least $2.41 trillion, of which ~$1.52 trillion is accumulated technical debt.
The evidence is clear. The next step is to measure your own.


