RESEARCH

30 Stats Every Engineering Leader Should Know About AI-Era Software Quality

Hundreds of reports exist on AI and software quality, most published by people selling a fix for it. Here are 30 independant reports that you should know about AI-Era Software Quality.

Most statistics in this field do not survive being looked up. The claim that a production bug costs 100 times more to fix than one caught in design traces back not to research but to IBM training-course notes from 1981, repeated for forty years as though it were a study. Every figure below names its source, its year and its sample size, and links to somewhere you can check it.

The AI code-quality gap

93% of software quality leaders have already adopted AI coding tools. Survey of 273 CTOs, CIOs, developers and QA directors, January 2026. SmartBear.

40% now generate more than 40% of their code with AI, and 60% expect to within a year. Same 273-respondent survey. SmartBear.

AI-co-authored pull requests carry 1.7 times more issues than human ones. Across 470 open-source pull requests: 10.83 issues per PR against 6.45. CodeRabbit inferred authorship rather than confirming it. The Register.

46% of developers distrust the accuracy of AI output, against 33% who trust it. Up from 31% distrust a year earlier. 33,244 answered that question, of 49,009 respondents in 177 countries. Stack Overflow 2025.

"Almost right, but not quite" is developers' single biggest frustration with AI tools. Chosen by 66% of 31,476 developers, the most-selected of seven options. Stack Overflow 2025.

45.2% say debugging AI-generated code is more time-consuming. The runner-up frustration in the same survey. Stack Overflow 2025.

A third of companies could not tell whether AI code caused their own incident. 34% of those that had an incident could not establish it within 24 hours. Harris Poll, 1,528 respondents, six countries. GitLab.

What the tests themselves cost

84% of Google's pass-to-fail test transitions involved a flaky test. Against a steady background rate of 1.5% of all test runs. Google Testing Blog, 2016.

End-to-end tests flake 28 times more often than unit tests. Across Google's 4.2 million tests in one week: 0.5% of small tests against 14% of large ones. Android emulator tests reached 25.46%. Google Testing Blog, 2017.

Only 1.23% of Google's test executions ever found a real breakage. In a month covering 4 billion test results, 2.07% of targets had both passed and failed; 1.23% survived flaky filtering. Memon et al., ICSE-SEIP 2017.

Up to half of all builds contain a flaky failure. Between 14% and 52% depending on the project, across five Microsoft projects over 30 days. 4.6% of individual tests were flaky. Lam et al., ISSTA 2019.

86% of flaky tests only flake in CI. Of 315 collected in a single day, most stopped behaving non-deterministically when re-run 100 times locally. Lam et al., ISSTA 2019.

Two thirds of Mozilla's job failures were not real failures. 25,871 of 38,596 failures across nearly 2 million runs, each classified by hand rather than by heuristic. Lampel et al., ESEC/FSE 2021.

Flaky tests outnumbered genuine failures four to one at Facebook. In the CI data sampled for the study, where almost 99.9% of dependency-selected targets passed. Machalica et al., ICSE-SEIP 2019.

Main-branch build success has fallen to 70.8%, a five-year low. Average daily workflow runs rose 59% year on year, across 28.7 million workflows in September 2025. CircleCI.

What failure costs

Poor software quality costs the US at least $2.41 trillion. Plus $1.52 trillion of accumulated technical debt, reported separately. A modelled estimate from 2022, with no later edition. CISQ.

The "100x" cost of fixing a bug late is nearer 30x, depending who you ask. Boehm's 1976 figures give 0.5 in design against 15 at installation. The same table's Baziuk column reaches 470 to 880 times. NIST/RTI, 2002.

The average US data breach now costs a record $11.5 million. Global average $4.99 million, across 602 organisations breached to February 2026. IBM calls its own sample "nonstatistical". IBM.

Downtime costs Global 2000 companies $600 billion a year. Up 50% in two years. Roughly $300 million per company, or $15,000 for every minute one is down. 2,000 executives, 20 countries. Splunk and Oxford Economics.

One bad update cost the Fortune 500 an estimated $5.4 billion. CrowdStrike, July 2024, excluding Microsoft. A weighted average of $44 million across 125 firms, modelled by an outage insurer. Parametrix.

Delta lost around $500 million from that single incident. $380 million of revenue and $170 million of non-fuel cost, offset by $50 million of unburned fuel, across 7,000 cancellations in five days. Delta, SEC filing.

UK banks logged 803 hours of unplanned outage in two years. More than 33 days, across at least 158 incidents at nine institutions, on their own returns to Parliament. Treasury Committee, 2025.

Customer and revenue impact

A 31% faster page lifted Vodafone's sales by 8%. A genuine 50/50 A/B test, roughly 34,000 visits per arm per day, all paid media traffic. Google web.dev.

A tenth of a second of mobile speed was worth 8.4% more retail conversions. And 9.2% higher order value, across 20.5 million sessions at 15 brands. A regression on natural variation, not an experiment. Deloitte for Google.

A 1.09% crash rate is enough for Google Play to bury your app. Above that share of daily users an app is "likely to be less discoverable", with a possible store-listing warning. Google Play Console.

Shoppers cut spending after 47% of bad experiences. Down 8 points year on year, rising to 58% in online retail and 62% in fast food. 20,001 consumers, 14 countries, Q3 2025. Qualtrics XM Institute.

Market adoption and investment

Gartner expects 70% of enterprises to have AI-augmented testing tools by 2028. Up from about 20% in early 2025, revised down from an earlier forecast of 80% by 2027. Magic Quadrant, 6 October 2025, ID G00828088. Gartner.

64% of technology executives plan to deploy agentic AI within 24 months. 2,501 CIOs and technology executives, fielded May and June 2025. Gartner.

Only 15% of companies have scaled AI in quality engineering enterprise-wide. 43% are still experimenting and 30% at operational deployment. Non-adopters rose from 4% to 11%. 2,000 interviews, 23 countries. World Quality Report 2025-26.

86% are raising testing budgets by more than 11% this year. And 92% see autonomous testing as a quality improvement. Same 273-respondent survey. SmartBear.

Ship Faster,
Ship Fearlessly

Product Verification Infrastructure

Ship Faster,
Ship Fearlessly

Product Verification Infrastructure