Much more than a pilot project in your team’s innovation roadmap, AI is now writing production code at most companies you'd recognize. We partnered with Wakefield Research to survey 400 U.S. executives and engineering leaders and find out what's actually happening once that code ships. The answer isn't the clean AI success story told in a lot of leadership decks.
An immense 83% of organizations now deploy production code that's more than 10% AI-generated. At least some AI code generation is the baseline, not an edge case, and it lines up with what researchers have been finding elsewhere: A 2026 NBER working paper tracking more than 100,000 GitHub developers found AI coding tools drove a 741% increase in code output but only a 20% bump in actual software releases.
Code generation sprinted ahead, but verification lags way behind.
And that’s what we’re calling the code verification crisis: AI generating code faster than any human or automated system can confirm it's actually safe to ship. Read on to learn what the report’s data shows once that gap starts showing up in production.
Believe it or not, 80% of organizations surveyed have traced a real production incident, outage, or customer-impacting defect directly to AI-generated code. Seventeen percent said it's happened more than once. Only 19% said it's never happened to them at all.
And the people closest to the code see it more clearly than the people managing it from above. Engineering leaders report these incidents at a higher rate than executives (82% vs. 78%), which tracks since they're the ones getting paged.
The financial numbers back it up, with 65% of organizations saying their most significant software quality incident in the past year cost $500,000 or more, 41% crossing the $1 million mark, and 17% going past the $5-million mark. Not hypothetical what-ifs buried in a risk assessment but financial losses that already happened.
Even worse, the ramifications rarely stay contained to engineering. An almost unfathomable 95% of organizations that traced an incident back to AI-generated code went on to experience a legitimate business consequence, whether a regulatory inquiry, a lost contract, negative press, a stock hit, or an executive departure. Among organizations that haven't traced an incident, that number drops to 54%. Clearly, the technical problem doesn't stay technical for long.
Here's where the stats become uncomfortable: 92% of respondents expressed at least some concern that their organization's current safeguards wouldn't catch an AI-driven quality or safety issue before it reached users. Forty-one percent said they're extremely or very concerned — split evenly between executives and engineering leaders, indicating shared alarm across the org chart.
Worse yet, 66% of respondents said their teams have compromised on quality or testing standards to hit a release deadline in the past year, and 53% have shipped to production with known, unresolved issues at least some of the time. Put plainly: The people writing and shipping this code know the safety net has gaps, yet they're still climbing over it to hit the date.
The widely held belief is that AI will shrink QA headcount. It hasn't, as 64% of organizations report their dedicated QA or testing headcount actually grew this past year, and only 13% saw it shrink. But don't necessarily read that as an industrywide vote of confidence in human oversight. Among organizations that have already traced an incident to AI code, 67% grew their QA team, compared to 51% of those who haven't. Teams aren't hiring more testers because leadership trusts the process more. There's simply just more code to verify, and they don’t believe there’s a faster way to do it.
At the same time, 84% of organizations have eliminated or significantly reduced at least one role due to AI adoption. The cuts are hitting junior talent hardest, with 53% percent reducing or eliminating entry-level developer roles. Forty-two percent have cut manual QA testers. Thirty-four percent have reduced technical writers. Thirty-one percent have cut QA managers. Entry-level hiring is down at 42% of organizations overall. The org chart is adding senior oversight while closing the door on the training that used to develop the next generation.
On the bright side, 89% of organizations report positive ROI from their AI testing tools — 30% call it significantly positive. That figure is a resounding win on paper. But the same organizations reporting that ROI are, in large numbers, the same ones tracing incidents to AI code, worrying about their safeguards, and shipping with known defects.
We believe that pattern suggests a lot of ROI math is measuring speed and output — more code shipped, faster releases — rather than the outcomes that actually protect the business: fewer incidents, lower defect costs, a smaller blast radius when something breaks.
None of this means AI-generated code is a mistake. However, most organizations are running a verification process built for a slower, more human-paced era of software development, and asking it to keep up with the exponential volume of code that AI generates. That math demands infrastructure built for the scale at which AI actually operates.
That's the gap Sauce Labs built AURA to close. Customers using our AI-Unified Release Assurance platform have seen incident rates drop by more than 90%, releases move 47% faster, and 38% of engineering capacity reclaimed for their engineering teams — outcomes independently validated and worth a look regardless of where your organization sits on the data above.
The full report goes deeper into every stat here, plus the trust gap between executives and engineering leaders, the sectors respondents would trust software developed and tested entirely by AI, whether deploying AI too quickly or not deploying quickly enough poses the greater risk to quality, and where organizations are placing their bets for 2027.