
Stephan Schmidt, a developer with forty years of experience, described what AI-assisted coding feels like from the reviewer’s seat: “Again and again I feel fatigue. I finish a feature, and another feature, concentrate on reviewing the code the AI generated, and fix a bug and finish a feature with such velocity and I feel fatigue after some hours — sometimes as soon as one hour.”
The pace is the problem. “My brain does not get the baking time to mentally process architecture, decisions and edge cases the AI creates,” Schmidt wrote. “I’m running a marathon at the pace of a sprint — speeds don’t match.”
Gabriel Scherer, an OCaml maintainer, put it more bluntly when rejecting a massive AI-generated pull request: reviewing AI-generated code is “more taxing than reviewing human-written code.” And the volume keeps climbing.
Nobody announced this transition. No memo. No team meeting. No CTO stood on stage and said, “We’re converting our most experienced builders into full-time auditors.” It just happened, one AI-generated pull request at a time.
Who Absorbs the Output
The numbers tell a clean story until you think about who absorbs the output.
GitHub’s 2025 Octoverse report logged 518.7 million merged pull requests, a 29% year-over-year increase. Faros AI’s telemetry analysis of over 10,000 developers found that individual developers completed 21% more tasks and merged nearly twice as many PRs. AI made the pipeline faster. Everyone celebrated.
But the same analysis found that code review time increased 91%. PR sizes grew 154%. And organizational delivery metrics, the ones that actually measure whether software ships and works, stayed flat.
The SmartBear/Cisco study, one of the largest empirical studies on code review effectiveness, established clear limits: reviewers catch 70–90% of defects only when reviewing under 400 lines of code per session at under 500 lines per hour. Beyond that threshold, effectiveness collapses. After 60 to 90 minutes of continuous review, defect detection drops sharply. Their researchers put it bluntly: above 1,000 lines per hour, “you can probably conclude that the reviewer isn’t actually looking at the code at all.”
Those cognitive limits haven’t changed. PR volume has nearly doubled. PR sizes have more than doubled. The math doesn’t resolve in favor of quality.
Fastly’s 2025 developer survey quantified who bears the burden. Senior developers with ten or more years of experience ship 2.5 times more AI-generated code than juniors. And they’re the ones actually scrutinizing it: nearly 30% of seniors significantly edit AI output before shipping, compared to 17% of juniors. The experience that makes them better reviewers is exactly what makes the work slower: they catch problems juniors miss, so every review takes longer.
The pipeline got faster. The bottleneck became human.
Almost Right Is the Most Expensive Kind of Wrong
Broken code is easy to reject. A function that crashes, a query that returns nothing, a build that fails: quick decisions, low cognitive cost.
AI doesn’t produce broken code. It produces code that almost works.
The Stack Overflow 2025 Developer Survey asked 49,000 developers what frustrated them most about AI tools. The top answer, at 66%: “AI solutions that are almost right, but not quite.” Close behind, 45% said debugging AI-generated code takes longer than fixing what a human wrote.
Here’s why “almost right” is so expensive. You’re reading a function that looks correct. The logic flows. The tests pass. But something is off, and you can’t skim your way to it. You have to hold the intended behavior in your head, line by line, and spot where the AI’s version quietly diverges. That’s not reviewing. That’s reverse-engineering someone else’s confidence.
CodeRabbit’s 2025 analysis of 470 open-source pull requests found that AI-generated PRs contain 1.7 times more issues than human-written ones. 75% more logic errors. Security issues up to 2.7 times higher. Not catastrophic failures. Subtle ones. The kind that pass CI and haunt you in production six weeks later.
Simon Willison, co-creator of Django, named the dynamic plainly: “Almost anyone can prompt an LLM to generate a thousand-line patch and submit it for code review. That’s no longer valuable.” What’s valuable now is the person who can verify the output actually works.
The role didn’t change its title. Just its daily reality.
The Flow State Nobody Mourned
Software engineers know a state that psychologists spent decades trying to name. The hours that vanish while solving a hard problem. The architecture that reveals itself as you build it. The satisfaction of writing something elegant that works on the first run. For many, it’s why they chose this profession.
Mihaly Csikszentmihalyi called it flow: intense concentration, a merging of action and awareness, a sense of control, intrinsic enjoyment. His central finding was that flow happened most often during work, not leisure, and required a specific condition: voluntarily stretching your abilities toward something difficult and worthwhile.
Code review systematically destroys the conditions for flow.
Gloria Mark’s research at UC Irvine established that it takes an average of 23 minutes and 15 seconds to regain deep focus after a single interruption. Chris Parnin’s study of 10,000 programming sessions at Georgia Tech found that developers typically get one uninterrupted two-hour block per day. After an interruption during a method edit, only 10% of programmers resumed within a minute. Most needed 10 to 15 minutes to start editing code again.
Review work is reactive by definition. You are evaluating someone else’s output on someone else’s timeline. Each PR is a context switch. Each context switch resets the clock on deep focus. A senior engineer reviewing a dozen PRs a day isn’t experiencing a dozen reviews. They’re experiencing a dozen context resets. Each costs ten to twenty-three minutes of cognitive recovery. And most workdays contain only one reliable window for uninterrupted thought.
The industry measures task completion going up. What it doesn’t capture: whether the tasks senior engineers are completing are the ones that require their expertise, or the ones the review queue assigned them by default.
Flow requires stretching your abilities toward something difficult. Auditing whether a machine stretched correctly is not the same thing.
The Identity Underneath the Role
Annie Vella named it directly: “We’re shifting from creators to orchestrators, from builders to overseers. From engineers to something that looks suspiciously like… managers.”
The identity issue is specific, and nobody talks about it. People became software engineers because they loved building things. Not managing things. Not reviewing things. Building. That craft, the late nights debugging a system until it finally clicks, the satisfaction of an architecture that holds under load, wasn’t just a job. It was how they understood themselves. A decade of deliberate practice, of stubborn debugging, of the particular satisfaction that comes from shipping something that works. That practice built expertise and identity simultaneously.
AI didn’t eliminate the expertise. It redirected it. The same deep knowledge that made them valuable as builders is exactly what makes them indispensable as reviewers. They’ve seen bad abstractions rot a codebase. They know where subtle bugs hide because they’ve been burned by those bugs before. The system needs their judgment more than ever.
But judgment deployed as oversight feels nothing like judgment deployed as creation. Building something is not the same as catching what someone else built wrong.
The Stack Overflow 2024 survey found that only 20% of professional developers reported being happy at work. A third were actively unhappy. The rest described themselves as complacent — not engaged, not leaving, just there. The LeadDev 2025 report found 22% of engineering leaders at critical burnout, 40% saying their teams are less motivated than a year ago. Sixty percent say AI hasn’t significantly boosted team productivity.
Nobody is asking whether the work changed. They’re measuring whether the output increased.
What the Dashboards Don’t Show
Here is the organizational paradox: every metric the AI era optimizes for suggests improvement. More PRs merged. More tasks completed. More code shipped. The dashboards are green.
GitClear found duplicated code blocks rose eightfold in their dataset of 211 million lines. Copy-pasted code surpassed refactored code for the first time. The codebase is growing faster than ever. Its quality is degrading at the same rate.
Addy Osmani, engineering leader at Google Chrome, identified the shift: “The bottleneck moved from writing code to proving it works.” He predicted the emergence of a formal role, “AI code auditor,” to handle the growing volume.
The prediction reveals the problem. When the solution to a broken workflow is creating a new job title for the people trapped in it, the workflow isn’t being fixed. It’s being institutionalized.
More code is being written. Less of it is being understood by the people responsible for shipping it. The senior engineers who could close that gap are spending their days in the review queue instead of at the whiteboard. And the juniors who should be learning from those seniors are generating more code for the queue instead of learning why the existing code works.
The Job That Changed Without a Memo
Nobody told the senior engineers they’d become auditors. There was no announcement, no transition plan, no conversation about whether this was the best use of their expertise. The tools got faster, the pipeline got wider, and the most experienced people in the organization became the last checkpoint before production.
Some are adapting. They’re learning to triage reviews, to set boundaries on queue depth, to carve out protected time for the architectural work that still needs human judgment. Others are burning out quietly, their dissatisfaction invisible inside dashboards that only track throughput.
The question isn’t whether AI is useful. It’s whether anyone asked the senior engineers if this is the job they wanted.