Skip to content

So you approved a pull request last week that you did not read. Not entirely.

It’s fine. “Everyone does it,” you tell yourself. The CI was green. The tests passed. You gave it a scroll, a nod, an LGTM. Your name is on the commit now. You told yourself you reviewed it. Move on to the next item on your plate.

This is not a story about bad engineers. It is a story about a phrase.

“Human in the loop” used to name a specific act. A senior engineer reading a diff, understanding the change, pushing back when the change was wrong. That is not what happened in your PR. What happened was a ritual that kept the shape of review and lost the function. The ritual is still called “human in the loop” because the phrase is comforting. The phrase is comforting because admitting what you actually did would mean admitting something harder.

The Psychological Comfort Blanket

“Human in the loop” has quietly shifted from a genuine safety mechanism to a psychological comfort blanket. The loop implies a human is actively deciding. In practice, the human is confirming. The loop is a confirmation click. The phrase exists because removing it would feel worse than failing to fulfill it.

Sonar’s 2026 developer survey is the most honest document the industry produced this year. 96% of developers say they do not fully trust that AI-generated code is functionally correct. 48% say they always verify it before committing.

The 48-point gap is the article.

It is the distance between what developers admit they should do and what they actually do, collected by a company with a financial stake in highlighting the gap. That commercial interest does not make the number wrong. It makes the number reported against interest. Sonar sells code verification tools. Their target users, in their own survey, publicly admit that only 48% always verify before committing. The company whose product exists to close that gap is telling you the gap is open.

Formal Authority vs Functional Authority

There is a name for what is happening in those reviews. Not laziness, not recklessness, but a gap between formal authority and functional authority.

Formally, you approved the change. Functionally, the AI decided.

The branch your name is on represents a decision you did not make, shipped through a ritual that used to mean something. The organization sees a reviewed PR. Your git blame shows your commit. The only person in the loop who knows what actually happened is you.

Every approval layer above engineering has a version of this problem, and every one of them has a way out. A CEO formally approves decisions functionally made by VPs, but the CEO can ask “walk me through the reasoning.” A VP can say “what did you consider and reject.” The formal approver, at every layer, can interrogate the functional decider. That is what the word “approval” has always meant.

An AI does not stand in the room. An agent’s reasoning is not preserved past the token it generated. The approver in an AI PR review is the first formal approver in any approval chain who cannot interrogate the decider, because the decider is not available for questioning.

You did not review a decision. You reviewed an artifact.

Those are different things, and the phrase “human in the loop” keeps both of them behind the same curtain.

LGTM Is Older Than AI. What Changed Is Volume.

Rubber-stamp review is not an AI problem. It predates Copilot by twenty years. Everyone agrees code review is essential. Under deadline pressure, “LGTM” becomes the default. The process exists. The substance erodes.

What changed is volume.

When humans wrote the code, the rate at which a team produced diffs was rate-limited by human typing speed. Skipping review on one PR left the next PR still human-authored. The rubber-stamp rate was a tax on careful engineering, annoying but bounded. Now an agent can open ten pull requests in the time it takes to read one. The tax became a structural floor.

A 2026 empirical study of 11,048 agentic pull requests found that 15.4% required explicit reviewer involvement through feedback or direct commits. Another 5.5% showed no visible interaction trace at all between author and reviewer. The rest sit in the middle. Approvals. Green checks. Two-word comments. Scrolls.

The pattern varies by tool in a way that should be uncomfortable if you use any of them. PRs from Copilot and Devin tended to show reviewer-mediated workflows. PRs from Codex and Cursor typically merged with minimal interaction. The rubber stamp is not evenly distributed. It tracks which tool is authoring the code more than which team is reviewing it.

The paper is careful about what those numbers do and do not mean. “No explicit involvement” is not “no review.” Someone may have read. Someone may have thought. The defining act of software engineering now requires a “may have.” That should bother you more than the percentage.

The next study is worse.

A 2026 METR study tested this gap directly. Four active maintainers of scikit-learn, Sphinx, and pytest reviewed 296 AI-generated pull requests across five frontier models, with 47 human-written patches mixed in as controls. They did not know which was which. Their merge rate on the AI patches ran 24 percentage points below what those same patches scored on SWE-bench, the AI coding benchmark. SWE-bench said the patches passed. The humans who would have to live with the code said they would not merge. That 24-point gap is the space between what automated benchmarks catch and what a careful human reviewer catches. It is also the space a rubber-stamp review leaves behind.

The Cognitive Cost

Underneath all of this is a math problem that goes unnamed.

Genuine review of AI-generated code is more expensive than writing the code would have been. To review it, you have to build the mental model the generator built and discarded. You load the library it chose, read the test it wrote, and do the thinking the author did in reverse, starting from the output. Then you render a binary judgment at the end of a nonlinear cognitive chain.

Approving is one click.

The cognitive cost of genuinely evaluating AI output often exceeds the cost of just trusting it. Every individual review decision is made under time pressure, with seven other PRs in the queue. Each decision is rational for the individual. The aggregate is what breaks. Ten thousand individually rational approvals become an organization that stopped reviewing code.

The Sonar survey has a quieter finding inside it that makes this worse. Junior developers report the highest productivity gains from AI (40%). They also report the highest concern that AI produces code that looks correct but isn’t reliable: 66% of juniors agree, compared to 48% of seniors with twenty-plus years of experience. The people getting the most speed from the tool are the same people most worried the output they are shipping is not dependable. They ship it anyway.

Seniors are less worried on every axis. Reliability, volume, codebase understanding: every concern drops as years of experience rise. That is not seniority making people calmer. That is two groups with two different rates of contact with AI output. The people with the judgment to catch subtle problems are not writing most of the code. The people writing most of the code do not yet have the judgment. That is not a generational complaint. That is a staffing model that does not resolve.

Vibe Coding Got a Dictionary Entry. The Caveat Did Not.

The cultural moment this phrase is trying to govern has a specific origin. In February 2025, Andrej Karpathy wrote the sentence that named a generation of workflows: “I ‘Accept All’ always, I don’t read the diffs anymore.”

6.8 million views. Merriam-Webster added “vibe coding” to the dictionary the same year.

What Merriam-Webster did not add was the caveat Karpathy included later in the same post. He said the mode was for “throwaway weekend projects.” He said he was “not really coding, just seeing stuff, saying stuff, running stuff, and copy pasting stuff.” He said “it mostly works” with the tone of someone who knew that “mostly works” was the ceiling of this mode, not the floor of it.

The industry kept the phrase “Accept All.” It left the caveat behind. A specific admission about personal weekend side projects became a general permission for production code review. A tweet became a culture. The warning label did not make the import.

This is the pattern “human in the loop” is running at a larger scale. The words survive. The conditions the words assumed drop away. What stays is the label. What leaves is the substance.

Open your company’s AI coding policy. Or your compliance page. Or the all-hands slide where leadership first mentioned Accept All. The Karpathy quote might be in there. The next sentence in the original tweet, the one about throwaway weekend projects, is not.

Dependency Has a Midpoint. Most of Us Crossed It.

Dependency is not binary. It is a spectrum, and most of us crossed the midpoint a while ago without marking the moment. We just noticed one day that we could not draft without asking first. We could not debug without pasting the stack trace in first. We could not think through the shape of a problem without having the tool think through it alongside us.

The review step was the first honest casualty of this shift. We stopped reviewing because we stopped wanting to. The cost of evaluating was higher than the cost of trusting, and we paid in trust without tracking the balance.

The phrase “human in the loop” stayed because the phrase let us not say out loud that the loop had already opened.

I use these tools every day. I draft documentation with them. I think through technical decisions with them. I let them take first passes on code that eventually ships. I am not writing this from outside the loop about the people inside it. I have caught myself many times tempted to accept output without questioning the implementation or its side effects, because understanding fully is expensive and approving is free.

The Phrase Stays Because We Need It To

“Human in the loop” does the work the loop no longer does.

It gives the compliance team a control to point at. It gives the auditor a checkbox. It gives the manager a clean line between the organization’s AI policy and any future incident. It gives you a story to tell when the thing you approved breaks: there was a human in the loop, the human was me, the loop included a review step. The substance did not show up to that review. The phrase did.

The honest version is that we kept the word because losing it would mean admitting something we are not ready to say. That we work alongside a tool we do not fully understand. That we ship decisions we did not make. That the part of the job called “review” has quietly become the part of the job called “accepting.” That the dependency crossed a line a while back, the line was not marked, and we are past it now.

You can still call it human in the loop if you want. There is still a human. Technically, there is still a loop.

But somewhere in the last eighteen months, for most of us, on most PRs, the loop stopped closing. The word stayed, because the word is what we had left when the thing the word described walked out of the room.

Your name is on the commit. Check it.

—Viz

Next: What Happens When AI Reads a Page

Read next