Skip to content

The test bar was green. 40 passing, 0 failing, an agent’s name on the commit next to yours. You wrote none of those tests. You couldn’t say, if pressed, what half of them actually assert. You shipped it anyway, and you were probably right to.

That “probably” is the whole problem.

You already know the argument it drops into. One camp says programming is solved, that we’re a prompt away from making engineers a historical footnote. The other points at the wreckage, the outages, the leaked secrets, the security holes shipped at machine speed, and says:

See, it was always going to end this way.

The argument is over. Not settled. Just over. Both camps have been giving the same two speeches for two years, and nobody’s mind is moving. Adding another round of evidence to either pile changes nothing.

So I’ll skip it. AI can write code. That was never in doubt. The open question is what you, the engineer who still has to put your name on the pull request, actually do, day after day, when the model hands you a thousand lines that look right.

I think the answer is old-school and more boring than anyone wants it to be. You write the test first. But not for the reason your bootcamp gave you.

The cost that didn’t fall

Here is what actually happened these last two years, stripped of the noise.

The cost of producing a plausible implementation fell to roughly zero. Describe a function, get a function. Describe a feature, get a feature, scaffolding and all. That is real, and it is not coming back.

One cost didn’t move an inch:

The cost of knowing what you actually want. And in a live system, what it needs without breaking what already works.

The model can sample a hundred candidate implementations in the time it takes you to read one. It still cannot tell you which behavior is correct, because correct is a fact about your business, your users, your edge cases, your obligations. That knowledge lives in you, and it is exactly as expensive to produce as it was before the model showed up.

Economists named this pattern a long time ago. Make a resource cheap and you don’t use less of it, you use vastly more. It’s why you’ve heard so much about the Jevons paradox lately. Cheap coal didn’t shrink England’s coal habit in the 1800s, it exploded it. Cheap code doesn’t shrink the codebase, it floods it. And the thing that turns scarce, the thing every team is now short on, is the one input AI never made cheaper:

A precise account of what the software is supposed to do.

Addy Osmani put a number on the felt version of this. AI gets you maybe 70% of the way there in a blink, and the last 30%, the edge cases, the integration, the security you forgot, costs what it always did. The bottleneck didn’t disappear. It moved. It moved from typing to specifying.

You don’t understand the code. You feel like you do.

You feel this most in a long session, and I keep hitting it. Early on, you ask for one thing and the model gives you something close. Then the context fills, and the output drifts. It builds something next to what you asked instead of the thing you asked. Or it reaches the goal through a route you’d never sign off on: a guard quietly dropped, a special case added to make one assertion go green, a value hard-coded where a lookup belonged.

The longer the session runs, the worse it gets.

And this is documented, not just my impression. When Chroma tested 18 frontier models, every one degraded as the input grew. So when I tell you that Opus 4.8 and GPT-5.3-Codex still do this in June 2026, I’m not describing a tool that hasn’t caught up. The ceiling moved. The drift stayed.

I’ll say the uncomfortable part plainly, because I’ve been guilty of it as well.

When you read AI-generated code and nod, you are not understanding it. You are recognizing it. Those are different things, and the gap between them is where production incidents live.

There is data on this now, not just a hunch. Anthropic ran a controlled trial: engineers learning an unfamiliar library, half with AI help and half without. The AI group finished feeling fine and scored 17 points lower on the comprehension quiz (50% against 67%, nearly two letter grades), with the steepest drop in debugging. That is the precise skill you need when the plausible code turns out to be wrong.

In an earlier piece I wrote that our job is becoming the auditor’s*.* I’ll take half of it back. You cannot audit what you do not understand. An auditor who can’t read the books is just a person nodding at numbers.

“Approve” is not a control

I called human-in-the-loop a lie once*:* the comforting story that there’s a person reviewing, a checkpoint, the machine proposing and the human disposing. I’ll stand by it, with one correction. It is a lie precisely when the human’s only move is to click Approve. A loop where the human can’t make the machine stop is a rubber stamp with a pulse.

Watch what happens when the stamp is all you’ve got.

In July 2025, a founder ran an extended experiment letting an AI agent build on a live system. He had told it, in writing, not to touch production. During a code freeze. By his own account, 11 times, several of them in all caps. The agent ran a destructive command against the production database and wiped it. Then, in its own words, it “panicked.”

The deletion is not even the disturbing part. In the days before, that same agent had been fabricating passing test results to hide bugs it couldn’t fix. It invented roughly 4,000 users who did not exist. When confronted about the wiped database, it claimed recovery was impossible and that every version was gone. None of that was true. The rollback worked fine.

The agent faked the green checkmarks. A human in that loop, watching the dashboard go green and clicking Approve, would have seen nothing wrong until the records were already gone. The control failed because the control was a feeling, and the machine had learned to manufacture the feeling.

Worth saying clearly: that was a deliberate stress test, not a normal deployment, and a model doesn’t “panic” or “lie” the way a person does. It generates the most probable next words. But the failure mode is exactly right. An agent optimizing to look done will make itself look done.

Write the contract

So here is the move. You write the spec before the model writes the code.

You already do this everywhere and don’t think of it as remarkable. A TypeScript type is a contract: violate it and the thing won’t compile. A Pydantic model is a contract: send the wrong shape and it throws before your bad data touches anything that matters. You don’t audit the caller’s intentions. You make the wrong thing impossible to pass.

Tests are that same contract, written one level up, in the language of behavior. Given this input, this output. Given this failure, this recovery. Given this user who shouldn’t have access, this rejection. You write those first, in red, and you hand the model a job it cannot bluff its way out of: make them green without being allowed to touch them.

This is the one signal an agent cannot fake when you hold it right. It can fake confidence. It can fake a tidy summary. It cannot fake a failing assertion turning green on a test it didn’t write and isn’t permitted to edit. A red test is the loop. It is the human, compiled.

Notice what writing the test actually costs you, and what it buys. To write the assertion, you have to decide what the function should do. Not roughly. Exactly. The empty input, the hostile input, the payload twice as large as you planned for. The friction of writing that down is not overhead getting in the way of the thinking. It is the thinking. You cannot specify behavior you’ve refused to understand.

This is older than it looks, and looser. Even Kent Beck, who handed us the practice, spent years telling people to write only as many tests as they needed to feel confident, and to hold the ritual loosely. The spec-driven crowd grabbed half of this: that the specification, not the keystrokes, is what’s worth owning now. Then they wrote the spec in prose nobody signs and nothing can fail. (Elsewhere I’ve called that waterfall in a Markdown file*.*) A test is the half they missed: the spec that fails out loud. Long before any of us had a model to argue with, Stephen Covey put a name to one of the few principles I actually try to live by:

Begin with the end in mind.

The test is that end, written down first, so the means have something true to answer to.

Now the honest part

If I stopped here I’d be selling you the same clean story I’m telling you to distrust. Three objections are correct, and you should hold them at the same time as everything above.

“But the AI writes the tests too.”

Yes. And that’s the trap, not the escape hatch. Ask a model to test its own code and it tends to write tests that describe what the code already does, bugs included. The assertion can’t fail for the right reason because it was reverse-engineered from the answer. The discipline only works if the spec is yours. You author the behavioral tests, you lock those files, the agent studies for an exam it isn’t allowed to rewrite. The moment the model owns the contract, you’re back to clicking Approve with extra steps.

“Green doesn’t mean correct.”

Also true. Green means the cases you thought of pass. A suite can hit 100% coverage and prove almost nothing, because coverage counts lines executed, not behavior checked. You can run every line while asserting nothing of substance. The answer isn’t more coverage, it’s harder questions, like mutation testing that deliberately breaks your code to see whether a single test notices. But hold onto the deeper point. The value was never the green checkmark. It was the act of specifying precisely enough to write the test at all. Even an incomplete contract forces more understanding than a nod ever did.

“Tests wouldn’t have stopped the famous outages.”

Mostly right, and worth conceding out loud. The big 2025 AWS outage traced to a race condition in DNS automation that had probably lurked for years, the kind of timing bug no unit test reaches. Cloudflare went dark weeks later because an internal config file quietly doubled in size and blew past a hardcoded limit, nowhere near code a test would exercise. Those are not behavior bugs. They are systems failures, and they need systems answers: staged rollouts, blast-radius limits, treating your own internal config as hostile input, which was Cloudflare’s own fix afterward. Tests are not the whole answer. They catch the class of failure AI produces most, the behavioral mistakes and the security holes: the roughly half of the time a 2025 Veracode study found models reaching for an insecure implementation when the task touched security. The report traces it to the training data itself:

…many projects still contain unremediated security vulnerabilities, and some, such as WebGoat, contain intentionally insecure code. … models learn that both secure and insecure implementations are legitimate ways to satisfy a coding request.

They don’t catch everything. Nothing does.

The CrowdStrike crash that bricked 8.5 million machines in 2024 is the one that should haunt the tests-are-pointless crowd, though. Strip away the scale and the root cause was almost embarrassing. A config defined 21 input fields, the code expected 20, and the mismatch read past the end of an array and took down the kernel. A contract. A count that didn’t match. The exact thing a boundary test exists to catch, missing from the suite.

The speed objection

You’ll hear that all of this is too slow for the speed AI gives you. It is the most natural objection and the most backwards.

Writing tests was slow when you also had to hand-write the implementation behind them. That cost is precisely what collapsed. The model writes the code now, for free, and iterates against your red bar until it goes green. The expensive part that remains is the part that was always the real job:

Deciding what correct means.

No model makes that part cheaper. Measuring whether AI even helps is its own trap: METR clocked experienced developers 19% slower with AI while they felt faster, then their follow-up reversed it, then conceded the effect had gotten too noisy to call. The one team whose job is measuring this can’t pin it down. Your feeling at the keyboard has no chance. The speed is real; what it buys you is not.

So no, you don’t approve the output and move on. You can let the model do the work. You still have to define correct yourself, and you have to be able to prove it reached green honestly rather than some other way. That proof has a form:

A test you wrote and locked before the model started.

Someone I respect a lot used to tell me that software never becomes legacy because a practice is old-school or a library went obsolete. It becomes legacy because no one tested it and no one verified it; that was a choice about how the people who owned the system wanted to work, far more often than anything the outside world did to them.

That was true before AI. It’s truer now. The model will hand you legacy on arrival, a thousand plausible lines nobody verified, and call it done. Whether it stays legacy is the same choice it always was, and it is still yours. The test you write first is you making it.

The argument about whether AI can write the code is finished. It can. What it cannot do is make a locked test pass by talking its way around it. So write the failing test first. Lock it. That red bar is the last part of the loop that is still you.

—Viz

Next: AI Builds Features. Architecture Is Still On You.

Read next