
Your eval dashboard says 0.84. It said 0.84 last week too.
Then a partner at the firm you sell to forwards a screenshot: your assistant told their analyst the target’s margins were improving, the quarter they collapsed. You go looking for the eval that should have caught it. It passed. It passed because the scorer checked whether the answer mentioned “operating margin,” and it did.
That eval is a unit test holding an eval’s job title. The two look identical in your test runner. Both take an input, run your code, check the output. They answer completely different questions, and once you confuse the questions, every number on the dashboard inherits the confusion.
A unit test asks: did this deterministic function return the one correct output? Run it a thousand times, get the same answer a thousand times.
An eval asks: how well does this probabilistic system do a job, across the range of inputs it will actually see? There is no single correct output. The system gives you a different answer each time you ask.
The score you get back isn’t a fact. It’s one draw from a distribution whose variance you never measured.
So here’s the problem. Most teams shipping LLM features do not have evals. They have unit tests with an eval’s job title, and that gap is why the dashboard says 0.84 while the users say it’s broken.
An eval is two parts, and both can be wrong
Strip away the tooling (Inspect, RAGAS, promptfoo, LangSmith, whatever you run) and every eval is two pieces:
- A dataset (the inputs you feed the system)
- A scorer (the logic that turns an output into a number).
The tooling is plumbing. The measurement is your dataset plus your scorer.
It lives or dies on a few properties that measurement people have fought over for decades. The one this whole piece runs on, construct validity, comes from Cronbach and Meehl writing about psychological tests in 1955. They defined it against the older idea of criterion validity, whether a score predicts an outcome. Reliability is older still, and sensitivity is just statistical power. They were measuring intelligence and anxiety, not language models. The problem is the same one you have: you can’t see the thing you care about directly, so you build an instrument and hope the number it prints matches the thing.
Construct validity: are you measuring the capability you claim, or an artifact? The “operating margin” check has none. It measures word presence and calls it correctness.
Reliability: run the instrument twice under the same conditions. Same reading? LLM evals fail this constantly. Rephrase a prompt, reorder the options, bump the temperature, and the score moves.
Sensitivity: can the instrument tell two systems apart? A 50-example eval cannot tell you whether version B is two points better than version A. It doesn’t have the resolution.
And the one hiding inside your 0.84 story: criterion validity, whether the eval score predicts the production outcome you care about. Your 0.84 didn’t predict the margins your assistant got backwards. It had no criterion validity. It was measuring something, just not the thing that bit you.
So what should the scorer actually be? Reach for the cheapest instrument that catches the failure. When the task has a known right answer (a route, a label, a JSON shape), exact match or a schema check is the whole scorer: deterministic, free, and it never drifts. The substring check at the top didn’t fail because it was code instead of a model. It failed because it was the wrong exact check. You reach for a model to judge only when the target is genuinely subjective, like tone or grounding, and that is where the real trouble starts.
Your judge is an instrument inside your instrument
When you do need a model to score, you’ve nested one probabilistic system inside another, and you almost certainly haven’t calibrated the inner one.
Model judges have biases, and they’re defaults, not edge cases. The first one named was position: show a judge two answers and ask which is better, swap the order and ask again, and in the original 2023 study the verdict flipped about a third of the time even for GPT-4, more for weaker judges, purely from where the answer sat. Strong judges have since largely tamed that one. The other two proved stubborner. Judges still reward padding and surface polish. Judges still rate their own outputs higher, and if your judge and your generator share a base model, that self-preference runs through your whole eval in a direction the eval can’t see from the inside.
So you validate the judge against human labels. And you do not use percent agreement to do it.
import numpy as np
from sklearn.metrics import cohen_kappa_score
human = np.array([1]*95 + [0]*5) # 95% PASS
judge = np.ones(100, dtype=int) # a judge that always says PASS
(judge == human).mean() # 0.95 the "agreement" that fools you
cohen_kappa_score(human, judge) # 0.00 the information it actually carries
Percent agreement is inflated when one class dominates. A judge that rubber-stamps everything looks 95% accurate while carrying zero information. Kappa is literally zero, because a constant answer has no variance to correlate with anything. Cohen’s kappa corrects for the agreement you’d get by chance, and the gap between those two numbers is the one that gets teams. (Hamel Husain and Shreya Shankar have a whole workflow for this: one domain expert labels real outputs with a written reason, the reasons become the judge’s rubric, and you ship the judge only when its kappa against a held-out set clears your bar.)
A judge is trustworthy not because it’s a big model. It’s trustworthy because you measured its agreement with someone who knows the domain.
The public benchmarks won’t save you
A reflex when evals feel like work is to point at a public number. MMLU, SWE-bench, the leaderboard of the week. Skip it. Those numbers are contaminated.
When a benchmark’s answers leak into training data, the benchmark stops measuring capability and starts measuring memorization. The score goes up; the capability doesn’t. MMLU has an estimated 6.49% of its questions containing outright errors. SWE-bench is worse: one audit found that a third of the patches models got credit for (32.67%) had the answer sitting in the issue itself, and another 31% passed only because the tests were too weak to catch a wrong fix. This is a treadmill, not a one-off: by early 2026 OpenAI had stopped reporting SWE-bench Verified, its own cleaned subset, on the grounds that contamination and weak tests meant the score no longer tracked real coding ability.
A model’s MMLU score tells you almost nothing about whether it’ll do your job. So you build the dataset yourself, from production traces, not from imagination.
The method is error analysis, and it’s dull on purpose. Pull real traces. Read them. Write one plain sentence per failure. Cluster the sentences. Count the clusters. The failure mode in 40% of your broken traces is your week; the elegant edge case that shows up once is not. You stop somewhere around 100 traces, when new ones stop producing new categories.
Every trace you haven’t read is instrument debt: a failure mode your eval can’t see because you never taught it the category exists. It comes due in production.
One real task, start to finish
Make it concrete. Say you’re routing support tickets to billing, bug, or account. Your first eval is the substring check from the top of this piece, and it ships a bug: “card declined” tickets keep landing in the engineering queue.
So you do the boring thing. You pull 100 real tickets and read the misroutes. A pattern falls out fast: 78% of the errors are payment failures (“my card was declined,” “payment failed at checkout”) routed to “bug,” because the classifier sees the word “declined” before it sees the billing context. The other 22% are account tickets worded like errors. Two categories, counted, ranked. The 78% one is your week.
Measured honestly, that classifier scores 0.82, with a 95% interval from 0.74 to 0.89. Not “82%.” A reading with a margin.
You fix the payment-failure flaw and the score jumps to 0.96, a real 14-point gain. A 100-example eval can just barely see that: its noise floor is around 13 points. If your fix had moved things only 3 points, those same 100 examples could not have told you whether it was real. Same afternoon of work, same dataset. The difference is that now the green number means something.
Every number in that paragraph comes from a fixed toy classifier with a planted flaw, not from hand-waving. That’s the whole point: a measurement is reproducible, or it isn’t a measurement.
How many examples you actually need
Sample size is where the comfortable evals fall apart.
To reliably detect a 2-point accuracy improvement around the 80% mark, you need on the order of 6,000 examples. A 4-point delta needs about 1,450. The relationship is roughly quadratic: halve the difference you want to catch, quadruple the examples.
Now read that against the 50-example eval making your ship decisions. It can’t see a 2-point change. It can’t reliably see even a 10-point change; the smallest difference 50 examples can catch is about 17 points.
When your 50-example eval says version B improved three points, that three points is inside the noise. You’re making a product call on a coin flip you’re admiring.
This doesn’t mean every eval needs thousands of examples. It means you size the instrument to the decision. Catching only catastrophic regressions? A few hundred examples is fine. For a launch decision between two close versions, you need the resolution, or you accept you’re guessing. (One cheap trick that buys resolution for free: run both versions on the same questions and compare per-question, not the two averages. Shared difficulty cancels, and your error bars shrink without a single extra API call.)
Your eval is probably lying to you if
- You’ve never run it twice and compared the two scores.
- Your scorer checks for a substring and calls it correctness.
- Your dataset is examples you wrote by hand, not traces you pulled from production.
- You report a single number with no interval, and nobody’s asked for one.
- Your suite has 50 examples and you use it to choose between model versions.
- The score went up and users noticed nothing. Or it held steady and users complained.
- You use a model judge and have never compared it to a human.
- You compared judge to human with percent agreement, not kappa, and the high number reassured you.
The pattern under all of these: a green number you trusted without checking the instrument that produced it.
Build the instrument first
Deterministic software earned a testing culture because the cost of a bug was visible and the correct output was knowable. Probabilistic software needs a measurement culture, with the seriousness a physicist brings to a noisy detector. Calibrate it. Characterize its error. Never confuse the reading with the truth.
Once you see an eval as an instrument instead of a test, a lot of confusing weeks resolve at once. The dashboard that said 0.84 while users complained was measuring the wrong thing. The improvement that didn’t show up in production was inside the error bars. The judge that loved every answer had a kappa of zero. None of these are mysteries. They’re what happens when you trust a reading without auditing the instrument.
This only gets more true as the systems get harder to measure. An agent taking ten steps, a conversation running twenty turns, a tool call buried deep in a trajectory: more surface to measure, the same question of whether you’re measuring it. The instrument discipline is what carries over.
Build the instrument first. The number is only as honest as the thing that made it.