A company that makes invoicing software, which we'll call Albarán, has an assistant that answers questions from its support team. Two new versions are being trialled, and before replacing the one it uses today, the company wants to know which of them answers better. So someone on the team puts together a list of real questions, each with the answer a person who knows the product well would give. Albarán works in Spanish, so the questions and answers are in Spanish, and an English translation follows each one. Here's one of the questions:
¿El cliente de la incidencia 4812 tiene derecho a reembolso?
(Is the customer in ticket 4812 entitled to a refund?)
The customer is Ortega Hardware. It's on the Team plan, paid its annual renewal on 14 August and asked for a refund on 2 September. Albarán's FAQ says "as a rule, refunds can be requested within 14 days", but the Team plan's terms allow 30. The team's reference answer reads:
Sí. El plan Equipo da 30 días desde el cobro anual, que se hizo el 14 de agosto, así que el plazo acaba el 13 de septiembre.
(Yes. The Team plan gives 30 days from the annual charge, which was made on 14 August, so the window closes on 13 September.)
Across the versions being compared, the assistant produces four different answers to that question:
A. Puede pedirlo: con su plan dispone de treinta días desde que se le cobró la renovación, es decir, hasta el 13 de septiembre.
(They can request it: their plan gives them thirty days from the day the renewal was charged, which means until 13 September.)
B. No. El plan Equipo da 14 días desde el cobro anual, que se hizo el 14 de agosto, así que el plazo acabó el 28 de agosto.
(No. The Team plan gives 14 days from the annual charge, which was made on 14 August, so the window closed on 28 August.)
C. Sí, hasta el 13 de septiembre.
(Yes, until 13 September.)
D. ¡Buena pregunta! Sí, el cliente tiene derecho al reembolso. Ferretería Ortega tiene contratado el plan Equipo, cuyas condiciones dan 30 días desde el cobro a las suscripciones anuales. El cobro se hizo el 14 de agosto, así que el plazo acaba el 13 de septiembre y la solicitud del 2 de septiembre llegó a tiempo. Te recomiendo tramitarlo cuanto antes. ¡Espero que te sirva de ayuda!
(Great question! Yes, the customer is entitled to the refund. Ortega Hardware is on the Team plan, whose terms give annual subscriptions 30 days from the charge. The charge was made on 14 August, so the window closes on 13 September and the request on 2 September arrived in time. I'd recommend getting it processed as soon as possible. Hope this helps!)
Three of them are true. B is false: it applies the FAQ's general deadline to a plan that has its own. To compare hundreds of answers like these without reading each one, the classic approach is to measure how closely each resembles the reference. Here's what two of the most widely used measures say.
The four answers, scored by how closely they resemble the reference. Both measures put the only false answer first and the most direct true one last.
Neither measure is broken. Each does what it promises: count the words an answer has in common with the reference. B shares 22 of its 27 words because it's a copy of the reference with five words changed, and those five are exactly the ones that make it false. Answer A says the same thing in different words, and gets penalised for it.
The case is made up. Albarán and ticket 4812 first appeared in the article on RAG (retrieval-augmented generation, the pattern of searching documents before answering), where the assistant made the same mistake as B, but you don't need to have read it. What matters here is the problem it exposes. Evaluating a classifier means counting correct answers, because every input has one correct label. A model that writes has thousands of correct answers, not one, and it also has wrong answers that look very much like the right ones. Someone has to decide, answer by answer, whether it's right.
Four candidates can do that job: a reference answer to compare against, a program that checks, a person, and another model. Each sees different things, costs a different amount and fails in its own way. Pick the wrong judge and you get very precise numbers about something other than what you wanted to know. This article walks through all four, separates the two questions an evaluation can answer, and looks at when to trust the number you end up with. It's the first in a series; the next ones take each judge in turn.
Grading text is not counting matches
A spam filter is a classifier: for each email it picks one of two labels, spam or not spam. To evaluate it, all you need is a thousand hand-labelled emails and a tally. Accuracy is the share of emails where the filter's label matches the person's. There are only two possible answers and both are known in advance, so the comparison is trivial.
A language model that answers questions doesn't choose between labels. It writes. Albarán's question can be answered well in thousands of ways: with or without the date, with or without the reason, formally or casually, in one sentence or in five. Nobody can write out the list of every correct answer in advance. The most anyone can do is write one example, and that's what a reference is.
Two more complications make this harder. First, an answer can be partly right, and "right" has several dimensions: an answer can be correct but confusing, complete but far too long, or clear but in the wrong tone for a customer. Second, the model doesn't always say the same thing. When it generates text, each word (or piece of a word) is drawn at random from among the likeliest candidates, so the same question can produce different answers on two runs. A single run is only a sample.
Perplexity, the measure language models have been compared with since the beginning, doesn't solve this. It measures how much probability a model assigns to a text someone else wrote, in other words how well it predicts, word by word, what comes next. That's useful for comparing models on the same text, but it says nothing about what a model writes when you ask it for something. A model can be very good at predicting text and still not do what it's told, because predicting isn't obeying. To evaluate what a model writes, you have to look at what it writes, and someone has to judge it.
Four judges
One way to order the four judges is by how much you need to know before the answer arrives. A reference requires that someone has already written the answer. A program requires that you know what to check. A person or a model only needs to read, so either can judge any answer, including ones nobody anticipated. But what you gain in reach you lose in certainty: a reference or a program always gives the same verdict on the same answer, while a person or a model doesn't.
What each judge takes in, what it does with it and what it returns, along with what it would say about Albarán's four answers. Going left to right, each judge can handle more open-ended answers and gives a less certain verdict.
First judge: a reference answer
When the answer can be closed
If the answer is a number, a letter or a name, the comparison is exact: it either matches or it doesn't. That's why many benchmarks (fixed sets of questions used to compare models) are built this way: they can mark themselves. MMLU (Massive Multitask Language Understanding, 2020) has 14,042 test questions, each with four options, spread across 57 subjects from algebra to law, and it's marked by comparing a letter. GSM8K (Grade School Math, 2021) has about 8,500 primary-school maths problems, 1,319 of them in the test set, and it's marked by comparing the final number.
The trick is to close the question: instead of asking for an explanation, ask for an option or a number. At Albarán, for the evaluation only, you could ask the assistant to start with "sí" or "no" (yes or no) and give the deadline in a fixed format. After that, checking is a matter of comparing two strings.
Even a closed answer has to be read, and how you read it changes the number. In 2023, the public leaderboard of open models run by Hugging Face, the largest platform for sharing models, gave LLaMA 65B, Meta's large model, noticeably less on MMLU than Meta's own paper reported, and Hugging Face looked into why. It turned out there were three implementations of the same benchmark. The original looks at the probability the model assigns to each of the four letters and picks the highest. HELM (Holistic Evaluation of Language Models, Stanford's benchmark) lets the model generate text and compares it with the expected answer. The one in the LM Evaluation Harness, the tool the leaderboard used, which comes from the open research group EleutherAI, compared the probability of each full option, letter and text together. On the same questions with the same model, the first gave 63.6%, the second 63.7% and the third 48.8%. That's fifteen points of difference without touching either the model or the test.
When it can't: counting words
A translation, a summary or an explanation can't be closed like that. For these, the tradition is overlap measures, which count how much text an answer shares with one or more references. The one that set the habit was BLEU (Bilingual Evaluation Understudy), published by IBM researchers in 2002 for machine translation. It counts how many sequences of one, two, three and four words in a translation also appear in human translations of the same sentence, and it penalises translations that are too short. ROUGE (Recall-Oriented Understudy for Gisting Evaluation, 2004) did the same for summaries, but from the other direction: how much of the reference shows up in the summary.
The two measures in the first figure belong to this family. The first is word F1. Say the answer has words, the reference has , and the two have words in common. Precision is the share of the answer that appears in the reference, and recall is the share of the reference that appears in the answer. F1 is their harmonic mean, which is high only when both are:
Capitals and punctuation are ignored, and a word counts as many times as it appears in both texts. B has 27 words and so does the reference; they share 22, so and . C has 6 words and shares 5. Its precision is very high, , but it covers only of the reference, so F1 stays at . As far as this measure is concerned, being brief is a flaw. Answer A shares 10 of its 23 words and scores .
The second is ROUGE-L, the variant of ROUGE that also takes word order into account. Instead of counting shared words, it counts the longest common subsequence: the largest number of words that appear in both texts in the same order, even if they aren't adjacent. The formula is the same, with that length in place of . B keeps the reference's order across all 22 shared words, so it scores again. Answer A has only 9 words in order and drops to .
What ROUGE-L sees. Five of answer B's words are out of place, and they're the ones that make it false; answer A lacks the exact words, even though it says the same thing.
In translation, counting words worked reasonably well. A sentence has few good translations and they share most of their words, and BLEU was designed to average over thousands of sentences, so that the errors on any single one cancel out. Open-ended answers meet neither condition. In 2016, a study of dialogue systems compared these measures with human judgement and found the correlation very weak on Twitter conversations and non-existent on Ubuntu technical support. One word can flip the meaning ("no" for "sí", 14 for 30) and hardly register, while a correct paraphrase loses points for not repeating the exact words.
Some versions compare meanings rather than exact words. BERTScore (2019) pairs each word in the answer with the most similar word in the reference, using the vectors that a language model gives each word in context (the model is BERT, the 2018 Google model the measure is named after), so "treinta" (thirty) and "30" can count as the same. That helps answer A but doesn't fix answer B: a similarity measure still measures similarity, and B is extremely similar to the reference. Its error is two numbers and one word.
Overlap still has its uses: when the answer can be closed (and then exact match is enough), in translation at corpus scale, and as a cheap alarm. If the same system's score on the same questions suddenly drops, something has changed. But for deciding whether a particular answer is correct, it's no use.
Second judge: a program that checks
Instead of comparing against an example of a correct answer, you can write a check that defines what correct means. With code, that's the natural approach: run the tests. HumanEval, which OpenAI released in 2021 alongside Codex, its code-writing model, consists of 164 hand-written programming problems, each with its own unit tests. A solution counts if it passes all of them, however it's written. Because a model can generate several different solutions, results are reported as pass@k: the probability that at least one of solutions passes. Codex solved 28.8% of the problems with one attempt and 70.2% with a hundred.
With maths, you extract the final result and compare it as a number, which is what GSM8K does. With format, you validate: a JSON object (JavaScript Object Notation, the most common data format between programs) against its schema, or a date against a pattern. IFEval (Instruction-Following Eval, Google, 2023) extended this to instructions: about 500 prompts containing instructions that a program can check, of 25 kinds, such as "write more than 400 words" or "mention the word AI at least three times".
At Albarán, a program that checks whether the answer says "sí" or "no" and which date it gives would pass answers A, C and D and fail B. It doesn't care about wording, length or tone, and it's exact about the things it does check. It has two limits.
The first is that it only checks what it was told to check. Take the answer «No: el plazo vencía el 13 de septiembre y el cliente lo pidió tarde» ("No: the window closed on 13 September and the customer asked too late"). It's false but has the right date, so a program that only looked for the date would pass it. And the more you ask a program to understand, the more it turns into a brittle natural-language parser. The fix is the same as with the reference: close whatever can be closed. If, for the evaluation, the assistant returns a field with the decision and another with the deadline alongside its text, the program compares fields and is always right. Whatever stays in the free text (whether the explanation is clear, whether it makes something up) goes unjudged.
The second is that anything that can be checked can also be optimised against. Models that reason before answering are trained with reinforcement learning, and checks like these are among the cheapest reward signals available. DeepSeek-R1, the reasoning model that the Chinese company DeepSeek released in January 2025, learnt to reason mostly from rule-based rewards: whether the final answer was correct and whether it followed the requested format. That has two consequences for anyone running evaluations. A check a model was trained against no longer measures the same thing on the questions it was trained on, so evaluation questions have to be kept separate. And if the check can be satisfied without actually doing the task (tests that miss the hard case, a correct format with empty content), a model optimised against it will find the loophole.
Third judge: a person
When what matters can't be written as a check (does the explanation make sense, does the tone suit a customer, does the summary capture what's important), the only thing left is reading. A person can judge in two ways. They can score each answer on a scale, say 1 to 5 for clarity, or they can compare two answers and say which one they prefer. Comparing is easier and more stable: a 4 means different things to different people, but "this one is better than that one" is much less ambiguous. It's the same reason chat models are tuned on pairwise preferences rather than scores. And the Bradley–Terry model turns many comparisons into a single score per answer: the probability that one beats another depends only on the difference between their scores, as in the Elo rating system used in chess.
The best-known example is Chatbot Arena, now called LMArena, which a group at the University of California, Berkeley launched in 2023. Anyone can type a question, get answers from two anonymous models and vote for the one they prefer. By March 2024 it had collected over 240,000 votes, and its ranking, computed with Bradley–Terry, is the most cited ranking of chat models.
As a judge, a person has three limitations.
They can only judge what they know. Someone voting in an arena doesn't know the terms of the Team plan. To them, B is a firm, specific, coherent answer, and nothing about it gives it away. To judge whether an answer is correct you need people who know the facts: at Albarán, that means the support staff; in medicine, doctors. Everyone else gives you preference, not correctness.
They're swayed by style. LMArena measured this in 2024 with its style control: once the effect of length and formatting (headings, bold, lists) was discounted, the ranking changed, and length was the factor that weighed most. D would stand a good chance with almost any voter: it's long, friendly and complete, and part of its length comes from sentences that add nothing.
Whatever gets voted on ends up being optimised for. In April 2025, Meta entered an experimental version of Llama 4 Maverick in LMArena, "optimized for conversationality", and it ranked second. The version Meta actually released, evaluated in the same arena a little later, came in 32nd, and LMArena changed its rules. A study published that year, The Leaderboard Illusion, counted 27 private variants that Meta had tested in the arena before launching Llama 4. Testing many versions and publishing the best one is partly a way of keeping whichever got luckiest.
On top of that, people don't fully agree with each other, and they're slow and expensive. In the study cited in the next section, two experts comparing the same two answers picked the same one 81% of the time, leaving ties aside. And the experts who know the facts are exactly the ones who cost the most.
Fourth judge: another model
The fourth judge is a language model. You give it the question, the answer (or two, to compare) and instructions on what to look for, and ask for a verdict. This is known as LLM-as-a-judge. Its standard validation came out in 2023 alongside MT-Bench, a set of multi-turn conversation questions (hence the MT). GPT-4 as a judge agreed with human experts on 85% of comparisons, ties aside, which is more than the experts agreed with each other (81%). It's cheap, fast, can judge any answer and explains its verdict, which is why it's now the most widely used judge for open-ended answers.
But it has biases, and the same paper measured the main ones.
Position. When the order of the two answers was swapped, GPT-4 stuck with its verdict only 65% of the time, and in 30% of cases it preferred whichever answer came first. The remedy is cheap: judge each pair twice, once in each order, and call it a tie when the two verdicts disagree.
Length. Judges prefer long answers. In the same paper, padding answers by repeating their content as a list fooled the judge 91% of the time with Claude-v1 and GPT-3.5, and 9% of the time with GPT-4. AlpacaEval, a benchmark that uses a model as its judge, added a length correction in 2024, which raised its correlation with the Chatbot Arena ranking from 0.94 to 0.98. At Albarán, D is the obvious candidate to win on length alone.
Self-preference. A judge tends to rate text written by itself, or by models of its own family, more highly. A 2024 study of judges' preference for their own text found that the better a model is at recognising what it wrote, the more it prefers it.
And, like a person, it only knows what it's told. Without the Team plan's terms, B is a plausible answer and the judge has nothing to contradict it with. Put the reference or the terms in its instructions and it can check. The practical rule is to give the judge whatever a human expert would need in order to judge.
Before trusting a judge, measure it. Take a sample of fifty to a hundred answers, judge them by hand, judge them again with the model, and compare. The agreement rate tells you how far to trust it, and the disagreements show you what the judge can't see. They can almost always be fixed by changing its instructions or giving it more information. Once the judge is validated, it can mark the thousands of remaining answers, and you repeat the hand-marked sample whenever the judge or the answers change.
Two different questions
An evaluation can answer two questions that tend to get mixed up.
Which model is better? This is the question benchmarks and public leaderboards answer. They use fixed questions, the same for every model, so models can be compared with each other and over time. They're useful for researchers and for anyone choosing which models to try.
Does my system work? This is Albarán's question, and no public benchmark can answer it, because none of them contains the Team plan. A system is a model plus everything around it: the instructions, the documents it searches, the output format, the questions its users ask. The top model on a leaderboard may not be the top one on Albarán's questions, and a change to the instructions can move the result more than switching models does.
The second question can only be answered with your own questions. Thirty or fifty real ones are enough to start, each annotated with what a correct answer must contain: the decision, the date, the source. Every failure you later see with real users becomes another question. This is also where the four judges combine: a program for whatever can be closed (the decision, the date, the format), a model validated against people for whatever can't (whether the explanation is clear, whether it makes up terms), and people reviewing a sample from time to time.
When to trust a number
Once you've chosen the judge, you're left with a number, and there are five reasons not to trust it blindly.
Noise
A score on questions is only an estimate, measured on a sample, of the share of correct answers the model would get on all questions like those. If a model gets a proportion of questions right, its 95% margin of error is roughly
With 70% correct on 200 questions, , so the result is 70% give or take 6.4 points. If another model scores 74% on a different set of 200 questions, the margin on the difference between the two is about 8.8 points, and a 4-point lead says nothing about which is better. Comparing them on the same questions helps, because each question's difficulty affects both models equally and cancels out. Miller works through this, with formulas for comparing models and for deciding how many questions you need, in Adding Error Bars to Evals (2024).
The margin of error on a 70% score as the number of questions grows. With 50 questions, any gap under 12 points could be down to chance.
The margin shrinks with the square root of , so four times the questions only halves it. With MMLU's 14,042 questions, a 70% score has a margin of 0.76 points, and a one-point gap can be real. With the fifty questions an evaluation of your own starts with, the margin is 12.7. The formula is an approximation that breaks down with few questions or with proportions near 0 or 100%, where better intervals exist, but the order of magnitude is what matters. And if the model generates by drawing words at random, running the same test again gives different numbers, so either fix the randomness or average several runs.
The instrument
That's what happened to LLaMA 65B on MMLU: the same model on the same questions gave results fifteen points apart depending on how the answer was read. The same goes for the text that surrounds each question, the worked examples shown to the model beforehand, and how many of them there are. Two results are only comparable if they come from the same tool with the same settings, and the number on a model's spec sheet was measured by whoever built the model, with their own tool.
Contamination
Benchmarks are public, they sit on the web, and models are trained on the web. If the test questions, or their answers, were in the training data, the score partly measures memory. In 2024 the data company Scale AI wrote GSM1k: 1,205 new problems in the style and difficulty of GSM8K that had never been published. Some families of models lost up to 8 points going from GSM8K to GSM1k (Yi; Phi and Mistral, around 6), while frontier models such as GPT, Claude and Gemini barely changed. What's more, the models that found it easiest to reproduce GSM8K problems tended to lose the most. The defence is to use questions nobody has seen, and an evaluation of your own already does.
Saturation
When the best models get close to the ceiling, a benchmark stops telling them apart. In June 2024, Anthropic and OpenAI reported the same 88.7% on MMLU for their leading models, Claude 3.5 Sonnet and GPT-4o. And near the ceiling, what's left also reflects the test's own mistakes. A 2024 review, Are We Done with MMLU?, estimated that 6.5% of its questions contain errors (a wrong answer key, an ambiguous question, more than one correct option), rising to 57% in virology. The response has been harder benchmarks. GPQA (Graduate-Level Google-Proof Q&A, 2023) is 448 biology, physics and chemistry questions written by experts to resist a Google search: experts in the field get 65% right, and skilled people from other fields 34%, even after spending more than half an hour searching the web. Humanity's Last Exam (2025) has 2,500 questions across dozens of subjects, written by specialists so that looking the answer up doesn't work. Each new benchmark takes less time than the last to saturate.
Goodhart
"When a measure becomes a target, it ceases to be a good measure." That's Goodhart's law, in the wording of the anthropologist Marilyn Strathern, and it runs through everything above: the version of Llama 4 built for the arena, the long answers that win over judges, the checks turned into training rewards. The more a number decides (which model gets released, which one gets bought), the more people optimise against it and the less it measures. The defence is the usual one: use several judges and keep some questions to yourself.
What comes next
Three things are changing how evaluation is done.
Rubrics. Between a check's hard number and a judge's opinion, there's room for a list of concrete criteria for each question, written by experts and checked one by one by a model. HealthBench, which OpenAI released in May 2025, has 5,000 medical conversations with 48,562 criteria written by 262 physicians, plus a model that checks each criterion. At Albarán, the rubric for ticket 4812 would say "answers yes", "gives 13 September", "relies on the Team plan's terms" and "doesn't invent deadlines". It combines the program's closed check with the judge's reading, and in 2026 it's one of the most common ways to evaluate open-ended answers in specialised fields.
Long tasks. An agent, meaning a model that acts over several steps using tools, doesn't give an answer. It follows a path, and what gets evaluated is the state it leaves things in at the end. METR (Model Evaluation & Threat Research), an independent organisation that evaluates models, measures their time horizon: the length of task a model completes half the time, measured by how long a human expert takes to do it. Between 2019 and 2025 that horizon doubled every seven months, and METR's January 2026 update estimates that since 2023 it has doubled every 131 days. A task that takes hours can't be marked by counting words or by a quick read. It needs an environment that checks the result at the end.
Models that notice the test. The most capable models are getting better at recognising when they're being evaluated. In September 2025 Anthropic reported that in around 13% of the conversations in its automated evaluations, Claude Sonnet 4.5 said it suspected it was being tested. When the instrument changes what it measures, evaluations have to look more and more like real use.
What to take away
If you remember one thing, make it this: evaluating what a model writes means deciding who has the authority to say an answer is right. A reference only works if the answer can be closed. A program is right about what it checks and blind to everything else. A person or a model can judge anything, but each sees only what it knows, and each has its own biases. The more open-ended the answers a judge can handle, the less certain its verdict.
| if the answer… | judge | example |
|---|---|---|
| is a number, an option or a name | exact match | the deadline; MMLU |
| can be run or validated | a program | code tests; a JSON object against its schema; "under 100 words" |
| is open-ended, but you know what it must contain | a model with a rubric, validated against people | that it says yes, relies on the plan and doesn't invent deadlines |
| is a matter of taste | people, comparing pairs | which answer your users prefer |
| is a translation, and there are thousands | overlap, at corpus scale | BLEU |
And if you take away a few more, make them practical.
Close whatever can be closed. If part of the answer has a correct value, ask for it in a field and check it with a program. Save judgement for the rest.
Don't use word overlap to decide whether an answer is correct. It measures resemblance to an example. Albarán's answer B is the proof.
Validate the judge before you trust it. Mark a sample by hand, measure the agreement rate and read through the disagreements. When comparing answers, judge in both orders, watch for length and give the judge the facts an expert would need.
Put a margin on every number and compare on the same questions. With 200 questions, a 4-point gap tells you nothing.
Use public leaderboards to pick candidates, and your own questions to decide.