Every major model release now follows the same choreography: a blog post, a technical report, and a table of benchmark scores with the new model sitting comfortably at the top. The scores are real. The benchmarks are real. But the selection of which benchmarks get highlighted has quietly become a strategic decision, not a scientific one.

This isn’t a conspiracy - it’s a structural problem. There’s no independent body that mandates which evaluations must appear in a model release. So companies run dozens of benchmarks internally and publish the ones that tell the most flattering story. A model that underperforms on, say, long-context retrieval tasks might dominate on coding benchmarks. Publish the coding numbers, bury the rest. Nothing about this is technically dishonest, which is exactly what makes it hard to push back on.

The Benchmarks Themselves Are Compromised

Beyond cherry-picking, there’s a more corrosive issue: contamination. Several widely-used benchmarks - including MMLU and HumanEval - have been around long enough that their contents appear in training data scraped from the web. A model scoring 90% on MMLU might be partially recalling answers rather than demonstrating reasoning. Researchers have pointed to this problem for years. Model developers acknowledge it in footnotes. And then everyone proceeds to cite MMLU scores anyway.

ARC-AGI was introduced partly as a response to this - designed to be resistant to memorisation. When OpenAI’s o3 scored highly on it in late 2024, it was treated as a landmark. But even ARC-AGI’s creator, François Chollet, was careful to note that the score didn’t mean what the press coverage implied it meant. The benchmark moved faster than the understanding of it.

What Actually Gets Measured

The benchmarks that survive long enough to matter tend to measure what’s easiest to measure: multiple-choice question accuracy, code that either compiles or doesn’t, math problems with unambiguous answers. What they don’t measure well is the thing most users actually care about - whether the model is reliable under pressure, whether it hallucinates plausibly, whether it knows when to say it doesn’t know.

Those qualities are harder to quantify, so they largely don’t appear in release tables.

The result is an evaluation ecosystem that’s optimised for producing publishable numbers rather than useful signal. Labs aren’t lying, but they are, collectively, managing perception. And at this point, reading a benchmark table without knowing which benchmarks were omitted is less informative than it looks.