Every major AI lab releases a new model and leads with the same thing: benchmark scores. MMLU, HumanEval, MATH, GPQA - a column of numbers, a comparison table with competitors, and a press release written to make it look like a moon landing. The model aced the bar exam. It outperforms PhD students on biology questions. It scores 90-something percent on reasoning tasks that were supposed to separate humans from machines.
And then you use it, and it confidently tells you something wrong.
This isn’t an anecdote problem. It’s a structural one. The benchmarks being used to evaluate frontier AI models have been so thoroughly studied, discussed, and trained-on that performance on them has decoupled from real-world capability. When a model scores well on MMLU - a multiple-choice dataset covering 57 subjects - it doesn’t necessarily mean it understands those subjects. It may mean it’s seen enough MMLU-adjacent data to pattern-match its way to the right answer. That’s a meaningful distinction that benchmark tables make invisible.
The contamination issue has been documented by researchers. Training datasets large enough to cover most of the internet will inevitably contain benchmark questions, answer keys, and discussion threads about both. Labs test for this with varying rigor, and the honest ones acknowledge it. But the press release rarely leads with the caveat.

The Score Always Goes Up
What makes this particularly hard to ignore is the one-directionality of it. Benchmark scores trend upward, almost without exception, as new models replace old ones. If those scores accurately reflected capability, the experience of using AI tools would be improving at roughly the same rate. For some narrow tasks - code generation in common languages, structured data extraction - it genuinely has. For open-ended reasoning, nuanced writing, and anything requiring consistent judgment across a long context, improvement has been far more uneven.
The labs are not exactly hiding this. Anthropic, OpenAI, and Google DeepMind all publish model cards that include known limitations. But those limitations appear in fine print while the benchmark numbers go in the headline.
The real problem is that there’s no agreed replacement. Benchmarks that are kept private to avoid contamination can’t be verified by outside researchers. Human evaluation is expensive, slow, and inconsistent. Vibes-based leaderboards like Chatbot Arena are interesting but reflect user preference, not accuracy.
So the industry keeps publishing the numbers it has, even when everyone quietly agrees those numbers mean less than they used to.