AI Rookies

Benchmark contamination

Fact

A model saw the test questions during training, so its score looks too high.

In Plain Words

It is like finding tomorrow's spelling test in your backpack. Suddenly you look like a genius, not a wizard.

It can puff up scores above real skill. You see it on public leaderboards, model tests, and model face-offs.

Related Concepts

Pretraining
If test data slips into pretraining, later scores stop being trustworthy.

LLM
Benchmark contamination can make an LLM look smarter than it is.

AGI
Benchmark contamination can make AGI progress look too rosy.

AI-regulation
Polluted tests can give AI-regulation a shaky base.