A state where model scores crowd so near perfect that the benchmark can no longer tell them apart.
It is like a spelling quiz after everyone got the answer sheet. The straight-A kid and the sneaky crammer both get 100.
The leaderboard becomes less useful. Test makers need harder, newer questions.
Benchmark contamination
Test questions in training can make saturation arrive sooner.
Leaderboard
Crowded scores make the leaderboard worse at showing real gaps.
Evaluation harness
The evaluation harness needs new questions to keep models apart.