A ranked list of AI models based on the same tests.
A model leaderboard is class test results pinned in the hallway. The top score gets stared at like it brought pizza.
People use it to pick models and spot trends. But a high score does not always mean it works best for you.
Benchmark contamination
Leaked tests can make leaderboard scores look fake.
Third-party AI evaluation
Third-party eval adds another view beyond one score.
LLM-as-a-judge
Some leaderboards use LLM-as-a-judge to help grade answers.
Frontier model
Frontier models often race each other on leaderboards.