A test that checks AI reasoning with very hard, new math problems.
FrontierMath is the final-boss math test for AI. No dusty answer sheet can save it.
It shows if top models can really reason. Its scores help rankings reward skill, not memorized questions.
Leaderboard
Its scores help compare models on hard math skills.
Benchmark contamination
New questions make it harder to score from memorized answers.
Reasoning-model
It tests how well a model reasons through hard math problems.