A benchmark that tests abstract reasoning with few-example grid puzzles.
ARC-AGI is like giving AI a puzzle placemat at a diner. No answer key. No sneaky peek at the back.
It uses tiny colored grids to test fresh rule-finding. It catches models that just crammed the old practice sheet.
AGI
ARC-AGI uses small grid puzzles to probe the core of AGI.
Reasoning-model
A Reasoning-model does better when it can infer a new rule.
Benchmark contamination
ARC-AGI uses new few-shot tasks, so memorized answers help less.