A test that judges LLMs by making them play games.
Talk is cheap. Put the AI in Mario Kart and let banana peels grade it.
It checks planning and memory. It also checks if the AI can use a game screen. You often see the scores on leaderboards.
Leaderboard
Game scores often go into leaderboards to compare models.
Agent harness
Many game benchmarks use an agent harness to run tests again and again.
Computer use
Game tasks often make the model watch the screen and take actions.
Benchmark contamination
If training data included the game tasks, the score may look too high.