AIME 2025
Artificial Analysis
Artificial Analysis' AIME 2025 evaluation measures olympiad-level mathematical reasoning with exact integer answers, directly fitting multi-step reasoning under constraints. | test metric name: test metric value test unit | Model: Test | Superduper good first evidence
Linked to
AI can correctly understand, reason through, and plan difficult tasks
Can it solve hard problems that require several correct steps while following constraints?
May 29, 2026