AIME: 2026 AI Leaderboard
The American Invitational Math Exam, used as a rolling frontier-math benchmark.
What it tests
AIME is a 15-question annual math competition for top US high school students, with integer answers 0-999. LLM labs now run fresh AIME problems each year as a contamination-resistant reasoning benchmark because the questions are published publicly only after the competition.
How it is scored
Accuracy over 15 problems, reported per year (AIME 2024, AIME 2025, etc.). Each year reruns against the current model lineup.
Why it matters
AIME is the cleanest year-over-year math-reasoning signal. A model that scores 99%+ on AIME 2024 but 60% on AIME 2026 is almost certainly training-data-contaminated; fresh-year scores are the honest test.
Leaderboard (7 models)
Sorted by AIMEscore. Tier column shows the tool's overall AIToolTier rank, which blends this benchmark with pricing, features, and real-world usability.
| # | Model | Tier | AIME score | Variant | Overall |
|---|---|---|---|---|---|
| 1 | Claude (Anthropic) Claude Fable 5 (launched 2026-06-09; suspended 2026-06-12 by US gov order; RESTORED globally 2026-07-01 after controls lifted 6/30) -- vendor SWE-Bench Pro 80.3% (vs GPT-5.5 58.6%); #1 LMArena Elo 1510 and #1 Artificial Analysis Index 65 as of 6/11. New default Sonnet 5 (2026-06-30): OSWorld-Verified 78.5%. Legacy Opus-line reasoning-suite scores shown below as baseline pending full third-party suites | A | 99.8% | AIME 2024 | 8.5/10 |
| 2 | GLM / Z.ai (Zhipu AI) GLM-5.2 (~753B MoE, launched 2026-06-13) -- vendor-published; third-party verification still settling | A | 99.2% | AIME 2026 | 8.0/10 |
| 3 | Microsoft MAI-Thinking-1 MAI-Thinking-1 (vendor-published 2026-06-02; third-party verification pending) | B | 97% | AIME 2025 | 7.5/10 |
| 4 | Gemma 4 (Google) Gemma 4 31B | A | 89.2% | AIME 2026 | 8.3/10 |
| 5 | Qwen (Alibaba) Qwen3.5-397B MoE | A | 87% | AIME 2025 | 8.8/10 |
| 6 | Nemotron (Nvidia) Llama-Nemotron Ultra 253B (prior gen -- Nemotron 3 Ultra 550B third-party scores pending) | B | 84.5% | AIME 2025 | 7.8/10 |
| 7 | ChatGPT GPT-5.5 (launched 2026-04-23; scores below are the GPT-5.4 baseline -- GPT-5.5 launch benchmarks per OpenAI are logged in Known Issues, pending third-party verification) | A | 83.3% | AIME 2024 | 8.8/10 |
About AIME
- Creator
- Mathematical Association of America
- Unit
- % (max 100)
- Official source
- https://www.maa.org/math-competitions/aime