Math

AIME: 2026 AI Leaderboard

The American Invitational Math Exam, used as a rolling frontier-math benchmark.

What it tests

AIME is a 15-question annual math competition for top US high school students, with integer answers 0-999. LLM labs now run fresh AIME problems each year as a contamination-resistant reasoning benchmark because the questions are published publicly only after the competition.

How it is scored

Accuracy over 15 problems, reported per year (AIME 2024, AIME 2025, etc.). Each year reruns against the current model lineup.

Why it matters

AIME is the cleanest year-over-year math-reasoning signal. A model that scores 99%+ on AIME 2024 but 60% on AIME 2026 is almost certainly training-data-contaminated; fresh-year scores are the honest test.

Leaderboard (7 models)

Sorted by AIMEscore. Tier column shows the tool's overall AIToolTier rank, which blends this benchmark with pricing, features, and real-world usability.

#ModelTierAIME score
1Claude (Anthropic)
Claude Opus 5 (launched 2026-07-24) is the current flagship and the default on Max -- Anthropic released COMPARATIVE claims only, no absolute scores: more than doubles Opus 4.8 on Frontier-Bench, within 0.5% of Fable 5 on CursorBench 3.2 at half the cost, 3x the next-best model on ARC-AGI 3, ~1.5x on Zapier AutomationBench, and beats Fable 5's best OSWorld 2.0 result at just over a third of the cost. Fable 5 (2026-06-09) holds vendor SWE-Bench Pro 80.3% (vs GPT-5.5 58.6%), #1 LMArena Elo 1510 and #1 Artificial Analysis Index 65 as of 6/11. Sonnet 5 (2026-06-30): OSWorld-Verified 78.5%. Legacy Opus-line reasoning-suite scores shown below as baseline pending third-party suites for Opus 5
A99.8%
2GLM / Z.ai (Zhipu AI)
GLM-5.2 (~753B MoE, launched 2026-06-13) -- vendor-published; third-party verification still settling
A99.2%
3Microsoft MAI-Thinking-1
MAI-Thinking-1 (vendor-published 2026-06-02; third-party verification pending)
B97%
4Gemma 4 (Google)
Gemma 4 31B
A89.2%
5Qwen (Alibaba)
Qwen3.5-397B MoE
A87%
6Nemotron (Nvidia)
Llama-Nemotron Ultra 253B (prior gen -- Nemotron 3 Ultra 550B third-party scores pending)
B84.5%
7ChatGPT
GPT-5.5 (launched 2026-04-23; scores below are the GPT-5.4 baseline -- GPT-5.5 launch benchmarks per OpenAI are logged in Known Issues, pending third-party verification)
A83.3%

About AIME

Creator
Mathematical Association of America
Unit
% (max 100)

Other benchmarks