MMLU: 2026 AI Leaderboard
The 57-subject knowledge test that became the default LLM benchmark.
What it tests
MMLU (Massive Multitask Language Understanding) is a 14,000-question multiple-choice exam spanning 57 subjects from elementary math to professional law. It measures how much a language model actually knows, not how well it reasons.
How it is scored
Models answer four-choice questions in a zero-shot or few-shot setting. The reported score is average accuracy across all subjects. Scores above 85% are considered strong; humans average roughly 89% on this test.
Why it matters
MMLU is the most widely-reported LLM benchmark, which makes it the easiest point of apples-to-apples comparison across vendors. Its weakness is saturation -- frontier models now cluster in the upper 80s and 90s, so small differences are statistical noise. Use it to rule out weak models, not to pick a winner among strong ones.
Leaderboard (9 models)
Sorted by MMLUscore. Tier column shows the tool's overall AIToolTier rank, which blends this benchmark with pricing, features, and real-world usability.
| # | Model | Tier | MMLU score | Variant | Overall |
|---|---|---|---|---|---|
| 1 | Claude (Anthropic) Claude Opus 5 (launched 2026-07-24) is the current flagship and the default on Max -- Anthropic released COMPARATIVE claims only, no absolute scores: more than doubles Opus 4.8 on Frontier-Bench, within 0.5% of Fable 5 on CursorBench 3.2 at half the cost, 3x the next-best model on ARC-AGI 3, ~1.5x on Zapier AutomationBench, and beats Fable 5's best OSWorld 2.0 result at just over a third of the cost. Fable 5 (2026-06-09) holds vendor SWE-Bench Pro 80.3% (vs GPT-5.5 58.6%), #1 LMArena Elo 1510 and #1 Artificial Analysis Index 65 as of 6/11. Sonnet 5 (2026-06-30): OSWorld-Verified 78.5%. Legacy Opus-line reasoning-suite scores shown below as baseline pending third-party suites for Opus 5 | A | 91.3% | MMLU | 8.5/10 |
| 2 | ChatGPT GPT-5.5 (launched 2026-04-23; scores below are the GPT-5.4 baseline -- GPT-5.5 launch benchmarks per OpenAI are logged in Known Issues, pending third-party verification) | A | 91% | MMLU | 8.8/10 |
| 3 | DeepSeek DeepSeek V4-Pro (SWE-bench + Arena Elo third-party verified post-launch; knowledge rows are V3.x baseline pending V4 figures) | A | 90.8% | MMLU | 8.0/10 |
| 4 | Muse Spark (Meta) Muse Spark | A | 89% | MMLU | 8.8/10 |
| 5 | Grok Grok 4.20 (baseline -- Grok 4.5 launched 2026-07-08; vendor-reported 4.5 scores in Known Issues pending third-party verification) | B | 88.5% | MMLU | 7.5/10 |
| 6 | Nemotron (Nvidia) Llama-Nemotron Ultra 253B (prior gen -- Nemotron 3 Ultra 550B third-party scores pending) | B | 88.4% | MMLU (Llama-Nemotron 70B) | 7.8/10 |
| 7 | Mistral AI Mistral Medium 3.5 (vendor-published; third-party verification pending) | B | 86% | MMLU | 7.5/10 |
| 8 | Gemma 4 (Google) Gemma 4 31B | A | 83% | MMLU | 8.3/10 |
| 9 | Falcon (TII) Falcon 3 10B | B | 73.1% | MMLU | 7.1/10 |
About MMLU
- Creator
- Hendrycks et al., 2020 (UC Berkeley)
- Unit
- % (max 100)
- Official source
- https://arxiv.org/abs/2009.03300