GPQA Diamond: 2026 AI Leaderboard
Graduate-level physics, biology, and chemistry written to defeat Google-search.
What it tests
GPQA (Graduate-level Google-Proof Q&A) Diamond is the hardest subset of a 448-question multiple-choice set written by PhDs in physics, biology, and chemistry. Questions are deliberately designed so that searching the web does not yield the answer.
How it is scored
Four-choice accuracy. Domain PhDs with unlimited internet access score about 65%; non-expert humans with search score roughly 34%. Frontier models in 2026 are hitting the 80s and 90s -- a major inflection.
Why it matters
GPQA Diamond is the most cited reasoning benchmark for frontier LLMs precisely because it resists memorization. A high score implies the model can synthesize knowledge, not just recite training data.
Leaderboard (11 models)
Sorted by GPQA Diamondscore. Tier column shows the tool's overall AIToolTier rank, which blends this benchmark with pricing, features, and real-world usability.
| # | Model | Tier | GPQA Diamond score | Variant | Overall |
|---|---|---|---|---|---|
| 1 | ChatGPT GPT-5.5 (launched 2026-04-23; scores below are the GPT-5.4 baseline -- GPT-5.5 launch benchmarks per OpenAI are logged in Known Issues, pending third-party verification) | A | 92.8% | GPQA Diamond | 8.8/10 |
| 2 | Claude (Anthropic) Claude Fable 5 (launched 2026-06-09; suspended 2026-06-12 by US gov order; RESTORED globally 2026-07-01 after controls lifted 6/30) -- vendor SWE-Bench Pro 80.3% (vs GPT-5.5 58.6%); #1 LMArena Elo 1510 and #1 Artificial Analysis Index 65 as of 6/11. New default Sonnet 5 (2026-06-30): OSWorld-Verified 78.5%. Legacy Opus-line reasoning-suite scores shown below as baseline pending full third-party suites | A | 91.3% | GPQA Diamond | 8.5/10 |
| 3 | GLM / Z.ai (Zhipu AI) GLM-5.2 (~753B MoE, launched 2026-06-13) -- vendor-published; third-party verification still settling | A | 91.2% | GPQA Diamond | 8.0/10 |
| 4 | Muse Spark (Meta) Muse Spark | A | 86% | GPQA Diamond | 8.8/10 |
| 5 | Grok Grok 4.20 (baseline -- Grok 4.5 launched 2026-07-08; vendor-reported 4.5 scores in Known Issues pending third-party verification) | B | 85% | GPQA Diamond | 7.5/10 |
| 6 | Gemma 4 (Google) Gemma 4 31B | A | 84.3% | GPQA Diamond | 8.3/10 |
| 7 | DeepSeek DeepSeek V4-Pro (SWE-bench + Arena Elo third-party verified post-launch; knowledge rows are V3.x baseline pending V4 figures) | A | 79.9% | GPQA Diamond | 8.0/10 |
| 8 | Qwen (Alibaba) Qwen3.5-397B MoE | A | 78.2% | GPQA Diamond | 8.8/10 |
| 9 | Nemotron (Nvidia) Llama-Nemotron Ultra 253B (prior gen -- Nemotron 3 Ultra 550B third-party scores pending) | B | 70.5% | GPQA Diamond | 7.8/10 |
| 10 | Llama 4 (Meta) Llama 4 Maverick (17B/400B MoE) | B | 69.8% | GPQA Diamond | 7.9/10 |
| 11 | Falcon (TII) Falcon 3 10B | B | 42.5% | GPQA Diamond | 7.1/10 |
About GPQA Diamond
- Creator
- Rein et al., 2023 (NYU/Cohere/Anthropic)
- Unit
- % (max 100)
- Official source
- https://arxiv.org/abs/2311.12022