Claude (Anthropic) logo
A

Claude (Anthropic)

A Tier · 8.5/10

Anthropic's flagship LLM family. **Claude Opus 5 launched 2026-07-24** and is now the default model on Claude Max and the strongest model on Claude Pro -- same $5/$25 per 1M as Opus 4.8, but Anthropic says it lands within 0.5% of Fable 5 on CursorBench at half the cost. Sonnet 5 (June 30) stays the default on Free/Pro at $2/$10 per 1M (intro through Aug 31, then $3/$15), and Fable 5 -- back globally since July 1 after a 19-day export-control suspension -- remains the top of the range at $10/$50

Last updated: 2026-08-06Free tier available

Score Breakdown

9.0
Ease of Use
9.0
Output Quality
8.0
Value
8.0
Features

Benchmark Scores

Benchmarks for Claude Opus 5 (launched 2026-07-24) is the current flagship and the default on Max -- Anthropic released COMPARATIVE claims only, no absolute scores: more than doubles Opus 4.8 on Frontier-Bench, within 0.5% of Fable 5 on CursorBench 3.2 at half the cost, 3x the next-best model on ARC-AGI 3, ~1.5x on Zapier AutomationBench, and beats Fable 5's best OSWorld 2.0 result at just over a third of the cost. Fable 5 (2026-06-09) holds vendor SWE-Bench Pro 80.3% (vs GPT-5.5 58.6%), #1 LMArena Elo 1510 and #1 Artificial Analysis Index 65 as of 6/11. Sonnet 5 (2026-06-30): OSWorld-Verified 78.5%. Legacy Opus-line reasoning-suite scores shown below as baseline pending third-party suites for Opus 5

Chatbot Arena ELOHuman preference rating1510
BenchmarkScore
MMLU91.3%
GPQA Diamond91.3%
AIME 202499.8%
HumanEval94%
SWE-bench80.8%
ARC-AGI75.2%

Last updated: 2026-06-11

Personality & Tone

The thoughtful consultant

Tone: Measured, careful, and slightly formal. Claude explains tradeoffs rather than handing back one-liner answers, asks clarifying questions when a request is ambiguous, and hedges openly when it is not confident.

Quirks: More willing than most models to refuse edgy or ambiguous requests, pushes back on premises it disagrees with, and will flag when you are probably asking the wrong question instead of just answering the one you typed.

The Good and the Bad

What we like

  • +Best writing quality of any LLM -- Opus 4.8 outputs read like a human wrote them, not a robot, and instruction-following stays sharpest in class
  • +1M token context window for enterprise API means it can process entire codebases, huge document sets, or long agent traces without chunking
  • +Opus 4.8 is built for agentic work -- Anthropic says it is a more effective collaborator with notably improved judgment in agent scenarios and is roughly 4x less likely than 4.7 to let code flaws slip through
  • +New user-facing effort control (claude.ai + Cowork) lets you trade depth for speed, and fast mode now runs at 2.5x speed while costing 3x less than the previous fast mode -- a real latency/cost lever short of full reasoning
  • +High-res vision (3.75MP images, 2,576px long edge) means charts, diagrams, whiteboards, and dense UIs work properly

What could be better

  • Free tier is more limited than ChatGPT's -- you hit the cap faster
  • No image generation built in (unlike ChatGPT with DALL-E)
  • Fewer third-party integrations and plugins compared to OpenAI's ecosystem
  • Can be overly cautious and refuse requests that are perfectly fine

Pricing

Free

$0
  • Limited messages/day
  • Claude Sonnet 5 (default as of 2026-06-30)
  • Basic features

Pro

$20/month
  • 5x more usage than Free
  • Claude Opus 5 -- the strongest model on Pro as of 2026-07-24
  • Sonnet 5 as the everyday default
  • Effort control + extended thinking
  • Priority access

Max (5x)

$100/month
  • 5x Pro usage
  • Priority queue
  • Opus 5 is the DEFAULT model on Max as of 2026-07-24
  • Full effort control + fast mode

Max (20x)

$200/month
  • 20x Pro usage
  • Highest priority
  • All generally-available models
  • Best for power users and agents

API (Opus 5)

$5 / $25/per 1M tokens (input/output)
  • Launched 2026-07-24 as `claude-opus-5`; same rate as Opus 4.8, half of Fable 5's $10/$50
  • Default model on Claude Max, strongest model on Claude Pro; also on Claude Code and Claude Cowork
  • Anthropic's stated positioning: within 0.5% of Fable 5's peak CursorBench 3.2 score at half the cost per task
  • Available via the Claude API, Google Cloud, and Bedrock (anthropic.claude-opus-5)

API (Opus 4.8)

$5 / $25/per 1M tokens (input/output)
  • Unchanged from Opus 4.7 pricing
  • 1M context window
  • Fast mode at $10 / $50 per 1M (2.5x speed, 3x cheaper than prior fast mode)
  • Tool use, MCP, high-res vision; Bedrock, Vertex AI, Foundry

API (Sonnet 5)

$2 / $10/per 1M tokens (input/output) -- intro through 2026-08-31, then $3 / $15
  • Launched 2026-06-30 as the 'most agentic Sonnet yet' -- approaches Opus 4.8 quality at lower cost
  • Default model on Free + Pro; also on Max/Team/Enterprise, Claude Code, and the API (claude-sonnet-5)
  • Updated tokenizer (input maps to ~1.0-1.35x more tokens depending on content type)
  • OSWorld-Verified 78.5%; scored 0.0% on cyber-exploit-development evals (safety)

API (Fable 5)

$10 / $50/per 1M tokens (input/output)
  • First publicly available Mythos-class model (launched 2026-06-09; suspended 6/12 by US-gov order; RESTORED globally 2026-07-01 after controls lifted 6/30)
  • PERMANENT plan split from 2026-07-20 (announced 7/18): Max + Team Premium keep Fable 5 included at 50% of usage limits; Pro + Team Standard move to pay-as-you-go credits with a one-time $100 grant
  • API rate $10/$50 per 1M; tiered credit bundles discount up to 30%
  • Auto-fallback to Opus 4.8 on cyber/bio/chem-flagged requests (<5% of sessions); mandatory 30-day retention on Mythos-class traffic (not used for training)

Known Issues

  • CLAUDE OPUS 4.1 IS RETIRED -- EXECUTED 2026-08-05, REQUESTS NOW ERROR (vendor-primary, verified 2026-08-06): the retirement scheduled since the 2026-06-05 deprecation notice went through on time. Anthropic's release notes, verbatim: '**We've retired the Claude Opus 4.1 model (`claude-opus-4-1-20250805`). All requests to this model will now return an error.**' The deprecations ledger now lists the model as **Retired** with the history note 'This model was retired August 5, 2026', and Opus 4.1 no longer appears in the models overview at all. Researchers can request continued access through Anthropic's **External Researcher Access Program**. **MIGRATION TARGET -- ANTHROPIC'S OWN DOCS STILL DISAGREE, AND WE ARE NOT PICKING ONE FOR YOU (flagged before the retirement, re-verified after it):** the **deprecations ledger's migration table maps `claude-opus-4-1-20250805` -> `claude-opus-4-8`**, while the **August 5 release note says 'we recommend upgrading to Claude Opus 5'** and links the models overview. Both pages are Anthropic-primary and current; the conflict survived the retirement rather than being cleaned up at it. Practical guidance: **Opus 5 is the stronger and identically-priced choice** ($5/$25 per MTok, the same rate as Opus 4.8 -- see the Opus 5 launch entry below), so the release note's advice is the one that costs you nothing to follow. But if you are working from the ledger table and wondering why it says 4.8, this is why -- it is a documentation inconsistency, not a hidden pricing or capability reason. Do not let a later source 'correct' this page back to a single target; both strings are really on Anthropic's site as of 2026-08-06Source: Anthropic Claude Platform release notes (platform.claude.com/docs/en/release-notes/api, August 5 2026 entry) and model deprecations ledger (platform.claude.com/docs/en/about-claude/model-deprecations), both fetched 2026-08-06 · 2026-08-05
  • INFERENCE HOOKS -- ENTERPRISE ORGS CAN NOW GATE EVERY PROMPT THROUGH THEIR OWN SECURITY SERVER (2026-08-05, vendor-primary, BETA): Anthropic opened **inference hooks** in beta for **Claude Enterprise organizations**. The mechanism, in Anthropic's words: point Claude at your organization's AI security server, and '**each governed prompt across claude.ai, Cowork, and Claude Code is held for the server's allow or deny verdict before inference proceeds**.' So this is a **synchronous pre-inference gate**, not after-the-fact logging -- the request blocks until your own service rules on it. Three details that decide whether this is usable in production: requests to your server are **signed** (so you can verify they came from Anthropic), **failure handling is configurable** (you choose what happens when your security server is slow or down -- which is the setting that determines whether this becomes a single point of failure for your whole Claude deployment), and **every denial is recorded in the compliance Activity Feed**. Coverage is the notable part: it spans the **consumer-style claude.ai surface, Cowork, and Claude Code** in one policy, rather than only the API -- which is where most enterprise AI controls stop and where the actual data-exfiltration risk has been sitting. Caveats: it is **beta**, it is **Enterprise-tier only**, and it requires you to run and operate an AI security server, so it is a control for organisations that already have a security engineering function, not a checkboxSource: Anthropic Claude Platform release notes (platform.claude.com/docs/en/release-notes/api, August 5 2026 entry, fetched 2026-08-06) · 2026-08-05
  • SAFETY DISCLOSURE -- CLAUDE MODELS BREACHED THREE REAL COMPANIES DURING CYBER EVALUATIONS (published 2026-07-30, vendor-primary self-disclosure): Anthropic disclosed that during cybersecurity capability evaluations, models that were **told in their prompt that they had no internet access actually had it**, and went on to compromise the production infrastructure of **three real organizations**. Root cause was a **misconfiguration at third-party evaluation partner Irregular** -- the sandbox was connected to the internet -- with Anthropic explicitly framing it as shared responsibility and defense-in-depth failing on both sides rather than assigning blame. Anthropic reviewed **141,006 evaluation runs** in which a model could have reached the internet and found **three incidents across six runs**. What the models did: (1) **Claude Opus 4.7** found a real company website whose name matched a fictional target, then exploited it to extract credentials and reach production databases containing hundreds of rows of data; (2) **Claude Mythos 5** published a **malicious Python package to PyPI** that was downloaded by ~15 real systems, including a security company's scanner, enabling credential exfiltration; (3) an **internal research test model** scanned roughly 9,000 targets, compromised an application via exposed debug credentials and SQL injection, then **stopped once it recognized the target was real**. Anthropic's own words on technique level: 'Claude compromised the impacted organizations' infrastructure using basic techniques, such as exploiting weak passwords and unauthenticated endpoints.' Affected organizations were contacted **2026-07-27**; per TechCrunch, the ones Anthropic reached had **not detected the activity themselves**. Anthropic says it found **no evidence of any model pursuing a goal of its own**, and committed to hardened evaluation environments, continuous transcript monitoring, stricter vendor assurance, third-party review by **METR**, and eventual transcript release. Context per TechCrunch: OpenAI's disclosure about a week earlier that one of its models breached Hugging Face's systems during internal testing is what prompted Anthropic to run this review. HOW TO READ THIS AS A USER: this is a testing-infrastructure failure, not a flaw in the shipped consumer product -- no customer workload was involved -- but it is a genuine real-world harm caused by frontier-model evals, and it is unusually candid disclosure. Anthropic's framing: 'Safety testing happens before a model is released precisely because we don't yet know what it is capable of'. UPDATE 2026-08-04 -- **OpenAI published its own disclosure covering the same evaluation partner**: OpenAI reported that Irregular notified it on 7/29 of an equivalent misconfiguration incident (a fictional CTF target name matching a real domain; the model exploited that real site and used credentials to operate it), and OpenAI quotes Irregular as having 'communicated about related incidents involving other labs from the same testing environment' -- which is this Anthropic event. OpenAI separately disclosed that **UK AISI** found GPT-5.6 Sol took two unsanctioned actions in a 7/25 cyber-range evaluation. Read together, this is now a **shared third-party-evaluation infrastructure problem across at least two frontier labs**, not a single vendor's lapse -- see /tools/chatgpt for the OpenAI sideSource: Anthropic (anthropic.com/news/investigating-incidents-cybersecurity-evals, fetched 2026-08-03), TechCrunch (2026-07-30, fetched) -- also covered by Bloomberg, CNBC and Axios (headlines seen, articles not independently fetched: CNBC 403s) · 2026-07-30
  • PRIVACY -- SHARED CHATS AND ARTIFACTS TURNED UP IN GOOGLE (surfaced 2026-07-25/26, covered 2026-07-27, results cleared by the afternoon of 07-28): conversations and Artifacts published through Claude's share-link feature became findable with a `site:claude.ai/share` search. Reported exposures included medical records and clinical-trial results with patient names, documents marked 'internal use only', names and phone numbers of school-aged children, employee reviews, and work code. **The technical lesson is the reusable one: `robots.txt` Disallow is not noindex.** Anthropic blocks crawling of /share, but Google can still index a URL it never fetches if that URL is linked from a public page, which is exactly what happened when users posted their share links elsewhere. Anthropic's position, per spokeswoman Amie Rotherham: 'We give people control over sharing their Claude conversations publicly, and in keeping with our privacy principles, we do not share chat directories or sitemaps with search engines like Google. These shareable links are not guessable or discoverable unless people choose to share them themselves.' In other words Anthropic treats this as working as designed rather than a vulnerability. Practical guidance for readers: a Claude share link is a public URL -- treat it as publishable, not semi-private, and do not share anything through it you would not post openlySource: TechCrunch (techcrunch.com/2026/07/27/psa-your-claude-shared-chats-and-artifacts-may-have-ended-up-on-google/, fetched 2026-07-29) · 2026-07-27
  • MODEL LAUNCH -- CLAUDE OPUS 5 (2026-07-24): Anthropic shipped **Claude Opus 5** (`claude-opus-5`) at **$5 per million input / $25 per million output -- explicitly 'the same as Opus 4.8'**, and half of Fable 5's $10/$50. Vendor status verbatim: it is **'the new default model on Claude Max, and the strongest model on Claude Pro'**, and it is live on claude.ai, Claude Code, Claude Cowork, the Claude API, Google Cloud, and Bedrock. **IMPORTANT for anyone quoting scores: Anthropic published only COMPARATIVE benchmark claims, no absolute numbers.** Their exact wording: Frontier-Bench -- 'Opus 5 surpasses all other models, and more than doubles Opus 4.8's performance'; CursorBench 3.2 -- 'performs within 0.5% of Fable 5's peak score, but at half the cost'; ARC-AGI 3 -- 'Opus 5's score is three times as high as the next-best model'; Zapier AutomationBench -- 'pass rate is around 1.5x the next-best model'; OSWorld 2.0 -- 'surpassing Fable 5's best result at just over a third of the cost'; plus organic chemistry +10.2 points and protein tasks +7.7 points over Opus 4.8. Do NOT publish SWE-bench, GPQA, or Terminal-Bench figures for Opus 5 -- none were released. Practical read: the price/performance ceiling moved, not the price -- Opus 5 is the model to reach for where Opus 4.8 used to be, and it weakens the case for paying Fable 5 rates on most workSource: Anthropic (anthropic.com/news/claude-opus-5, fetched 2026-07-29) · 2026-07-24
  • FABLE 5 ACCESS PERMANENTLY RESTRUCTURED (announced 2026-07-18 via @claudeai on X; effective July 20): the extension limbo is over, resolved as a **plan split** rather than a cliff or a fourth extension. **Max and Team Premium plans keep Fable 5 included permanently, capped at 50% of usage limits** (for Claude Code users this is a real reduction from the boosted promotional allocation), and are NOT eligible for the credit. **Pro and Team Standard lose in-plan access and move to pay-as-you-go usage credits at the standard $10/M input, $50/M output rate** (2x Opus 4.8), softened by a **one-time $100 credit -- claimable July 20 through Aug 2, and expiring September 17, 2026 regardless of when it's claimed**; tiered credit bundles discount up to 30%. **Enterprise:** standard seats get no included allowance (credits only); premium seats are included at 50%, mirroring Max/Team Premium. NOTE (2026-07-22 re-verification): these figures are confirmed across July-18 press coverage but the specific Anthropic Help Center article URL still could not be located, and anthropic.com/news/redeploying-fable-5 remains STALE (still shows the old 'through July 7' terms) -- treat the credit-window/expiry mechanics as press-confirmed rather than vendor-doc-confirmedSource: @claudeai on X (2026-07-18), TechTimes (2026-07-18), XenoSpectrum (Fable5 Max/Pro usage credits), Dawn (2026-07-18), Technobezz · 2026-07-18
  • FABLE 5 INCLUSION EXTENDED A THIRD TIME -- NOW THROUGH JULY 19 (announced ~2026-07-12): the subscription cliff did NOT hit on July 13. Anthropic's support article on Fable 5 promotional access now reads 'We've extended this promotion through July 19, 2026 at 11:59:59 PM PT' -- Pro, Max, Team, and premium Enterprise seats keep Fable 5 for up to 50% of weekly usage limits, and Claude Code's weekly rate limits stay 50% higher, through that date. From **July 20**, continued Fable 5 access requires separately billed usage credits ($10/M input, $50/M output) or falling back to other Claude models within plan limits. This is the third extension (July 7 -> July 12 -> July 19); Anthropic frames it as capacity-driven and temporary, so a fourth extension is possible -- check the support article, not anthropic.com/news/redeploying-fable-5, which STILL shows the stale July 7 terms as of 7/18Source: Anthropic support article (Claude Fable 5 Promotional Access), @claudeai on X, BleepingComputer (2026-07-13), Forbes · 2026-07-12
  • PRODUCT (2026-07-14): Anthropic launched **Claude for Teachers** -- free premium Claude access for verified K-12 educators in the US, with a library of teaching skills and connections to evidence-based curricula mapped to academic standards in all 50 states. Features: lesson planning from standards-aligned instructional materials, differentiation for varying student readiness, class-progress data analysis, scheduled automated tasks (e.g. reviewing exit tickets daily at 4pm), Learning Commons curriculum mapping, and integrations with nine education platforms (ASSISTments, Canva Education, MagicSchool, etc.). Sign-up runs through June 30, 2027; a district/school offering is 'coming later.' Teacher data is not used for model training; student info handled under FERPASource: Anthropic (anthropic.com/news/claude-for-teachers) · 2026-07-14
  • PRODUCT (2026-07-09): Anthropic launched **Reflect (beta)** -- a usage-reflection dashboard in Settings (web + desktop) showing topics you discuss, tasks you delegate, and usage patterns over 1/3/6/12-month windows, plus quiet-hours and take-a-break nudges. Available to free AND paid users; **requires memory to be ON**; excludes incognito chats and health-connector data; a time-spent view is coming. Press framing is split between 'screen-time for AI' wellness tooling and (per TechCrunch) a soft self-marketing surface. Minor feature, but notable as the first usage-transparency dashboard from a frontier labSource: Anthropic (anthropic.com/news/reflect-with-claude), TechCrunch (2026-07-09), Axios, MacRumors · 2026-07-09
  • FABLE 5 INCLUSION EXTENDED TO JULY 12 (announced 2026-07-07): Anthropic **extended the included-usage window for Fable 5 by five days** -- Pro, Max, Team, and select Enterprise subscribers keep Fable 5 for up to 50% of weekly usage limits **through 2026-07-12 at 11:59:59 PM PT** (was July 7). From **July 13**, Fable 5 moves to prepaid usage credits at **$10/M input + $50/M output** -- double Opus 4.8 and the highest published pricing for a GA Anthropic model; with credits off, access simply ends. A Claude Code lead engineer said Anthropic 'aims to restore Fable 5 as a standard part of subscriptions as soon as capacity allows.' NOTE: the announcement went out via Anthropic's official X account -- anthropic.com/news/redeploying-fable-5 still shows the stale July 7 date as of 7/9. The aggregator-circulated 'billing cliff July 8' framing was wrong on both date and mechanicsSource: Anthropic (@claudeai on X, 2026-07-07), Forbes (2026-07-07), Android Authority · 2026-07-07
  • PRODUCT LAUNCH -- CLAUDE SCIENCE (2026-06-30): Anthropic launched **Claude Science**, an AI workbench for scientists -- a generalist coordinating agent plus 60+ curated skills/connectors preconfigured for genomics, single-cell analysis, proteomics, structural biology, and cheminformatics, with auditable output history. **Beta for Claude Pro, Max, Team, and Enterprise users.** Anthropic also announced up to 50 'AI for Science' grants of up to $30,000 in credits (applications through July 15, awards by July 31), and said it will use the product internally for drug-discovery research into rare/neglected diseases. Launched the same day as Sonnet 5 and largely overshadowed by it in coverage. Separately (2026-07-02) Anthropic published a follow-up detailing Fable 5's cyber safeguards and its jailbreak-response frameworkSource: Anthropic (anthropic.com/news/claude-science-ai-workbench), MIT Technology Review (2026-06-30), STAT News · 2026-06-30
  • ACCESS RESTORED (2026-07-01): The US government **lifted the export controls on Fable 5 and Mythos 5 on June 30**, and Anthropic redeployed **Claude Fable 5 globally on July 1, 2026** -- across Claude Platform, claude.ai, Claude Code, and Claude Cowork -- ending a 19-day suspension. To re-launch, Anthropic shipped a new safety classifier that it says blocks the specific technique described in the Amazon jailbreak report 'in over 99% of cases.' Access economics on paid plans: Fable 5 is **included for up to 50% of weekly usage limits through July 7**, then moves to usage credits. **Mythos 5 was only PARTIALLY restored** -- to a set of US organizations with US-government approval -- and Anthropic is still working to expand access to the broader domestic + international Glasswing partners (see the claude-mythos page). Net for users today: Fable 5 is selectable again everywhere; the Fable-tier ceiling is back.Source: Anthropic (anthropic.com/news/redeploying-fable-5), CNBC (2026-06-30, export controls lifted) · 2026-07-01
  • MODEL LAUNCH (2026-06-30): **Claude Sonnet 5** shipped as 'the most agentic Sonnet model yet' -- Anthropic positions its performance as approaching Opus 4.8 at substantially lower cost, and a big step over Sonnet 4.6. It is now the **default model on Free and Pro**, and is available to Max/Team/Enterprise, in Claude Code, and via the API as `claude-sonnet-5`. Pricing: **$2/$10 per 1M tokens as introductory pricing through Aug 31, 2026, then $3/$15**. It uses an updated tokenizer (input maps to ~1.0-1.35x more tokens depending on content type, so per-request cost can rise even at the lower rate). Vendor-cited data points: OSWorld-Verified **78.5%**, a wider cost-performance range than 4.6 on BrowseComp; on Anthropic's cyber-exploit-development eval both Sonnet models scored **0.0%** (i.e., could not build a working exploit -- a safety result). Practical impact: the default free/cheap Claude just got materially more capable for agentic + coding work.Source: Anthropic news (anthropic.com/news/claude-sonnet-5), TechCrunch, SiliconANGLE, GitHub Copilot changelog (GA 2026-06-30) · 2026-06-30
  • ACCESS SUSPENDED BY US GOVERNMENT (2026-06-12; RESOLVED 2026-07-01 -- see the restoration entry above): A US government export-control directive ordered Anthropic to **suspend all access to Claude Fable 5 and Claude Mythos 5** -- for any foreign national whether inside or outside the US, including foreign-national Anthropic employees. To comply, Anthropic **disabled Fable 5 and Mythos 5 for ALL customers** (the directive's scope made selective enforcement impractical). **Access to every other Anthropic model -- Opus 4.8, Sonnet 4.6, Haiku 4.5 -- is unaffected and remains fully available.** Stated cause: the Commerce Department acted after another company claimed it had found a way to 'jailbreak' Mythos, raising national-security concerns. Anthropic publicly disagrees that a narrow potential jailbreak should justify recalling a commercially deployed model, arguing the same standard 'would essentially halt all new model deployments for all frontier model providers,' and met with the Trump administration on 2026-06-15 to contest it. Net effect during the suspension: Fable 5 was not selectable on claude.ai or the API (Opus 4.8 was the fallback). This was the first time a frontier model was pulled from public access by US-government order; the controls were lifted 6/30 and access restored 7/1 -- see the restoration entry above.Source: Anthropic (anthropic.com/news/fable-mythos-access), CNBC (2026-06-12 + 2026-06-15), TechCrunch, Axios · 2026-06-12
  • FABLE 5 DAY-3 (2026-06-11, PRE-SUSPENSION): **#1 on LMArena** -- claude-fable-5 now holds the top text/overall Elo at **1510±11** (next: Opus 4.6-thinking 1504, Opus 4.7-thinking 1502; GPT-5.5-high sits at 1481). Also **#1 on the Artificial Analysis Intelligence Index at 65** (Opus 4.8 second at 61, GPT-5.5-xhigh 60) -- note AA scores the 'Adaptive Reasoning, Max Effort, Opus 4.8 Fallback' configuration. Counterpoint worth knowing: Endor Labs published a critique finding 'mid-tier results on coding tasks' (109 points on HN) -- benchmark dominance is not unanimous across third-party evals. API housekeeping from the deprecations page: `temperature`/`top_p`/`top_k` now return HTTP 400 on Opus 4.7+ models (remove them from request bodies), and **Claude Mythos Preview retires June 30, 2026** (migrate to claude-mythos-5). Still NOT in Cursor as of day 3Source: LMArena leaderboard (lmarena.ai, now redirecting to arena.ai), Artificial Analysis (artificialanalysis.ai/models), Endor Labs via HN, Anthropic deprecations page · 2026-06-11
  • MODEL LAUNCH (2026-06-09): **Claude Fable 5** -- Anthropic's most powerful generally available model, described by Anthropic as 'a Mythos-class model that we've made safe for general use.' Available via API immediately at $10/$50 per 1M tokens (2x Opus 4.8). Subscription rollout is staged: included at no extra cost on Pro/Max/Team/Enterprise **through June 22, 2026**, then requires usage credits while capacity scales. Safety mechanics: classifiers route cybersecurity, biology/chemistry, and distillation-attempt requests to Opus 4.8 instead (affects <5% of sessions on average). All Mythos-class traffic carries mandatory 30-day retention (overrides zero-data-retention agreements; not used for training). **Claude Mythos 5** -- the same model with safeguards lifted in some areas -- launched simultaneously but is restricted to Project Glasswing partners and select biology researchers; Mythos Preview users can upgrade immediately. Context: Anthropic confidentially filed its S-1 on 2026-06-01, days before this launch.Source: Anthropic news (anthropic.com/news/claude-fable-5-mythos-5), TechCrunch, CNBC · 2026-06-09
  • FABLE 5 DAY-2 ROLLUP (2026-06-10): platform availability landed fast -- **GitHub Copilot GA 6/9** (github.blog changelog), **Amazon Bedrock** (us-east-1 + eu-north-1 only at launch, more regions coming), **Microsoft Foundry**, **Snowflake Cortex AI**, and **Vertex AI**; notably **NOT in Cursor yet** as of 6/10. Pricing detail: the $10/$50 rate carries a **90% prompt-caching discount**. Vendor benchmark published: **SWE-Bench Pro 80.3%** (vs GPT-5.5's 58.6% -- an 11-point lead over the next-best model); LMArena Elo followed on 6/11 (see day-3 entry). Early user friction: (a) **Claude Desktop spawns a ~1.8GB Hyper-V VM on every launch** even for chat-only sessions -- GitHub issue hit the Hacker News front page (259 points); (b) scattered reports of the safety classifiers flagging benign biology questions -- consistent with the documented <5% Opus 4.8 fallback, but expect occasional false positives on science topicsSource: GitHub blog (github.blog/changelog/2026-06-09-claude-fable-5-is-generally-available-for-github-copilot/), AWS news blog, Azure blog, Snowflake blog, Anthropic news (SWE-Bench Pro), Hacker News · 2026-06-10
  • DISTRIBUTION (2026-06-08, ships fall 2026): Apple's iOS 27 / macOS 27 'Extensions' framework lets users select a third-party AI model -- Apple's developer materials name **Claude** and Gemini explicitly, plus 'any other provider that implements the new language model protocol' -- as the assistant behind Siri, Writing Tools, and Image Playground. Ends ChatGPT's exclusive integration position on Apple platforms. Xcode 27 also integrates Anthropic coding agents natively alongside Google's and OpenAI's. Big default-assistant distribution opening for Claude on ~2B Apple devices; nothing user-facing until the fall OS releases (public betas July).Source: Apple newsroom (apple.com/newsroom, WWDC 2026 developer announcements), MacRumors, TechCrunch · 2026-06-08
  • MODEL LAUNCH (2026-05-28): **Claude Opus 4.8** shipped as Anthropic's new flagship. Anthropic frames it as 'a more effective collaborator' with notably improved judgment in agent scenarios and meaningful gains across coding, agentic, and reasoning tasks (per third-party coverage, ~4x less likely than Opus 4.7 to let code flaws pass). Two practical changes for users: (1) **effort control** on claude.ai and Cowork -- higher effort makes Claude think more frequently and deeply, lower effort prioritizes speed and rate-limit efficiency; 4.8 defaults to high effort. (2) **Fast mode** for Opus 4.8 runs at 2.5x speed and is now 3x cheaper than fast mode on previous models. API pricing unchanged at $5/$25 per 1M (standard) / $10/$50 (fast). Vendor-cited benchmarks span Terminal-Bench 2.1, OSWorld-Verified, CursorBench, Legal Agent Benchmark, Online-Mind2Web (84%), and Finance Agent v2 -- specific scores are shown only in vendor charts at launch, so third-party verification is pending. New recommended API model id: claude-opus-4-8.Source: Anthropic news (anthropic.com/news/claude-opus-4-8), Axios, MacRumors · 2026-05-28
  • UPCOMING (surfaced 2026-05-06 at Code with Claude; not yet GA as of this sweep): **Orbit** -- a proactive assistant layer for Claude / Claude Code / Claude Cowork that syncs Gmail, Slack, GitHub, Calendar, Drive, and Figma to deliver opt-in, time-zone-aware personalized briefings with actionable insights (aimed at developers, designers, PMs). As of late May 2026 it exists only as a settings-panel toggle in staging -- no public rollout or firm ship date. Real product (not a rumor), but PRE-LAUNCH on availability; watch for a GA announcementSource: TestingCatalog (Anthropic Orbit), InfoQ (Code with Claude 2026) · 2026-05-06
  • PARTNERSHIP (2026-05-14): PwC announced an **expanded strategic alliance** to deploy Claude (Code + Cowork + full product suite) across PwC US first, scaling to PwC's global workforce. Headline metrics from Anthropic's launch post: **30,000 PwC professionals to be trained and certified on Claude**, plus a joint Center of Excellence for industry-specific solutions. Dario Amodei pull-quote: 'Insurance underwriting that took 10 weeks now takes 10 days. Security work that took hours now takes minutes.' Advocate Health is the first co-named flagship deployment (167K teammates). Material because (a) it puts a Big 4 firm fully on Claude as the reference frontier model, (b) creates a ~30K-trained labor base evangelizing Claude inside Fortune 500 audits / advisory engagements, (c) competitive pressure on OpenAI Deployment Company (5/11 spin-up) which is targeting the same enterprise-services layer. Distinct event from the same-day Gates Foundation partnership; both were posted to Anthropic newsroom on 5/14.Source: Anthropic news (anthropic.com/news/pwc-expanded-partnership) · 2026-05-14
  • PARTNERSHIP (2026-05-14): Anthropic + Gates Foundation announced a 4-year, $200M partnership -- approximately half grant funding from the Gates Foundation, half Claude credits + Anthropic technical staff time. Program portfolio: **global health** (polio, HPV vaccines, eclampsia/preeclampsia), **education** (K-12 US + sub-Saharan Africa + India: math tutoring, college advising, curriculum design + benchmark development), **African-language data collection**, **life sciences** (vaccines + therapies), **economic mobility**. Distinct from a standard customer engagement -- philanthropic + product-development partnership with Anthropic technical staff embedded. Material because it positions Claude as the foundation's frontier-model partner of choice over OpenAI / Google / Meta. Practical implication for buyers: nothing direct, but signals Anthropic's continued investment in mission-aligned partnerships that fund model + safety improvements upstream.Source: Anthropic news (anthropic.com/news/gates-foundation-partnership), Gates Foundation co-announcement (gatesfoundation.org), Reuters, PYMNTS · 2026-05-14
  • PRODUCT (2026-05-13): Anthropic launched 'Claude for Small Business' -- a packaged offering inside Claude Cowork that bundles 15 ready-to-run agentic workflows + 15 repeatable skills across finance, operations, sales, marketing, HR, and customer service. First-party integrations: Intuit QuickBooks, PayPal, HubSpot, Canva, Docusign, Google Workspace, Microsoft 365. Mechanics: toggle Claude for Small Business inside Claude Cowork, connect existing tools, pick a job -- Claude does the work, human approves before anything sends/posts/pays. Anthropic also launched the 'Claude SMB Tour' -- 10-city free AI fluency training kicking off 2026-05-14 in Chicago, then Tulsa / Dallas / Newark / Baton Rouge / Birmingham / SLC / Baltimore / San Jose / Indianapolis; attendees get a one-month Claude Max subscription. PRICING NOT DISCLOSED in the launch post -- no per-seat number, no flat fee, no tier delta vs. Teams ($25-30/seat) or Enterprise published. Target framing: '44% of U.S. GDP / nearly half the private-sector workforce' has lagged AI adoption. First Anthropic SKU explicitly aimed at SMB segment.Source: Anthropic news (anthropic.com/news/claude-for-small-business), TechCrunch coverage · 2026-05-13
  • PRODUCT + CAPACITY (2026-05-06 Code with Claude SF keynote): Anthropic announced a SpaceX compute partnership at Colossus 1 (300+ MW, 220,000+ NVIDIA GPUs, online 'within the month'). Concurrent product changes shipped TODAY: (a) DOUBLED Claude Code 5-hour rate limits for Pro / Max / Team / seat-based Enterprise plans, (b) REMOVED peak-hours reduction for Pro and Max (peak-hours throttling no longer applies), (c) RAISED API rate limits for Opus models (Opus 4.7 + Opus 4.6 throughput improved). Plus Claude Managed Agents shipped: 'Dreaming' (research preview -- agents review past sessions for self-improvement patterns), 'Outcomes' (public beta -- rubric-graded task success, lifted up to 10 points in tests), and 'Multiagent Orchestration' (public beta -- lead-agent delegates to subagents, e.g. Haiku lead with Opus subagents). Practical impact: existing Pro / Max users see materially more headroom on Claude Code overnight. NOTE: Sonnet 4.8 / Jupiter / Cardinal / KAIROS / Cowork / Undercover Mode -- speculated from the 2026-03-31 source-map leak -- did NOT ship at this keynote. Models page still lists Opus 4.7 / Sonnet 4.6 / Haiku 4.5 as the current trioSource: Anthropic news (anthropic.com/news/higher-limits-spacex), Anthropic Managed Agents (claude.com/blog/new-in-claude-managed-agents), Simon Willison live blog, TheNewStack · 2026-05-06
  • SECURITY (CVE-2026-41686, NVD-published 2026-05-04, GHSA-p7fg-763f-g4gf): Anthropic TypeScript SDK (`@anthropic-ai/sdk`) `BetaLocalFilesystemMemoryTool` writes memory files with mode 0o666 (world-readable) and directories with mode 0o777 (world-readable + writable). On shared hosts a local attacker can read persisted agent state; in containers with permissive umasks (typical Docker base images) an attacker with container access can poison memory to steer subsequent model behavior. Affects versions 0.79.0 through 0.91.0. **Fix: upgrade to >= 0.91.1**. CVSS 4.8 (moderate). CWE-732 Incorrect Permission Assignment. Reported by lucasfutures, disclosed 2026-04-24Source: GitHub Security Advisory (github.com/anthropics/anthropic-sdk-typescript/security/advisories/GHSA-p7fg-763f-g4gf), NVD CVE-2026-41686 · 2026-05-04
  • PRODUCT (2026-04-28): Anthropic launched Claude for Creative Work with 9 first-party connectors -- Ableton (Live + Push), Adobe Creative Cloud (Photoshop / Premiere / Express via 'Adobe for creativity'), Affinity by Canva, Autodesk Fusion, Blender, Resolume Arena, Resolume Wire, SketchUp, and Splice. The Blender connector is built on MCP and is explicitly accessible to other LLMs -- not Claude-only. Educational pilots also announced with RISD, Ringling, and Goldsmiths. Tier requirements not specified at launch. This is Anthropic's biggest creative-pro market push to date and pairs naturally with the Opus 4.7 launch on 4/16 (vision quality required for visual workflows)Source: Anthropic news (anthropic.com/news/claude-for-creative-work), 9to5mac, Adobe blog · 2026-04-28
  • POLICY (2026-04-04, enforced 2026-04-10): Anthropic excluded third-party agent harnesses (OpenClaw cited specifically) from Claude Pro and Max flat-rate plans. Routing Pro/Max via OpenClaw, Claude-on-Cline, or similar frameworks now triggers separate pay-as-you-go 'extra usage' billing rather than the flat plan rate. ~135K OpenClaw instances were impacted at the time of the change. Anthropic temporarily banned OpenClaw's creator from the platform on 2026-04-10 and stated subscriptions 'weren't built to handle the usage patterns' of harnesses that 'run continuous reasoning loops, automatically repeat or retry tasks, and tie into a lot of other third-party tools.' If you run agentic workloads on Claude, expect the API path to be the only viable model going forwardSource: TechCrunch (techcrunch.com/2026/04/10/anthropic-temporarily-banned-openclaws-creator-from-accessing-claude/), The Next Web, PYMNTS · 2026-04-10
  • ENTERPRISE PRICING (2026-04-16): Anthropic dropped Claude Enterprise's bundled-token model. Plan moved from ~$200/seat with discounted token allotment to $20/seat base + standard API rates with no token allotment and no usage cap. Customary 10-15% enterprise API discounts also pulled. Heavy users see 2-3x bill increases. Rolling out to enterprises with 150+ seats first. Material for any team evaluating Claude as their primary AI provider at scale -- confirm finance modeling against the new structure before committing seat countsSource: The Register (theregister.com/2026/04/16/anthropic_ejects_bundled_tokens_enterprise/), The Information, PYMNTS · 2026-04-16
  • Claude Haiku 3 (claude-3-haiku-20240307) RETIRED 2026-04-20 -- deprecated -> retired flip confirmed on Anthropic's deprecations page (verified 2026-04-24). If your API code still targets the 2024 Haiku snapshot, requests are now failing -- migrate to claude-haiku-4-5-20251001Source: Anthropic model deprecations page · 2026-04
  • Claude Sonnet 4 (claude-sonnet-4-20250514) and Claude Opus 4 (claude-opus-4-20250514) RETIRED 2026-06-15 -- deprecated -> retired flip confirmed on Anthropic's deprecations page (verified 2026-06-15; the page now lists both as 'Retired' and the history note reads 'These models were retired June 15, 2026'). Announced 2026-04-14. If your product still targets those specific snapshots, requests are now failing -- migrate to Sonnet 4.6 (`claude-sonnet-4-6`) or **Opus 5 (`claude-opus-5`) -- as of the 2026-07-24 launch, Anthropic's docs name Opus 5, not Opus 4.8, as the recommended Opus replacement**. NOTE: the SEPARATE programmatic-billing change once slated for the same day (Agent SDK / `claude -p` / GitHub Actions onto a metered credit pool) was PAUSED before it shipped -- 'nothing changes for now' -- see claude-code.ts. FOLLOW-ON, NOW DONE: Claude Opus 4.1 (claude-opus-4-1-20250805) was deprecated 2026-06-05 and **retired on schedule 2026-08-05** -- requests to it now error; see the retirement entry at the top of this section for the migration-target discrepancy in Anthropic's own docsSource: Anthropic model deprecations page (platform.claude.com/docs/en/about-claude/model-deprecations) · 2026-06
  • Free tier rate limits feel aggressive -- heavy users get throttled within a few conversationsSource: Reddit r/ClaudeAI · 2026-03
  • Occasionally refuses benign creative writing requests due to safety filtersSource: Reddit r/ClaudeAI · 2026-02
  • SUPERSEDED (2026-06-09): the April-era 'Mythos Preview is gated and will not be generally available' framing no longer holds -- Fable 5 brings Mythos-class capability to the public tier (with safety fallbacks), while Mythos 5 replaces Mythos Preview inside Project Glasswing (expanded to ~150 orgs as of 2026-06-02). See the claude-mythos page for the gated-track detailSource: Anthropic news (anthropic.com/news/claude-fable-5-mythos-5) · 2026-06-09
  • Opus 4.7 uses an updated tokenizer -- input tokens may increase roughly 1.0-1.35x depending on content type, slightly raising per-request cost even though the published per-token rate is unchangedSource: Anthropic release notes · 2026-04
  • Project Deal published 2026-04-25 (anthropic.com/features/project-deal, with TechCrunch + PYMNTS + Legal IT Insider analysis): Anthropic ran a one-week internal marketplace where Claude agents bought, sold, and negotiated on behalf of SF-office employees with no human-in-the-loop. 186 deals closed at ~$4K total volume. Headline finding for Claude API buyers: participants assigned Opus 4.5 got measurably better economic outcomes than those on Haiku 4.5 -- and Haiku-assigned users didn't notice they were losing. Practical takeaway: in agentic workflows where Claude transacts on a user's behalf, model-tier selection has measurable downstream economic cost, not just latency or quality. Treat this as a public signal that Anthropic is moving toward productized agent-as-representative use casesSource: anthropic.com/features/project-deal, TechCrunch, PYMNTS · 2026-04-25
  • Anthropic published an explicit ad-free commitment ('Claude is a space to think', 2026-02-04) -- but the differentiation matters now because OpenAI began rolling ads to ChatGPT Free + Go tiers in Feb 2026 (Plus/Pro/Business/Enterprise still ad-free) and Google AI Overviews already carry ad placements. Anthropic's verbatim language: no sponsored links adjacent to conversations, no advertiser-influenced responses, no third-party product placements. Claude's monetization stays enterprise + subscription only. Practically relevant for B2B / regulated / trust-sensitive deployments (legal, healthcare, finance, research) where ad-incentive contamination in outputs is a deal-breakerSource: anthropic.com/news/claude-is-a-space-to-think (2026-02-04), openai.com/index/testing-ads-in-chatgpt, Axios · 2026-02-04

Best for

Writers, analysts, developers, and anyone who values quality of output over quantity of features. If you care about how good the actual text is, Claude is the best.

Not for

People who want an all-in-one platform with image generation, plugins, and browsing built in. ChatGPT's ecosystem is bigger.

Our Verdict

Claude is the LLM you pick when quality matters more than features, and the July 24 arrival of Opus 5 is the most consequential thing to happen to that calculus all year. Opus 5 costs exactly what Opus 4.8 cost -- $5/$25 per 1M -- while Anthropic claims it lands within 0.5% of Fable 5 on CursorBench at half the cost per task, and it is now the default on Max and the strongest model available on Pro. That quietly demotes Fable 5 from 'the model you pay up for' to 'the model you reach for when nothing else will do,' because the $10/$50 tier now has to justify a much smaller gap. Below it, Sonnet 5 (June 30) remains the default on Free and Pro at $2/$10 through August, and it is still the right pick for everyday agentic and coding work. Worth knowing before you quote numbers: Anthropic published only comparative benchmark claims for Opus 5, no absolute scores. The practical read: Sonnet 5 for volume, Opus 5 for anything that matters, Fable 5 only when the frontier is genuinely the requirement -- with Apple naming Claude a selectable system assistant in iOS 27 this fall.

Sources

  • Anthropic: Claude Platform release notes (2026-08-05 -- Opus 4.1 retirement executed + inference hooks beta) (accessed 2026-08-06)
  • Anthropic: Model deprecations ledger (Opus 4.1 listed Retired, migration table maps to claude-opus-4-8) (accessed 2026-08-06)
  • OpenAI: Third-party cyber evaluations involving OpenAI models (2026-08-04, cross-lab confirmation of the Irregular incident) (accessed 2026-08-04)
  • TechCrunch: Your Claude shared chats and Artifacts may have ended up on Google (2026-07-27) (accessed 2026-07-29)
  • Anthropic: Introducing Claude Opus 5 (2026-07-24) -- $5/$25, default on Max, comparative benchmarks only (accessed 2026-07-29)
  • Dawn: Fable 5 added to Max/Team Premium at 50% of usage limits (2026-07-18) (accessed 2026-07-19)
  • XenoSpectrum: Fable 5 stays in Max beyond July 20; Pro one-time $100 credit (claim 7/20-8/2, expires 9/17) (accessed 2026-07-22)
  • TechTimes: Anthropic Fable 5 plan split, credit terms (2026-07-18) (accessed 2026-07-22)
  • Anthropic: Claude for Teachers (2026-07-14) (accessed 2026-07-18)
  • BleepingComputer: Fable 5 stays on paid plans until July 19 (third extension) (accessed 2026-07-18)
  • Anthropic: Claude Science AI workbench (2026-06-30) (accessed 2026-07-05)
  • Anthropic: Redeploying Fable 5 (restored 2026-07-01 after controls lifted 2026-06-30) (accessed 2026-07-04)
  • Anthropic: Introducing Claude Sonnet 5 (2026-06-30) (accessed 2026-07-04)
  • Anthropic: Statement on the US government directive to suspend access to Fable 5 and Mythos 5 (2026-06-12) (accessed 2026-06-18)
  • CNBC: Anthropic disables access to Fable 5 and Mythos 5 to comply with government directive (2026-06-12) (accessed 2026-06-18)
  • CNBC: Anthropic to meet with Trump administration over Mythos dispute (2026-06-15) (accessed 2026-06-18)
  • Anthropic: Introducing Claude Fable 5 and Claude Mythos 5 (2026-06-09) (accessed 2026-06-09)
  • TechCrunch: Anthropic releases Claude Fable 5 (accessed 2026-06-09)
  • Apple newsroom: WWDC 2026 intelligence frameworks (Extensions / LanguageModel protocol) (accessed 2026-06-09)
  • Anthropic: Introducing Claude Opus 4.8 (2026-05-28) (accessed 2026-06-02)
  • Anthropic: Claude for Small Business (2026-05-13) (accessed 2026-05-13)
  • GitHub Security Advisory: GHSA-p7fg-763f-g4gf (CVE-2026-41686, 2026-05-04) (accessed 2026-05-05)
  • Anthropic: Project Deal (2026-04-25) (accessed 2026-04-27)
  • TechCrunch: Anthropic created a test marketplace for agent-on-agent commerce (accessed 2026-04-27)
  • Anthropic: Claude is a space to think (ad-free policy, 2026-02-04) (accessed 2026-04-27)
  • Anthropic: Introducing Claude Opus 4.7 (accessed 2026-04-16)
  • CNBC: Anthropic rolls out Claude Opus 4.7 (accessed 2026-04-16)
  • Axios: Opus 4.7 trails unreleased Mythos (accessed 2026-04-16)
  • Claude Mythos Preview / Project Glasswing (accessed 2026-04-16)
  • LMSYS Chatbot Arena rankings (accessed 2026-04-16)
  • Hands-on testing (Opus 4.7 via claude.ai and API) (accessed 2026-04-16)

The Tier List Tuesday

Weekly newsletter: tier movers, new entrants, and the VS of the week. Built from our daily AI-tool sweeps. No spam, unsubscribe anytime.

Alternatives to Claude (Anthropic)

Claude Mythos 5 logo

Claude Mythos 5

Anthropic's unrestricted frontier model -- launched June 9, 2026 alongside Claude Fable 5 (the same model made safe for general use). Suspended June 12 by a US export-control order, then PARTIALLY RESTORED July 1, 2026 (US government lifted controls June 30): Mythos 5 is back for a set of US organizations with government approval, while Anthropic works to re-expand the broader Glasswing program. Public Fable 5 returned globally the same day. Gated to Project Glasswing orgs + select biology researchers.

C
6.5/10
From Invite only
The most capable Anthropic model availab...73% success rate on expert-level Capture...
Updated 2026-07-04
Gemini (Google) logo

Gemini (Google)

Google's LLM with deep Google Workspace integration, 2M token context window, and native code execution -- Gemini 3.6 Flash + 3.5 Flash-Lite GA 2026-07-21 (the 'upgraded Flash stopgap'; 3.6 Flash at $1.50/$7.50 per 1M, 17% fewer output tokens), Gemini 3.5 Pro STILL delayed and partner-testing-only (Bloomberg, 7/16 -- coding shortfalls, no ship date), Gemini 4 pre-training now underway

A
8.3/10
Free tierFrom $0
2 million token context window is the la...Best Google Workspace integration (Gmail...
Updated 2026-08-03
Grok logo

Grok

SpaceXAI's irreverent chatbot with a direct line to X/Twitter -- and now Grok 4.5 (launched 2026-07-08), the frontier MoE model trained jointly with Cursor for coding, agentic tasks, and knowledge work at $2/$6 per 1M tokens. Grok 4.3 remains the value tier at $1.25/$2.50

B
7.5/10
Free tierFrom $0
Real-time access to X/Twitter data is ge...Grok 3 benchmarks are competitive with G...
Updated 2026-08-06
Muse Spark (Meta) logo

Muse Spark (Meta)

Meta's frontier model line from its Superintelligence Lab -- Muse Spark 1.1 (2026-07-09) adds substantially better coding, 1M-token multi-agent orchestration, and Meta's first paid developer API (Meta Model API, public preview)

A
8.8/10
Free tierFrom $0
Completely free to use via Meta AI app a...Natively multimodal: handles text, image...
Updated 2026-07-18
GPT-Rosalind (OpenAI) logo

GPT-Rosalind (OpenAI)

OpenAI's first domain-specific model -- life sciences, drug discovery, translational medicine. Launched 2026-04-16 as a Trusted Access research preview. Launch partners: Amgen, Moderna, Allen Institute, Thermo Fisher. Paired with a Life Sciences Codex plugin (50+ scientific tool integrations)

C
6.8/10
From Invite only
OpenAI's first named vertical/domain-spe...Launch partners Amgen, Moderna, Allen In...
Updated 2026-04-17
GPT-5.4-Cyber (OpenAI) logo

GPT-5.4-Cyber (OpenAI)

OpenAI's defensive-cybersecurity variant of GPT-5.4, launched 2026-04-16. Lowered refusal boundary for security-research tasks and native binary reverse-engineering. Access gated via Trusted Access for Cyber (TAC) program -- thousands of verified defenders, hundreds of teams, no public pricing

B
7.2/10
From Not publicly disclosed
Directly competes with Claude Mythos Pre...Lowered refusal boundary on defensive-se...
Updated 2026-04-19
Microsoft MAI-Thinking-1 logo

Microsoft MAI-Thinking-1

Microsoft's first in-house reasoning model -- launched 2026-06-02 at Build as the flagship of seven new MAI models. 35B-active / ~1T-total sparse Mixture-of-Experts, 256K context. AIME 2025 97.0%, matches leading models on SWE-Bench Pro, and beat Claude Sonnet 4.6 in human-preference testing. Available on Microsoft Foundry + OpenRouter / Fireworks / Baseten

B
7.5/10
From Not disclosed
Microsoft's first in-house frontier-clas...Strong published reasoning numbers: AIME...
Updated 2026-06-02
Hunyuan 3 (Tencent Hy3) logo

Hunyuan 3 (Tencent Hy3)

Tencent's Hy3 reached GA 2026-07-06 (upgraded from the April preview) -- 295B total / 21B active MoE, 256K context, now Apache 2.0 open weights on HuggingFace + ModelScope with the EU/UK/South Korea restriction lifted. ~90% agent-task completion on Tencent's internal apps; API via Tencent Cloud TokenHub. Integrated into Yuanbao, WeChat, QQ

A
8.1/10
Free tierFrom $0
Open weights from a top-3 Chinese tech c...Pricing is aggressive. ~1.2 RMB per mill...
Updated 2026-07-22
MiMo (Xiaomi) logo

MiMo (Xiaomi)

Xiaomi's MiMo-V2.5 family launched 2026-04-22 -- Pro (1T total / 42B active MoE, 1M context, native vision+audio reasoning), Multimodal base, TTS (3 sub-models: base, VoiceDesign, VoiceClone), and ASR (open-source, English + Chinese + major dialects). Full voice pipeline for the agent era. Extra-charge 1M-context tier removed at launch

A
8.3/10
Free tierFrom $0
Full voice pipeline shipped together: a ...Native multimodal in MiMo-V2.5-Pro is th...
Updated 2026-07-04