Mistral AI logo
B

Mistral AI

B Tier · 7.5/10

**Mistral now powers Firefox Smart Window (2026-09-16)** -- Mozilla's beta AI browsing assistant runs on Mistral models for France and North America, UK and Germany later this year, with zero data retention. **Mistral raised EUR 3B in a Series D at a post-money valuation above EUR 21B (2026-09-08)** -- led by Samsung Electronics with EQT's Scaleup Europe Fund and PSG Equity, the largest equity round ever by a European tech company per Mistral. European AI lab with open and commercial models -- Le Chat is now **Vibe** (May 28 2026): one agent across Work Mode + Code Mode with a VS Code extension and CLI, powered by Mistral Medium 3.5 (128B dense, 256k context, 77.6% SWE-Bench Verified). Newest release: **Shieldstral 1.0** (Aug 4 2026), a 3B Apache 2.0 multimodal safety classifier that runs on one 16GB GPU. Earlier 2026 line: Small 4 (119B MoE Apache 2.0), Voxtral TTS. **Mistral Medium 3.1 and Medium 3 were both retired 2026-08-31** -- Medium 3.5 is the vendor-listed replacement for each

Last updated: 2026-09-17Free tier available

Score Breakdown

6.0
Ease of Use
8.0
Output Quality
9.0
Value
7.0
Features

Benchmark Scores

Benchmarks for Mistral Medium 3.5 (vendor-published; third-party verification pending)

BenchmarkScore
MMLU86%
HumanEval92%
MATH69%
SWE-Bench Verified77.6%

Last updated: 2026-04-29

Personality & Tone

The European pragmatist

Tone: Efficient, terse, and slightly blunt. Mistral answers in fewer words than Claude or ChatGPT, especially on factual questions, and rarely hedges or softens its take.

Quirks: Trained with less Anglocentric data than Llama, so it handles French, German, and Spanish notably better than US-origin models. Refusal rates are lower than ChatGPT or Gemini on most gray-area prompts.

The Good and the Bad

What we like

  • +Mistral Medium 3.5 (April 29 2026) is Mistral's first 'flagship merged' model -- 128B dense, 256k context, 77.6% on SWE-Bench Verified, in public preview at $1.5/$7.5 per million tokens. Closes most of the coding-benchmark gap to Claude Opus / GPT-5.5 at materially lower API cost
  • +Vibe Remote Agents (also 4/29) lets you launch cloud-based coding sessions that run asynchronously and in parallel via CLI or Le Chat -- file diffs, tool calls, and the ability to teleport a local session to the cloud while preserving history and approval state. Unique in the category as of today
  • +Le Chat Work Mode (4/29) is the first agentic mode shipped at the consumer-chat tier -- multi-step task completion, cross-tool workflows, research synthesis, inbox triage, with explicit approval gates for sensitive operations
  • +Mistral Small 4 (March 2026) unifies the previously-split Small/Magistral/Pixtral/Devstral lines into one 119B MoE Apache-2.0 model. Voxtral TTS (March 2026) fills the speech gap with a competent open-source 4B-param model that runs on consumer hardware
  • +Extremely competitive API pricing remains the moat -- Small 4 at $0.20/1M tokens, Medium 3.5 at $1.5/$7.5 per million tokens, against frontier-class quality

What could be better

  • −Le Chat web interface is bare-bones compared to ChatGPT or Claude
  • −Smaller ecosystem -- fewer integrations and community resources
  • −Less brand recognition means less community help when you get stuck
  • −Documentation could be better, especially for newer models

Pricing

Le Chat (Free)

$0
  • ✓Web chat interface with Mistral models
  • ✓Mistral Small 4 + Medium 3.5 available (Medium 3 retired 2026-08-31)
  • ✓Basic features, limited rate

API (Mistral Small 4)

$0.20/per 1M tokens
  • ✓119B MoE, Apache 2.0 open-weight
  • ✓Unifies Small/Magistral/Pixtral/Devstral into one model
  • ✓Fast, efficient, 128K context

API (Mistral Medium 3.5)

$1.5 / $7.5/per 1M tokens (input/output)
  • ✓Public preview SHIPPED 2026-04-29 -- Mistral's first 'flagship merged' model
  • ✓128B dense, 256k context, 77.6% SWE-Bench Verified
  • ✓Underlies new Vibe Remote Agents + Le Chat Work Mode

API (Mistral Medium 3 -- RETIRED 2026-08-31)

$1/per 1M tokens -- no longer available
  • ✓Launched April 9, 2026
  • ✓EU AI Act compliance metadata
  • ✓**RETIRED 2026-08-31 alongside Medium 3.1** -- `mistral-medium-2505` was deprecated 2026-05-22 and reached retirement today; the vendor-listed replacement is **Mistral Medium 3.5**

API (Mistral Large 3)

$2/per 1M tokens
  • ✓Flagship sparse MoE
  • ✓256K context
  • ✓MRL license (paid for commercial self-hosting)

Voxtral TTS

$0
  • ✓4B-param open-source speech model, March 2026
  • ✓9 languages, runs on consumer hardware
  • ✓Apache 2.0

System Requirements

Hardware needed to self-host. Min = smallest viable setup (usually heavy quantization). Max = full-precision / production-grade.

Model variantMinMax
Mistral Small 3 / Devstral 2 (24B dense, Apache 2.0)10 GB VRAM (Q4)1× A100 40 GB FP16
Mistral 14B / 8B / 3B (Apache 2.0)6 / 4 / 2 GB VRAM (Q4)24 / 16 / 8 GB VRAM FP16
Mixtral 8x22B (legacy)64 GB RAM + 24 GB GPU (Q3)2× A100 80 GB FP16
Mistral Large 3 (flagship)Not self-hostable under free terms -- MRL licenseRequires paid commercial license to self-host

Known Issues

  • MISTRAL POWERS FIREFOX SMART WINDOW -- FIRST CONSUMER-BROWSER DISTRIBUTION FOR MISTRAL MODELS, WITH ZERO DATA RETENTION (2026-09-16, vendor-primary): Mistral and Mozilla announced a partnership under which **Firefox Smart Window (beta), Mozilla's AI browsing assistant, 'is now powered by Mistral models'**. Smart Window summarises complex searches, recalls things you clicked away from and sources information from your open tabs. **Regions: France and North America now; the United Kingdom and Germany 'expected to follow later this year'.** Privacy terms as stated: conversations 'aren't saved on Mozilla's servers by default, and partners like Mistral agree to zero data retention'. Mistral frames it as 'two open source advocates' pairing open weights with open distribution, and says the models are fine-tuned on regional languages and dialects so the browser assistant 'feels native'. **What it changes for this page:** Mistral's consumer reach has historically been Vibe (ex-Le Chat) only; Firefox is the first third-party consumer surface at scale, and the zero-retention clause is a concrete privacy commitment worth quoting against Chrome's Gemini integration. Not stated: which Mistral model, whether inference is local or cloud, or any commercial terms.Source: Mistral (mistral.ai/news/mistral-x-mozilla/, RSS pubDate 'Wed, 16 Sep 2026 12:00:00 GMT', on-page 'September 16, 2026', by-line 'Mistral and Mozilla') -- fetched 2026-09-17 via curl with browser UA · 2026-09-16
  • EUR 3 BILLION SERIES D AT MORE THAN EUR 21 BILLION POST-MONEY, LED BY SAMSUNG -- THE LARGEST EUROPEAN TECH EQUITY ROUND ON RECORD, PER MISTRAL (2026-09-08, vendor-primary): Mistral announced 'it has raised **EUR 3 billion in a Series D funding round at a post-money valuation of more than EUR 21 billion**, the largest equity fundraising round ever completed by a European technology company, three years after the company's launch.' **Samsung Electronics led**, with co-leads **Scaleup Europe Fund (managed by EQT)** and existing investor **PSG Equity**. **Stated use of funds:** expanding frontier research, scaling compute for training, building out infrastructure (the European sovereign-compute programme announced 8/11), and commercial and international growth; Mistral says it now operates across **20 countries** and supports **125+ global enterprises** including Airbus, ASML and HSBC. **WHY IT MATTERS ON THIS PAGE:** this roughly doubles the valuation from the September 2025 round and puts a strategic hardware investor at the top of the cap table -- Samsung is a memory and foundry supplier to every frontier lab, so the tie-up reads as a compute-and-distribution partnership as much as a financing. It also funds the open-weight strategy explicitly ('sovereign, open-weight AI'), which is the commercial bet that distinguishes Mistral from the closed US labs it competes with on Vibe and the API. **Same window, minor:** a **Cloudera partnership** (9/10) to run Mistral models inside customers' Cloudera environments for regulated industries, and a solutions post (9/09) on migrating 40,000 lines of Fortran 77 to C++ with agents for a European energy operator. Neither changes pricing or product availability. **No model releases from Mistral this window** -- newest model post remains Shieldstral (8/04).Source: Mistral (mistral.ai/news/mistral-makes-sovereign-open-weight-ai-to-frontier/, RSS pubDate 'Tue, 08 Sep 2026 12:00:22 GMT'); Mistral (mistral.ai/news/mistral-x-cloudera/, RSS pubDate 'Thu, 10 Sep 2026 10:42:55 GMT') -- both fetched 2026-09-14 · 2026-09-08
  • TWO MISTRAL MEDIUM GENERATIONS RETIRE TODAY -- MEDIUM 3.1 AND MEDIUM 3 BOTH HIT END-OF-LIFE 2026-08-31 (verified on Mistral's model lifecycle table): both models now sit in the **'Deprecated & retired models'** section of Mistral's models overview with a retirement date of **8/31/2026**. Specifically: **Mistral Medium 3.1** (`mistral-medium-2508`, v25.08) -- deprecated **5/22/2026**, retired **8/31/2026**; and **Mistral Medium 3** (`mistral-medium-2505`, v25.05) -- deprecated **5/22/2026**, retired **8/31/2026**. **The vendor-listed alternative for both is Mistral Medium 3.5.** **Why this matters more than a routine EOL:** Medium is Mistral's volume commercial tier, and this retires **two consecutive generations on the same day**, roughly 100 days after a single joint deprecation notice. If you pinned a model string rather than tracking an alias, a `mistral-medium-2508` or `mistral-medium-2505` call breaks today, and the migration target is a different model with a different price -- **Medium 3.5 is $1.5/$7.5 per 1M**, against the **$1 per 1M** this page recorded for Medium 3. **So for anyone still on Medium 3 the retirement is also a price increase, which the lifecycle table does not tell you.** Note Mistral's deprecation-to-retirement gap here was about **three months**, shorter than the twelve months OpenAI gave the Assistants API -- worth knowing when you plan around Mistral model lifetimes generally. Same-day cohort from the same 5/22 notice: **Devstral 2**, **Magistral Medium 1.2**, **Magistral Small 1.2** and **Mistral Nemo 12B** were all retired on 7/31, so this 8/31 pair is the tail of that wave rather than a new announcement.Source: Mistral (docs.mistral.ai/getting-started/models/models_overview/, 'Deprecated & retired models' table -- rows 'Mistral Medium 3.1 | 25.08 | mistral-medium-2508 | 5/22/2026 | 8/31/2026 | Mistral Medium 3.5' and 'Mistral Medium 3 | 25.05 | mistral-medium-2505 | 5/22/2026 | 8/31/2026 | Mistral Medium 3.5') -- fetched 2026-08-31 via curl · 2026-08-31
  • AGENTIC SEARCH -- MISTRAL'S RETRIEVAL LAYER, WITH UNUSUALLY LARGE VENDOR BENCHMARK DELTAS (2026-08-20, vendor-primary): Mistral shipped **Agentic Search**, a multi-step retrieval loop positioned explicitly against one-shot RAG. Rather than handing a model a fixed set of retrieved chunks, it **navigates documents using five tools -- search, open, navigate, read and grep** -- and builds on your existing search index rather than requiring a new one. **AVAILABILITY:** through the **Mistral Search Toolkit**, and built into **Libraries in both Studio and Vibe**. Mistral stresses the deployment story for regulated buyers: 'portable and open tooling helps you unlock value from your data **without crossing your isolation boundaries** in the cloud or on-premises.' **THE VENDOR NUMBERS, QUOTED AND LABELLED AS VENDOR-MEASURED -- NO THIRD-PARTY VERIFICATION EXISTS YET:** on **FinanceBench**, correctness on financial filings goes **from 26.7% to 86%** (Mistral describes this as 'up to 3x correctness'); on **OfficeQA Pro**, table-heavy multi-document questions gain **+45.6 points, from 6.3% to 51.9%**; **p90 latency falls up to 39.6%** and token consumption drops '**by up to one-third**' from fewer repeated searches. **HOW TO READ THOSE DELTAS HONESTLY:** a 26.7% baseline is a one-shot RAG configuration Mistral chose, and the size of the gain is a statement about how badly chunk-based RAG performs on dense financial filings as much as about Agentic Search itself. The efficiency claims (latency, tokens) are the more transferable ones because they are architectural -- targeted navigation genuinely issues fewer retrieval calls. **NO PRICING IS PUBLISHED IN THE POST** and it does not state whether Agentic Search is metered separately from Search Toolkit usage; treat cost as unknown until Mistral's pricing page reflects it. **STRATEGIC READ:** this extends the 5/28 Search Toolkit into the agentic layer and pairs with the 8/11 in-region inference push -- Mistral is assembling a sovereignty-first enterprise stack where the differentiator is that the retrieval runs inside your boundary, not that the model is the smartest availableSource: Mistral AI (mistral.ai/news/agentic-search/, RSS pubDate 'Thu, 20 Aug 2026 12:00:17 GMT', fetched 2026-08-20 via curl) · 2026-08-20
  • REGIONAL ENDPOINTS GO GA, A PRIORITY TIER ENTERS PREVIEW, AND MISTRAL STARTS HOSTING SOMEONE ELSE'S OPEN MODEL (2026-08-11, vendor post): three changes, two of them concrete product news. **(1) Mistral Regional Endpoints are now generally available** -- customers 'choose whether their inference runs in Europe or the US', aligning inference location with data-residency, regulatory and latency requirements. Read the carve-out Mistral itself states rather than the headline: processing happens in the selected region **'subject to limited, safeguarded transfers to sub-processors that may occur outside that region'** -- so this is regional control, not an absolute guarantee that no byte leaves. **(2) Mistral Priority Tier, public preview** -- committed service levels for mission-critical workloads, custom rate limits, backed by an **uptime SLA**. Mistral's claim: it is 'the only European AI lab to offer both' region choice and a committed SLA-backed service level. No pricing published for either. **(3) The genuinely surprising one: Mistral's platform will serve third-party open models, starting with Z.ai's GLM-5.2**, running on 'the same infrastructure, regional controls, and service commitments as Mistral models.' A frontier lab reselling a Chinese lab's open weights under its own sovereignty guarantees is a real strategic shift -- it concedes that customers want model choice more than they want Mistral-only, and it turns Mistral's European infrastructure into the product rather than the models alone. **USEFUL CROSS-CHECK: this vendor page names GLM-5.2 as the current Z.ai model**, which independently corroborates our standing decision not to ship the unsourced 'GLM-5.5' claim that has circulated on aggregators. **(4) Company/roadmap, not a product:** Mistral is assembling an anchor group of enterprises (**ASML, Amadeus** among the quoted participants) whose multi-year commitments fund European infrastructure, sold as **European Compute Units (ECUs)**, targeting **up to 1 GW of capacity by 2030**. That is a financing structure and a 2030 ambition -- no capacity exists today on the strength of this post, and it should not be read as shipped infrastructureSource: Mistral AI (mistral.ai/news/regional-inference-open-models-new-compute/, RSS pubDate Tue, 11 Aug 2026 12:00:27 GMT, on-page date August 11, 2026, fetched 2026-08-13) · 2026-08-11
  • SHIELDSTRAL RELEASED (2026-08-04, vendor-primary): Mistral shipped **Shieldstral 1.0**, a **3B-parameter, Apache 2.0 open-weights, policy-adaptive multimodal safety classifier** -- a guard model that screens prompts, responses, and **images** for harmful content. The design point is that it treats moderation as question-answering: you supply your **safety policy in plain language at inference time**, with **no retraining or fine-tuning**, and it returns a **calibrated yes/no probability from a single forward pass**. Mistral's claim is that it **matches or beats guard models up to 7x its size** on text safety, refusal detection, policy adaptability, and multimodal safety. Practical appeal: it runs on **a single 16GB NVIDIA GPU**, so self-hosters and small teams can put a real moderation layer in front of an open model without renting a second big box. Weights are on Hugging Face as **mistralai/Shieldstral-1.0-3B**; no pricing (open weights, free to download). Caveats worth knowing before you deploy it: **all comparative benchmark numbers are vendor-published and not yet third-party verified**, and **multilingual coverage is listed as future work** -- Mistral did not claim non-English safety performance at launch, which is a notable gap for a lab whose main differentiator is multilingual strength. Context: guard models are becoming table stakes as the **EU AI Act's Article 50 transparency duties went enforceable 2026-08-02**, and an EU-hosted, open-weights, self-deployable classifier is a pointed answer to US-hosted moderation APIsSource: Mistral AI (mistral.ai/news/shieldstral/), Mistral AI news RSS (pubDate 2026-08-04), Hugging Face (mistralai/Shieldstral-1.0-3B) · 2026-08-04
  • MICROSOFT PARTNERSHIP EXPANDED -- MULTIBILLION-DOLLAR DEAL (2026-07-21, Microsoft newsroom): Microsoft and Mistral announced a major expansion of their strategic partnership aimed at enterprises and regulated industries (finance, healthcare, manufacturing). Terms: **thousands of NVIDIA Vera Rubin GPUs** allocated for EU-based compute, and **Mistral Medium 3.5 + Mistral OCR 4 now available in Microsoft Foundry and Copilot Studio**. The pitch is control/sovereignty -- cloud, Azure Local, and fully air-gapped/disconnected deployment so regulated customers can run frontier Mistral models on their own terms. Strategically this deepens Mistral's distribution on Azure (Microsoft is also a Mistral investor) and gives European enterprises a non-US-lab frontier option inside the Microsoft stack. Not a new base model -- an availability + partnership expansionSource: Microsoft (news.microsoft.com/source/2026/07/21/microsoft-and-mistral-expand-strategic-partnership-to-give-enterprises-and-regulated-industries-frontier-ai-they-can-control/) · 2026-07-21
  • ROBOSTRAL NAVIGATE (2026-07-08, vendor-primary): Mistral's first **embodied-AI navigation model** -- an 8B-param model, built in-house and trained entirely in simulation (400K trajectories across 6K simulated environments, RL via CISPO), that guides **wheeled, legged, and flying robots** using just a single RGB camera + a plain-language instruction (no LiDAR/depth sensors). Vendor benchmarks: **R2R-CE 79.4% success (seen) / 76.6% (unseen)** -- +9.7 pts over the best single-camera approach and +4.5 over the best depth/multi-camera system. Caveats: all results are simulation-only, the pointing-based approach can't handle targets outside the camera's field of view, and **no weights, API, or license were published** -- access is 'talk with our team.' A research/enterprise play, not a product you can use today; notable as Mistral's entry into robotics. Same week (7/9): **Prompt & Skills Management** shipped in Mistral Studio -- a versioned system-of-record for prompts and skillsSource: Mistral AI (mistral.ai/news/robostral-navigate/), Mistral AI news (Studio prompt management, 2026-07-09) · 2026-07-08
  • LEANSTRAL 1.5 RELEASED (2026-07-02, hit #1 on Hacker News 7/4): a formal-verification / Lean 4 theorem-proving model -- 119B total / 6B active params, Apache 2.0 open weights (mistralai/Leanstral-1.5-119B-A6B on Hugging Face) plus a FREE API endpoint (leanstral-1-5). Vendor-reported results: saturates miniF2F (100%), 587/672 on PutnamBench, 87% FATE-H / 34% FATE-X. Niche (math/proof engineering) but notable as a genuinely open frontier release in a specialty domain where closed labs dominateSource: Mistral AI blog (mistral.ai/news/leanstral-1-5) · 2026-07-02
  • REBRAND (2026-05-28): **Le Chat is now 'Vibe'** -- Mistral merged its consumer chat product into a single agent brand spanning Work Mode and Code Mode, with a new VS Code extension and CLI. Mistral Medium 3.5 (public preview since 4/29, broader rollout 5/22) is the default model powering Vibe's remote coding agents. Adjacent late-May moves: Emmi AI (physics/industrial simulation, via acquisition) added to the enterprise platform and a new Search Toolkit (5/28). If you bookmarked chat.mistral.ai as 'Le Chat,' it's the same product under the new name -- pricing tiers unchangedSource: Mistral AI blog (mistral.ai/news/vibe-agent, mistral.ai/news/vibe-remote-agents-mistral-medium-3-5) · 2026-05-28
  • ENTERPRISE PRODUCT (2026-04-28 public preview): Mistral Workflows -- a Temporal-powered durable orchestration engine for AI workloads. Built on the same Temporal core that backs Netflix / Stripe / Salesforce, with Mistral-added streaming, payload handling, multi-tenancy, and observability. Python SDK v3.0, Helm-deployable workers, customer-perimeter data residency. Human-in-the-loop approvals via simple Python (wait_for_input()), full execution tracking in Studio, deploys cloud / on-prem / hybrid. Distinct from Vibe Remote Agents (the consumer-facing async coding sessions); Workflows is the enterprise infra layer that makes them and other AI workloads durable at scale. Live customers cited at preview: ASML, ABANCA, CMA-CGM, France Travail, La Banque Postale, Moeve. Pricing during preview not disclosedSource: Mistral AI blog (mistral.ai/news/workflows) · 2026-04-28
  • Mistral Medium 3.5 SHIPPED 2026-04-29 in public preview, accompanied by two net-new agentic offerings: Vibe Remote Agents (cloud-based coding sessions, async + parallel, CLI or Le Chat entry) and Le Chat Work Mode (agentic chat for multi-step tasks across tools). The model is 128B dense, 256k context, and posts 77.6% on SWE-Bench Verified. Pricing is $1.5/$7.5 per million tokens (input/output). 'Flagship merged' framing means Medium 3.5 supersedes Medium 3 for new workloads -- existing Medium 3 deployments continue to workSource: Mistral AI blog (mistral.ai/news/vibe-remote-agents-mistral-medium-3-5) · 2026-04-29
  • Le Chat occasionally slower than competitors during European business hoursSource: Reddit r/MistralAI · 2026-03
  • Voxtral TTS English output is competent but trails ElevenLabs v3 on expressiveness -- it's positioned as an open-source alternative, not a quality leaderSource: TechCrunch Voxtral coverage · 2026-03

Best for

Developers who want cheap, high-quality API access. Also strong for multilingual applications and European companies that prefer an EU-based AI provider for data residency.

Not for

Non-technical users looking for a polished chat experience. ChatGPT and Claude are much better as consumer products.

Our Verdict

Mistral is the scrappy underdog that keeps surprising people. Their models are impressively efficient -- you get near-GPT-4 quality at a fraction of the API cost. But the consumer experience (Le Chat) is rough. This is primarily a developer's tool. If you're building AI applications on a budget, Mistral should be on your shortlist.

Sources

  • Mistral: Mistral and Mozilla are bringing open, private and multilingual AI to your web browser -- Firefox Smart Window (2026-09-16) (accessed 2026-09-17)
  • Mistral: Mistral raises EUR 3B to make sovereign, open-weight AI the technology frontier -- Series D, >EUR 21B post-money, Samsung-led (2026-09-08) (accessed 2026-09-14)
  • Mistral: Cloudera and Mistral partner for sovereign enterprise AI (2026-09-10) (accessed 2026-09-14)
  • Mistral docs: models overview + deprecated/retired model lifecycle table -- Medium 3.1 and Medium 3 both retired 2026-08-31, replacement Medium 3.5 (verified 2026-08-31) (accessed 2026-08-31)
  • Mistral AI: Introducing Agentic Search -- FinanceBench 26.7% to 86%, OfficeQA Pro +45.6pt (2026-08-20) (accessed 2026-08-20)
  • Mistral AI: In-region inference, open models, and new European infrastructure for sovereign AI -- Regional Endpoints GA, Priority Tier preview, GLM-5.2 hosting, ECUs (2026-08-11) (accessed 2026-08-13)
  • Mistral AI: Introducing Shieldstral (2026-08-04) (accessed 2026-08-04)
  • Hugging Face: mistralai/Shieldstral-1.0-3B (accessed 2026-08-04)
  • Microsoft: Microsoft and Mistral expand strategic partnership (2026-07-21) (accessed 2026-07-22)
  • Mistral AI: Leanstral 1.5 (2026-07-02) (accessed 2026-07-05)
  • Mistral AI: Workflows public preview (2026-04-28) (accessed 2026-05-04)
  • Mistral AI: Vibe Remote Agents + Mistral Medium 3.5 (2026-04-29) (accessed 2026-04-30)
  • Mistral AI official site (accessed 2026-04-30)
  • TechCrunch: Mistral releases Voxtral TTS (accessed 2026-04-16)
  • SiliconANGLE: hardware-efficient language models (accessed 2026-04-16)
  • LMSYS Chatbot Arena rankings (accessed 2026-04-16)
  • API testing (accessed 2026-04-16)

The Tier List Tuesday

Weekly newsletter: tier movers, new entrants, and the VS of the week. Built from our daily AI-tool sweeps. No spam, unsubscribe anytime.

Alternatives to Mistral AI

Llama 4 (Meta) logo

Llama 4 (Meta)

Meta's open-weights family -- Scout (10M context), Maverick (multimodal 400B MoE). NOTE: Meta's frontier work moved to the proprietary Muse Spark line in April 2026; Llama remains downloadable and supported but is effectively in maintenance mode

B
7.9/10
Free tierFrom $0
Llama 4 Scout has a 10M token context wi...Llama 4 Maverick is natively multimodal ...
Updated 2026-06-09
DeepSeek logo

DeepSeek

**DeepSeek-V4.1-Flash shipped 2026-09-10**: a 552B MoE on a new Causal Encoder-Decoder architecture (8B active for input, 16B for output), natively multimodal, MIT-licensed on HF, and priced BELOW the V4-Flash it replaces -- $0.15/$0.60 off-peak, $0.30/$1.20 peak per 1M, with cache-hit input cut to $0.003/$0.006. V4-Flash and V4-Flash-Vision-Exp are retired (aliases route to V4.1-Flash). DeepSeek said V4-Pro would be phased out on 9/14, then **reversed that four days later** -- V4-Pro stays on the API at unchanged $0.66/$1.98 off-peak, $1.32/$3.96 peak until V4.1-Pro lands. V4.1-Flash's vendor benchmarks beat V4-Pro (DeepSWE 74.2, Terminal-Bench 2.1 90.6, GPQA 90.9)

A
8.0/10
Free tierFrom $0
Pricing is absurdly cheap compared to GP...DeepSeek-R1 reasoning model genuinely co...
Updated 2026-09-14
Gemma 4 (Google) logo

Gemma 4 (Google)

Google DeepMind's open-weights model family -- multimodal, 256K context, runs on edge devices

A
8.3/10
Free tierFrom $0
Apache 2.0 license -- truly permissive, ...Multimodal: handles text + image input (...
Updated 2026-04-19
DiffusionGemma (Google) logo

DiffusionGemma (Google)

Google DeepMind's experimental open-weights TEXT-DIFFUSION model (June 10, 2026) -- 26B MoE (3.8B active), Apache 2.0, generates 256-token blocks in parallel with bidirectional attention for up to 4x faster output (1,000+ tok/s on H100). Trades some quality vs Gemma 4 for raw speed

C
6.8/10
Free tierFrom $0
Text diffusion instead of autoregression...First open-weights text-diffusion model ...
Updated 2026-06-10
Qwen (Alibaba) logo

Qwen (Alibaba)

**Qwen-Image-2.1 went open-weight (weights 2026-09-14, blog 2026-09-20)** -- a unified text-to-image and editing model with a 7B visual-generation component, native transparent (RGBA) output, up to 10 reference images and mask-guided local edits -- but under a **research-only licence, not Apache**: commercial use needs a separate licence from Alibaba. Alibaba's open-weights + API family, and August 2026 was its biggest open-release month yet: **Qwen3.8-2.4T-A95B -- the Max-class flagship -- went open-weight on 2026-08-08** (custom Qwen3.8-Max licence), **Qwen3.8-27B** dense landed Apache 2.0 (8/05), and **Qwen3.8-Flash-Next** (8/24 weights, 8/26 blog) previews the Qwen4 architecture: 125B main + 51B n-gram embeddings with only 6B active, served as Qwen3.8-Flash at $0.15/$0.47 per 1M. Qwen 3.7 Max remains the GA API flagship ($2.50/$7.50)

A
8.8/10
Free tierFrom $0
Qwen 3.6-Plus (launched Mar 30 2026) is ...Qwen3.5 Small (0.8B / 2B / 4B / 9B) is t...
Updated 2026-09-21
GLM / Z.ai (Zhipu AI) logo

GLM / Z.ai (Zhipu AI)

**GLM-5.3 and GLM-5.3-Flash released 2026-08-25** (weights on Hugging Face): 5.3 is a post-training-only upgrade of the 5.2 base that Z.ai calls 'the most capable open-weights model for coding' (Terminal Bench 2.1 88.2, DeepSWE 66.9, CyberGym 84.5), under a custom GLM-5.3 licence at $1.40/$4.40 per 1M; **GLM-5.3-Flash** is a new 320B/18B-active natively multimodal model under MIT at $0.15/$0.50. Zhipu AI's open-weights flagship -- GLM-5.2 (launched 2026-06-13) is a ~753B-parameter MoE with a 1M-token context and the new IndexShare sparse-attention architecture (~2.9x lower per-token FLOPs at 1M context), MIT licensed. Vendor benchmarks put SWE-Bench Pro at 62.1 (up from GLM-5.1's 58.4) and it tops the Artificial Analysis open-weights Intelligence Index; VentureBeat reports it beats GPT-5.5 on several long-horizon coding benchmarks at roughly 1/6 the cost. Drop-in for Claude Code / Cline / OpenCode. Still trained outside the Nvidia stack on Huawei Ascend silicon

A
8.0/10
Free tierFrom $0
GLM-5.1 (2026-04-07) topped SWE-Bench Pr...First frontier model trained entirely on...
Updated 2026-09-14
Kimi K3 (Moonshot) logo

Kimi K3 (Moonshot)

Moonshot's 2.8T-parameter Kimi K3 (launched 2026-07-16/17) is the largest open-weight model ever released -- 1M context, multimodal, $3/$15 per 1M via API, ranked best-available on Arena.AI at launch. WEIGHTS SHIPPED ~2026-07-26/27 on Hugging Face (2.8T total / 104B activated, safetensors) under a custom Kimi K3 License, not the Modified MIT of the K2 line

A
8.1/10
Free tierFrom $3 / $15
Frontier-tier performance -- Elo 1309 on...Beats Claude Opus 4.5 on several coding ...
Updated 2026-09-05
Nemotron (Nvidia) logo

Nemotron (Nvidia)

Nvidia's open-weights family -- hybrid Mamba-Transformer MoE architecture, optimized for efficient reasoning on Nvidia hardware. Nemotron 3 Ultra (550B total / 55B active) shipped 2026-06-04 as the family flagship, joining Super (120B/12B, March) and Nano

B
7.8/10
Free tierFrom $0
Hybrid Mamba-Transformer architecture dr...Nemotron 3 Super activates only 12B of 1...
Updated 2026-09-07
MiniMax M3 logo

MiniMax M3

**MiniMax-H3's weights are public on Hugging Face (repo 2026-07-28, updated 2026-08-13) under a community licence whose open-weight grant is limited to the US, EU, UK and South Korea**, and **MiniMax Music 3 (2026-08-07)** ships open weights for five-minute full-song generation with a prominent-attribution commercial clause. MiniMax's coding/agent flagship -- M3 (June 1 2026): 1M-token context, MSA sparse attention (>15x decoding speedup at long context), SWE-Bench Pro 59.0%, Terminal-Bench 66.0%. OPEN WEIGHTS LIVE on HuggingFace since June 12 (~428B total / ~23B active, native multimodal, minimax-community license)

A
8.4/10
Free tierFrom $0
229B/10B-active MoE delivers Tier-1 agen...Sparse MoE design: ~10B active params du...
Updated 2026-09-17
Falcon (TII) logo

Falcon (TII)

UAE's Technology Innovation Institute open-weights family -- Falcon 3 optimized for efficient sub-10B deployment on consumer hardware

B
7.1/10
Free tierFrom $0
Apache 2.0 license -- fully permissive f...Sub-10B sizes run on any consumer GPU or...
Updated 2026-04-13
gpt-oss (OpenAI) logo

gpt-oss (OpenAI)

OpenAI's FIRST open-weight models -- gpt-oss-120b (single 80GB GPU, near parity with o4-mini on reasoning) and gpt-oss-20b (runs on 16GB edge devices). Apache 2.0. Launched 2025-08-05. gpt-oss-safeguard ships in 2026 as the safety-tuned variant

A
8.1/10
Free tierFrom $0
First-ever OpenAI open-weight release --...gpt-oss-120b approaches o4-mini on reaso...
Updated 2026-04-17
IBM Granite 4.0 logo

IBM Granite 4.0

IBM's enterprise-focused open-weight family -- Granite 4.0 hybrid Mamba-2 + transformer architecture (70-80% memory reduction vs pure transformer), 3B to 32B sizes, Apache 2.0. First open model family to secure ISO 42001 certification. Nano 350M runs on CPU with 8-16GB RAM. 3B Vision variant landed 2026-04-01

A
8.2/10
Free tierFrom $0
Hybrid Mamba-2 + transformer architectur...Granite 4.0 Nano (350M and 1.5B) is genu...
Updated 2026-04-17
Arcee Trinity-Large-Thinking logo

Arcee Trinity-Large-Thinking

Arcee AI's US-made open-weight frontier reasoning model -- launched 2026-04-01. 398B total params, ~13B active. Sparse MoE (256 experts, 4 active = 1.56% routing). Apache 2.0, trained from scratch. #2 on PinchBench trailing only Claude 3.5 Opus. ~96% cheaper than Opus-4.6 on agentic tasks

A
8.1/10
Free tierFrom $0
Rare US-made frontier-tier open-weight r...Trained from scratch (not a fine-tune) a...
Updated 2026-04-17
Olmo 3 (AI2) logo

Olmo 3 (AI2)

Allen Institute for AI's fully-open frontier reasoning models -- Olmo 3 family (2025-11-20) includes 7B and 32B sizes, four variants (Base, Think, Instruct, RLZero). Apache 2.0 with fully open data + checkpoints + training logs. Olmo 3-Think 32B matches Qwen3-32B-Thinking at 6x fewer training tokens

B
7.9/10
Free tierFrom $0
FULLY OPEN is a different category than ...Olmo 3-Think 32B matches Qwen3-32B-Think...
Updated 2026-04-17
AI21 Jamba2 logo

AI21 Jamba2

AI21 Labs' hybrid SSM-Transformer (Mamba-style) open-weight family -- Jamba2 launched 2026-01-08. Two sizes: 3B dense (runs on phones / laptops) and Jamba2 Mini MoE (12B active / 52B total). Apache 2.0, 256K context, mid-trained on 500B tokens

A
8.0/10
Free tierFrom $0
Hybrid SSM-Transformer (Mamba-style) arc...Jamba2 3B dense runs realistically on iP...
Updated 2026-04-17
StepFun Step 3.7 Flash logo

StepFun Step 3.7 Flash

StepFun's (China) agent-focused open-weight family -- Step 3.7 Flash (May 28 2026): 198B sparse MoE vision-language model, ~11B active, 256K context, Apache 2.0, ~400 tok/s, SWE-Bench Pro 56.3. Supersedes Step 3.5 Flash (Feb 2026) as the flagship

B
7.8/10
Free tierFrom $0
Step 3.5 Flash at 196B total / 11B activ...Agent-focused tuning explicitly -- tool ...
Updated 2026-06-10
Cohere Command A logo

Cohere Command A

Cohere's enterprise-multilingual flagship -- 111B params, 256K context, runs on 2x H100. 23 languages. CC-BY-NC 4.0 on weights (research / non-commercial), commercial requires Cohere enterprise contract. Follow-ups: Command A Reasoning + Command A Vision

B
7.5/10
Free tierFrom $0
Best-in-class multilingual open-weight m...Runs on just 2x H100 at FP16 for the ful...
Updated 2026-04-17
LongCat-2.0 (Meituan) logo

LongCat-2.0 (Meituan)

Meituan's open-source 1.6T-parameter MoE (~48B active) with native 1M-token context, MIT license -- trained entirely on domestic Chinese AI ASICs and revealed as the stealth 'Owl Alpha' model that had been topping OpenRouter

B
7.9/10
Free tierFrom $0
Genuinely open frontier scale: 1.6T tota...Native long context: LongCat Sparse Atte...
Updated 2026-07-05
Inkling (Thinking Machines Lab) logo

Inkling (Thinking Machines Lab)

Mira Murati's $12B lab ships its first model (2026-07-15): a 975B/41B-active open-weights MoE that reasons natively over text, images, and audio with a 1M-token context -- positioned not as the strongest model, but as the best starting point for fine-tuning via Tinker

A
8.0/10
Free tierFrom $0
First US frontier-adjacent open-weights ...Natively multimodal INPUT across text, i...
Updated 2026-07-18
Bonsai 27B (PrismML) logo

Bonsai 27B (PrismML)

The first 27B-class model that runs on a phone (2026-07-14) -- ternary and 1-bit quantizations of Qwen3.6 27B squeeze a multimodal, tool-calling, 262K-context model into 3.9-5.9GB under Apache 2.0

B
7.9/10
Free tierFrom $0
Genuine first: a 27B-class model running...Keeps the grown-up capabilities: multi-s...
Updated 2026-07-18