MiMo (Xiaomi)
A Tier · 8.3/10
Xiaomi's MiMo-V2.5 family launched 2026-04-22 -- Pro (1T total / 42B active MoE, 1M context, native vision+audio reasoning), Multimodal base, TTS (3 sub-models: base, VoiceDesign, VoiceClone), and ASR (open-source, English + Chinese + major dialects). Full voice pipeline for the agent era. Extra-charge 1M-context tier removed at launch
Score Breakdown
Personality & Tone
Xiaomi's voice-first agentic stack
Tone: Direct, multimodal-aware. MiMo-V2.5-Pro is comfortable mixing image, audio, and text inputs in a single turn -- it's been trained for that, not retrofitted to it.
Quirks: Voice-pipeline orientation makes MiMo unusually expressive when audio is in the loop -- TTS variants (VoiceDesign, VoiceClone) and ASR are surfaced as first-class products, which most Chinese frontier vendors haven't done. PRC content filters apply on chat surfaces.
The Good and the Bad
What we like
- +Full voice pipeline shipped together: a frontier reasoning model (Pro), a multimodal base, a TTS family, and an open-source ASR -- Xiaomi positions MiMo-V2.5 as 'voice for the agent era,' which is rare in 2026 (most vendors ship one of these and integrate the others later)
- +Native multimodal in MiMo-V2.5-Pro is the differentiator -- vision and audio reasoning in one model, not bolted on after the fact. Closer to the Gemini 2.5 / GPT-5.5 design than to text-first models with separate vision adapters
- +Removing the surcharge for the full 1M-context tier at launch is a real value move -- Alibaba, Anthropic, and OpenAI all charge meaningfully more per token for full-context windows. Xiaomi flattening this lowers the barrier to long-document and agentic workloads
- +Open-source MiMo-V2.5-ASR is the practical takeaway for privacy-sensitive teams. Cohere Transcribe + Whisper had been the open-ASR options through 2025; MiMo-V2.5-ASR adds a Chinese-dialect-strong third entry
- +Listed on Artificial Analysis at launch -- third-party verification path is open, even if scores are still being filled in
What could be better
- −Third-party benchmarks are still developing as of launch week -- Xiaomi's own published numbers are the dominant evidence, which warrants the usual self-reporting discount
- −PRC content filters apply on Pro and Multimodal -- the same regulated-topic refusals that Hy3, DeepSeek, and Qwen exhibit. ASR is less affected by content filtering since it's transcription, not generation
- −English creative-writing polish lags Western frontier models -- pick MiMo for Chinese-language work, multimodal reasoning, or voice pipelines first, English prose second
- −Geo-availability: API access for non-Chinese developers may require an Xiaomi developer account and KYC; check the docs before assuming OpenAI-style account-creation friction
Pricing
Free (consumer)
- ✓Xiaomi consumer device integration (HyperOS, Mi AI)
- ✓Web chat at mimo.xiaomi.com
- ✓Basic usage limits apply
API (MiMo-V2.5-Pro)
- ✓1T total / 42B active MoE
- ✓Native 1M context window with NO extra-charge tier (Xiaomi removed the surcharge for the full window at launch)
- ✓Native multimodal: vision and audio reasoning in one model
- ✓OpenAI- and Anthropic-API-compatible endpoints (the standard pattern Chinese frontier models adopted in 2025-26)
API (MiMo-V2.5 multimodal base)
- ✓Image + audio + video + text in a single API call
- ✓Cheaper than Pro for workloads that don't need 1M context or 42B-active capacity
MiMo-V2.5-TTS (3 sub-models)
- ✓Base TTS (general voice synthesis)
- ✓VoiceDesign (designed-from-scratch synthetic voices)
- ✓VoiceClone (replicate a target voice from a sample)
MiMo-V2.5-ASR (open-source)
- ✓Open-source under a permissive license
- ✓English + Mandarin Chinese + major Chinese dialects (Cantonese, Shanghainese, etc.)
- ✓Self-hostable for privacy-sensitive transcription workloads
System Requirements
Hardware needed to self-host. Min = smallest viable setup (usually heavy quantization). Max = full-precision / production-grade.
| Model variant | Min | Max |
|---|---|---|
| MiMo-V2.5-Pro (1T total, 42B active MoE)API-only flagship pattern matches Qwen 3.6-Max-Preview and DeepSeek V4-Pro positioning. Xiaomi may or may not open Pro weights later | API-only at launch -- weights not released for Pro | API-only -- weights not released |
| MiMo-V2.5-ASR (open-source)Open-source under a permissive license per Xiaomi's launch comms. Self-hostable for privacy-sensitive transcription. Strong specifically on Chinese dialects (Cantonese, Shanghainese) | 8 GB VRAM (RTX 3060 tier) for English + standard Mandarin | 1× A100 40 GB for full dialect coverage at production throughput |
Known Issues
- OLDER SERIES RETIRED (2026-06-30): Xiaomi fully deprecated the prior **MiMo-V2 series** (V2-Pro / V2-Omni / V2-Flash) on 2026-06-30 -- migrate to the MiMo-V2.5 family this page covers, which is the designated replacement. If you pinned a `mimo-v2-*` model id, it's now failing; the V2.5 line is current. (Reinforces V2.5 as the live family; no change to our review scope.)Source: Xiaomi MiMo docs (mimo.mi.com/docs updates), 36Kr, Hacker News · 2026-06-30
- NEW PRODUCT (2026-06-11): **MiMo Code v0.1.0** -- Xiaomi's open-source terminal AI coding agent (MIT license, TypeScript, built on OpenCode tech), 4,300+ GitHub stars on launch day and a top Hacker News story (357 points). MiMo-V2.5 model built in FREE; also supports DeepSeek and Kimi models. Differentiators: persistent memory system (project memory + session checkpoints + task progress), a '/dream' weekly self-maintenance agent, Compose mode (idea → design/plan/code/test/review pipeline), voice control, and curl/npm install (`mimo` command, Windows via npm). Vendor benchmark claims (UNVERIFIED third-party): SWE-Bench Pro 62% vs Claude Code 57%, Terminal-Bench 73% vs 68%. It's a v0.1.0 -- expect churn -- but a free, open-source, frontier-backed CLI agent is direct pressure on Claude Code / Codex CLI / Grok Build pricing. Standalone review page candidate once it stabilizes past 0.xSource: GitHub (github.com/XiaomiMiMo/MiMo-Code), mimo.xiaomi.com/mimocode, Hacker News, Gizmochina · 2026-06-11
- MIMO CODE DAY-2 CAVEATS (2026-06-12, press analysis): the launch benchmarks deserve skepticism -- they're self-reported, and the 'beats Claude Code' comparison pits MiMo Code + MiMo-V2.5 against Claude Code running **Sonnet 4.6**, not the flagship models. VentureBeat's hands-on found genuine strength on ultra-long 200+ step tasks. Privacy note: the 'free for a limited time' model routes your code through Xiaomi servers -- a data-residency consideration for commercial codebases. v0.1.0 remains the only release (no hotfixes yet)Source: TechTimes (2026-06-12), VentureBeat (2026-06-11), GitHub releases · 2026-06-12
- MiMo-V2.5 family launched 2026-04-22 with four product lines released in parallel: Pro (1T/42B MoE, 1M context, native vision+audio), Multimodal base, TTS (Base + VoiceDesign + VoiceClone), and open-source ASR. This is Xiaomi's first explicit 'voice for the agent era' positioning and the first time it has shipped frontier-class reasoning + voice in a single coordinated launchSource: Xiaomi product site (mimo.xiaomi.com), Gizmochina, Artificial Analysis listing · 2026-04-22
- 1M-context surcharge removed at launch on Pro -- Xiaomi explicitly priced parity between short-context and full-context calls. Watch whether they reintroduce a tier later as adoption scales; the no-surcharge stance is unusual at this scaleSource: Xiaomi launch comms, Gizmochina · 2026-04
- ASR is open-source; TTS is API-only. If you need a fully self-hostable voice pipeline you can use ASR locally + a different TTS (ElevenLabs, Cohere, Murf) on top, or wait for Xiaomi to potentially open the TTS weights laterSource: Xiaomi announcement · 2026-04
- PRC content filtering applies on the reasoning/chat surfaces. Same regulated-topic pattern as Hy3, DeepSeek, Qwen, Kimi, GLMSource: Pattern across Chinese frontier APIs · 2026-04
Best for
Teams building voice-first agentic products that need a coordinated reasoning + TTS + ASR stack from a single vendor. Also Chinese-market builders and developers who need strong multimodal (vision + audio) inputs in one API call without stitching three providers together. The no-surcharge 1M-context stance makes MiMo-V2.5-Pro especially attractive for long-document agentic workloads.
Not for
English-first creative writing (Claude / GPT-5.5 still lead), regulated geographies that block Chinese AI APIs, or teams whose only voice need is English TTS (ElevenLabs is more mature). Also not the right fit if you need fully proven third-party benchmark verification today -- that takes weeks post-launch.
Our Verdict
MiMo-V2.5 is Xiaomi treating voice as a first-class agentic surface, not an after-the-fact integration. Shipping Pro + Multimodal + TTS + open-source ASR together -- with native vision and audio reasoning baked into the flagship and the 1M-context surcharge removed -- is the most coordinated voice-stack launch from a Chinese frontier vendor in 2026. The benchmark story will fill in over the next few weeks; for now, treat MiMo as a serious option for voice-pipeline builds, multimodal Chinese-language workloads, and self-hosted dialect-strong ASR. For text-only English-first work, Claude / GPT / Gemini still lead and DeepSeek is still the cheapest text-first frontier alternative.
Sources
- Xiaomi MiMo docs: MiMo-V2 series deprecation (2026-06-30) (accessed 2026-07-04)
- Xiaomi: MiMo-V2.5-Pro product page (accessed 2026-04-25)
- Gizmochina: Xiaomi introduces MiMo-V2.5 TTS and ASR full voice pipeline (accessed 2026-04-25)
- Artificial Analysis: MiMo-V2.5-Pro listing (accessed 2026-04-25)
- The Asian Mirror: Xiaomi MiMo V2.5 voice AI launch (accessed 2026-04-25)
Explore more MiMo (Xiaomi) rankings
Deeper leaderboards, benchmarks, task-specific tier lists, and status/pricing pages for MiMo (Xiaomi).
The Tier List Tuesday
Weekly newsletter: tier movers, new entrants, and the VS of the week. Built from our daily AI-tool sweeps. No spam, unsubscribe anytime.
Alternatives to MiMo (Xiaomi)
Claude (Anthropic)
**Life Sciences Verification Program opened 2026-09-17** -- verified life-science teams get Mythos 5.1, Opus 5 and Sonnet 5 with the biology safeguards relaxed (Standard Use grants renewed yearly; High-risk Use grants per project every six months, Opus 5 and Sonnet 5 today, Mythos still US-government-gated), on top of the 2026-08-27 Claude team plan for scientists (10,000 seats, standard free, premium $15/mo). Anthropic's flagship LLM family. **Claude Opus 5 launched 2026-07-24** and is now the default model on Claude Max and the strongest model on Claude Pro -- same $5/$25 per 1M as Opus 4.8, but Anthropic says it lands within 0.5% of Fable 5 on CursorBench at half the cost. **Sonnet 5's $2/$10 per 1M is now permanent** -- Anthropic cancelled the 2026-09-01 rise to $3/$15 and made the launch rate standard -- and it stays the default on Free/Pro. **Claude Fable 5.1 and Mythos 5.1 launched 2026-09-01** and now top the range -- same $10/$50 per 1M as Fable 5, with the saving delivered entirely through a 4x cheaper cache read ($1 -> $0.25/MTok), so it is 25-45% cheaper only if your workload reuses cached context. **From 2026-08-14 future Claude models watermark their text output globally** (SynthID-Text; no extra tokens, no price or speed change, no identifying information -- detector API not shipped yet), and the **legacy Workbench plus the experimental prompt-tools APIs retired 2026-08-17**
Claude Mythos 5.1
Anthropic's trusted-access frontier model. **Mythos 5.1 launched 2026-09-01** alongside Fable 5.1, and Anthropic now states outright that they are **the same model with different safeguards** -- the 5.1-cycle gap is 60.9% vs 55.8% on Terminal-Bench 4.0, which Anthropic attributes to safeguard interventions rather than capability and expects to shrink. Originally launched June 9, 2026 alongside Claude Fable 5. Suspended June 12 by a US export-control order, then PARTIALLY RESTORED July 1, 2026 (US government lifted controls June 30): Mythos 5 is back for a set of US organizations with government approval, while Anthropic works to re-expand the broader Glasswing program. Public Fable 5 returned globally the same day. Gated to Project Glasswing orgs + select biology researchers.
Gemini (Google)
**Gemini 3.8 Live and 3.8 Live Extended Thinking launched 2026-09-15** -- native speech-to-speech models in the Gemini API at $0.005/min audio in and $0.018/min audio out (audio-in a tenth of GPT-Live-1's $0.05/min voice layer, and on the same price row as the older 3.1 Flash Live), with Extended Thinking taking the #1 spot on Artificial Analysis' Speech to Speech Quality Index (82.6) and rolling into Gemini Live, Search Live and Workspace Docs/Gmail/Keep. **Gemini app for Windows shipped 2026-09-10** (Alt+Space overlay, Google-app connectors, Gemini Spark on desktop) and the Sept 9 AI-plan refresh added Google Pics and Sheets canvas for AI Pro/Ultra plus Gemini Spark hooks into Chrome and Google Photos. Google's LLM with deep Google Workspace integration, 2M token context window, and native code execution -- **Gemini 3.8 Flash launched 2026-09-02** -- the third Flash release in six weeks -- at the SAME $0.75/$3.75 per 1M intro rate as 3.7 Flash and, critically, **on the same 2026-12-31 expiry** (then $1.50/$7.50), so the discount window did not reset. HLE-Verified 54.9%; Google warns it uses more tokens per task because it 'works harder', so cost per task can rise at an unchanged rate. **Gemini 3.8 Flash Cyber** is trusted-defenders-only via the new **Fairwind Program**. Gemini 3.7 Flash (2026-08-13) remains fully supported for efficiency-first work. Gemini 3.5 Pro STILL delayed and partner-testing-only (Bloomberg, 7/16 -- coding shortfalls, no ship date), so Flash keeps shipping while the Pro line stalls. Gemini 4 pre-training underway. **Two new models in Aug 2026: Gemini 3.5 Transcribe (2026-08-26) speech-to-text and Gemini Omni 1.1 Flash (2026-08-27) production video with 40s scene extension, start/end frames and 4K** -- Transcribe's rate card has since appeared (about $0.005/min blended, verified 9/17); Omni 1.1 Flash's still has not
Grok
SpaceXAI's irreverent chatbot with a direct line to X/Twitter -- and now **Grok 4.6 (launched 2026-08-12)**, focused on long-running agents and interactive/visual work, at the same $2/$6 per 1M as Grok 4.5 (fast variant 2x). xAI says it matches GPT-5.6 Sol on the AA Intelligence Index (61), though Sol still leads it on DeepSWE and Terminal-Bench. **Grok Bot** (8/11, early beta) adds always-on agents with their own cloud computer. **Grok 4.6 reached GitHub Copilot on 2026-08-14** across all five paid Copilot SKUs (off by default for orgs), adding to its day-one Cursor availability. Grok 4.3 remains the value tier at $1.25/$2.50
Muse Spark (Meta)
Meta's frontier model line from its Superintelligence Lab -- Muse Spark 1.1 (2026-07-09) adds substantially better coding, 1M-token multi-agent orchestration, and Meta's first paid developer API (Meta Model API, public preview)
GPT-Rosalind (OpenAI)
OpenAI's first domain-specific model -- life sciences, drug discovery, translational medicine. Launched 2026-04-16 as a Trusted Access research preview. Launch partners: Amgen, Moderna, Allen Institute, Thermo Fisher. Paired with a Life Sciences Codex plugin (50+ scientific tool integrations)
GPT-5.6-Cyber / GPT-5.4-Cyber (OpenAI)
**OpenAI is retiring GPT-5.4-Cyber (notice 2026-09-11, removed from the API 2026-10-01); its replacement is `gpt-5.6-cyber`**, an alias for OpenAI's most advanced purpose-trained cyber models, gated behind the Daybreak program and priced at $12.50/$75 per 1M (first published price for any OpenAI cyber model). Original page: OpenAI's defensive-cybersecurity variant of GPT-5.4, launched 2026-04-16. Lowered refusal boundary for security-research tasks and native binary reverse-engineering. Access gated via Trusted Access for Cyber (TAC) program -- thousands of verified defenders, hundreds of teams, no public pricing. On **2026-08-17 OpenAI published its first dedicated post on the Hugging Face incident**, conceding it 'underestimated the real-world cyber capabilities of our AI models' and confirming it now releases cyber capabilities only to trusted defenders. **On 2026-09-01 OpenAI confirmed GPT-6 Astra meets the Critical cyber threshold -- the first model it has ever designated at that level -- and on 2026-09-03 committed $1B to Daybreak for Frontline Defenders**
Microsoft MAI-Thinking-1
Microsoft's first in-house reasoning model -- launched 2026-06-02 at Build as the flagship of seven new MAI models. 35B-active / ~1T-total sparse Mixture-of-Experts, 256K context. AIME 2025 97.0%, matches leading models on SWE-Bench Pro, and beat Claude Sonnet 4.6 in human-preference testing. Available on Microsoft Foundry + OpenRouter / Fireworks / Baseten
Hunyuan 3 (Tencent Hy3)
Tencent's Hy3 reached GA 2026-07-06 (upgraded from the April preview) -- 295B total / 21B active MoE, 256K context, now Apache 2.0 open weights on HuggingFace + ModelScope with the EU/UK/South Korea restriction lifted. ~90% agent-task completion on Tencent's internal apps; API via Tencent Cloud TokenHub. Integrated into Yuanbao, WeChat, QQ