Muse Spark (Meta)
A Tier · 8.8/10
Meta's frontier model line from its Superintelligence Lab -- Muse Spark 1.1 (2026-07-09) adds substantially better coding, 1M-token multi-agent orchestration, and Meta's first paid developer API (Meta Model API, public preview)
Score Breakdown
Benchmark Scores
Benchmarks for Muse Spark
| Benchmark | Description | Score | |
|---|---|---|---|
| MMLU | Knowledge across 57 subjects | 89% | |
| GPQA Diamond | Graduate-level science questions | 86% | |
| HumanEval | Python code generation | 91% | |
| Humanity's Last Exam | Frontier difficulty questions | 58% | |
| HealthBench Hard | 42.8% |
Last updated: 2026-04-19
The Good and the Bad
What we like
- +Completely free to use via Meta AI app and Meta.ai -- no subscription required for the full model
- +Natively multimodal: handles text, images, and audio in a single architecture, not bolted-on adapters
- +Contemplating mode orchestrates multiple agents reasoning in parallel -- competitive with Gemini Deep Think and GPT-5.4 Pro
- +Runs on 10x less compute than comparable frontier models according to Meta's benchmarks
- +260K token context window -- competitive with most frontier models
- +Scores 52 on Artificial Analysis Intelligence Index and 58% on Humanity's Last Exam in Contemplating mode
- +HealthBench Hard 42.8 -- a real differentiator. Far ahead of Claude Opus 4.6 (14.8), Gemini 3.1 Pro (20.6), and GPT-5.4 (40.1). Meta trained Muse Spark with 1,000+ licensed physicians in the loop, which shows up in medical-reasoning evaluations
- +Distribution advantage: will reach billions of users across Facebook, Instagram, and WhatsApp
What could be better
- −Meta Model API is public-preview only, with no published pricing yet -- production commitments are premature
- −Locked into Meta's ecosystem -- you need Meta accounts and their apps to access it
- −Early reviews say it's 'competitive but not leading' -- doesn't clearly beat GPT-5.4 or Claude on any benchmark
- −Privacy concerns given Meta's data practices across its social platforms
- −No fine-tuning, no custom instructions, no equivalent of Custom GPTs or Claude Projects
Pricing
Free (Meta AI app)
- ✓Full Muse Spark access
- ✓Text, image, audio input
- ✓Available on Meta.ai, Facebook, Instagram, WhatsApp
- ✓Contemplating mode
Meta Model API (public preview)
- ✓Launched 2026-07-09 with Muse Spark 1.1 -- Meta's first paid model API
- ✓Public preview for developer access
- ✓Early partners: Replit, Cline, Box, OpenClaw Foundation
- ✓Pricing not yet disclosed in the launch post
Known Issues
- MODEL + API LAUNCH (2026-07-09): **Muse Spark 1.1** shipped via Meta Superintelligence Labs -- available immediately in 'Thinking' mode in the Meta AI app and meta.ai, and launched alongside the **Meta Model API public preview**, Meta's first paid model API (resolving the long-standing 'no public API' gap). Vendor-cited improvements: substantially better coding on large, complex codebases (bug diagnosis, feature implementation, code migrations); stronger perception/multimodal reasoning/tool use across images, video, and audio; multi-agent orchestration where Spark 1.1 acts as a main agent delegating to subagents while managing a **1M token context window**; and computer use across multiple applications. Early API partners: Replit, Cline, Box, OpenClaw Foundation. NOTE: no pricing was disclosed in the launch post -- aggregator claims of '75% below rivals' are unverified; wait for the vendor price sheetSource: Meta AI blog (ai.meta.com/blog/introducing-muse-spark-meta-model-api/), NYT (2026-07-09), CNBC, Axios · 2026-07-09
- SUPERSEDED (2026-07-09): the April-era 'no public API, private preview only' status no longer holds -- the Meta Model API entered public preview with Muse Spark 1.1 (see entry above).Source: Meta AI blog (ai.meta.com/blog/introducing-muse-spark-meta-model-api/) · 2026-07-09
- Contemplating mode adds significant latency (30+ seconds) as multiple agents reason in parallel before respondingSource: DataCamp review, Reddit r/artificial · 2026-04
Best for
Anyone who wants frontier-level AI for free. If you use Meta's apps (Facebook, Instagram, WhatsApp) already, Muse Spark is the most accessible high-quality LLM with zero cost.
Not for
Developers who need API access for production apps (not available yet). Also not ideal for enterprise users who need data privacy guarantees given Meta's data handling practices.
Our Verdict
Muse Spark is Meta's strongest statement yet that frontier AI should be free. The model is genuinely competitive with GPT-5.4 and Claude on benchmarks, and Contemplating mode's multi-agent reasoning is a real differentiator. The catch? No API, no customization, and you're locked into Meta's ecosystem. For casual users who just want a great free chatbot, this is arguably the best deal in AI right now. For developers or enterprise users, you'll be waiting.
Sources
- Meta AI blog: Introducing Muse Spark 1.1 + Meta Model API (2026-07-09) (accessed 2026-07-18)
- Meta AI official blog (accessed 2026-04-13)
- TechCrunch coverage (accessed 2026-04-13)
- Artificial Analysis benchmarks (accessed 2026-04-13)
- DataCamp review (accessed 2026-04-13)
Explore more Muse Spark (Meta) rankings
Deeper leaderboards, benchmarks, task-specific tier lists, and status/pricing pages for Muse Spark (Meta).
The Tier List Tuesday
Weekly newsletter: tier movers, new entrants, and the VS of the week. Built from our daily AI-tool sweeps. No spam, unsubscribe anytime.
Alternatives to Muse Spark (Meta)
Claude (Anthropic)
Anthropic's flagship LLM family. **Claude Opus 5 launched 2026-07-24** and is now the default model on Claude Max and the strongest model on Claude Pro -- same $5/$25 per 1M as Opus 4.8, but Anthropic says it lands within 0.5% of Fable 5 on CursorBench at half the cost. **Sonnet 5's $2/$10 per 1M is now permanent** -- Anthropic cancelled the 2026-09-01 rise to $3/$15 and made the launch rate standard -- and it stays the default on Free/Pro. **Claude Fable 5.1 and Mythos 5.1 launched 2026-09-01** and now top the range -- same $10/$50 per 1M as Fable 5, with the saving delivered entirely through a 4x cheaper cache read ($1 -> $0.25/MTok), so it is 25-45% cheaper only if your workload reuses cached context. **From 2026-08-14 future Claude models watermark their text output globally** (SynthID-Text; no extra tokens, no price or speed change, no identifying information -- detector API not shipped yet), and the **legacy Workbench plus the experimental prompt-tools APIs retired 2026-08-17**
Claude Mythos 5.1
Anthropic's trusted-access frontier model. **Mythos 5.1 launched 2026-09-01** alongside Fable 5.1, and Anthropic now states outright that they are **the same model with different safeguards** -- the 5.1-cycle gap is 60.9% vs 55.8% on Terminal-Bench 4.0, which Anthropic attributes to safeguard interventions rather than capability and expects to shrink. Originally launched June 9, 2026 alongside Claude Fable 5. Suspended June 12 by a US export-control order, then PARTIALLY RESTORED July 1, 2026 (US government lifted controls June 30): Mythos 5 is back for a set of US organizations with government approval, while Anthropic works to re-expand the broader Glasswing program. Public Fable 5 returned globally the same day. Gated to Project Glasswing orgs + select biology researchers.
Gemini (Google)
Google's LLM with deep Google Workspace integration, 2M token context window, and native code execution -- **Gemini 3.8 Flash launched 2026-09-02** -- the third Flash release in six weeks -- at the SAME $0.75/$3.75 per 1M intro rate as 3.7 Flash and, critically, **on the same 2026-12-31 expiry** (then $1.50/$7.50), so the discount window did not reset. HLE-Verified 54.9%; Google warns it uses more tokens per task because it 'works harder', so cost per task can rise at an unchanged rate. **Gemini 3.8 Flash Cyber** is trusted-defenders-only via the new **Fairwind Program**. Gemini 3.7 Flash (2026-08-13) remains fully supported for efficiency-first work. Gemini 3.5 Pro STILL delayed and partner-testing-only (Bloomberg, 7/16 -- coding shortfalls, no ship date), so Flash keeps shipping while the Pro line stalls. Gemini 4 pre-training underway. **Two new models in Aug 2026: Gemini 3.5 Transcribe (2026-08-26) speech-to-text and Gemini Omni 1.1 Flash (2026-08-27) production video with 40s scene extension, start/end frames and 4K** -- neither shipped with a public rate card
Grok
SpaceXAI's irreverent chatbot with a direct line to X/Twitter -- and now **Grok 4.6 (launched 2026-08-12)**, focused on long-running agents and interactive/visual work, at the same $2/$6 per 1M as Grok 4.5 (fast variant 2x). xAI says it matches GPT-5.6 Sol on the AA Intelligence Index (61), though Sol still leads it on DeepSWE and Terminal-Bench. **Grok Bot** (8/11, early beta) adds always-on agents with their own cloud computer. **Grok 4.6 reached GitHub Copilot on 2026-08-14** across all five paid Copilot SKUs (off by default for orgs), adding to its day-one Cursor availability. Grok 4.3 remains the value tier at $1.25/$2.50
GPT-Rosalind (OpenAI)
OpenAI's first domain-specific model -- life sciences, drug discovery, translational medicine. Launched 2026-04-16 as a Trusted Access research preview. Launch partners: Amgen, Moderna, Allen Institute, Thermo Fisher. Paired with a Life Sciences Codex plugin (50+ scientific tool integrations)
GPT-5.4-Cyber (OpenAI)
OpenAI's defensive-cybersecurity variant of GPT-5.4, launched 2026-04-16. Lowered refusal boundary for security-research tasks and native binary reverse-engineering. Access gated via Trusted Access for Cyber (TAC) program -- thousands of verified defenders, hundreds of teams, no public pricing. On **2026-08-17 OpenAI published its first dedicated post on the Hugging Face incident**, conceding it 'underestimated the real-world cyber capabilities of our AI models' and confirming it now releases cyber capabilities only to trusted defenders. **On 2026-09-01 OpenAI confirmed GPT-6 Astra meets the Critical cyber threshold -- the first model it has ever designated at that level -- and on 2026-09-03 committed $1B to Daybreak for Frontline Defenders**
Microsoft MAI-Thinking-1
Microsoft's first in-house reasoning model -- launched 2026-06-02 at Build as the flagship of seven new MAI models. 35B-active / ~1T-total sparse Mixture-of-Experts, 256K context. AIME 2025 97.0%, matches leading models on SWE-Bench Pro, and beat Claude Sonnet 4.6 in human-preference testing. Available on Microsoft Foundry + OpenRouter / Fireworks / Baseten
Hunyuan 3 (Tencent Hy3)
Tencent's Hy3 reached GA 2026-07-06 (upgraded from the April preview) -- 295B total / 21B active MoE, 256K context, now Apache 2.0 open weights on HuggingFace + ModelScope with the EU/UK/South Korea restriction lifted. ~90% agent-task completion on Tencent's internal apps; API via Tencent Cloud TokenHub. Integrated into Yuanbao, WeChat, QQ
MiMo (Xiaomi)
Xiaomi's MiMo-V2.5 family launched 2026-04-22 -- Pro (1T total / 42B active MoE, 1M context, native vision+audio reasoning), Multimodal base, TTS (3 sub-models: base, VoiceDesign, VoiceClone), and ASR (open-source, English + Chinese + major dialects). Full voice pipeline for the agent era. Extra-charge 1M-context tier removed at launch