Grok logo
B

Grok

B Tier · 7.5/10

SpaceXAI's irreverent chatbot with a direct line to X/Twitter -- and now **Grok 4.6 (launched 2026-08-12)**, focused on long-running agents and interactive/visual work, at the same $2/$6 per 1M as Grok 4.5 (fast variant 2x). xAI says it matches GPT-5.6 Sol on the AA Intelligence Index (61), though Sol still leads it on DeepSWE and Terminal-Bench. **Grok Bot** (8/11, early beta) adds always-on agents with their own cloud computer. Grok 4.3 remains the value tier at $1.25/$2.50

Last updated: 2026-08-13Free tier available

Score Breakdown

7.0
Ease of Use
7.5
Output Quality
7.5
Value
8.0
Features

Benchmark Scores

Benchmarks for Grok 4.20 (baseline -- Grok 4.5 launched 2026-07-08; vendor-reported 4.5 scores in Known Issues pending third-party verification)

Chatbot Arena ELOHuman preference rating1420
BenchmarkScore
MMLU88.5%
GPQA Diamond85%
HumanEval90%
Humanity's Last Exam50.7%

Last updated: 2026-04-13

Personality & Tone

The irreverent contrarian

Tone: Casual, jokey, and willing to swear. Grok takes strong positions without hedging, leans into an edgy 'based' persona, and cracks jokes far more often than Claude, ChatGPT, or Gemini.

Quirks: Engages with topics other chatbots refuse, pulls live context from X so it reflects whatever is trending that hour, and will freely mock things -- including itself. In SuperGrok's multi-agent mode it can sound like several personalities arguing with each other.

The Good and the Bad

What we like

  • +Real-time access to X/Twitter data is genuinely useful for tracking breaking news and trending topics
  • +Grok 3 benchmarks are competitive with GPT-4o and Claude 3.5 -- this is not a vanity project anymore
  • +The personality is refreshing if you're tired of overly cautious AI assistants -- it'll actually joke around
  • +DeepSearch mode does solid multi-step research, pulling from web and X data simultaneously

What could be better

  • The snarky personality gets old fast when you're trying to get serious work done
  • Tied to the X ecosystem -- you need an X account, and the real-time data skews toward X's user base
  • SuperGrok at $30/mo is steep when Claude Pro and ChatGPT Plus are $20 with arguably better core models
  • Image generation is no longer the weak spot it was -- Imagine Image 2.0 (2026-08-07) added region editing, segmentation, background removal and 5-image multi-ref, and xAI places it #2 worldwide on both Arena image boards; the remaining gap vs ChatGPT/Gemini is analysis and document understanding, not generation

Pricing

Free

$0
  • ~10 prompts per 2 hours
  • Basic Grok access
  • Requires X account

X Premium

$8/month
  • Higher query limits
  • Grok 4.20 access
  • Bundled X social features

X Premium+

$40/month
  • Higher Grok 4.20 access
  • Ad-free X
  • Priority responses

SuperGrok

$30/month
  • Full Grok 4.20 (4-agent multi-agent system)
  • DeepSearch mode
  • Highest rate limits
  • Think mode
  • $300/yr option (16% off)

SuperGrok Heavy

$300/month
  • Grok 4 Heavy model
  • Highest priority
  • Multi-agent at scale
  • Note: Grok 4.3 beta-gating ended 2026-05-02

API (Grok 4.5, launched 2026-07-08)

$2 / $6/per 1M tokens (input/output)
  • SpaceXAI + Cursor joint frontier MoE model
  • Faster variant $4/$18 via Cursor
  • Available in Grok Build, SpaceXAI console/API, Cursor (all plans)
  • NOT available in EU at launch
  • Vendor benchmarks: Terminal-Bench 2.1 83.3%, SWE-Bench Pro 64.7% (third-party verification pending)

API (Grok 4.3)

$1.25 / $2.50/per 1M tokens (input/output)
  • Production launch 2026-05-02 (~40% input / ~60% output price cut vs 4.20)
  • 1M context window
  • Reasoning tokens billed at output rate
  • Native video input + PDF/PPT/spreadsheet output
  • Custom Voices voice cloning free on console (80+ presets, 28 languages)
  • Imagine Agent Mode (creative workflow agent, beta)

Known Issues

  • GROK 4.6 SHIPPED (2026-08-12, vendor-primary) -- AND IT ARRIVED FIVE DAYS AFTER THE DATE WE RECORDED AS MISSED. Read this alongside the 7/16 entry below, which correctly stated on 8/10 that 4.6 had not shipped: the Musk-stated 8/7 target did pass without a release, but xAI then published **x.ai/news/grok-4-6, dated Aug 12, 2026**. xAI's framing: 4.6 'builds on Grok 4.5 with a particular focus on long-running agents and more ambitious interactive and visual work,' staying with tasks 'across many steps.' **PRICING: $2/M input and $6/M output** -- identical to Grok 4.5, so this is a capability bump at flat price, not a price move. **There is also a 'fast variant which is twice the price'** ($4/$12 by that arithmetic; xAI states the multiplier, not the absolute figures, so we do not print derived numbers as vendor facts). **VENDOR-PUBLISHED EVAL TABLE (Grok 4.6 High vs Grok 4.5 High):** AA Intelligence Index **61 (from 56)**, GDPVal-AA v2 **1753 (from 1526)**, CursorBench v3.2 **69.9% (from 66.7%)**, DeepSWE v1.1 **65.9% (from 54%)**, FrontierCode v1.1 Extended **61.3% (from 56.6%)**, APEX-Agents **57.5% (from 47.1%)**, Terminal-Bench v3.0 **26% (from 15.7%)**, APEX-SWE **56.4% (from 53.6%)**, AA-Briefcase **1577 (from 1313)**, Harvey LAB **15.8% (from 12.9%)**. **THE HEADLINE CLAIM, STATED PRECISELY:** xAI says 4.6 'matches GPT-5.6 Sol on the Artificial Analysis Intelligence Index' -- both at **61** in xAI's own table, with **Fable 5 Max at 62**. **BUT THE SAME TABLE SHOWS WHERE IT LOSES**, and this is the part marketing coverage will drop: on **DeepSWE v1.1 GPT-5.6 Sol scores 73% against Grok 4.6's 65.9%**, and on **Terminal-Bench v3.0 Sol scores 34.6% and Fable 5 Max 34.1% against Grok 4.6's 26%**. So 'matches Sol' is true of one composite index and not of the agentic-coding benchmarks underneath it. All competitor figures are xAI-selected from rivals' published system cards or leaderboards; third-party verification pending. **AVAILABILITY DAY ONE:** Cursor and Grok Build, the xAI API, and partners **OpenRouter, Vercel and Cloudflare**, with **2x included usage inside Grok Build and Cursor for the first week**. **METHOD NOTE FOR FUTURE SWEEPS: this is the resolution of a pending dated flip that had already been marked 'did not ship' once. A missed vendor date is not a cancelled product -- keep checking the slug index.**Source: SpaceXAI (x.ai/news/grok-4-6, on-page date Aug 12, 2026, fetched 2026-08-13 via curl -- x.ai 403s WebFetch); corroborated by Cursor's own post (cursor.com/blog/grok-4-6, dated Aug 12, 2026, fetched 2026-08-13) · 2026-08-12
  • GROK BOT -- AN AGENT PRODUCT WITH ITS OWN COMPUTER, AND AN UNUSUAL CROSS-VENDOR ENTITLEMENT (2026-08-11, vendor-primary, EARLY BETA): xAI launched **Grok Bot**, described as 'your team of always-on agents' -- 'AI teammates you can give real work to.' The design claim that distinguishes it from chat-style agents: **bots get their own cloud computer**, sign into the tools you already use, and 'work across apps, inboxes, and more,' including **'platforms with no clean API or MCP'** -- i.e. it drives real UIs rather than requiring integrations. They 'finish jobs end to end, and only come back when something needs your approval,' retain conversation memory, and 'learn how you like things done.' xAI says it began as an internal prototype used for sales outbound, marketing campaigns, office operations and bug fixes. **AVAILABILITY IS THE NOTEWORTHY PART, QUOTED EXACTLY:** 'Grok Bot is in beta and available today for **SuperGrok Heavy, Cursor Ultra, and Cursor Teams Premium subscribers** on desktop and iOS,' with **enterprise access by waitlist**. That means an xAI product ships as an entitlement of a *Cursor* subscription -- which we verified is coherent rather than a typo: **Cursor's own pricing page lists Ultra and Teams Premium as real tiers and advertises 'Generous limits for Grok' on every paid plan, with Grok as a top-level nav item**. This follows SpaceX's Anysphere/Cursor acquisition and is the clearest sign yet that xAI treats Cursor's subscriber base as a first-party distribution channel. **Caveats to hold:** it is explicitly **early beta**, no pricing is published for standalone access, and no benchmark or reliability data accompanies the autonomy claims -- an agent with persistent credentials into your live SaaS tools is a materially larger blast radius than a chat assistant, and xAI publishes no security model for it in this postSource: SpaceXAI (x.ai/news/introducing-grok-bot, on-page date Aug 11, 2026, fetched 2026-08-13 via curl); tier names cross-checked against cursor.com/pricing (fetched 2026-08-13) · 2026-08-11
  • IMAGINE IMAGE 2.0 IS GA -- PRECISE EDITING ARRIVES, AND GROK IS NOW #2 ON BOTH IMAGE ARENAS (2026-08-07, vendor-primary): xAI shipped **Imagine Image 2.0 as the new Quality Mode** on grok.com/imagine and the iOS and Android apps. The pitch is explicitly utilitarian -- 'make images you can use in real work' -- with the model planning typography and layout 'the way a designer would' so dense multi-part visuals hold together and small text stays sharp. **The substantive change is that editing is now first-class, not a re-roll.** Four tools: the **magic wand** edits only the region you point at and leaves the rest untouched; **segmentation** selects precise areas to change; **background removal** exports any subject on transparency; and **multi-ref editing accepts up to 5 input images in a single generation**, which removes a manual compositing step. **Smart resize** refills the frame across ten ratios (1:2, 9:16, 2:3, 3:4, 1:1, 4:3, 3:2, 16:9, 2:1). xAI also added **templates** -- prepackaged workflows for photo editing, product shots, headshots, icons, game assets, mascots, e-commerce and UGC photos, emoji and merch. **RANKING, QUOTED PRECISELY:** xAI says Image 2.0 'ranks second in the world in both text-to-image generation and image editing' on Arena overall Elo **as of Aug 7, 2026** -- note this is a vendor-stated placement citing the Arena leaderboards, third-party re-verification pending, and that **xAI models are listed on Arena under 'SpaceXAI'**, which is where to look if you go check it yourself. This directly retires the long-standing con on this page that Grok's image generation lags ChatGPT and Gemini -- on the vendor's own leaderboard reading it no longer does, though the con predates the Imagine line and is kept below as historical context for Grok 3-era usersSource: SpaceXAI (x.ai/news/grok-imagine-image-2, on-page date Aug 7, 2026, fetched 2026-08-10 via curl -- x.ai 403s WebFetch) · 2026-08-07
  • IMAGINE VIDEO 1.5 GETS REFERENCES, TEXT-TO-VIDEO AND NATIVE 1080p (2026-07-31, vendor-primary -- this page had only ever recorded the 6/3 Imagine 1.5 *preview*, so the shipped feature set was several steps behind): xAI substantially expanded its video model. **Three things landed.** (1) **Multi-reference conditioning** -- pass reference images and 'each reference image locks one thing in place: a face, a product, a location', so you can keep a character and swap the scene, keep the scene and swap the character, or hold both and change only the action. **Up to seven references per generation.** (2) **Voice reference / voice consistency** -- supply a character image plus a voice reference and 'both hold: the same face and the same voice in every scene', which is the piece most competing video models still cannot do without a separate dubbing pass. (3) **Text-to-video and native 1080p** -- generation from a prompt with no starting image (xAI describes it as pairing their image generation with image-to-video), and 1080p output for both text-to-video and image-to-video, up from the 720p ceiling the 6/3 preview shipped with. **AVAILABILITY IS SPLIT AND WORTH READING CAREFULLY:** text-to-video and native 1080p are **generally available** on grok.com/imagine, iOS and Android. Image and voice references started **2026-07-31 in the US only, for SuperGrok Heavy and SuperGrok Plus**, on grok.com/imagine and iOS, with xAI saying they roll out to all tiers 'over the next few days' -- so by now that gate has probably widened, but the vendor post is the last first-party statement of scope. **In the API**, image references, text-to-video and native 1080p are live under the model id **`grok-imagine-video-1.5`**; **voice reference support is 'available on request'**, i.e. not self-serve. No pricing was published in the postSource: xAI/SpaceXAI (x.ai/news/grok-imagine-video-1-5-references, datePublished 2026-07-31T00:00:00Z, fetched 2026-08-06 via curl -- x.ai 403s WebFetch) · 2026-07-31
  • PRODUCT PAIR (2026-07-15/16, vendor posts on x.ai/news): (1) **Grok Automations** (7/16) -- scheduled and autonomous recurring tasks in the Grok app (standing queries, monitoring, repeat jobs), SpaceXAI's answer to ChatGPT's Scheduled Tasks. (2) **Grok Build open-sourced** (7/15) -- the terminal coding agent's source is now public, relevant if you want to audit or extend the harness. Roadmap noise, clearly labeled: Musk indicated a **Grok 4.6** is in the pipeline (~7/17-18, no date, no specs -- aggregator-grade until a vendor post exists) and Grok 5's timeline has slid repeatedly. **GROK 4.6 UPDATE -- RESOLVED, SEE THE 2026-08-12 ENTRY ABOVE.** The history is worth keeping because it shows how these dates actually behave: the Musk-stated **8/7 target did pass with nothing shipped** (verified 8/10 by re-enumerating the full x.ai/news index -- no post, no slug), and then **Grok 4.6 launched on 2026-08-12**, five days late. So the 8/10 reading was accurate at the time and the product was not cancelled, merely slipped. Standing rule this produced: treat 'Grok X is out' as unsourced until a vendor post exists, but do **not** conclude from a missed date that a release is dead. EU AVAILABILITY STILL UNRESOLVED: the 'mid-July' EU promise for Grok 4.5 has conflicting reports (one aggregator claims EU access landed ~7/16; another says still unavailable) and x.ai's Grok 4.5 page wouldn't load for verification -- we are NOT flipping EU availability until vendor wording confirms; re-check next sweepSource: SpaceXAI (x.ai/news/grok-automations, x.ai/news/grok-build-open-source -- both listed on the x.ai news index, scraped in-session) · 2026-07-16
  • GROK 4.5 LAUNCHED (2026-07-08 developer surfaces, public rollout reported 7/9): SpaceXAI shipped **Grok 4.5, 'our smartest model built for coding, agentic tasks, and knowledge work'** -- a mixture-of-experts frontier model **trained jointly with Cursor** (SpaceX closed its Anysphere/Cursor acquisition in June) on trillions of tokens of Cursor data. Musk's framing: 'Opus-class, but faster, more token-efficient and lower cost.' **API pricing: $2/M input + $6/M output** (a faster variant at $4/$18 is offered through Cursor). Day-one availability: **Grok Build, the SpaceXAI console/API, and Cursor on all plans**; press reports public access via grok.com and the X app from 7/9; **NOT available in the EU at launch**. Vendor-reported benchmarks (charts, third-party verification pending): **Terminal-Bench 2.1 83.3%, SWE-Bench Pro 64.7%, #1 on Harvey's Legal Agent Benchmark**, with standout token efficiency (~16K output tokens per SWE-Bench Pro task vs ~67K for Opus 4.8 max per launch charts); Artificial Analysis measured 91.3 tok/s on the API. CONSUMER ACCESS CONFIRMED (7/10 press): available to **X Premium and SuperGrok subscribers** in the Grok interface, and it's the **default model in Grok Build** (beta, SuperGrok + X Premium+). API also live on OpenRouter, Vercel, Cloudflare, Snowflake, Databricks. **INDEPENDENT BENCHMARKS PUBLISHED (Artificial Analysis, 7/9-10): Intelligence Index 54** -- AA's launch article places it 4th among frontier models behind Fable 5, GPT-5.5, and Opus 4.8 (the model page ranks it #8 of 188 counting all model variants); **GDPval-AA v2 Elo 1543**; 92.9 output tok/s measured; 500K context window per AA. AA calls the 16-point gen-over-gen jump xAI's largest ever. Press caveat: reviewers note a higher hallucination rate than frontier peers. Aggregator claims of a '1.5T-param V9 foundation' remain UNVERIFIED -- treat as rumorSource: SpaceXAI (x.ai/news/grok-4-5), Cursor blog (cursor.com/blog/grok-4-5), Artificial Analysis (artificialanalysis.ai/models/grok-4-5), Axios (2026-07-08) · 2026-07-10
  • GROK BUILD PLUGIN MARKETPLACE (2026-06-11, vendor-primary): xAI launched a built-in **plugin marketplace for Grok Build** -- plugins bundle skills, slash commands, agents, hooks, MCP servers, and LSPs; installs are commit-SHA-pinned for supply-chain safety; the catalog is open to community submissions via PR. Launch partners: MongoDB, Vercel, Sentry, Chrome DevTools, Cloudflare. Mirrors the plugin/extension pattern Claude Code and Gemini-CLI-era tooling established -- Grok Build is maturing fast for a product still labeled betaSource: xAI news (x.ai/news/grok-plugin-marketplace), GitHub (github.com/xai-org/plugin-marketplace) · 2026-06-11
  • JUNE CLUSTER (2026-06, all vendor-primary on x.ai/news): **Grok Imagine 1.5 Preview** (6/3) -- image-to-video generation up to 720p, available as an API preview. **Composer 2.5** (6/1) -- xAI's 'fast, SOTA model for long-running tasks,' now selectable in the Grok Build /models menu for SuperGrok and X Premium+ subscribers (NOT related to Cursor's Composer line despite the name). **Grok Build 0.1 on the API** (5/29) -- the coding-agent model behind Grok Build became directly callable via the xAI API: 256K context, always-on reasoning, text + image input. Grok Build itself ('Introducing Grok Build,' 5/25) is in early beta for ALL SuperGrok and X Premium+ subscribers -- broader than the original Heavy-tier-only gate. Also: Grok voice now powers Vapi (6/3) and Gopuff's 'Go' shopping agent (6/9). NOTE: 'Grok 5' / 'V9-Medium mid-June' claims remain aggregator-only with zero vendor signal -- not real until x.ai posts itSource: xAI news (x.ai/news/grok-imagine-1-5, x.ai/news/composer-2-5, x.ai/news/grok-build-0-1, x.ai/news/grok-build-cli), x.ai/build/changelog · 2026-06-03
  • PRODUCT (2026-05-18): xAI shipped **Grok Skills** -- a persistent-memory Skills layer on Grok 4.3. Skills are user-defined named capabilities Grok carries across sessions on web / iOS / Android (recipe collection, code-review checklist, study-habits coach, etc.). Each Skill is a stored prompt + behavioral pattern Grok consults when invoked by name. Per-user storage; not shared across accounts. Differentiates Grok from ChatGPT Memory (passive recall) toward configurable named tools. Pairs with the 5/14 Grok Build CLI ship -- Skills are the consumer-facing persistent-state layer, Build CLI is the developer-facing one. Material in the 'agent goes where you go' competitive narrative alongside Codex on mobile (5/14) + Cursor Jira integration (5/19) + Devin Windows VMs (5/21).Source: xAI news (x.ai/news), xAI release notes (docs.x.ai/developers/release-notes) · 2026-05-18
  • MODEL LINEUP CONSOLIDATION (2026-05-15, went live 12:00 PT): xAI auto-redirected **8 deprecated model slugs** to grok-4.3 (or grok-imagine-image-quality for the image model). Affected slugs: grok-4-1-fast-reasoning, grok-4-1-fast-non-reasoning, grok-4-fast-reasoning, grok-4-fast-non-reasoning, grok-4-0709, grok-code-fast-1, grok-3, grok-imagine-image-pro. All requests now silently bill at grok-4.3 rates ($1.25 input / $2.50 output per 1M tokens). Anyone with these slugs pinned in production or referenced inside a Copilot/Cursor/Codex multi-model selector now pays the new rate without any code change. Migration path: explicitly switch to grok-4.3 in your model selector and audit token-spend after 5/15 since the new rate may differ from what each deprecated slug was previously billed at. The grok-code-fast-1 slug retirement is the same event that took the model off GitHub Copilot's Chat/inline/agent surfaces on 5/15Source: xAI docs (docs.x.ai/developers/migration/may-15-retirement) · 2026-05-15
  • PRODUCT (2026-05-14): xAI launched **Grok Build CLI** in early beta -- an agentic terminal-native CLI for coding, app development, and workflow automation. Spawns up to **8 concurrent agents** in parallel. Powered by Grok 4.3 beta with a 16-agent Heavy architecture and **2M token context window**. Vendor-primary launch posts at x.ai/news/grok-build-cli and x.ai/cli, plus Musk's public invitation to wider beta testers on X. **Access gate**: launched first to SuperGrok Heavy tier ($299/mo, intro offer $99/mo for 6 months) -- not yet available to standard Premium / SuperGrok subscribers. Positions Grok as a direct competitor to Claude Code, Codex CLI, and Cursor CLI for terminal-first agentic coding workflows. The 8-agent parallelism + 2M context is the differentiating feature -- single longest context window of any production coding CLI as of todaySource: xAI news (x.ai/news/grok-build-cli), xAI product page (x.ai/cli), Musk on X · 2026-05-14
  • xAI joined SpaceX on 2026-02-02 -- SpaceX acquired xAI. Procurement, billing, and compliance workflows now route through SpaceX's vendor pipeline. For regulated industries (healthcare, finance, US government) this may require re-qualifying xAI as a vendor even if Grok itself was previously approvedSource: xAI announcement (x.ai/news/xai-joins-spacex), SpaceX updates · 2026-02
  • Grok Speech (STT + TTS) APIs launched 2026-04-17 as separate products from the chatbot -- see /tools/grok-voice on this site. Built on the same stack Grok Voice uses. Not included in Premium/SuperGrok consumer tiers; billed separately at $0.10/hr STT batch and $4.20/1M char TTSSource: xAI Grok STT/TTS announcement · 2026-04
  • Real-time X data can surface misinformation from viral posts without adequate fact-checkingSource: Reddit r/artificial · 2026-02
  • Free tier rate limits are aggressive -- many users report hitting caps within a few queriesSource: X/Twitter user reports · 2026-03
  • Grok 4.20's 4-agent system (Grok, Harper, Benjamin, Lucas) can take 30+ seconds for complex queries as agents debate internally. Grok 4.20 Beta 2 (landed ~2026-04-07) improved instruction-following, reduced hallucinations, better LaTeX and image search -- partially addresses the slowness and reliability complaints from early 4.20 feedbackSource: Reddit r/grok, IBTimes · 2026-04
  • PRODUCTION LAUNCH (2026-05-02): Grok 4.3 went broadly available beyond the SuperGrok Heavy beta. New consumer + API features: **Custom Voices voice cloning suite** (clone voice from ~1 minute of speech in <2 minutes, two-stage passphrase + speaker-embedding consent gate, 80+ preset voices, 28 languages, free on console); **Imagine Agent Mode** (creative production workflow agent, beta); native video input + reasoning-by-default; native PDF / PowerPoint / spreadsheet output. **API pricing: $1.25 input / $2.50 output per 1M tokens** -- ~40% input cut + ~60% output cut vs Grok 4.20. 1M context window. Reasoning tokens billed at output rate. xAI's pattern is silent ship via grok.com model selector + console UI rather than vendor blog post -- vendor-primary verification through grok.com itself plus 4+ tier-1 press sources (VentureBeat, Winbuzzer, The Decoder, Phemex)Source: VentureBeat (venturebeat.com/technology/xai-launches-grok-4-3-at-an-aggressively-low-price-and-a-new-fast-powerful-voice-cloning-suite), Winbuzzer 2026-05-03, The Decoder, grok.com console · 2026-05-02
  • Grok 4.3 Beta dropped 2026-04-17 as a SuperGrok Heavy exclusive ($300/mo tier). Elon Musk clarified on 2026-04-18 that the live checkpoint is ~0.5T params; the full 1T version is ~5 days from finishing training. Beta gating ENDED 2026-05-02 with broader rollout (see entry above)Source: PiunikaWeb, BuildFastWithAI, xAI release notes, Musk posts on X (2026-04-18) · 2026-04

Best for

People who live on X/Twitter and want an AI that can tap into that data in real-time. Also good for users who find mainstream chatbots too sanitized and want something with more personality.

Not for

Enterprise users who need reliable, consistent outputs. Also not the best pick if you don't use X -- the real-time data advantage disappears and you're left with a solid-but-not-best-in-class LLM.

Our Verdict

Grok has come a long way from being dismissed as Elon's pet project. The Grok 3 models are legitimately competitive, and the real-time X integration is a unique differentiator that no other chatbot can match. But the value proposition gets muddier when you strip away the X angle -- at $30/mo for SuperGrok, you're paying a premium for personality and Twitter data. If those matter to you, Grok is great. If not, Claude or ChatGPT give you more for less.

Sources

  • SpaceXAI: Introducing Grok 4.6 -- $2/$6, full eval table, Cursor + Grok Build day one (2026-08-12) (accessed 2026-08-13)
  • Cursor blog: Grok 4.6 (corroborates launch date, pricing and 2x first-week usage) (accessed 2026-08-13)
  • SpaceXAI: Introducing Grok Bot -- always-on agents, SuperGrok Heavy + Cursor Ultra/Teams Premium (2026-08-11) (accessed 2026-08-13)
  • SpaceXAI: Imagine Image 2.0 -- precise editing, multi-ref, smart resize, #2 on both Arena image boards (2026-08-07) (accessed 2026-08-10)
  • SpaceXAI: Imagine Video 1.5 with References -- multi-reference, voice consistency, text-to-video, native 1080p (2026-07-31) (accessed 2026-08-06)
  • SpaceXAI: Grok Automations (2026-07-16) (accessed 2026-07-18)
  • SpaceXAI: Grok Build open source (2026-07-15) (accessed 2026-07-18)
  • SpaceXAI: Introducing Grok 4.5 (2026-07-08) (accessed 2026-07-09)
  • Cursor blog: Introducing Grok 4.5 (joint training details) (accessed 2026-07-09)
  • Axios: SpaceXAI launches new model, Grok 4.5 (accessed 2026-07-09)
  • xAI May 15 model retirement docs (accessed 2026-05-19)
  • VentureBeat: xAI launches Grok 4.3 with voice cloning (2026-05-02) (accessed 2026-05-05)
  • Winbuzzer: xAI Grok 4.3 + Custom Voices (2026-05-03) (accessed 2026-05-05)
  • xAI official site (accessed 2026-04-17)
  • xAI Grok 4.20 announcement (accessed 2026-04-17)
  • IBTimes: Grok 4.20 Beta 2 April 2026 (accessed 2026-04-17)
  • BuildFastWithAI: Grok 4.3 Beta 2026-04-17 (accessed 2026-04-17)
  • Artificial Analysis: Grok 4.20 (accessed 2026-04-17)
  • Reddit r/grok, r/artificial (accessed 2026-04-17)

The Tier List Tuesday

Weekly newsletter: tier movers, new entrants, and the VS of the week. Built from our daily AI-tool sweeps. No spam, unsubscribe anytime.

Alternatives to Grok

Claude (Anthropic) logo

Claude (Anthropic)

Anthropic's flagship LLM family. **Claude Opus 5 launched 2026-07-24** and is now the default model on Claude Max and the strongest model on Claude Pro -- same $5/$25 per 1M as Opus 4.8, but Anthropic says it lands within 0.5% of Fable 5 on CursorBench at half the cost. Sonnet 5 (June 30) stays the default on Free/Pro at $2/$10 per 1M (intro through Aug 31, then $3/$15), and Fable 5 -- back globally since July 1 after a 19-day export-control suspension -- remains the top of the range at $10/$50

A
8.5/10
Free tierFrom $0
Best writing quality of any LLM -- Opus ...1M token context window for enterprise A...
Updated 2026-08-10
Claude Mythos 5 logo

Claude Mythos 5

Anthropic's unrestricted frontier model -- launched June 9, 2026 alongside Claude Fable 5 (the same model made safe for general use). Suspended June 12 by a US export-control order, then PARTIALLY RESTORED July 1, 2026 (US government lifted controls June 30): Mythos 5 is back for a set of US organizations with government approval, while Anthropic works to re-expand the broader Glasswing program. Public Fable 5 returned globally the same day. Gated to Project Glasswing orgs + select biology researchers.

C
6.5/10
From Invite only
The most capable Anthropic model availab...73% success rate on expert-level Capture...
Updated 2026-07-04
Gemini (Google) logo

Gemini (Google)

Google's LLM with deep Google Workspace integration, 2M token context window, and native code execution -- **Gemini 3.7 Flash launched 2026-08-13** at intro pricing of $0.75/$3.75 per 1M (doubling to $1.50/$7.50 on 2027-01-01), superseding Gemini 3.6 Flash after just three weeks and posting large coding/agent gains (DeepSWE 65.3% vs 49.0%). Gemini 3.5 Pro STILL delayed and partner-testing-only (Bloomberg, 7/16 -- coding shortfalls, no ship date), so Flash keeps shipping while the Pro line stalls. Gemini 4 pre-training underway

A
8.3/10
Free tierFrom $0
2 million token context window is the la...Best Google Workspace integration (Gmail...
Updated 2026-08-13
Muse Spark (Meta) logo

Muse Spark (Meta)

Meta's frontier model line from its Superintelligence Lab -- Muse Spark 1.1 (2026-07-09) adds substantially better coding, 1M-token multi-agent orchestration, and Meta's first paid developer API (Meta Model API, public preview)

A
8.8/10
Free tierFrom $0
Completely free to use via Meta AI app a...Natively multimodal: handles text, image...
Updated 2026-07-18
GPT-Rosalind (OpenAI) logo

GPT-Rosalind (OpenAI)

OpenAI's first domain-specific model -- life sciences, drug discovery, translational medicine. Launched 2026-04-16 as a Trusted Access research preview. Launch partners: Amgen, Moderna, Allen Institute, Thermo Fisher. Paired with a Life Sciences Codex plugin (50+ scientific tool integrations)

C
6.8/10
From Invite only
OpenAI's first named vertical/domain-spe...Launch partners Amgen, Moderna, Allen In...
Updated 2026-04-17
GPT-5.4-Cyber (OpenAI) logo

GPT-5.4-Cyber (OpenAI)

OpenAI's defensive-cybersecurity variant of GPT-5.4, launched 2026-04-16. Lowered refusal boundary for security-research tasks and native binary reverse-engineering. Access gated via Trusted Access for Cyber (TAC) program -- thousands of verified defenders, hundreds of teams, no public pricing

B
7.2/10
From Not publicly disclosed
Directly competes with Claude Mythos Pre...Lowered refusal boundary on defensive-se...
Updated 2026-04-19
Microsoft MAI-Thinking-1 logo

Microsoft MAI-Thinking-1

Microsoft's first in-house reasoning model -- launched 2026-06-02 at Build as the flagship of seven new MAI models. 35B-active / ~1T-total sparse Mixture-of-Experts, 256K context. AIME 2025 97.0%, matches leading models on SWE-Bench Pro, and beat Claude Sonnet 4.6 in human-preference testing. Available on Microsoft Foundry + OpenRouter / Fireworks / Baseten

B
7.5/10
From Not disclosed
Microsoft's first in-house frontier-clas...Strong published reasoning numbers: AIME...
Updated 2026-06-02
Hunyuan 3 (Tencent Hy3) logo

Hunyuan 3 (Tencent Hy3)

Tencent's Hy3 reached GA 2026-07-06 (upgraded from the April preview) -- 295B total / 21B active MoE, 256K context, now Apache 2.0 open weights on HuggingFace + ModelScope with the EU/UK/South Korea restriction lifted. ~90% agent-task completion on Tencent's internal apps; API via Tencent Cloud TokenHub. Integrated into Yuanbao, WeChat, QQ

A
8.1/10
Free tierFrom $0
Open weights from a top-3 Chinese tech c...Pricing is aggressive. ~1.2 RMB per mill...
Updated 2026-07-22
MiMo (Xiaomi) logo

MiMo (Xiaomi)

Xiaomi's MiMo-V2.5 family launched 2026-04-22 -- Pro (1T total / 42B active MoE, 1M context, native vision+audio reasoning), Multimodal base, TTS (3 sub-models: base, VoiceDesign, VoiceClone), and ASR (open-source, English + Chinese + major dialects). Full voice pipeline for the agent era. Extra-charge 1M-context tier removed at launch

A
8.3/10
Free tierFrom $0
Full voice pipeline shipped together: a ...Native multimodal in MiMo-V2.5-Pro is th...
Updated 2026-07-04