AI Tool Tier List · Updated Daily

AI tools, ranked S to F.

Every tool tested, scored, and placed in its tier. We report the bugs, show the pricing traps, and tell you what actually works. No sponsored placements.

S
9.0+
A
8.0
B
7.0
C
6.0
D
5.0
F
<5
142
Tools ranked
23
Categories
10,011+
Comparisons
Daily
Updates

01 The Tier List

Top picks across 5 categories.

142 tools ranked
Last updated: Aug 31, 2026

03 Latest Reviews

Recently reviewed and updated.

See trending →
Claude (Anthropic) logo

Claude (Anthropic)

Anthropic's flagship LLM family. **Claude Opus 5 launched 2026-07-24** and is now the default model on Claude Max and the strongest model on Claude Pro -- same $5/$25 per 1M as Opus 4.8, but Anthropic says it lands within 0.5% of Fable 5 on CursorBench at half the cost. **Sonnet 5's $2/$10 per 1M is now permanent** -- Anthropic cancelled the 2026-09-01 rise to $3/$15 and made the launch rate standard -- and it stays the default on Free/Pro. Fable 5 -- back globally since July 1 after a 19-day export-control suspension -- remains the top of the range at $10/$50. **From 2026-08-14 future Claude models watermark their text output globally** (SynthID-Text; no extra tokens, no price or speed change, no identifying information -- detector API not shipped yet), and the **legacy Workbench plus the experimental prompt-tools APIs retired 2026-08-17**

A
8.5/10
Free tierFrom $0
Best writing quality of any LLM -- Opus ...1M token context window for enterprise A...
Updated 2026-08-31
Grok logo

Grok

SpaceXAI's irreverent chatbot with a direct line to X/Twitter -- and now **Grok 4.6 (launched 2026-08-12)**, focused on long-running agents and interactive/visual work, at the same $2/$6 per 1M as Grok 4.5 (fast variant 2x). xAI says it matches GPT-5.6 Sol on the AA Intelligence Index (61), though Sol still leads it on DeepSWE and Terminal-Bench. **Grok Bot** (8/11, early beta) adds always-on agents with their own cloud computer. **Grok 4.6 reached GitHub Copilot on 2026-08-14** across all five paid Copilot SKUs (off by default for orgs), adding to its day-one Cursor availability. Grok 4.3 remains the value tier at $1.25/$2.50

B
7.5/10
Free tierFrom $0
Real-time access to X/Twitter data is ge...Grok 3 benchmarks are competitive with G...
Updated 2026-08-31
GPT-5.4-Cyber (OpenAI) logo

GPT-5.4-Cyber (OpenAI)

OpenAI's defensive-cybersecurity variant of GPT-5.4, launched 2026-04-16. Lowered refusal boundary for security-research tasks and native binary reverse-engineering. Access gated via Trusted Access for Cyber (TAC) program -- thousands of verified defenders, hundreds of teams, no public pricing. On **2026-08-17 OpenAI published its first dedicated post on the Hugging Face incident**, conceding it 'underestimated the real-world cyber capabilities of our AI models' and confirming it now releases cyber capabilities only to trusted defenders

B
7.2/10
From Not publicly disclosed
Directly competes with Claude Mythos Pre...Lowered refusal boundary on defensive-se...
Updated 2026-08-31
Mistral AI logo

Mistral AI

European AI lab with open and commercial models -- Le Chat is now **Vibe** (May 28 2026): one agent across Work Mode + Code Mode with a VS Code extension and CLI, powered by Mistral Medium 3.5 (128B dense, 256k context, 77.6% SWE-Bench Verified). Newest release: **Shieldstral 1.0** (Aug 4 2026), a 3B Apache 2.0 multimodal safety classifier that runs on one 16GB GPU. Earlier 2026 line: Small 4 (119B MoE Apache 2.0), Voxtral TTS. **Mistral Medium 3.1 and Medium 3 were both retired 2026-08-31** -- Medium 3.5 is the vendor-listed replacement for each

B
7.5/10
Free tierFrom $0
Mistral Medium 3.5 (April 29 2026) is Mi...Vibe Remote Agents (also 4/29) lets you ...
Updated 2026-08-31
GitHub Copilot logo

GitHub Copilot

AI code assistant that lives in your editor -- autocomplete on steroids, now with the broadest model picker of any coding tool. **Gemini 3.7 Flash landed 2026-08-13 and Grok 4.6 on 2026-08-14**, both across Pro, Pro+, Max, Business and Enterprise and both still off by default for orgs. **Claude Opus 5 landed 2026-07-24** (Pro+, Max, Business, Enterprise) and **Grok 4.5 on 2026-07-28** (all five paid SKUs, up to 500K context, text and image input, low/medium/high reasoning effort). Both bill usage-based at provider list price, and both are off by default for Business/Enterprise until an admin enables the policy -- though that posture flips on **2026-08-26**, when new GA models covered by GitHub's data-retention agreement start auto-enabling for orgs. **MAI-Code-1.1-Flash landed 2026-08-11** (native vision, 0.25x premium-request multiplier, 73% below the model it replaces; automatic for Free/Student, manual elsewhere, off by default for orgs) and **MAI-Code-1-Flash retires 2026-09-10**. **GitHub Models (the separate free playground) was fully retired 2026-07-30.** Usage-based billing went live 2026-06-01 with AI Credits and token metering; code completions are still free; new signups for Student/Pro/Pro+/Max remain PAUSED

A
8.3/10
Free tierFrom $0
Inline code completions feel magical -- ...Works directly in VS Code, JetBrains, Ne...
Updated 2026-08-31
Cursor logo

Cursor

AI-native code editor, agent-first in Cursor 3 -- **now a SpaceX subsidiary as of 2026-08-14, when the April acquisition option closed**. **OpenAI is cutting it off: on 2026-08-28 OpenAI notified SpaceX it will wind down the contract supplying OpenAI models to Cursor, proposed shutoff 2026-11-12, and says it is providing no future models to Cursor in the meantime** -- so the model-neutrality pitch is narrowing. Still home to **Grok 4.6 on day one (2026-08-12)**, following the Grok 4.5 model Cursor trained jointly with SpaceXAI on trillions of Cursor tokens (still $2/$6 per 1M, fast variant 2x, desktop/web/iOS/CLI/SDK), with Composer 2.5 as the fast lower-cost tier. Since 2026-08-11 the **Cursor Ultra and Teams Premium tiers also entitle you to xAI's Grok Bot**

A
8.3/10
Free tierFrom $0
Cursor 3's agent-first redesign (April 2...Composer 2 is Cursor's own frontier codi...
Updated 2026-08-31

04 Reviews you can actually trust

Every review is based on hands-on testing, cross-referenced user sentiment from G2, Reddit, and Capterra, and real pricing data. We report known bugs. We don't do paid placements.

No paid placements
Real bug reports
Updated daily