AI Catchup

Our Top AI Model

Our top AI model is Claude Fable 5, run at medium reasoning effort, with GPT-5.6 Sol at xhigh effort as the co-pick for terminal-heavy, long-horizon coding agents. Fable 5 Max leads the Artificial Analysis Intelligence Index at 62; Sol Max leads DeepSWE v1.1 (73%) and Terminal-Bench v3.0 (34.6%). Grok 4.6 is the price-to-capability pick at $2 per million input tokens.

The Picks

  1. #1

    Claude Fable 5

    Best overall (run at medium effort)

    Anthropic's Mythos-class model made safe for general use, addressable as claude-fable-5 at $10/M input and $50/M output tokens. SpaceXAI's August 12, 2026 launch table shows Fable 5 Max leading the Artificial Analysis Intelligence Index at 62 and CursorBench v3.2 at 70.5%, and Anthropic cites Cognition's FrontierCode evaluation placing Fable 5 highest among frontier models even at medium effort. Run it at medium as your default and step up for the hardest work. Know the caveats: safety classifiers fall back a small share of requests to Opus (under 5% of sessions, per Anthropic), and all Mythos-class traffic carries a mandatory 30-day retention policy.

  2. #2

    GPT-5.6 Sol

    Best for long-horizon coding agents (run at xhigh)

    OpenAI's flagship, described by OpenAI as a step function better than GPT-5.5. In SpaceXAI's August 12 launch table, Sol at Max effort leads DeepSWE v1.1 at 73% and Terminal-Bench v3.0 at 34.6%, the two rows that matter most if your workload is long-horizon software engineering in a terminal. API pricing is $5/M input and $30/M output tokens at short context. Sol exposes reasoning effort from none through max; run agent work at xhigh, and hold max for when quality outweighs latency and cost.

  3. #3

    Grok 4.6

    Best price-to-capability

    SpaceXAI's August 12, 2026 release matches GPT-5.6 Sol Max on the Artificial Analysis Intelligence Index (61 vs 61) at $2/M input and $6/M output tokens below 200,000 prompt tokens. It is live in Cursor and Grok Build with 2x included usage for the first week. Terminal-Bench v3.0 (26%) is the caution row: bake it off on your own terminal-heavy tasks before making it a default.

  4. #4

    Gemini 3.1 Pro

    Best long context

    Handles massive documents and codebases with multi-million-token windows. Still leads ARC-AGI-1 Verified (98.0%) and BrowseComp (85.9%). Ideal for tasks that require ingesting and reasoning over entire repositories or long documents.

  5. #5

    Claude Sonnet 5

    Best value

    Anthropic's mid-tier and its most agentic Sonnet yet: it plans, drives browsers and terminals, and runs autonomously at a level Anthropic says recently required larger models. Pricing is permanent at $2/M input and $10/M output tokens, and it is the default model for Free and Pro plans. The smart choice for high-volume work where cost and latency matter more than peak capability.

  6. #6

    DeepSeek R1

    Best open-weight

    Competitive reasoning at lower cost, fully open. The leading option for teams that need to self-host or want complete transparency into model weights.

How They Compare

EvalClaude Fable 5 (Max)GPT-5.6 Sol (Max)Grok 4.6 (High)
AA Intelligence Index626161
CursorBench v3.270.5%67.2%69.9%
DeepSWE v1.170%73%65.9%
FrontierCode v1.1 (Extended)63.6%60.6%61.3%
APEX-Agents59.2%56.7%57.5%
Terminal-Bench v3.034.1%34.6%26%

Figures are from SpaceXAI's Grok 4.6 launch table (August 12, 2026), which reports the best of self-reported or publicly available results for competitors. Scores are at each model's listed effort setting, not our recommended day-to-day settings.

Changelog

  • August 12, 2026: Claude Fable 5 promoted to the top pick, recommended at medium reasoning effort, with GPT-5.6 Sol at xhigh as the co-pick for terminal-heavy, long-horizon coding agents. Benchmarks re-based on SpaceXAI's August 12 Grok 4.6 launch table. Grok 4.6 added at #3 for price-to-capability; Claude Sonnet 5 replaces Sonnet 4.6 as the value pick; GPT-5.5, GPT-5.5 Pro, and Claude Opus 4.7 rotate out, superseded by GPT-5.6 Sol and Fable 5.
  • April 23, 2026: GPT-5.5 promoted to top overall model on its April 23 launch. State-of-the-art on Terminal-Bench 2.0 (82.7%), GDPval (84.9%), OSWorld-Verified (78.7%), FrontierMath Tier 4 (35.4%), and CyberGym (81.8%). Matches GPT-5.4 per-token latency while delivering higher intelligence and using fewer tokens on equivalent Codex tasks. Claude Opus 4.7 moves to #2 as the pick for SWE-Bench Pro-style work. GPT-5.5 Pro added as a separate tier for the hardest questions.
  • April 16, 2026: Claude Opus 4.7 promoted to top model on its April 16 launch. Notable gains on the hardest software engineering tasks, state-of-the-art on Finance Agent and GDPval-AA at that time, plus higher-resolution vision (2,576 pixels on the long edge) and a new xhigh effort level. Pricing unchanged from Opus 4.6 at $5/M input and $25/M output.
  • March 2026: Initial picks published. Claude Opus 4.6 selected as top model.

Frequently Asked Questions

Why is Claude Fable 5 the top pick?

It leads the Artificial Analysis Intelligence Index at 62 and CursorBench v3.2 at 70.5% in SpaceXAI's August 12, 2026 launch table, and Anthropic cites Cognition's FrontierCode evaluation placing it highest among frontier models even at medium effort. That efficiency at medium effort is what wins it the default slot.

When should I use GPT-5.6 Sol instead?

When your workload is long-horizon software engineering or terminal-heavy agent loops. Sol at Max effort leads DeepSWE v1.1 at 73% and Terminal-Bench v3.0 at 34.6% in the same launch table. We recommend running it at xhigh effort, reserving max for accuracy-critical work.

What reasoning effort should I set?

Run Fable 5 at medium: Cognition's FrontierCode evaluation shows it highest among frontier models even there, and higher effort costs more tokens and latency. Run GPT-5.6 Sol at xhigh for agent work; Sol's effort scale goes none, low, medium, high, xhigh, then max.

How often does the top model change?

We re-evaluate whenever a major model release changes the landscape. GPT-5.5 held the top slot from April 23 to August 12, 2026, when Claude Fable 5 took it, with GPT-5.6 Sol named as the agent co-pick in the same review.

Get the weekly AI Catchup

Tools, practices, and what matters, in your inbox every week.