AI Catchup

The AI Catchup - September 22, 2026

By 12 min read

Welcome back. Both labs shipped a new model today. Anthropic released Claude Opus 5.5, the first model in its Claude 5.5 family. The same day, OpenAI released GPT-6 Sol and GPT-6 Luna, the mid-tier and budget models under GPT-6 Astra. All three launches are pitched on the same promise: frontier-class results for much less money per task.

So we did the comparison the launch posts avoid. Both vendors publish cost-per-task charts at every effort level, and we pulled every data point from both pages and lined them up. Below: the price sheet, the benchmarks, our verdict on performance per dollar, and how each new model stacks up against GPT-6 Astra and Claude Fable 5.1.

Let us get into it.

At a Glance

What changedWhat you can do now that you could not beforeCan you use it today?
Claude Opus 5.5 is live as claude-opus-5-5Get a model that beats Fable 5.1 on every row of Anthropic's table, at $4 and $20 per million tokens instead of $10 and $50Yes, on the Claude Platform, AWS, Google Cloud, and Azure, and as the Opus default in Claude Code
GPT-6 Sol replaces GPT-5.6 Sol at half the priceRun OpenAI's mid-tier model at $2 in and $10 out, with half the factual errors of its predecessorYes, as gpt-6-sol in the API, and in ChatGPT Work and Codex on paid plans, rolling out through the day
GPT-6 Luna replaces GPT-5.6 Luna at half the priceRun a model that scores 66.6% on DeepSWE for about $0.22 a taskYes, as gpt-6-luna; Free and Go users get it in the desktop app
Claude Code limits go upPro, Max, and Team five-hour limits rise 20%, and you get a rate-limit reset you can bankYes. Use the reset from Settings, Usage on the web or desktop app
GPT-6 prompt caching improvesChange reasoning effort or toggle tools mid-conversation without losing the cache, and set explicit cache breakpointsYes, in the OpenAI API

Our Verdict: Who Wins on Performance Per Dollar

Opus 5.5 at medium effort is the best value at the frontier. GPT-6 Sol is the best value for high-volume work. GPT-6 Luna is the cheapest model that can still write real code. That is how we read every cost-per-task data point the two vendors published today.

The case for Opus 5.5 is that you rarely need to turn it up. At medium effort, its default in Claude Code, it scores 1576 Elo on GDPval-AA v2.1 at an estimated $0.86 a task. That beats GPT-6 Astra at max effort, which scores 1542 at $4.53. On FrontierCode, which grades whether a coding agent's changes are ready to merge, Opus 5.5 at medium scores 54.6% at $0.80 a task. That is above Astra's best (53.3%) and above GPT-6 Sol's best (49.3%, at $2.14 by OpenAI's estimate). On Terminal-Bench 4.0, medium Opus 5.5 (57.6%, $2.94) matches Astra at high effort (57.9%, $7.21).

The case for Sol is that cheap tokens stay cheap when the job is routine. On AutomationBench, which tests business workflows across 47 tools, Sol at high effort scores 31.2% for $0.24 a task. Opus 5.5 at high effort scores 32.0% for $0.70 a task. That is the same result for about a third of the cost. Opus 5.5 goes further at max effort (40.0% at $1.37), so Sol wins on value, not on ceiling.

Luna is a different tier altogether. At max effort it scores 66.6% on DeepSWE, which tests long software-engineering tasks in real codebases, for about $0.22 a task. Sol at max effort scores 68.8% for $2.74 a task. Luna is 2.2 points behind for about a twelfth of the price.

One caveat changes how you should read the headlines. Each vendor estimates cost per task with its own method, and on one benchmark the two estimates disagree badly. OpenAI's chart puts Claude Opus 5 at max effort on AutomationBench at $3.05 a task. Anthropic's chart, using Zapier's leaderboard, puts the same run at $1.27. That gap is where OpenAI's claim that Sol costs "9% of Opus 5's cost per task" comes from. On Anthropic's figures it is about 21%. The two vendors do agree on GPT-6 Astra, within 5% on both benchmarks they both chart. Treat any cross-vendor cost multiple as a range, not a number.

The Price Sheet

Per-token price is where Sol and Luna win outright. Opus 5.5 costs twice as much as Sol per token, but both charge $0.20 per million for cached input, which is most of the bill in long agent sessions.

ModelInput ($/M)Cached input ($/M)Output ($/M)API ID
GPT-6 Luna$0.10$0.01$0.50gpt-6-luna
GPT-6 Sol$2$0.20$10gpt-6-sol
Claude Opus 5.5$4$0.20$20claude-opus-5-5
GPT-6 Astra$10separate rate$50gpt-6-astra
Claude Fable 5.1$10$0.25$50claude-fable-5-1

OpenAI says Sol and Luna are 50% cheaper than GPT-5.6 Sol ($4 and $20) and GPT-5.6 Luna ($0.20 and $1.20), and the cached rate is a 90% discount on input. Opus 5.5 is 20% cheaper than Opus 5 on input and output, and 60% cheaper on cache reads; Anthropic says that works out to about 40% less per task on typical workloads. Opus 5.5 Fast mode costs $8 and $40 for up to 2.5x speed. Cache writes on Opus 5.5 cost $5 per million.

The Benchmarks, Side by Side

Anthropic and OpenAI mostly publish different benchmarks, so there are two tables here. Both show the best score each vendor reported for each model, and every number is vendor-reported.

Anthropic's table (Opus 5.5 at max effort unless noted, with production safeguards on):

BenchmarkOpus 5.5Fable 5.1GPT-6 AstraGPT-5.6 Sol
Terminal-Bench 4.066.4% (xhigh)55.8%57.9% (high)37.3%
FrontierCode v1.1 Main54.4%50.3%53.3%47.5%
CursorBench 4.057.8%51.8%not reported41.7%
GDPval-AA v2.1 (Elo)1846173515421588
AutomationBench40.0%31.4%41.4%28.8%
Humanity's Last Exam, with tools67.7%65.6%57.2%not reported
Terminal-Bench-Science 0.158.7%52.6%64.6%22.4%
OSWorld 2.0, partial81.8%80.7%not reportednot reported
Chartography, with tools89.0%88.4%not reportednot reported

OpenAI's charts (best score per model, with OpenAI's cost-per-task estimate):

BenchmarkGPT-6 AstraGPT-6 SolGPT-6 LunaGPT-5.6 Sol
AutomationBench 1.0.641.4% ($1.73)33.2% ($0.27)20.7% ($0.04)28.8% ($0.67)
Agents' Last Exam V159.3% ($6.23)56.4% ($2.93)50.9% ($0.15)53.6% ($5.08)
FrontierCode 1.1 Main53.3% ($4.59)49.3% ($2.14)42.4% ($0.11)47.5% ($5.19)
DeepSWE 1.174.1% ($4.43)68.8% ($2.74)66.6% ($0.22)72.7% ($6.46)
OSWorld 2.0 offline, partial73.5% ($9.07)64.4% ($3.25)52.7% ($0.27)66.2% ($7.71)
Factual error rate (lower is better)3.9%4.5%7.6%8.4%

Our full Opus 5.5 article has its score and cost at every effort level. Sol and Opus 5.5 appear together on only two benchmarks, FrontierCode and AutomationBench, and Opus 5.5 scores higher on both (54.4% vs 49.3%, and 40.0% vs 33.2%). Nobody has published Sol or Luna on Terminal-Bench 4.0, CursorBench, or GDPval, or Opus 5.5 on DeepSWE or Agents' Last Exam. The two OSWorld numbers use different settings, so do not compare them.

A detail OpenAI's post does not mention: at top effort, GPT-6 Sol scores below GPT-5.6 Sol on DeepSWE (68.8% vs 72.7%) and on OSWorld offline (64.4% vs 66.2%). Sol's upgrade is price, reliability, and business-workflow scores, not raw coding ceiling. Below the top it is much cheaper for the same score: on OSWorld offline, Sol at xhigh scores 60.5% for $2.21 a task, where GPT-5.6 Sol needed $5.93 for 60.9%. That is a real gain, but it is not the same as a smarter model.

Opus 5.5 vs GPT-6 Astra and Fable 5.1

Opus 5.5 beats Astra on four of the six benchmarks they share, at 40% of Astra's per-token price. The clear wins are Terminal-Bench 4.0 (66.4% vs 57.9%), GDPval-AA (1846 vs 1542), and Humanity's Last Exam (67.7% vs 57.2%). FrontierCode is close (54.4% vs 53.3%). Astra keeps AutomationBench (41.4% vs 40.0%) and Terminal-Bench-Science (64.6% vs 58.7%), though Anthropic lists the science benchmark's standard error at 3.5 to 5 points. OpenAI still calls Astra "the world's best model for computer use" and leads its own DeepSWE, OSWorld, and Agents' Last Exam charts. Opus 5.5 has not been scored on any of those three, so that question stays open.

Against Fable 5.1, Opus 5.5 wins every row at 40% of the price. Anthropic says Opus 5.5 "performs at the level of Claude Fable 5.1 on most work" and adds that in its own use the gap is narrower than the scores suggest. The cost difference is not narrow. In Anthropic's internal test translating HAProxy from C to Rust, both models passed nearly all of HAProxy's regression tests. Opus 5.5 finished in 9.5 hours against 12 for Fable 5.1, at 51% less cost. On Terminal-Bench 4.0, medium Opus 5.5 (57.6%, $2.94) beats Fable 5.1 at max effort (55.8%, $19.50). If you moved to Fable 5.1 on September 1 for coding, run your workload on Opus 5.5 this week.

Partner reports agree. Optiver says Opus 5.5 "matched Opus 5's quality in about half the turns," cutting that workload's cost by 40 to 50%. Deloitte says it caught 72% of known bugs at its lowest effort, against 56% for Opus 5 at high effort. Box measured a third of the tokens Opus 5 used.

GPT-6 Sol and Luna vs GPT-6 Astra

Astra is still OpenAI's best model on every chart, and Sol gets you most of the way for a fifth of the token price. Sol trails Astra by 3 points on Agents' Last Exam, 4 on FrontierCode, 5 on DeepSWE, 8 on AutomationBench, and 9 on OSWorld offline. On factual errors the two are close: 4.5% for Sol against 3.9% for Astra, both at the top of their range.

Where Sol clearly beats Astra is when the effort-cost trade-off favors it. On AutomationBench, Sol at xhigh (33.2%, $0.27) beats Astra at low (30.3%, $1.08) at a quarter of the cost. OpenAI's post says to "choose it [Astra] when you want the best results." We would put it more narrowly: use Astra for computer use and the hardest long tasks, and default to Sol for everything else.

Luna is the model the price cut matters most for. It costs a hundredth of Astra per token. OpenAI says that at higher effort it matches GPT-5.6 Sol on factuality for about a hundredth of the cost. And it scores within 8 points of Astra on DeepSWE. For subagents, bulk classification, and first-pass code changes, it is now the default to test first.

OpenAI also reports alignment gains carried over from Astra. In its coding-deception evaluation, deliberately built to make models dishonest, Sol's deception rate falls from 10.4% to 1.3%, and Luna's from 9.5% to 2.8%. Astra is at 0.5%.

What Changes in Your Tools

Claude Code. Opus 5.5 is now the default model on Pro, Max, Team, Enterprise, and API setups, and needs Claude Code 2.1.280 or later. Anthropic is raising five-hour limits on Pro, Max, Team, and seat-based Enterprise plans. ClaudeDevs puts the increase at 20% and says the lower model price makes limits go about 25% further. Thinking is always on, the default effort is Medium, and API users can switch with /claude-api migrate. Cybersecurity work beyond routine bug-fixing is rerouted to Opus 4.8, the same kind of safeguard Fable 5.1 uses. Anthropic says Claude Sonnet 5.5 and Haiku 5.5 follow in the coming weeks.

Codex and ChatGPT Work. Sol and Luna are rolling out today in ChatGPT Work and Codex for Plus, Pro, Business, Enterprise, and Edu, but not yet in Chat. OpenAI's developer announcement also mentions a banked reset for Plus, Pro, and Business accounts.

GitHub Copilot. Sol is rolling out on Pro+, Max, Business, and Enterprise, and Luna on those plus Pro, across VS Code, JetBrains, Xcode, the Copilot CLI, and the cloud agent.

The OpenAI API. GPT-6 prompt caching now survives changes in reasoning effort and turning tools on or off. Explicit breakpoints let you choose where a cached prefix ends. GitHub reports that the caching changes cut the share of Copilot prompt tokens needing fresh processing by more than 50%.

Ship It This Week

Move Opus 5 and Fable 5.1 coding work to claude-opus-5-5 at medium effort. Compare cost per completed task, not per token. Raise effort only on tasks that fail.

Route volume work to gpt-6-sol at high effort. For CRM updates, ticket triage, and multi-app workflows, it matches Opus 5.5 at high effort on AutomationBench for about a third of the cost.

Try gpt-6-luna as your subagent model. At $0.10 and $0.50 per million tokens, a failed run costs almost nothing, so let your evals decide.

Do not assume GPT-6 Sol beats GPT-5.6 Sol on hard coding. On DeepSWE and OSWorld, the old model at max effort still scores higher. Re-run your own coding evals before you switch a pinned model.

Spend the banked resets wisely. Use it from Settings, Usage on the web or desktop app, not the terminal. Save it for a migration or refactor day, not a routine one.

Quick Hits

  • Claude Code Projects run parallel cloud sessions. A project gives Claude one coordinating conversation that hands work to cloud sessions, which keep going after you close your laptop.
  • Claude Code reads AGENTS.md. Version 2.1.277 falls back to AGENTS.md when a project has no CLAUDE.md, which makes repos shared with Codex users easier.
  • Claude Design, Slides, and Docs work inside Claude Code. Point it at an RFC and ask for a review deck or UI mockup without leaving the terminal.
  • Cowork and chat are becoming one Claude. The merge starts with Pro and Max users on web, desktop, and mobile.
  • OpenAI launched Astra for Law. It is a legal build of GPT-6 Astra with a legal search index and 26 partner plugins, for selected U.S. firms.
  • OpenAI will publish misalignment reports. A new disclosure framework launched with six reports from training and evaluation.

That is it for this issue. Three models, one theme: frontier performance got cheaper overnight, from both labs, on purpose. The deciding question is no longer which model is smartest. It is which effort setting on which model clears your bar for the least money. If a teammate is still paying Fable 5.1 or Astra prices for everyday coding work, forward this.

Until next week, stay caught up.

Get the weekly AI Catchup

Tools, practices, and what matters, in your inbox every week.