GPT-6.1 Sol: Astra-Level Coding at a Fifth of the Price, With Cost at Every Effort Level
OpenAI released GPT-6.1 Sol on September 29, 2026, at $2 input and $10 output per million tokens, a fifth of GPT-6 Astra's price, with cached input at $0.10. On OpenAI's charts it edges past Astra's best DeepSWE score at high effort and trails Astra by 2.1 points on OSWorld offline. The API model ID is gpt-6.1-sol.
OpenAI released GPT-6.1 Sol at DevDay on September 29, 2026, one week after GPT-6 Sol. It keeps GPT-6 Sol's $2 input and $10 output price per million tokens and halves cached input to $0.10. OpenAI says it "nearly matches GPT-6 Astra's intelligence on agentic coding, computer use, and professional work at one-fifth of Astra's standard input and output token prices" (OpenAI). It is gpt-6.1-sol in the API and is in ChatGPT Work and Codex for Plus, Pro, Business, Enterprise, and Edu users.
The verdict: GPT-6.1 Sol is now our default pick among OpenAI models for coding agents, and GPT-6 Astra earns its price only for top-end computer use, long scientific tasks, and the hardest business workflows. On OpenAI's own charts, GPT-6.1 Sol at high effort scores 75.2% on DeepSWE for $0.65 a task, above Astra's best of 74.1% at $4.43. It trails Astra by 2.1 points on OSWorld offline, 5.3 points on AutomationBench, and 11.1 points on Terminal-Bench Science. If you run GPT-6 Sol, switch: GPT-6.1 Sol's best score is better on all six benchmark charts OpenAI published. Check two API changes first. The none effort level is gone, and Chat Completions no longer supports tool calling.
Key Takeaways
- Price. $2 input, $0.10 cached input, and $10 output per million tokens. Cached input is half GPT-6 Sol's $0.20 and a tenth of Astra's $1.
- Coding. GPT-6.1 Sol scores 75.2% on DeepSWE at high effort for $0.65 a task, 6.4 points above GPT-6 Sol's best and above Astra's 74.1%.
- Computer use. GPT-6.1 Sol reaches 71.4% on OSWorld 2.0 offline at max effort for $1.27 a task, 7 points above GPT-6 Sol and 2.1 below Astra, which costs $9.44 a task there.
- Effort. Use high for coding and documents, xhigh or max for business workflows, and max for computer use and science. On DeepSWE, xhigh and max score lower than high and cost more.
- API changes. Effort runs from
lowtomaxwithmediumas the default. Thenonelevel is gone, and Chat Completions works only without tool calling, so agents belong on the Responses API. - Availability. GPT-6.1 Sol is in ChatGPT Work and Codex for Plus, Pro, Business, Enterprise, and Edu users, and not yet in Chat. A GPT-6.1 Sol Ultrafast with up to 8x faster token generation in Codex is coming.
What to Do Now
| If you use | Do this |
|---|---|
| GPT-6 Sol in the API | Switch to gpt-6.1-sol. Replace any none effort setting and move tool calls from Chat Completions to the Responses API. |
| GPT-6 Astra for coding agents | Rerun your evals on GPT-6.1 Sol at high effort. OpenAI's DeepSWE chart puts it ahead of Astra for about a seventh of the cost per task. |
| GPT-6 Astra for computer use or science | Keep Astra where the last few points matter, and test GPT-6.1 Sol at max effort on the rest. |
| Codex or ChatGPT Work on a paid plan | Select GPT-6.1 Sol. Enterprise and Edu workspace owners enable it in Workspace settings > Permissions & roles. |
| Claude Opus 5.5 or Sonnet 5.5 | Compare on your own tasks. The labs share few benchmarks, and their numbers for Opus 5.5 differ. |
Pricing and API Specs
| Item | GPT-6.1 Sol | GPT-6 Sol | GPT-6 Astra |
|---|---|---|---|
| Input (per million tokens) | $2 | $2 | $10 |
| Cached input | $0.10 | $0.20 | $1 |
| Cache writes | $2.50 | $2.50 | $12.50 |
| Output | $10 | $10 | $50 |
| Context window / max input | 1,050,000 / 922,000 | 1,050,000 / 922,000 | 1,050,000 / 922,000 |
| Max output tokens | 128,000 | 128,000 | 128,000 |
| Knowledge cutoff | April 30, 2026 | April 20, 2026 | April 30, 2026 |
Sources: the GPT-6.1 Sol, GPT-6 Sol, and GPT-6 Astra model pages.
On GPT-6.1 Sol, prompts with more than 272K input tokens are priced at 2x input and cache rates and 1.5x output for the full request, so one oversized turn in an agent loop costs more than the headline rate. Batch and Flex cost 50% less than Standard, fast mode costs 2x Standard, and regional processing adds a 10% premium where available (OpenAI model page).
The model takes text and image input and returns text. In the Responses API it supports web search, file search, image generation, code interpreter, hosted shell, apply patch, skills, computer use, MCP, and tool search. It does not support fine-tuning or predicted outputs. Rate limits run from 500 requests and 500,000 tokens per minute at Tier 1 to 15,000 requests and 40,000,000 tokens per minute at Tier 5.
What Changed From GPT-6 Sol
The per-token price stays the same, so the quality gains below come at no extra rate. Four things change for developers, plus one limit to know:
- Cached input halves. It costs $0.10 per million tokens, 5% of the uncached input rate, against $0.20 and 10% on GPT-6 Sol. Agents that resend a long prefix every turn save the most.
- The
noneeffort level is gone. GPT-6.1 Sol supportslow,medium(default),high,xhigh, andmax; thenoneandminimalreasoning efforts are not supported. Change any request that sendsnone. - No tool calling in Chat Completions. GPT-6 Sol allowed function calling in Chat Completions only with
reasoning_effortset tonone. GPT-6.1 Sol supports Chat Completions only without tool calling, so move agent loops to the Responses API. - Newer knowledge. The knowledge cutoff moves from April 20, 2026 to April 30, 2026, the same as Astra.
- Data residency. GPT-6.1 Sol supports US and EU data residency, but fast mode is unavailable with EU data residency.
Benchmarks at Every Effort Level
OpenAI published a score-versus-cost chart for each evaluation, with one point per reasoning effort. The table lists every GPT-6.1 Sol point from the launch page's charts. Costs are OpenAI's estimates per task, and every score is vendor-reported (OpenAI).
| Benchmark | Low | Medium | High | Xhigh | Max |
|---|---|---|---|---|---|
| DeepSWE | 64.4% / $0.17 | 73.0% / $0.42 | 75.2% / $0.65 | 71.9% / $0.79 | 71.9% / $1.57 |
| OSWorld offline | 59.0% / $0.42 | 66.8% / $0.77 | 69.6% / $0.96 | 69.4% / $1.05 | 71.4% / $1.27 |
| AutomationBench | 24.7% / $0.16 | 31.7% / $0.19 | 33.2% / $0.23 | 35.5% / $0.25 | 36.1% / $0.30 |
| GDP.pdf | 27.0% / $0.33 | 30.0% / $0.34 | 32.0% / $0.35 | 31.8% / $0.37 | 31.0% / $0.42 |
| Terminal-Bench Science | 43.7% / $1.79 | 47.6% / $2.34 | 51.1% / $2.76 | 53.7% / $2.89 | 57.0% / $5.47 |
| Factual error rate (lower is better) | 7.7% | 6.3% | 4.5% | 4.1% | 4.6% |
DeepSWE tests long software-engineering tasks in real codebases. OSWorld offline scores long computer-use workflows. AutomationBench runs business workflows across sales, support, finance, and other functions. GDP.pdf asks professional questions about complex PDFs, and Terminal-Bench Science covers data analysis, simulations, and model fitting. The factuality test uses de-identified ChatGPT conversations where users had flagged an earlier model's error, so OpenAI says they are not representative of typical usage, where factual errors are more rare.
Best score for each model, from the same charts:
| Benchmark | GPT-6.1 Sol | GPT-6 Sol | GPT-6 Astra |
|---|---|---|---|
| DeepSWE | 75.2% ($0.65) | 68.8% ($2.74) | 74.1% ($4.43) |
| OSWorld offline | 71.4% ($1.27) | 64.4% ($3.37) | 73.5% ($9.44) |
| AutomationBench | 36.1% ($0.30) | 33.2% ($0.27) | 41.4% ($1.73) |
| GDP.pdf | 32.0% ($0.35) | 28.0% ($0.35) | 32.2% ($1.91) |
| Terminal-Bench Science | 57.0% ($5.47) | 27.6% ($12.18) | 68.1% ($23.80) |
| Factual error rate | 4.1% | 4.5% | 3.9% |
Which Effort Level to Use
Start at high, and go higher only for business workflows, factual answers, computer use, and science. The chart points make the case:
- Coding: high. DeepSWE peaks at high with 75.2% for $0.65. Xhigh and max both drop to 71.9% and cost up to 2.4 times as much.
- Documents: high. GDP.pdf also peaks at high with 32.0%, and every setting costs between $0.33 and $0.42 a task.
- Business workflows: xhigh or max. AutomationBench climbs to 35.5% at xhigh for $0.25 and 36.1% at max for $0.30.
- Computer use and science: max. OSWorld offline reaches 71.4% at max. Terminal-Bench Science reaches 57.0%, but max nearly doubles its cost, from $2.89 at xhigh to $5.47.
- Factual answers: xhigh. The error rate bottoms out at 4.1%.
The default, medium, is already strong. It scores 73.0% on DeepSWE for $0.42, above GPT-6 Sol's best score at any setting.
How Close GPT-6.1 Sol Gets to Astra
Close enough that Astra is hard to justify for coding or document work. On DeepSWE, GPT-6.1 Sol at high beats Astra's best score. On GDP.pdf it scores 32.0% for $0.35 a task, against Astra's 32.2% for $1.91. OpenAI says the factual error rate stays within 1.9 percentage points of Astra's at every tested setting.
The gap is real elsewhere. On OSWorld offline, OpenAI says GPT-6.1 Sol "comes within 2.1 percentage points of Astra's score at maximum reasoning effort at roughly one-seventh the cost per task." On AutomationBench, Astra's best is 41.4% against 36.1%. On Terminal-Bench Science, Astra leads 68.1% to 57.0%, and OpenAI says Astra "should be used for the most difficult scientific research tasks."
No GPT-6.1 Astra shipped at DevDay. TechCrunch reports that OpenAI is not launching one, citing a Wall Street Journal report that OpenAI scrapped the release over safety concerns raised during internal testing (TechCrunch).
GPT-6.1 Sol vs Claude Opus 5.5 and Sonnet 5.5
On OpenAI's charts, GPT-6.1 Sol beats Claude Opus 5.5 on cost everywhere and on score in most places, but OpenAI ran Opus 5.5 with fallback models, and Anthropic's own numbers differ. Opus 5.5 costs $4 per million input tokens, $20 per million output tokens, and $0.20 per million for cache reads, twice GPT-6.1 Sol's rates (Anthropic).
| Benchmark (OpenAI's chart) | GPT-6.1 Sol | Opus 5.5 with fallbacks |
|---|---|---|
| GDP.pdf, best | 32.0% ($0.35) | 28.8% ($0.83) |
| AutomationBench, medium | 31.7% ($0.19) | 29.5% ($0.65) |
| AutomationBench, max | 36.1% ($0.30) | 42.5% ($1.44) |
| Terminal-Bench Science, max | 57.0% ($5.47) | 63.3% ($23.21) |
Opus 5.5 still leads at the top of AutomationBench and on Terminal-Bench Science, at four to five times the cost per task. Anthropic's page reports Opus 5.5 at 40.0% on AutomationBench, from Zapier runs without fallback models that counted safeguard interventions as failures, and 58.7% on Terminal-Bench-Science 0.1 (Anthropic). Treat the cross-vendor numbers as ranges.
Claude Sonnet 5.5 costs the same as GPT-6.1 Sol per token, $2 input and $10 output, with cache reads at $0.20 per million, twice GPT-6.1 Sol's cached rate (Anthropic). Anthropic's Sonnet 5.5 results include Terminal-Bench 4.0, GDPval-AA, and OSWorld 2.1, and none of the benchmarks OpenAI charts for GPT-6.1 Sol appear on Anthropic's page, so no shared benchmark settles it. Run both on your own tasks.
Safety Evaluations
OpenAI says GPT-6.1 Sol shows lower failure rates than GPT-6 Sol on transparency about broken search tools, respecting explicit restrictions, and avoiding unauthorized outcomes during agentic tasks. At max effort, GPT-6.1 Sol failed to disclose a broken search tool 2.1% of the time, against 4.9% for GPT-6 Sol and 1.5% for Astra. It circumvented a warning 23.5% of the time, against 64.4% for GPT-6 Sol and 17.4% for Astra. On a computer-use safety stress test at xhigh, it failed 4.3% of the time, against 17.4% for GPT-6 Sol and 2.4% for Astra. OpenAI observed no attempts to bypass an automated safety reviewer. These tests deliberately target hard situations and do not measure failure rates in typical use (OpenAI).
Availability
- ChatGPT Work and Codex. OpenAI's announcement says GPT-6.1 Sol is available starting today to all Plus, Pro, Business, Enterprise, and Edu users in ChatGPT Work and Codex. The ChatGPT release notes describe a rollout that starts with Pro users and expands to Plus, Business, Enterprise, and Edu (ChatGPT release notes).
- Workspace controls. Enterprise and Edu workspace owners can enable model access in Workspace settings > Permissions & roles.
- Chat. GPT-6.1 Sol is not yet available in Chat.
- API.
gpt-6.1-solworks on the Responses API, Batch, and Chat Completions without tool calling. Realtime, Assistants, and fine-tuning are not supported.
GPT-6.1 Sol Ultrafast
A faster GPT-6.1 Sol is coming, but OpenAI has published no date beyond "in the coming days." The announcement says GPT-6.1 Sol Ultrafast will offer up to 8x faster token generation than standard speed in Codex. OpenAI's DevDay recap calls Ultrafast its premium speed tier, with up to 8x faster token generation in Codex, reaching 300 tokens per second, and up to 6x in the API (OpenAI DevDay recap). The recap says GPT-6.1 Sol Ultrafast is coming soon.
GPT-6 Astra Ultrafast is available today in the API and in ChatGPT Work and Codex on Pro 500 and Enterprise plans. Our Pro 500 and Astra Ultrafast coverage has the plan details. Neither the announcement nor the recap gives Sol Ultrafast's price or plan eligibility. In the API today, GPT-6.1 Sol's fast mode costs 2x Standard.
When to Choose GPT-6.1 Sol
Choose GPT-6.1 Sol at high effort for coding agents and PDF-heavy document work. It matches or beats Astra on those charts for a fraction of the cost per task.
Choose GPT-6.1 Sol at max effort for computer-use agents on a budget. It lands within 2.1 points of Astra on OSWorld offline for $1.27 a task.
Keep GPT-6 Astra for long scientific tasks and the hardest business workflows, where it leads by 11.1 and 5.3 points.
Retire GPT-6 Sol. GPT-6.1 Sol costs the same per token, less for cached input, and posts a better best score on all six benchmark charts OpenAI published. Keep GPT-6 Sol only if you depend on none effort or tool calls through Chat Completions.
Use GPT-6 Luna for bulk subagent work where cost per call matters more than the last few points. It costs $0.10 per million input tokens and $0.50 per million output tokens (GPT-6 Luna model page).
Choose Claude Opus 5.5 for business workflows at max effort, where OpenAI's own chart puts it ahead, if the higher cost per task is acceptable. Test Claude Sonnet 5.5 against GPT-6.1 Sol on your own tasks; they cost the same per token.
DevDay also brought dots, ChatGPT's always-on agents, and the Pro 500 plan with Astra Ultrafast.
Sources
- OpenAI, "Introducing GPT-6.1 Sol" (pricing, benchmark and cost charts, safety evaluations, availability, Ultrafast; September 29, 2026): https://openai.com/index/introducing-gpt-6-1-sol/
- OpenAI, GPT-6.1 Sol model page (effort levels, endpoints, tools, data residency, pricing, context, rate limits): https://developers.openai.com/api/docs/models/gpt-6.1-sol
- OpenAI, GPT-6 Sol model page (effort levels, Chat Completions limit, pricing, cutoff): https://developers.openai.com/api/docs/models/gpt-6-sol
- OpenAI, GPT-6 Astra model page (pricing, context, cutoff): https://developers.openai.com/api/docs/models/gpt-6-astra
- OpenAI, GPT-6 Luna model page (pricing): https://developers.openai.com/api/docs/models/gpt-6-luna
- OpenAI, "DevDay 2026 Recap" (Ultrafast speeds and availability; September 29, 2026): https://openai.com/index/devday-2026-recap/
- OpenAI Help Center, ChatGPT release notes (GPT-6.1 Sol rollout and workspace controls; September 29, 2026): https://help.openai.com/en/articles/6825453-chatgpt-release-notes
- Anthropic, "Introducing Claude Opus 5.5" (Opus 5.5 pricing and benchmarks): https://www.anthropic.com/claude-opus-5-5
- Anthropic, "Introducing Claude Sonnet 5.5" (Sonnet 5.5 pricing and benchmarks): https://www.anthropic.com/claude-sonnet-5-5
- TechCrunch, "OpenAI launches GPT-6.1 Sol, says it nearly matches GPT-6 Astra and costs less" (GPT-6.1 Astra report; September 29, 2026): https://techcrunch.com/2026/09/29/openai-launches-gpt-6-1-sol-says-it-nearly-matches-gpt-6-astra-and-costs-less/
Read next
Related coverage- OpenAI
GPT-6 Sol and Luna: Half the Price of GPT-5.6, With Benchmarks and Cost at Every Effort Level
OpenAI released GPT-6 Sol and GPT-6 Luna on September 22, 2026, at $2 and $10 and $0.10 and $0.50 per million tokens, half the price of GPT-5.6. Sol makes about half as many factual errors as GPT-5.6 Sol. Luna scores 66.6% on DeepSWE for about $0.22 a task. GPT-6 Astra still leads every OpenAI benchmark.
- OpenAI
OpenAI Launches GPT-6 Astra for Computer Use and Complex Work
OpenAI is rolling out GPT-6 Astra, a model focused on computer use, browsing, software engineering, cybersecurity, science, and complex professional workflows. Access starts with a limited set of organizations before expanding to ChatGPT plans and developer platforms.
- Claude
Claude Opus 5.5: Fable 5.1-Level Scores at $4 and $20, With Benchmarks, Effort Costs, and Migration Guide
Claude Opus 5.5, released September 22, 2026, beats Claude Fable 5.1 on every benchmark Anthropic published and costs $4 and $20 per million tokens, 60% less than Fable 5.1. It scores 66.4% on Terminal-Bench 4.0 against 57.9% for GPT-6 Astra. Thinking is always on, the default effort is medium, and four API changes can break Opus 5 code.
- Claude
Claude Sonnet 5.5 Launches With Lower Task Costs and Cursor Support
Anthropic introduced Claude Sonnet 5.5 on September 28 with model ID `claude-sonnet-5-5`. It keeps Sonnet 5's $2 input and $10 output price per million tokens; Anthropic reports output generation more than 30% faster and up to 30% lower cost per task. Cursor lists the model in Settings > Models.
Frequently Asked Questions
What is GPT-6.1 Sol?
GPT-6.1 Sol is OpenAI's upgrade to GPT-6 Sol, released September 29, 2026, at DevDay. OpenAI says it nearly matches GPT-6 Astra on agentic coding, computer use, and professional work at a fifth of Astra's token price. Its API model ID is gpt-6.1-sol.
How much does GPT-6.1 Sol cost?
GPT-6.1 Sol costs $2 per million input tokens, $0.10 per million cached input tokens, and $10 per million output tokens. Cached input is half GPT-6 Sol's rate. Prompts over 272K input tokens cost 2x on input and 1.5x on output, and Batch and Flex are 50% off.
Is GPT-6.1 Sol as good as GPT-6 Astra?
On coding and documents, roughly yes. OpenAI's charts show GPT-6.1 Sol at 75.2% on DeepSWE against Astra's best of 74.1%, and 32.0% against 32.2% on GDP.pdf. Astra still leads OSWorld offline by 2.1 points, AutomationBench by 5.3, and Terminal-Bench Science by 11.1.
What changed from GPT-6 Sol?
Best scores improve on all six benchmark charts OpenAI published, including 75.2% against 68.8% on DeepSWE and 71.4% against 64.4% on OSWorld offline. Cached input drops from $0.20 to $0.10 per million tokens. The none effort level is gone, and Chat Completions no longer supports tool calling.
Which reasoning effort should I use with GPT-6.1 Sol?
Use high for coding and document work, where DeepSWE and GDP.pdf peak. Use xhigh or max for business workflows, max for computer use and scientific tasks, and xhigh when factual accuracy matters most. Medium is the default and already scores 73.0% on DeepSWE for $0.42 a task.
Where can I use GPT-6.1 Sol?
In ChatGPT Work and Codex for Plus, Pro, Business, Enterprise, and Edu users, but not yet in Chat. In the API it runs on the Responses API, Batch, and Chat Completions without tool calling. Enterprise and Edu workspace owners enable it in workspace settings.
What is GPT-6.1 Sol Ultrafast?
It is a faster version that OpenAI says will arrive in the coming days, with up to 8x faster token generation than standard speed in Codex. The announcement does not give its price or which plans get it. GPT-6 Astra Ultrafast is already available on Pro 500 and Enterprise plans.