Prompting GPT-6 Astra and GPT-6.1 Sol: Effort, Migration, and Prompts That Keep It Working
GPT-6 Astra follows instructions more closely than earlier OpenAI models and asks more questions, so prompting it is mostly removing old scaffolding: grant permission to finish, trim skills and AGENTS.md, and stop ordering tests. In the API, replace none effort with low, move tool calls to Responses, and switch caching to prompt_cache_options.ttl.
GPT-6 Astra needs less prompting than GPT-5.6, not more. OpenAI says Astra is stronger at instruction following and more sensitive to instructions in skills and files such as AGENTS.md, and it is more likely to stop and ask. Most of the work is deleting instructions written for weaker models and telling Astra plainly when a task is finished. This guide distills OpenAI's Using GPT-6 guide and its skills and prompts post for Astra.
For the models themselves, see our coverage of the GPT-6 Astra launch and GPT-6.1 Sol. For cross-model fundamentals, start with our prompt engineering playbook.
Key Takeaways
- Replace
noneeffort withlow. GPT-6 Astra and GPT-6.1 Sol rejectnone; Astra answers it with an HTTP 400 error. - Set Astra's effort explicitly. OpenAI lists Astra's effort values but does not document a default.
- Move tool calls to the Responses API. Astra and GPT-6.1 Sol accept Chat Completions requests, but tool calling requires Responses.
- Switch prompt caching to
prompt_cache_options.ttl. Coming from GPT-5.5 or earlier, it replacesprompt_cache_retention. - Grant permission to finish. Astra asks more questions than earlier models; tell it which work needs no approval.
- Audit skills and
AGENTS.md. Astra takes them more literally, so cut long descriptions, recipes, and test mandates, and soften ask-first language. - Define completion up front. Astra may return after a first implementation unless you say what done means.
Which GPT-6 Model and Effort to Use
Start on GPT-6.1 Sol for coding agents and move a task to GPT-6 Astra only when your own evals show the gap is worth it. OpenAI describes GPT-6.1 Sol as near-Astra performance at a lower cost. Use GPT-6 Luna for focused, high-volume work. The API constraints differ by model:
| GPT-6 Astra | GPT-6.1 Sol | GPT-6 Luna | |
|---|---|---|---|
| Model ID | gpt-6-astra | gpt-6.1-sol | gpt-6-luna |
| Effort values | low, medium, high, xhigh, max | low, medium, high, xhigh, max | none, low, medium, high, xhigh, max |
| Default effort | Not documented | medium | medium |
| Tool calls in Chat Completions | No, use Responses | No, use Responses | Only with reasoning_effort: "none" |
| Fast mode with EU data residency | No | Yes | Yes |
| Ultrafast data residency | US and global processing only | US, EU, and global processing | Not listed for Luna |
| Input / output per 1M tokens | $10 / $50 | $2 / $10 | $0.10 / $0.50 |
Sources: OpenAI's GPT-6 Astra, GPT-6.1 Sol, and GPT-6 Luna model pages, the reasoning guide, and the Ultrafast guide.
Two details catch teams out. Astra's default effort is absent from its model page, the GPT-6 guide, and the reasoning guide, which gives medium as the default only for the other GPT-6 models: GPT-6.1 Sol, GPT-6 Sol, and GPT-6 Luna. Set reasoning.effort on every Astra request. And Fast mode for GPT-6 Astra does not include a latency SLA, so do not plan an EU or latency-bound deployment around Astra Fast.
Migrating From GPT-5.6 or GPT-6 Sol: The API Checklist
The migration is seven checks, and OpenAI's own Codex skill can apply them for you. Work through them in this order:
- Change the model. Set
modeltogpt-6-astra,gpt-6.1-sol, orgpt-6-luna. - Fix effort. On Astra and GPT-6.1 Sol, replace
nonewithlow. If you usedminimal, start withlowand compare results on representative tasks. The field isreasoning.effortin Responses andreasoning_effortin Chat Completions. - Move tool calls to Responses. Astra and GPT-6.1 Sol support Chat Completions only without tools.
- Drop sampling parameters. When effort is not
none, removetemperature,top_p, andtop_logprobs. In Chat Completions also removelogprobs; in Responses removemessage.output_text.logprobsfrominclude. - Update caching. Coming from GPT-5.5 or earlier, replace
prompt_cache_retentionwithprompt_cache_options.ttlset to"30m". That value is the only one supported and the default, and a cached prefix stays reusable for 30 minutes after its last write or reuse. Budget for cache writes, which cost 1.25 times the uncached input rate on GPT-5.6 and later. - Change effort without breaking the cache. If you raise or lower effort between turns, add a
configuration_updateinput item and leave request-levelreasoning.effortalone. OpenAI supports this in standard, single-agent mode, and it changes only reasoning effort. - Let Codex do it. OpenAI's guide gives a one-line Codex prompt,
$openai-docs migrate this project to the GPT-6 model family, which uses the OpenAI Docs skill. You can download the same skill for other coding agents.
If you are still on gpt-5.6-sol and depend on none effort, it remains a live model that accepts none; our GPT-5.6 prompting guide covers it.
Why Astra Stops to Ask, and How to Keep It Working
Astra asks more because it is designed to. OpenAI says the model is more likely to ask a question when extra input could materially change the result, which can make it stop where you expected it to make reasonable assumptions and persist. Fix this with permission, not pressure.
OpenAI's guide offers three starting prompts, which you should copy from the initiative and follow-through section and adjust:
- Infer intent and persist. Tell the model to work out the task's scope from the conversation, then keep going until the goal is met, taking reversible steps such as worktrees, merge-conflict fixes, and draft PRs on its own. OpenAI's version tells it to "bias towards action and carry the user's intended task to completion."
- Treat requests as instructions. A request phrased as a question about ability should start the work, not end with a plan, an offer, or a partial answer.
- Approve results, not intentions. Ask the model to finish everything already authorized so your approval is the last step before a deploy, merge, or publish, and to skip unsolicited warnings and checklists about hypothetical risk.
Tune all three to the autonomy your product allows. A customer-facing assistant may want the questions; a coding agent in a disposable worktree usually does not.
Define Completion Before You Start
Say what done looks like in the request itself. OpenAI's skills post warns that Astra can reach a first implementation and come back for review while work remains, and that a rule to stop for review after the first implementation pulls it toward stopping early. If the job includes running the result, inspecting it, and fixing what fails, put all three in the request.
Thomas Ricouard's architectural visualization post shows the shape. His request asked Astra to back up the scene, build the larger house in Blender, and keep iterating until the house, furniture, and garden were detailed and the geometry was correct, with the Unreal Engine 5 export as the named next step. That is a finish line, not a first draft. If you want exploration beyond a first pass, say what to explore and where to stop.
Skills and AGENTS.md: Audit What Astra Now Takes Literally
Audit every skill and instruction file your agent can read. OpenAI's guide says Astra can be more sensitive to instructions in skills and AGENTS.md, and that unclear or conflicting guidance in a skill can make it pause and block work early. OpenAI's skills post gives the specific cleanups:
| Old habit | What to do for Astra |
|---|---|
| Long skill descriptions that trigger on a whole domain | Keep descriptions as short as possible and say exactly when the skill applies. A migration skill should fire when you add or change a migration, not whenever you touch a database. |
One long SKILL.md covering several workflows | Make the root file a minimal router that points to supporting docs and scripts, so the model reads only what the task needs. |
| Step-by-step recipes | Cut them back. Overly specific guidance can now hinder results, and guidance written for Sol or Luna may overconstrain Astra. |
An AGENTS.md rule to read several docs before every edit | Point to each doc by situation: one for service boundaries, one for schema changes, one for deployments. |
| "Always run the tests" | Remove it. Astra checks its work on its own, so the old line can cause unnecessary testing. |
| Strong ask-first language for older models | Soften it, and grant explicit permission for workflows you know are safe, such as a local test suite on disposable fixtures. |
The skills post also notes that when you load too many skills, Codex starts shortening their descriptions to fit, so the model sees less of each one. Fewer, sharper skills beat a large library.
Two prompts from OpenAI's guide help when instructions still collide. State that the user's explicit instructions take precedence over a skill's guidelines. And ask the model, whenever a skill makes it pause, ask permission, or leave work unfinished, to name and link the exact SKILL.md file, quote the instruction, and say whether it is an explicit requirement or its own interpretation. That second prompt is the fastest way to find the silent conflicts in a large instruction set.
Writing Style and Formatting
Specify the output style you want, because Astra's default leans on lists, tables, and Markdown. OpenAI's guide says the model tends toward detailed, formatted responses and may reuse phrases across sessions. Its suggested prompts ask for concise paragraphs with one idea each, lists only for parallel or sequential items, the main point stated early, plain verbs over jargon, and no stock phrases or contrastive framing that introduces an alternative nobody asked about. Copy the full blocks from the writing style section.
Subagent Delegation
Tell Astra when to delegate, because it may do so less often than your harness expects. OpenAI says GPT-6 Astra is trained to divide work across parallel subagents and responds well to prompting about how and when to do it. Its sample prompt tells the model to delegate whenever parallel work would save time or improve quality, whether it is the root agent or a subagent. A second prompt asks for legible inter-agent messages with proper spacing, since messages between agents can contain grammar or spacing errors. For delegation patterns on the Claude side, see our subagent patterns guide.
Testing Without Over-Testing
Calibrate testing to the change instead of ordering more of it. OpenAI says that for coding tasks Astra tends to test thoroughly before calling a task complete, which can mean broader tests than a small task needs. Its suggested prompt says not to write tests for reversible, low-impact changes that mirror the implementation, to run the checks the change requires, and to broaden testing only when new changes, failures, or open concerns justify it.
For larger projects, give the model tools to verify with rather than instructions to verify more. In OpenAI's games post, the Void Explorer game exposes a small JavaScript state interface, window.__VOID_EXPLORER__, plus named test scenes for orbit, travel, descent, and landing that Playwright can load and check. Ricouard says he would build those tools early in his next game, because they let Astra investigate and test changes independently.
New API Features Worth Adopting
GPT-6 adds four capabilities that change how you build agents, beyond the prompting changes above:
- Async tool calling. Set
async: trueon a function or custom tool, and the model keeps reasoning or answers other parts of the request while your application runs the tool. Return the result with the originalcall_id. - Mid-turn steering. Send a correction or new requirement while the model is working. Over a WebSocket connection, the Responses API keeps completed work and folds the update into a continuation.
- Reasoning changes that keep the cache. The
configuration_updateitem from the checklist above lets you raise effort for a hard step and lower it for routine follow-ups. - Misalignment monitoring. For Astra, OpenAI says its systems asynchronously monitor for misalignment and trigger alerts when necessary.
For raw speed, Ultrafast mode for GPT-6 Astra and GPT-6.1 Sol is available to all API users: set service_tier to ultrafast. On GPT-6.1 Sol, Ultrafast costs 6x Standard. Our Pro 500 and Astra Ultrafast coverage covers the ChatGPT side.
Sources
- OpenAI, "Using GPT-6" (models, what's new, limitations, prompting best practices, migration quickstart): https://developers.openai.com/api/docs/guides/latest-model
- OpenAI, "Rethinking skills and prompts for GPT-6 Astra" (Eric Provencher, September 11, 2026): https://developers.openai.com/blog/rethinking-skills-and-prompts-for-gpt-6-astra
- OpenAI, "Architectural visualization with Astra" (Thomas Ricouard, September 4, 2026): https://developers.openai.com/blog/architectural-visualization-with-astra
- OpenAI, "Building games with Astra" (Thomas Ricouard, September 4, 2026): https://developers.openai.com/blog/how-to-build-games-with-astra
- OpenAI, GPT-6 Astra model page: https://developers.openai.com/api/docs/models/gpt-6-astra
- OpenAI, GPT-6.1 Sol model page: https://developers.openai.com/api/docs/models/gpt-6.1-sol
- OpenAI, GPT-6 Luna model page: https://developers.openai.com/api/docs/models/gpt-6-luna
- OpenAI, GPT-5.6 Sol model page: https://developers.openai.com/api/docs/models/gpt-5.6-sol
- OpenAI, Reasoning guide: https://developers.openai.com/api/docs/guides/reasoning
- OpenAI, Prompt caching guide: https://developers.openai.com/api/docs/guides/prompt-caching
- OpenAI, Fast mode guide: https://developers.openai.com/api/docs/guides/fast-mode
- OpenAI, Ultrafast mode guide: https://developers.openai.com/api/docs/guides/ultrafast-mode
Read next
More practices and workflows- OpenAI
GPT-6.1 Sol: Astra-Level Coding at a Fifth of the Price, With Cost at Every Effort Level
OpenAI released GPT-6.1 Sol on September 29, 2026, at $2 input and $10 output per million tokens, a fifth of GPT-6 Astra's price, with cached input at $0.10. On OpenAI's charts it edges past Astra's best DeepSWE score at high effort and trails Astra by 2.1 points on OSWorld offline. The API model ID is gpt-6.1-sol.
- OpenAI
OpenAI Launches GPT-6 Astra for Computer Use and Complex Work
OpenAI launched GPT-6 Astra on September 3, 2026, as a model for computer use, browsing, software engineering, cybersecurity, science, and complex professional workflows. Access started with a limited set of organizations; Astra is now available in ChatGPT Work and Codex, the OpenAI API, Microsoft Foundry, and Amazon Bedrock.
- OpenAI
Prompting GPT-5.6: Message Roles, Effort, and Agentic Prompts That Work
GPT-5.6 rewards precise, explicit prompts: structure developer messages as identity, instructions, examples, then context, keep stable content first for prompt caching, and pick reasoning effort deliberately (xhigh for complex multi-step work). Move saved prompt objects into code before OpenAI shuts down v1/prompts on November 30, 2026.
- Claude
Prompt Engineering in 2026: The Playbook That Works Across Claude and GPT
Prompt engineering in 2026 has two layers: cross-model fundamentals (clear direct instructions, motivated rules, 3-5 structured examples, XML boundaries, long documents before the question) and model-specific overrides that matter more than ever, because the newest models need instructions deleted as often as added. Start here, then apply the guide for your model.
- Claude Code
Claude Code Subagent Patterns: 10 Reusable Agent Definitions
A Claude Code subagent is a delegated worker with its own context window, defined as a markdown file with YAML frontmatter in .claude/agents/. Subagents can edit files when you grant Edit or Write, nest three layers deep by default, and run 20 at a time. These 10 definitions cover the highest-value delegations.
Frequently Asked Questions
Does GPT-6 Astra support reasoning effort none?
No. Setting none on GPT-6 Astra returns an HTTP 400 error, and GPT-6.1 Sol rejects both none and minimal. OpenAI's migration guide says to use low instead. GPT-6 Sol and GPT-6 Luna still accept none, so check which model each request targets.
What is GPT-6 Astra's default reasoning effort?
OpenAI does not state one. The Astra model page lists low, medium, high, xhigh, and max, while the reasoning guide gives medium as the default only for the other GPT-6 models: GPT-6.1 Sol, GPT-6 Sol, and GPT-6 Luna. Set effort explicitly on every Astra request rather than relying on an omitted value.
Why does GPT-6 Astra keep asking for permission?
OpenAI says Astra is more likely to ask when extra input could change the result, and it can stop where you expected it to assume and continue. Tell it to infer intent, finish reversible work without asking, and save approval for a concrete, reviewable result.
How should I change AGENTS.md and skills for GPT-6 Astra?
Audit them. Astra is more sensitive to instructions in skills and AGENTS.md, so shorten skill descriptions to when the skill applies, turn long skills into routers, point to docs by situation instead of before every edit, and soften test and ask-first rules written for older models.
Can I use tool calling with GPT-6 models in Chat Completions?
Not with GPT-6 Astra or GPT-6.1 Sol: both support Chat Completions, but tool calling requires the Responses API. GPT-6 Sol and GPT-6 Luna allow function calling in Chat Completions only with reasoning_effort set to none. Move agent loops to Responses.