Official GPT-6 Astra launch artwork. Source: .
The GPT-6 Astra API launched on September 3, 2026 with the stable model id gpt-6-astra. It offers a 1,050,000-token context window, a 922,000-token maximum input, a 128,000-token maximum output, text and image inputs, and tools including function calling, web search, computer use, code interpreter and MCP.
The release is less about another flagship number than a change in how OpenAI prices and behaves. OpenAI positions GPT-6 Astra for computer use, browsing, software engineering, science and professional work. It also says the model finishes several hard evaluations with fewer output tokens, which can lower estimated API cost per task even though the list price is well above GPT-5.6 Sol. Buyers should keep those two numbers apart: the published rates are first-party prices; the cheaper-per-task claim is an OpenAI evaluation estimate.
The ChatGPT change is easier to see. Official side-by-side captures show Astra asking a question that can change the result on a career site, a college shortlist and a grocery list. GPT-5.6 Sol more often answers immediately. That is product behavior you have to handle, not a slogan.
This guide separates model facts, OpenAI-reported evaluations, official UI demos and migration risks so you can decide what to test.
GPT-6 Astra API at a glance
| Item | Official specification | Deployment note |
|---|---|---|
| Release date | September 3, 2026 | This article was checked on September 6, 2026 |
| Stable model id | gpt-6-astra | The docs currently list this snapshot only |
| Context window | 1,050,000 tokens | Retrieval and context selection still matter |
| Maximum input | 922,000 tokens | Prompts above 272K input tokens use higher rates |
| Maximum output | 128,000 tokens | Put caps on agent turns, time and total output |
| Knowledge cutoff | April 30, 2026 | Fresh facts still need retrieval or web tools |
| Input format | Text and images | Output is text; image generation is a separate tool |
| Reasoning | low, medium, high, xhigh, max | none is unsupported; Fast mode is unavailable with EU data residency |
| Built-in tools | Web search, file search, computer use, code interpreter, MCP, skills, hosted shell | Tool calling should use the Responses API |
| Surfaces | OpenAI API, Azure, Bedrock | Chat Completions can generate text; tools require Responses |
The specification comes from OpenAI's . The model can inspect images, but it returns text. If a workflow needs pictures, connect an image-generation tool instead of expecting the model to draw.
A 1M context window helps with large repositories, long sessions and document packs. Production systems should still retrieve, cache and trim context, because longer inputs raise latency and cost. The 272K threshold matters more than the headline window: once a prompt crosses it, the whole request is billed at the long-context rates.
Where GPT-6 Astra is rolling out
The launch covers ChatGPT, the API and cloud platforms, but the schedule is not identical on every surface.
ChatGPT starts with a limited set of organizations and then reaches Plus, Pro, Business and Enterprise over the following days. Astra usage sits inside existing subscription allowances, with extra credits available for purchase. Pro, Business and Enterprise also get GPT-6 Astra Pro. Enterprise admins have to turn Astra on; it is off by default.
The OpenAI API uses the model name gpt-6-astra. OpenAI recommends the for new work. Chat Completions still works for text generation, but tool calling has to move to Responses.
Codex ships an updated harness with Astra. OpenAI says that, combined with the model's efficiency, task completion on Mind2Web is about 1.9x faster than the current GPT-5.6 Sol experience. Codex also adds an experimental way to keep and search notes across context windows, so a long debugging session does not lose the reason a fix failed every time the window is compacted.
Azure and AWS Bedrock will offer the same model. Teams already running on a cloud vendor should check residency, zero-data-retention and Fast mode limits in that region rather than copying the OpenAI.com price table.
For most product integrations, start with gpt-6-astra and Responses. ChatGPT's Astra Pro tier, Codex experiments and vendor-region limits need their own checks.
GPT-6 Astra API pricing
OpenAI published Standard list prices per million tokens. Cache writes are billed at 1.25x the uncached input rate.
| Meter | Short context | Above 272K input tokens |
|---|---|---|
| Uncached input | $10.00 | $20.00 |
| Cached input | $1.00 | $2.00 |
| Cache writes | $12.50 | $25.00 |
| Output | $50.00 | $75.00 |
Batch and Flex are 50% of Standard. Fast mode is 2x the applicable rates and has no latency SLA. Fast and Priority are unavailable with EU data residency. Regional processing endpoints add a 10% uplift.
Take a short-context job that uses 1 million uncached input tokens and 200,000 output tokens. The model charge is about $20: $10 for input and $10 for output. If the same request hits cached input, the input line can fall to $1, but the first cache write is still $12.50. This example excludes per-call tool fees such as search or computer use.
Against GPT-5.6 Sol's $4 / $20 rates, Astra's list price is higher. OpenAI's argument is that, in several evaluation setups, Astra finishes with fewer output tokens and a lower estimated API cost per task. Treat that as a billing hypothesis. Measure task success and the full invoice together before you cut over.
See the for every service tier. ChatGPT allowances, Codex credits and API invoices are not the same meter.
What the GPT-6 Astra benchmarks show
OpenAI's launch tables place GPT-6 Astra next to GPT-5.6 Sol, Claude Fable 5.1, Claude Opus 5 and Gemini 3.8 Flash across computer use, coding, science and safety evaluations. The numbers are OpenAI's best reported scores under specified harnesses, effort settings and tools. They are not Mengbi retests.
The computer-use and professional-work gaps are the clearest. On Agents' Last Exam, Astra scores 59.3%, Opus 5 scores 55.5% and Sol scores 53.6%; OpenAI says Astra also used about 65% fewer output tokens than Opus 5 at those settings. On the OSWorld 2.0 offline set, Astra scores 72.6% at roughly 40 minutes per task, versus 65.7% and about 75 minutes for Sol. BenchCAD geometric overlap is 95.9% for Astra, 83.3% for Sol and 84.3% for Fable 5.1.
Coding is not a sweep. Terminal-Bench 4.0 is 57.9% for Astra, 37.3% for Sol and 55.8% for Fable 5.1. DeepSWE v1.1 is 74.1% for Astra, 72.7% for Sol, 73.8% for Gemini 3.8 Flash and 73.7% for Opus 5. That is enough to put Astra in the same band as the current coding leaders, not enough to declare it first on every repository job.
Science and abstract reasoning take up the most launch copy. Terminal-Bench Science 0.1 is 64.6% for Astra, 52.6% for Fable 5.1 and 22.4% for Sol. GPQA Diamond is 96.0%. FrontierMath Tier 4 is 97.6%, also described in the post as about 98%. ARC-AGI-3 is 99.9%, with a footnote that two Responses API harness settings were changed. OpenAI also says Astra helped tighten a bound on small prime gaps from 240 to 186 and improved a large-gap term that had stood for more than 80 years. The proofs are official PDFs; they are research claims, not a product SLA.
Cybersecurity needs its own reading. OpenAI says Astra meets the Critical threshold in its Preparedness Framework. In a research setting without production safeguards, ExploitBench is 100% versus 78.5% for Sol, and ExploitGym is 42.4%. The version shipping today refuses more advanced offensive work such as writing proof-of-concept exploits. Defensive code review and patching are in scope; looser safeguards are planned through OpenAI Daybreak.
Use the tables to choose what to test. They mix tools, prompts, time budgets and no-safeguard research configurations, so they cannot replace your own jobs. A better method is a fixed set of 30 to 100 real tasks, with pass/fail, human edits, tool-call counts, total tokens, timeouts and the full bill.
Official GPT-6 Astra UI demos
OpenAI's announcement includes ChatGPT comparison screens and demos in electronics, architecture, games and CAD. They show when Astra stops to ask a question, and which professional apps computer use can enter. They are still curated launch clips.
Career site: ask the destination before publishing
Same prompt: build a personal site from LinkedIn. Sol publishes a portfolio; Astra asks about the target role. Source: .
Both models get a career-change request. Sol spends 13 minutes 15 seconds and returns a live portfolio. Astra pauses at 20 seconds to ask which job the user is moving into. OpenAI is arguing for judgment: without a destination, the site narrative changes. In a ChatGPT or agent product, treat that pause as default behavior. Prompt the model either to finish a reviewable draft first or to wait for an answer.
College search: is $80,000 a hard cap?
Same California senior prompt. Sol returns a school table; Astra splits the $80,000 budget into two interpretations. Source: OpenAI's announcement.
Sol quickly names Stony Brook, Pitt and other options. Astra asks whether $80,000 is a hard pre-aid limit or whether higher-list-price schools can count if aid brings them down. That question changes the shortlist. If your product cannot ask the user mid-task, write the default assumption into the system prompt, or the run will stop on a Question card.
Grocery list: household size, allergies and budget
Same prompt: a cheaper grocery list for next week, without cooking every night. Sol already has ingredients; Astra is still asking about household size. Source: OpenAI's announcement.
The same behavior shows up in a household task. After 12 seconds, Sol has a two-person list built around cooking twice. Astra uses the same interval to ask about people, allergies and a weekly budget. That can be safer in a consumer assistant. In a batch API it can become an unfinished tool loop. OpenAI's migration guide says the same thing in another form: if the model keeps asking for approval, use the initiative prompts so it first produces a reviewable result.
KiCad: schematic to a manufacturable board
OpenAI describes this as a 15-second condensed playback of Astra placing parts and routing copper in KiCad. Frame taken from the official demo video.
OpenAI uses this clip as the computer-use showcase: turn an electronics schematic into a manufacturable PCB. The frame shows traces, vias, pads and a board-level 3D view. It shows that the model can enter a specialist desktop app. It does not show that any schematic routes cleanly on the first pass, or that the result passed DRC or a fabricator.
Blender house and a playable city
First panel from the Solace / Garden House storyboard, labeled as a September 2026 matched-camera comparison. Source: OpenAI launch assets.
Used in the announcement to show playable scenes and motion, not only concept art. Source: OpenAI launch assets.
These stills sit in the professional-work and visual-output section. Astra models a house in Blender and, in the surrounding copy, moves the scene into Unreal Engine 5. Another demo is a city game scene. They are useful for asking whether the model can emit assets a designer can keep editing. They are not proof that a first pass will survive design review. ChatGPT Sites is mentioned in the same stretch: a prompt can create, host and share a site or a small game.
FreeCAD: a five-speed gearbox
The window title is Transmission_5Speed_GearMotion. The UI is a real FreeCAD desktop session. Source: OpenAI's announcement.
The FreeCAD shot moves computer use from browser forms into CAD. The model tree, task panel, 3D view and Python console are all visible. Use it as an example of continuing work inside a professional GUI. Acceptance still depends on dimensions, assembly constraints and whether an engineer can open the exported file.
How to call the GPT-6 Astra API
Create a key in and store it in the server environment variable OPENAI_API_KEY. The official Python example uses the Responses API. The request below has no tools, so you can first confirm auth, the model id and the response shape.
OpenAI() reads OPENAI_API_KEY. Do not put the key in a browser bundle, a public repository or a frontend environment variable. The server should also log request ids, token use, tool-call counts, latency and stop reasons, and cap retries plus total agent turns.
To raise reasoning strength, set reasoning.effort. Astra does not accept none. If you are coming from none or minimal on an older model, OpenAI suggests starting at low and comparing.
Chat Completions still works, but OpenAI is explicit: tool calling for GPT-6 Astra requires Responses. Projects that already use function calling, computer use or MCP should not only rename the model. Async tool calling, mid-turn steering and configuration_update also sit on Responses. See the and the .
A successful text request is not a production agent. Computer use, search and code interpreter can add per-call fees. The client still has to handle tool results, pending call_ids and failed retries.
Migrating from GPT-5.6 Sol to GPT-6 Astra
If the current request only sends model and a text prompt, start by swapping the model id. Requests that set sampling, tools, cache or Fast mode need a pass through OpenAI's .
| Current setting | What GPT-6 Astra expects |
|---|---|
gpt-5.6-sol or another 5.x model | Change to gpt-6-astra and keep the old model as a rollback group |
reasoning.effort: none or minimal | Start at low and compare; Astra does not support none |
| Chat Completions tool calling | Move to the Responses API |
temperature, top_p, top_logprobs | Remove them |
| Fast or Priority with EU data residency | Use Standard |
| Changing effort between responses | Use configuration_update and leave request-level effort unchanged so the cache prefix survives |
prompt_cache_retention from GPT-5.5 or earlier | Replace with prompt_cache_options.ttl, for example "30m" |
| Repeated approval pauses | Add initiative prompting so the model first finishes a reviewable result |
Do not move all traffic on day one. Start with 5% to 10% shadow or canary traffic. Record completion, output length, question count, tool errors and cost for the same jobs on Sol and Astra. Confirm rollback, timeouts and log fields before you raise the share.
One easy misread: Astra is more willing to stop at a fork that would change the answer. In ChatGPT that is a Question card. In the API it can become an unfinished turn waiting for input. Caution is not the same as a shorter job. Compare the cost and duration of a finished business result, not time-to-first-token alone.
GPT-6 Astra vs GPT-5.6 Sol
| Workload | Try GPT-6 Astra first | Keep GPT-5.6 Sol for now |
|---|---|---|
| Long coding and computer-use jobs | You need higher completion and less wasted exploration | The current flow is stable and a change is expensive |
| Documents, slides and specialist software | You need template following and GUI operation | The task is short and the output is simple |
| Cost-sensitive batch work | You will measure total tokens, not only the list price | Sol already finishes the job at a lower bill |
| Low-latency chat | You can test Fast mode or low effort | P95 is already fine and you must stay on EU residency |
| Defensive security | You need stronger review and patching help | You cannot absorb stricter refusals or false stops |
| Production cutover | You have canaries, rollback and per-task scoring | You still lack a test set and usage monitoring |
Sol remains the cheaper flagship tier. Astra is a better fit once you have shown that the work is hard enough to justify $10 / $50 rates, and that the product can handle questions, refusals and safety interrupts. If you are choosing an IDE or terminal agent to host the model, start with Mengbi's . If the work spans research, spreadsheets and everyday collaboration, the is closer to this ChatGPT and Codex shape. For Anthropic's contemporaneous flagship, see .
How to evaluate GPT-6 Astra before launch
A reusable scorecard should at least record:
- Whether the task finished and whether the final tests passed.
- How often the model asked a question, waited for approval or needed a human prompt.
- Uncached input, cached input, cache writes, output tokens and the full charge.
- Tool-call totals, failures, unfinished async calls and repeated calls.
- P50 and P95 latency for first token, full response and the whole job.
- Timeouts, rate limits, empty responses and structured-output parse failures.
- Safety refusals, false interrupts, factual errors and whether sourced answers can be checked.
Do not fill the set with only the easy wins. Include vague briefs, missing budgets, dead links, oversized files, tool errors, repeated function results and jobs that change requirements mid-run. Keep the input, acceptance script and human rubric fixed, or you will compare prompts instead of models.
OpenAI's system card and safety update also flag a monitoring gap: Astra's written reasoning is harder to watch. Complex tasks can still expose the necessary steps, but simpler tasks are less monitorable. Medical, legal, financial, privileged and cybersecurity jobs still need human review and system-level permissions. Misalignment monitoring can stop an API task outright; ChatGPT and Codex may ask the user to confirm before continuing.
Frequently Asked Questions
When was GPT-6 Astra released?
OpenAI published GPT-6 Astra on September 3, 2026. Developers can call the model id gpt-6-astra through the OpenAI API, Azure and AWS Bedrock. ChatGPT Plus, Pro, Business and Enterprise access follows over the next few days.
How much does the GPT-6 Astra API cost?
Standard short-context rates are $10 per million input tokens, $1 for cached input, $12.50 for cache writes and $50 per million output tokens. Above 272K input tokens, input and cache rates double and output is billed at 1.5x. Batch and Flex are half price. Fast mode is 2x.
What is the GPT-6 Astra context window?
The context window is 1,050,000 tokens, with a 922,000-token maximum input and a 128,000-token maximum output. Long context raises latency and cost, and crossing 272K input tokens also raises the rate for the whole request.
Can GPT-6 Astra inspect images?
Yes. It accepts image input and returns text. That is useful for screenshots, charts and document pages. It does not generate images directly; use an image-generation tool for that.
Does GPT-6 Astra require the Responses API?
Text generation still works on Chat Completions. Function calling, computer use, MCP and the new Astra features need Responses. New projects should start there.
Does GPT-6 Astra support Fast mode?
Yes, with no latency SLA, at 2x the applicable rates. Fast and Priority are unavailable with EU data residency; use Standard instead.
Should every GPT-5.6 Sol workload upgrade now?
No. Astra belongs in a test group for computer use, long coding jobs and professional artifacts. Sol is cheaper and can remain the right default for short or cost-sensitive traffic. Compare task success, question count, latency and the full bill on a canary first.
Sources and verification notes
Model specifications, prices, migration items and evaluation numbers in this article come from OpenAI's own materials:
Official UI images and demo frames come from OpenAI's announcement and its attached video. This article uses a limited set of frames for commentary and keeps the source links. Benchmarks are OpenAI's, or those of the evaluation authors it cites. Mengbi did not retest with a paid API key and did not take OpenAI sponsorship. Prices, model status, safeguards and preview features can change; re-check the official pages before a purchase or production cutover.
Last verified: September 6, 2026.

