Table of contents
- 1. Release overview: dates, specs, and who gets it
- 2. "I'm on Plus and it isn't there" — availability by plan
- 3. Pricing: the right way to ask whether $10/$50 is expensive
- 4. Benchmarks: why the side-by-side tables don't hold up
- 5. The first model to reach "Critical" cyber capability
- 6. For developers: what breaks when you migrate
- 7. Compared with Claude Opus 5 and Gemini 3.8 Flash
- 8. Caveats: the numbers this article left out
- 9. Picking by use case: when Astra is the right call
- FAQ
On September 3, 2026, OpenAI released its new flagship, GPT-6 Astra (OpenAI API documentation). The API is generally available with no gating, the model ID is gpt-6-astra, and the context window is 1,050,000 tokens.
That said, "it launched" and "your plan can use it" turned out to be two different things. OpenAI itself announced a staged rollout starting with a limited set of organizations, and on ChatGPT Plus it does not appear in the regular chat interface. Section 2 stands on its own for readers who came here only for that one answer.
There is one more way this article differs from other write-ups: it carries no benchmark score table. We list the metric names OpenAI cited, but we only print values we could trace back to the page that published them. Section 4 explains why — and the short version is that the comparisons circulating right now, of the "GPT-6 Astra 74.1% vs. Claude Opus 5 96%" variety, line up numbers that were measured with different rulers in the first place.
GPT-6 Astra
Released September 3, 2026 / API generally available
Sources: OpenAI
1. Release overview: dates, specs, and who gets it
OpenAI positions GPT-6 Astra as its most intelligent model to date, describing it as state of the art at computer use, browsing, software engineering, science, and professional work. The official phrasing is that it excels at carrying out multi-step workflows that span code, the browser, and business software — a description that puts the emphasis on work that runs for a long time rather than single question-and-answer turns (OpenAI model guidance).
| Item | GPT-6 Astra | Notes |
|---|---|---|
| Release date | September 3, 2026 | Rolling out from a limited set of organizations |
| Model ID | gpt-6-astra | API is ungated |
| Context window | 1,050,000 tokens | Input cap 922,000 |
| Max output | 128,000 tokens | — |
| Knowledge cutoff | April 30, 2026 | Details in our list of cutoff dates |
| Reasoning effort | low / medium / high / xhigh / max | none is not supported |
| Supported features | Streaming, structured outputs, function calling, file search, image input, web search, prompt caching | — |
| Rate limits | 500–15,000 RPM / 500K–40M TPM | Depends on your usage tier |
Inputs above 272,000 tokens carry a surcharge. Input and caching double, and output goes up 1.5x. A million-token context window is priced so that the more of it you use, the higher your unit rate climbs, so it is worth treating "what fits" and "what pays off" as separate questions.
2. "I'm on Plus and it isn't there" — availability by plan
This was by far the most common complaint in the days after launch. The short answer: Plus does give you Astra, but not in the regular chat interface.
What follows is based on the official announcement on the OpenAI developer community and on posts from OpenAI's official account.
OpenAI said on its official account that day one would cover a limited set of organizations, and that over the following days access would widen to Plus, Pro, Business, and Enterprise, plus the API and AWS. The next day it added that Pro, Enterprise, and Business Premium can use it in ChatGPT Work and Codex, that it is live in the API, and that the rollout to Plus and Business may take a few days.
Generally available. No gate, no waitlist — specify gpt-6-astra and it works. This is the one place you can reliably reach it today.
Available in ChatGPT Work and Codex. This is the scope OpenAI has explicitly described as already shipped to all users.
ChatGPT Work and Codex only. It does not appear in the model list of the regular chat interface. Look in the model picker inside Work or Codex instead.
In short, people are looking in the wrong place. Plenty of users open ChatGPT, scan the model list, see nothing, and stop there — but on Plus the way in is Work and Codex.
🟡 We cannot tell you it will show up in regular chat if you wait a few days. OpenAI's initial announcement said all Plus users within days, but as of September 8, 2026, five days after launch, it still has not arrived in regular chat on Plus. A thread asking for clarification is open on OpenAI's developer forum, pointing out that access was announced for all Plus users while Plus access is in fact limited to Work and Codex, and some reports say Pro is required to use it in regular chat.
The original announcement was a plan, not a confirmed availability date — what Plus reliably gives you today is Work and Codex. (As of September 8, 2026. We will update this if the situation changes.)
3. Pricing: the right way to ask whether $10/$50 is expensive
The API price is $10 in / $50 out per million tokens. That is exactly double Claude Opus 5 at $5 / $25, so on unit price alone it sits at the expensive end.
In the same documentation, though, OpenAI argues that because it reaches better results with fewer output tokens, the estimated API cost per task is actually lower. What a reasoning model bills you is unit price multiplied by the tokens it actually emitted, so a 2x unit price still comes out cheaper if the output is less than half as long.
Do not take that at face value. "Cheaper because it writes less" is OpenAI's own estimate, not a third-party measurement. It also flips easily depending on the workload — fire off large volumes of short answers and the unit price hits you directly, and settings that make it think longer (xhigh / max) push output tokens up.
Run ten of your own representative tasks and look at the bill — there is no other way to know.
4. Benchmarks: why the side-by-side tables don't hold up
These are the metrics on which OpenAI described Astra as achieving state-of-the-art results (official announcement on the OpenAI developer community). The metric names are public, but that announcement carries no table of scores.
One thing jumps out of that list. It is dominated by computer-use, terminal, and agent-flavored metrics, with almost nothing from the classic knowledge-and-trivia family. That matches the framing in section 1 about work that runs for a long time.
Why we left the numbers out
Write-ups elsewhere carry tables along the lines of "Astra 74.1% vs. Claude Opus 5 96%". This article does not, because the comparison does not actually hold.
For coding, OpenAI cites the DeepSWE family while Anthropic cites SWE-bench Verified. Those are different tests that merely sound alike, with different difficulty and different scoring. The moment you place one's number next to the other's, the table has stopped meaning anything.
Many of the circulating figures are roundups copying other roundups, and for some of them the same value is nowhere to be found once you open the publisher's page. Our policy is not to print a number we cannot trace to whoever published it.
On metrics where every top model is pinned to the high nineties, a one- or two-point gap means nothing. It is a gap you can call a win in print, but it will not show up in day-to-day work.
So what can you compare? Only figures where both models have published values on the same ruler. For example, on DeepSWE v1.1, which we covered in our Opus 5 article, Opus 5 scores 68.8% and GPT-5.6 Sol scores 72.7% — two values from the same test, so the comparison stands. We will add Astra's number on that metric as soon as we can confirm it on the publisher's page.
5. The first model to reach "Critical" cyber capability
This may be the bigger story of the two. OpenAI has classified GPT-6 Astra as the first model to reach Critical cybersecurity capability under its own Preparedness Framework (OpenAI Deployment Safety Hub, system card published September 3, 2026).
By OpenAI's own account, Astra can find previously unknown vulnerabilities and devise novel ways to exploit them against hardened systems without step-by-step human direction. In evaluations run with safety mitigations removed, it is documented as achieving a path to arbitrary code execution in a hardened browser and privilege escalation on a hardened operating system.
| Values published by OpenAI | GPT-6 Astra | GPT-5.6 Sol |
|---|---|---|
| ExploitBench (success on first attempt) | 88.0% | 55.9% |
| ExploitBench (success within four attempts) | 99.2% | 68.7% |
| Gray Swan IPI Arena attack success rate (lower is better defended) | 8.5% | 27.0% |
| HealthBench Professional (length-adjusted) | 63.4 | — |
| Severity-3-or-higher deviations in Codex simulations | 0.063% (34 out of 54,218) | — |
Offensive capability went up, but as a target it has become harder to break. The success rate of indirect prompt injection fell from Sol's 27.0% to 8.5%. If you point agents at external web pages, that is an improvement you will feel in production.
OpenAI's response was not to dial the capability back but to draw a line by use case. It helps with defensive work such as secure code review and patching, while refusing offense-leaning requests such as writing proof-of-concept code for a known vulnerability. On top of that come broader monitoring of all tool-using reasoning, user confirmation for consequential actions, and adjusted safety boundaries for under-18s, with a separate vetted limited-access track for biology research and cybersecurity research.
That way of drawing the line runs through the other new models released in September too. Every vendor fenced off its most dangerous capabilities and opened the gate only to qualified parties. The era of shipping everything to everyone is ending — of the three back-to-back releases, that is the shift I think is most widely overlooked.
6. For developers: what breaks when you migrate
If you are switching over existing code, some of it will error out as written. These are the changes spelled out in OpenAI's model guidance.
Do not pass temperature, top_p, top_logprobs and the like. On reasoning models they not only do nothing, they can be rejected outright.
Tool support in Chat Completions is limited. For agent work it is effectively mandatory.
none reasoning effortIf you were using none or minimal, OpenAI's instruction is to start from low and compare results.
prompt_cache_retention became prompt_cache_options.ttl. It is only a name change, but if it is silently ignored your cache stops working and only the bill goes up.
On behavior, OpenAI explicitly notes that Astra asks the user clarifying questions more readily than before. That is welcome when a human is sitting there in a conversation, but it becomes a reason batch jobs stall when nobody is watching. If you want autonomous execution, say so in the prompt.
One more thing: Fast mode is unavailable in EU data residency environments, and it comes with no latency SLA. Keep that in mind if your design assumes a given response speed.
7. Compared with Claude Opus 5 and Gemini 3.8 Flash
Between September 1 and 3, 2026, Anthropic, Google, and OpenAI shipped new models on three consecutive days. Here is the current lineup, restricted to the items where published values can be confirmed.
| Model | Released | Input/output (per 1M tokens) | Cutoff | Context |
|---|---|---|---|---|
| GPT-6 Astra | September 3, 2026 | $10 / $50 | April 30, 2026 | 1,050,000 |
| Claude Opus 5 | July 24, 2026 | $5 / $25 | May 2026 | 1,000,000 |
| Claude Fable 5.1 | September 1, 2026 | $10 / $50 | June 2026 | 1,000,000 |
| Gemini 3.8 Flash | September 2, 2026 | $0.75 / $3.75 (introductory price, through December 31, 2026) | March 2026 | 1,000,000 |
Three things stand out.
① The top tier has converged on $10/$50. Astra and Fable 5.1 charge the same. Meanwhile Anthropic has not changed its guidance that most use cases should start with Opus 5, so the picture in which the $5/$25 Opus 5 is where the real contest happens continues.
② Fable 5.1 has the freshest knowledge. June 2026, followed by Opus 5 (May), Astra (April 30), and Gemini 3.8 Flash (March). That said, if you pair the model with web search, the gap barely matters in practice. Cutoffs bite when search is turned off.
③ Price is the only dimension separated by an order of magnitude. At its introductory rate, Gemini 3.8 Flash costs roughly one-thirteenth of Astra on input. Line the models up by intelligence and they bunch together; line them up by cost and they split by an order of magnitude — that is the shape of autumn 2026, and the gap is starker than it was at the time of our previous-generation comparison.
8. Caveats: the numbers this article left out
To be upfront about it: there are figures in other write-ups that you will not find here.
Specific values for DeepSWE, FrontierMath, and ARC-AGI are omitted because we could not reach the page that published them. We will add them once the numbers are confirmed.
Figures of the "so many messages per five hours" kind are circulating, but we have not been able to confirm them in the official help pages. They are also the sort of value that shifts during a rollout.
Specs, pricing, and migration changes come from the API documentation and model guidance; the safety numbers come from the Deployment Safety Hub. Every one of them was checked on the publisher's own page.
And one more thing. Information right after a launch moves. Availability changes on a scale of days, and pricing and limits can be revised. This article reflects the state of things in September 2026.
9. Picking by use case: when Astra is the right call
The work runs autonomously for a long time. Multi-step tasks that cross the browser and business software, terminal work, and agents that read external web pages (thanks to the improved resistance to prompt injection).
When you want to halve the unit price. On writing quality and everyday coding, nothing published so far shows a decisive gap. If Opus 5 covers you, it is the cheaper answer.
Price is what moves by orders of magnitude. For pushing summaries, classification, and drafts through in bulk, the introductory rate on Gemini 3.8 Flash is overwhelming. See also our overview of Gemini.
Get a feel for them before you pay. There is no harm in working through each vendor's free tier before deciding. Here is our comparison of the three free tiers.
Let me close with the most reliable way to decide. Rather than studying benchmark tables, it is quicker to run ten of your own representative tasks through two or three models with identical inputs. Time taken, amount billed, and the quality of what came back. Look at those three yourself and you will get a more accurate answer than any table can give you.
FAQ
Q1. Can I use GPT-6 Astra yet?
The API has been generally available since September 3, 2026, with no gate and no waitlist. On the ChatGPT side the rollout is staged: Pro, Enterprise, and Business Premium can use it in ChatGPT Work and Codex. Plus and Business Standard are included as well, but OpenAI has said the rollout to them may take a few days.
Q2. I'm on ChatGPT Plus and it isn't in the model list.
You are most likely looking in the wrong place. On Plus it is available in ChatGPT Work and Codex, and it does not appear in the model list of the regular chat interface. Check the model picker there instead. Note that while OpenAI's initial announcement said all Plus users within days, it still had not reached regular chat as of September 8, 2026, five days after launch, and a thread asking for clarification is open on OpenAI's developer forum. Some reports say Pro is required to use it in regular chat, so we cannot state as fact that it will arrive if you wait.
Q3. Which is better, this or Claude Opus 5?
It depends on the use case. And at this point the published numbers cannot settle it, because OpenAI and Anthropic use different coding benchmarks (the DeepSWE family versus SWE-bench Verified), which means any table lining up the two companies' figures does not hold. On price, Astra is $10/$50 and Opus 5 is $5/$25, so Opus 5 is the cheaper of the two.
Q4. It costs twice as much as Opus 5 — will it really run up a bigger bill?
That depends on the task. OpenAI says the estimated cost per task is actually lower because it emits far fewer output tokens, but that is the vendor's own estimate. Workloads that fire off large volumes of short answers feel the unit price directly, and making it think longer with xhigh or max increases output. Measure the bill on your own representative tasks.
Q5. What breaks when I migrate existing code?
Four things mainly. Drop the sampling parameters such as temperature and top_p; use the Responses API instead of Chat Completions; note that the none reasoning effort is gone (OpenAI's instruction is to start from low and compare); and rename prompt_cache_retention to prompt_cache_options.ttl. Watch out for the last one in particular, because if it is silently ignored your cache stops working and only the bill goes up.
Q6. Does "Critical cyber capability" mean the model is dangerous?
It refers to a capability level in OpenAI's Preparedness Framework, not a verdict on danger. Astra is the first model to reach that level, which is why OpenAI is drawing a line by use case. It helps with defensive work such as secure code review and patching, but refuses offense-leaning requests such as writing proof-of-concept code. If anything it is harder to attack than before: the success rate of indirect prompt injection dropped from GPT-5.6 Sol's 27.0% to 8.5%.
Q7. Is a million-token context window a better deal the more of it I use?
No. Inputs above 272,000 tokens carry a surcharge, doubling input and caching and raising output by 1.5x. In many situations, selecting the parts you actually need is cheaper than dropping in an enormous context wholesale.
Sources
Every figure in this article was confirmed on the publisher's page listed below. Nothing was copied over from roundup articles.
- OpenAI — GPT-6 Astra Model (API documentation)
Model ID, context window (1,050,000 / input 922,000 / output 128,000), pricing ($10 / $50 / cache $1 and $12.50), the surcharge above 272K tokens, the April 30, 2026 cutoff, reasoning effort levels, supported features, rate limits - OpenAI — Model guidance
The positioning language, the claim that estimated cost per task is lower because output tokens are fewer, the four migration changes, the Fast mode restrictions, the tendency to ask clarifying questions - OpenAI — Deployment Safety Hub: GPT-6 Astra
Critical cyber capability under the Preparedness Framework, ExploitBench (88.0% / 99.2% against Sol's 55.9% / 68.7%), Gray Swan IPI Arena (8.5% against 27.0%), HealthBench, the 0.063% deviation rate, the mitigations and the refusal policy. System card published September 3, 2026 - OpenAI — GPT-6 Astra announcement (developer community)
The September 3, 2026 release date, availability by plan, the list of benchmark names described as state of the art - OpenAI developer forum — clarification thread on the scope of Plus access
The observation that access was announced for all Plus users while actual access is limited to Work and Codex. The basis for the 🟡 note in section 2 - OpenAI — official GPT-6 Astra announcement page
The primary announcement. Automated retrieval from our site is blocked, so we verified the contents through the announcement and documentation above
Related articles
- Claude Opus 5: the complete release guide — the current flagship it will be measured against
- GPT-5.6: the complete release guide — the previous generation and its Luna / Terra / Sol lineup
- Claude Fable 5.1: breaking changes and migration — the other new model released the same week
- Knowledge cutoff dates for the major AI models — how far each model's knowledge runs