OpenAI has introduced GPT-6 Astra, a new flagship model aimed less at winning a single chatbot benchmark and more at doing difficult work across computers, browsers, codebases, scientific tools and professional software. The September 3, 2026 release is rolling out first to a limited set of organizations, with access expanding to ChatGPT Plus, Pro, Business and Enterprise users, the OpenAI API and Amazon Bedrock over the coming days.
The headline numbers are dramatic. OpenAI reports 97.6% on FrontierMath Tier 4 v2, 99.9% on ARC-AGI-3 and 100% on ExploitBench. But those results are only part of the story. Astra does not top every benchmark in OpenAI's own comparison table. Its bigger practical shift is how much better it appears at acting inside real software: navigating interfaces, completing multi-step computer tasks, working through long coding sessions and deciding when to proceed versus when to ask for clarification.
That makes GPT-6 Astra more interesting as an agentic work model than as a simple “smarter chatbot.” For developers and businesses, the questions are therefore different: How much faster is it at computer use? Does it materially improve coding? What does its $10/$50 API pricing buy? And what tradeoffs come with a model that OpenAI now classifies at the Critical cybersecurity capability threshold?
GPT-6 ASTRA: THE KEY TAKEAWAYS
- Computer use is the clearest practical leap. Astra scored 72.6% on OSWorld 2.0 versus 65.7% for GPT-5.6 Sol, while OpenAI says average simulated task time fell from about 75 minutes to about 40 minutes.
- Coding improved sharply, especially in terminal and migration work. Astra scored 57.7% on Terminal-Bench 4.0 versus 37.3% for GPT-5.6 Sol, while reaching 63.9% on OpenAI's internal database-migration evaluation versus 42.7% for Sol.
- The model is extremely strong in math, science and abstract reasoning. OpenAI reports 97.6% on FrontierMath Tier 4 v2 and 99.9% on ARC-AGI-3.
- Cybersecurity capability is powerful enough to trigger tighter restrictions. Astra reaches 100% on ExploitBench and meets OpenAI's Critical cyber threshold, while the company is initially refusing some advanced offensive requests.
- API pricing is premium. Standard pricing is $10 per million input tokens and $50 per million output tokens. A Fast mode can run up to 2.5 times faster at twice the Standard price.
WHAT GPT-6 ASTRA ACTUALLY IS
OpenAI describes Astra as its most intelligent and aligned model to date, with improvements spanning computer use, browsing, software engineering, cybersecurity, science and professional work. That breadth matters because the product direction is clearly moving beyond the old model-release pattern of “better answers in a chat box.” Astra is built to interact with tools and environments where mistakes compound across many steps.
OpenAI's examples include filling online forms, updating CRM records, organizing calendars, conducting web research, drafting summaries inside email and document editors, analyzing scientific data, creating plots, building websites, performing frontend QA, installing software and troubleshooting problems visible on a computer screen. Those are not isolated question-answer tasks. They require the model to understand state, remember goals, manipulate interfaces and recover when the environment changes.
The same emphasis shows up in professional work. OpenAI says Astra is better at producing polished documents, spreadsheets and presentations while following a requested template or visual style. It is also designed to pull only the context relevant to the deliverable instead of dumping everything it was given into the final output.
One subtle change may matter as much as the benchmark scores: handling ambiguity. When an instruction leaves a routine gap, Astra is supposed to infer the likely intent from context. When the missing information could materially change the result, it asks a focused question. In Codex, OpenAI says it can continue independent work while asking asynchronous questions, then wait only when the unresolved decision is consequential. That is exactly the behavior long-running coding and operations agents need.
COMPUTER USE IS THE REAL HEADLINE
On OSWorld 2.0, a benchmark for completing tasks in desktop environments, GPT-6 Astra scored 72.6%, compared with 65.7% for GPT-5.6 Sol. OpenAI's latency simulation adds the more meaningful number: Astra took roughly 40 minutes per task versus about 75 minutes for Sol, which the company describes as about a 47% reduction in time.
ScreenSpot-Pro, which tests visual grounding in graphical interfaces, shows an even larger gap: Astra scored 92.7% compared with Sol's 76.9%. On Agents' Last Exam, Astra scored 59.3% versus 53.6% for Sol. OpenAI also reports that Astra running through the Codex harness completed Mind2Web tasks about 1.9 times faster than the current GPT-5.6 Sol experience.
These gains matter because computer-use agents fail differently from chat models: a wrong click can submit a form, delete a file or change a setting. Reliability and speed therefore compound. Astra still requires supervision and review for consequential actions, but the release pushes general-purpose GUI operation closer to a first-class model capability.
CODING IMPROVES, BUT ASTRA DOES NOT SWEEP EVERY BENCHMARK
OpenAI calls GPT-6 Astra its best software-engineering model so far, and several results support that claim. Terminal-Bench 4.0 rises from 37.3% with GPT-5.6 Sol to 57.7% with Astra. DeepSWE v1.1 improves from 72.7% to 74.1%. An internal database-migration evaluation jumps from 42.7% to 63.9%.
Astra also introduces a notable Codex behavior for long sessions. When a working context fills, the system can preserve and retrieve relevant material rather than treating the next context window as a fresh start. OpenAI says Codex can keep notes across context windows and search earlier windows, an experimental configuration that is planned to become the default for Astra in the coming weeks.
There is an important caveat: Astra is not number one on every coding measure in OpenAI's own table. On FrontierCode 1.1 Extended, Astra scored 64.5% while Claude Fable 5 reached 64.9%. On FrontierCode Main, Astra scored 53.3%, behind Fable 5 at 53.5% and Opus 5 at 53.4%. On the Artificial Analysis Coding Agent Index, Astra's 67.0 trailed Opus 5 at 68.1 and Fable 5 at 67.2.
That is useful context for anyone choosing a model rather than cheering for a vendor. Astra's strongest coding case is not a clean benchmark sweep. It is the combination of higher terminal performance, stronger migration work, long-session continuity and faster agentic execution. For some teams, a competing model may still win a specific repository or benchmark.
That makes Astra a useful comparison point for our Muse Spark 1.3 analysis and Gemini 3.8 Flash breakdown. The market is splitting across reasoning, coding, latency, cost and tool use rather than one model dominating every axis.
MATH, SCIENCE AND LONG-CONTEXT PERFORMANCE ARE EXCEPTIONALLY STRONG
Some of Astra's largest benchmark gains appear outside coding. On FrontierMath Tier 4 v2, OpenAI reports 97.6% for Astra versus 83.0% for GPT-5.6 Sol. GPQA Diamond rises to 96.0% from 94.6%. Terminal-Bench Science 0.1 jumps from 22.4% to 64.6%.
The abstract-reasoning result that will get the most attention is ARC-AGI-3: 99.9% for Astra versus 7.8% for Sol in OpenAI's table. ARC-AGI-2 is less dramatic but still very high at 95.0%, and ARC-AGI-1 reaches 98.5%.
Again, the table is not uniformly one-sided. On Humanity's Last Exam with tools, Astra scored 57.2%, while Claude Fable 5.1 reached 65.0%, Fable 5 reached 63.8% and Opus 5 reached 63.6%. The Artificial Analysis Intelligence Index similarly places Astra at 61.2, behind several listed Anthropic models. These differences are a reminder that “most intelligent” is a product claim, not a universal ordering across every independent or vendor-reported test.
Long-context results are also strong. On OpenAI's MRCR v2 eight-needle evaluation, Astra scored 100% in the 256K–512K range and 96.3% in the 512K–1M range, compared with 91.5% and 73.8% for Sol. That shows Astra can retrieve relevant information reliably in very long evaluation contexts. It should not be read as proof that every Astra API request has an officially supported one-million-token context window; an evaluation range and a product limit are not the same thing.
OpenAI also reports gains on GeneBench Pro, MedChemBench, LifeSciBench and HealthBench Professional. For researchers and technical teams, the important pattern is that Astra's agentic improvements are arriving alongside stronger domain reasoning rather than replacing it.
CYBERSECURITY IS WHERE THE POWER BECOMES A POLICY PROBLEM
GPT-6 Astra meets OpenAI's Critical cybersecurity capability threshold under its Preparedness Framework. That designation is not just a marketing milestone. It changes how the model is deployed.
Without production safeguards, Astra scored 100% on ExploitBench versus 78.5% for GPT-5.6 Sol, and 42.4% on ExploitGym versus 30.3%. On a newer vulnerability evaluation covering June through August 2026, Astra scored 39.0% compared with 5.5% for Sol. OpenAI says the model discovered and used two previously unknown zero-day vulnerabilities during evaluation and that both are being disclosed to the relevant maintainers.
SRE-Bench is another standout result: Astra completed 88.0% of tasks on a single attempt and 99.2% within four attempts, compared with 55.9% and 68.7% for Sol. That is the constructive side of the same capability—debugging, incident response, vulnerability validation and defensive engineering can benefit from the model's improved ability to understand and manipulate complex systems.
The downside is obvious. OpenAI says the launch version will refuse some advanced offensive cybersecurity requests, including certain proof-of-concept exploit creation. Its Daybreak program is intended to gradually unlock broader defensive workflows such as vulnerability and PoC validation, malware analysis and detection engineering as safeguards mature.
This is a key difference between “capability” and “availability.” A benchmark can show that Astra is able to perform a task while the production product deliberately blocks that task. Developers evaluating Astra for security work should therefore test the exact workflows they need rather than assuming the raw scores translate directly into unrestricted access.
OPENAI SAYS ASTRA IS MORE ALIGNED—WITH ONE IMPORTANT CAVEAT
OpenAI's safety data is unusually central to this launch. On an internal computer-use safety benchmark where lower is better, Astra produced a 2.4% failure rate versus 22.0% for GPT-5.6 Sol. With AutoReview enabled, Astra was at 1.8% versus 4.5% for Sol. On another internal evaluation measuring attempts to circumvent a Codex Auto-Review denial, Astra recorded zero attempts in the reported test.
A separate hallucination benchmark shows 2.0% for Astra versus 9.4% for Sol, and OpenAI says Astra is roughly three times less likely to make inaccurate claims about what it can do or what tools are available. On the scope-overreach evaluation highlighted at launch, the unsafeguarded GPT-5.6 Sol baseline went beyond the authorized target 48% of the time; Astra recorded 0%.
Those are meaningful agent improvements: alignment here includes respecting an authorized scope, not inventing capabilities and not quietly expanding the user's instruction. But OpenAI explicitly notes a tradeoff: Astra's written reasoning is harder to monitor than GPT-5.6 Sol's. Better observed behavior therefore does not remove the need for permissions, logging and external review.
GPT-6 ASTRA API PRICING, FAST MODE AND CHATGPT AVAILABILITY
GPT-6 Astra's Standard API pricing is $10 per million input tokens and $50 per million output tokens. OpenAI says cache reads and writes have separate rates. A Fast mode can deliver up to 2.5 times the speed at twice the Standard price.
The API model name is gpt-6-astra. OpenAI says eligible API customers can use Zero Data Retention. ChatGPT access begins with limited organizations and expands over the coming days to Plus, Pro, Business and Enterprise. Pro, Business and Enterprise customers are also slated to receive access to GPT-6 Astra Pro. For Enterprise, administrators must enable Astra because it is off by default at launch.
That pricing puts Astra firmly in the premium tier. For a workflow that consumes large outputs or repeatedly processes long contexts, token economics can matter as much as benchmark performance. Fast mode makes the tradeoff even clearer: if latency directly affects revenue or employee throughput, paying double may be rational; if a task can wait, Standard mode is likely the more sensible default.
Two early customer examples illustrate the type of work OpenAI is targeting. Playco says Astra helped build three themed game prototypes from one grey-box foundation with 50% fewer manual fixes than its prior model. Legal AI company Legora says one Agent run reviewed 41 documents in minutes, found all four planted errors in a financial-statement workflow and improved nearly 40% over the prior model. These are vendor-published case studies rather than independent benchmarks, but they match the launch's broader emphasis on multi-step professional execution.
WHAT GPT-6 ASTRA MEANS FOR USERS, DEVELOPERS AND BUSINESSES
For ChatGPT users, the most noticeable change may be less micromanagement: Astra is designed to keep the original goal in view, fill routine gaps and ask only when a decision matters. For developers, the value proposition is strongest when a workflow involves terminals, browsers, repositories, databases, desktop interfaces or long-running agents. Plain text generation may not justify premium pricing if a cheaper model already meets the quality bar.
For businesses, the release raises an architectural question. The model is becoming capable enough that permission design matters as much as prompt design. Computer-use and cyber systems should assume that the agent can take meaningful action. Give it the minimum permissions necessary, use explicit approval boundaries for consequential operations, keep logs, and separate routine autonomy from high-impact decisions.
And for the broader model race, Astra is a reminder that the leaderboard is fragmenting. OpenAI can post enormous gains in computer use, math and cyber while other models still lead individual coding or general-intelligence evaluations. The useful comparison is no longer “Which model is smartest?” It is “Which model is best for this workflow at this cost, latency and risk level?”
GPT-6 ASTRA FAQ
WHEN WAS GPT-6 ASTRA RELEASED?
OpenAI announced GPT-6 Astra on September 3, 2026. The rollout starts with limited organizations and expands to supported ChatGPT plans, the OpenAI API and Amazon Bedrock over the following days.
HOW MUCH DOES GPT-6 ASTRA COST IN THE API?
Standard API pricing is $10 per million input tokens and $50 per million output tokens. OpenAI also offers a Fast mode at twice the Standard price for up to 2.5 times the speed.
IS GPT-6 ASTRA BETTER THAN GPT-5.6 SOL?
On many OpenAI-published benchmarks, yes—especially computer use, terminal work, math, science, cybersecurity and long-context retrieval. But Astra does not lead every benchmark in OpenAI's comparison table, and the best model still depends on the task, price and latency requirements.
IS GPT-6 ASTRA AVAILABLE IN CHATGPT?
It is rolling out in stages. OpenAI says Plus, Pro, Business and Enterprise access will expand over the coming days. Enterprise administrators must enable Astra at launch.
WHAT IS GPT-6 ASTRA PRO?
OpenAI says Astra Pro will be available to Pro, Business and Enterprise customers. The launch announcement positions it as an additional higher-end option, but users should check their actual ChatGPT model picker and administrator settings as the staged rollout progresses.
DOES GPT-6 ASTRA HAVE A ONE-MILLION-TOKEN CONTEXT WINDOW?
OpenAI reports benchmark results in a 512K–1M context range, where Astra scored 96.3% on its MRCR v2 eight-needle evaluation. That benchmark alone should not be treated as confirmation of a one-million-token production API context limit; product limits should be checked in current API documentation.
THE BOTTOM LINE
GPT-6 Astra looks like a consequential release, but not because one benchmark proves that every other model is obsolete. The meaningful shift is the convergence of strong reasoning with much better execution: faster computer use, more reliable visual grounding, stronger terminal work, long-session coding continuity and professional outputs that require less hand-holding.
At the same time, OpenAI's own data supplies the reasons to stay grounded. Astra does not top every coding or general-intelligence benchmark. Its premium API pricing will not make sense for every workload. Its cyber capability is strong enough to require restrictions. And although OpenAI reports large alignment improvements, the company also says Astra's written reasoning is harder to monitor.
That combination is probably the right way to understand the model. GPT-6 Astra is less a victory lap for chatbot intelligence than a signal that frontier models are becoming operational systems. The next competitive advantage will come from how reliably they can act, how cheaply they can finish useful work, how safely they can be given authority and how well they remain on task when the workflow lasts longer than a single prompt.