GPT-6 Astra: What OpenAI's New Model Actually Does
OpenAI's new flagship posts big gains on computer use and science — but the benchmarks come with asterisks, a price premium, and a phased rollout.

OpenAI has started rolling out GPT-6 Astra, its newest flagship model, and the launch arrived with an unusually loud claim: company president Greg Brockman ended the press briefing with "Welcome to the AGI era." Behind the headline is a model built to do computer work — control apps, fill spreadsheets, build websites — rather than just describe how. This post breaks down what Astra actually does, how it benchmarks, what it costs, and the caveats worth knowing before you plan around it.
What is GPT-6 Astra?
Astra is OpenAI's latest and most capable model, positioned as a leap in autonomous computer use. OpenAI released it to a limited set of customers, labelling it the world's most intelligent AI system and the first model to trigger advanced internal safety protections given its cyber capabilities. The company says it sets a new high-water mark for autonomously controlling computer systems and performing tasks on a user's behalf — things like filling out spreadsheets and creating websites from scratch.
The framing from OpenAI is that this is a new mode of computing. Rather than requiring a dedicated API integration for every application, Astra is designed to navigate software much as a person does — working across browsers, spreadsheets, websites and desktop apps, producing finished documents, and carrying out multistep workflows instead of merely telling you how to do them.
That's the part that matters most for anyone building AI agents: a model that operates the same software people do is a general-purpose worker, not a chatbot with plugins.
The "AGI era" claim — how seriously to take it
OpenAI leaned hard into the AGI narrative, but it hedged the definition. Brockman said he personally believes OpenAI has reached AGI while leaving users to decide whether Astra meets the definition — "I think it might be about this model," he said, before closing with "Welcome to the AGI era."
Notably, the term has lost its old contractual weight. Asked whether OpenAI was formally declaring it had achieved AGI, Brockman said the term was no longer tied to a contractual trigger — a reference to its earlier agreement with Microsoft — and described it instead as a "mission concept or spiritual concept."
Read that as marketing framing over a technical milestone. The interesting story is in the capabilities and the caveats, not the label.
GPT-6 Astra benchmarks: where it leads, and the asterisks
Astra's real gains show up on computer use and science tasks more than on coding. It's worth reading these numbers with the footnotes OpenAI itself attached.
Coding is a modest step, not a blowout. On the DeepSWE v1.1 agentic coding test, Astra scored 74.1% versus 70.8% for GPT-5.6 Sol. But that's not a clear lead over rivals — Meta reported 75.4% for Muse Spark 1.3 at its maximum reasoning setting, and Gemini 3.8 Flash and Claude Opus 5 sit around 74% on the public leaderboard, with overlapping uncertainty ranges. In other words, no clear coding king.
Computer use is where the jump is real. On the OSWorld V2-Offline benchmark for desktop application work, OpenAI says Astra scored 72.6%, up from 65.7% for Sol, and cut average time per task from about 75 minutes to 40. That efficiency gain — same or better score in less time — is arguably more useful than a headline accuracy number.
Science and reasoning scores are high, with caveats. Astra posted a 98.6% on ARC-AGI-3 and a 97.6% on FrontierMath Tier 4. The asterisks matter: the ARC-AGI-3 result reflects Astra plus an agent harness that retains reasoning between turns, and OpenAI has exclusive access to part of the FrontierMath benchmark, whose developer it funded. Benchmarks that measure "the model plus OpenAI's scaffolding" aren't apples-to-apples with a raw model score.
The honest summary: Astra is genuinely strong on computer use and research-style tasks, roughly at parity on coding, and several of its flashiest numbers depend on OpenAI's own harness and benchmark access.
GPT-6 Astra pricing
Astra is a premium model. Once it's live in the API, Astra will cost $10 per million input tokens and $50 per million output tokens — 2.5 times Sol's current promotional price, though it matches Anthropic's pricing for Fable 5.1.
For context, that's well above the cheaper frontier options. Muse's standard price is $1.25 / $4.25 per million input/output tokens, and Google's introductory prices for Gemini 3.8 Flash are $0.75 / $3.75.
OpenAI's counter-argument is that per-token price isn't the same as per-task cost: a higher per-token rate doesn't necessarily mean a higher bill if the model finishes a job in fewer steps and needs fewer retries. The catch is that the launch data is too sparse to prove those savings offset the premium, so treat "cheaper per task" as a claim to verify on your own workloads.
For now the lineup is just Astra and Astra Pro — unlike GPT-5.6, OpenAI has not announced Luna, Terra, and Sol variants for this generation.
When can you use it? The phased rollout
You probably can't use Astra yet. The model is launching in phases, and OpenAI said a limited group of companies in its application-based cybersecurity program, Daybreak, get first access.
Broader access follows. Astra will be available "in the coming days" for ChatGPT Plus, Pro, Business and Enterprise customers and API developers. Pro, Business, and Enterprise users also get GPT-6 Astra Pro, and the model is slated to reach the OpenAI API and AWS.
Why cybersecurity gated the launch
The reason for the staged rollout is Astra's cyber ability. It's the first model OpenAI has designated as reaching its "critical" cybersecurity threshold under its preparedness framework — meaning it can potentially find and exploit previously unknown vulnerabilities across well-protected systems without step-by-step human guidance.
That's not hypothetical. In OpenAI's testing the model developed exploits for hardened browsers and operating systems and found two previously unknown vulnerabilities, which OpenAI says it is disclosing to maintainers. As a result, the standard-access version of Astra refuses some advanced cybersecurity work like exploit discovery, and API tasks that trip a cybersecurity check get stopped outright rather than paused for approval.
There's also an alignment wrinkle worth flagging: OpenAI disclosed that Astra's written reasoning was harder to monitor than Sol's in tests designed to elicit monitoring evasion — a reminder that more capability doesn't automatically mean more transparency.
What GPT-6 Astra means for AI agents
Strip away the AGI talk and the practical shift is this: frontier models are being built to operate software directly, filling spreadsheets and building sites rather than describing how. For teams building agent workflows, three takeaways stand out:
- Computer use is the new battleground. The biggest gains are in operating desktop and browser apps, not raw chat. Design workflows around agents that click, type, and verify — not just answer.
- Harness beats raw model. Several of Astra's best scores come from the scaffolding around it: persistent reasoning across context windows, compaction, and the ability to keep working while it asks you a question. The system matters as much as the weights.
- Cost and access are real constraints. A premium price and a gated, cyber-restricted rollout mean a single frontier model won't be the right fit for every task. Routing simpler steps to cheaper models keeps bills sane.
Put a computer-use agent to work today
You don't have to wait on a Daybreak invite to build agents that operate real software. Eigent is an open-source, local "Cowork" desktop app that runs a multi-agent workforce across your browser, files, and tools — the same "do the work, don't just describe it" pattern Astra is chasing, but model-agnostic and self-hosted, so you can plug in whichever model fits each task. See how a workforce automates multistep jobs on the Eigent workflows hub, or download Eigent and build your first computer-use agent this afternoon.
Recent Posts

Claude Fable 5.1 and Mythos 5.1: What's New, Explained
Claude Fable 5.1 and Mythos 5.1 explained: the same model in two safeguard tiers, with new benchmarks, roughly 25 to 45 percent lower cost, and access details.

Gemini 3.8 Flash: What's New for Coding and AI Agents
Gemini 3.8 Flash brings big coding and agentic-reasoning gains at the same low price, plus a new 3.8 Flash Cyber variant. Benchmarks, pricing, and how to use it.

Muse Spark 1.3: Meta's Frontier Coding and Agentic Model, Explained
Muse Spark 1.3 is Meta's new frontier model for coding and agents. See benchmarks, pricing, the xhigh vs max variants, what changed, and how it compares.