OpenAI released a new flagship AI model on September 3, 2026. It’s called GPT-6 Astra. President Greg Brockman told reporters it marks a “generational leap in capability” and called it the start of the “AGI era.”
That’s a big claim. Some of what backs it up is genuinely impressive. Some of it needs a closer look before you take the headline number at face value. Here’s what’s actually confirmed.
| QUICK ANSWER
GPT-6 Astra is OpenAI’s newest flagship model, released September 3, 2026. It’s built for autonomous computer use, operating a browser, editing files, writing code, and completing multi-step tasks with limited human direction. Pricing starts at $10 per million input tokens and $50 per million output tokens. Access is rolling out in stages, starting with enterprise and cybersecurity customers, then expanding to ChatGPT subscribers and API users. |
GPT-6 Astra is the successor to GPT-5.6 Sol, OpenAI’s previous flagship model. OpenAI describes it as its most intelligent and best-aligned model yet.
The core focus isn’t just answering questions better. It’s built to complete real, multi-step work: reading a screen, operating software, and finishing a task across multiple applications without a person guiding every step.
OpenAI’s own demonstrations showed Astra completing tasks that go well beyond a chat window:
OpenAI also upgraded its Codex coding tool alongside Astra. The update targets a real, common problem: AI coding agents losing track of earlier decisions during long tasks.
That kind of context loss is exactly the review-and-governance issue we covered in is vibe coding safe for production software: more autonomy makes it more important to know what the model actually did, and why.
|
Standard input |
$10 per million tokens |
|
Standard output |
$50 per million tokens |
| Cached input |
$1 per million tokens |
| Cache writes |
$12.50 per million tokens |
That’s roughly two and a half times the price of GPT-5.6 Sol. OpenAI isn’t charging more for general intelligence. It’s charging more for autonomy and reliability on real, complex tasks.
|
Context window |
1,050,000 tokens |
|
Max output |
128,000 tokens |
|
Knowledge cutoff |
April 30, 2026 |
| Input types |
Text and image |
| Reasoning levels |
Low, medium, high, Xhigh, max |
Full technical details are published directly on OpenAI’s developer documentation.
|
Benchmark |
GPT-6 Astra |
Comparison |
|
OSWorld 2.0 (computer use) |
72.6% | 65.7% for GPT-5.6 Sol |
|
FrontierMath Tier 4 v2 |
97.6% | 87.8% for Claude Fable 5.1 |
| Agents’ Last Exam | 59.3% |
55.5% for Claude Opus 5 |
| Humanity’s Last Exam (with tools) | Lower |
Claude Fable 5.1 scored higher |
Results are mixed, not a clean sweep. Astra leads on several benchmarks tied to computer use and agentic tasks. Claude Fable 5.1 came out ahead on at least one major reasoning benchmark. Anyone telling you one model wins across the board isn’t reading the full comparison.
| THE FINE PRINT
OpenAI’s headline 99.9% score on ARC-AGI-3 came from a custom testing setup called Provider Adapter, which keeps the model’s reasoning active between turns. Under the benchmark’s standard test harness, the same model scored 62.7% instead. Both numbers are real. They measure different things, and only one of them reflects how most people will actually use the model. |
This matters beyond one benchmark. It’s the same pattern we’ve flagged in other AI-adoption assets: a headline number that sounds definitive often has a specific, favorable setup behind it.
We covered a version of this same caution in top 10 AI tools for businesses, where picking a tool based on a marketing claim, instead of your actual use case, is how the wrong tool gets chosen.
The model is also reaching developers through cloud platforms, including Amazon Bedrock and Microsoft Azure. OpenAI’s own launch details are posted on its developer community forum.
OpenAI says Astra is the first of its models to reach what it calls “critical” cybersecurity capability. Under its own framework, that means the model can find previously unknown security flaws in well-defended systems and build ways to exploit them, without step-by-step human guidance.
That’s exactly the kind of capability that needs real governance before it needs adoption.
We covered this pattern directly in is agentic AI ready for regulated industries: the technology tends to arrive faster than the oversight built to manage it responsibly. A model this capable, deployed without a clear review process, is a governance gap waiting to be found the hard way.
If you’re weighing whether a model like GPT-6 Astra fits into your team’s workflow, and what guardrails it would actually need, that indeed needs a conversation. Book a Discovery Call and we’ll help you think it through.
AI & Automation, Ecommerce
Search “AI personalization ROI” and you’ll find the same handful of numbers everywhere. And trust us, those numbers are interesting at best, and misleading at worst. 40% more revenue. 400% ROI. 5–8x returns. They all sound impressive. They also contradict each other, sometimes within the same article. The problem isn’t that personalization doesn’t work. It’s […]
AI & Automation, AI & Enterprise Strategy, Compliance & Governance, Thought Leadership
Is agentic AI ready for healthcare and finance is the wrong framing at this point. It is already running in both, at real scale, making decisions that affect real patients and real credit applications. The question worth asking now is narrower and more useful: has the governance around these systems kept pace with how fast […]
AI & Automation, AI & Enterprise Strategy, Artificial Intelligence
Call center chatbot. Virtual assistant. AI agent. Copilot. Half the vendors selling you software right now use these words interchangeably. The other half use them to mean completely different things, sometimes on the same product page. That confusion is not just annoying marketing copy. It changes what you should actually build, what it should cost, […]