OpenAI Launches GPT-6 Astra to Scale Autonomous Knowledge Work

OpenAI Launches GPT-6 Astra to Scale Autonomous Knowledge Work

OpenAI Launches GPT-6 Astra to Scale Autonomous Knowledge Work

OpenAI introduced GPT-6 Astra today, a new model designed for advanced computer use and professional productivity. The release follows months of research into autonomous agent capabilities, positioning Astra as a state-of-the-art tool for software engineering, scientific analysis, and research tasks. Astra reached a 99.9% score on ARC-AGI-3 and a 98% score on FrontierMath Tier 4, exceeding performance milestones set by previous frontier models.

A person using a smartphone for digital tasks

A New Level of Autonomous Computer Control

Unlike its predecessors, GPT-6 Astra excels at navigating complex digital environments to perform tasks that typically require human intervention. Users can delegate workflows like filling out customer records in a CRM, organizing calendars, or installing and testing software. The model automates these processes by interacting directly with browser interfaces and OS-level applications. Performance tests on OSWorld 2.0 showed Astra completing real knowledge-work tasks with 47% less latency than the previous GPT-5.6 Sol model, effectively halving the time required for complex research and data extraction jobs.

Smartphone display showing social media application icons

A Step Change in Professional Work

GPT-6 Astra pairs advances in computer use with targeted training for professional environments, to help tackle complex work tasks. It combines the intelligence required for complex problems with the ability to carry out multistep workflows and produce polished documents, spreadsheets, and presentations. GPT-6 Astra is the best model for adhering to existing templates and producing slides that are well laid out and succinctly convey key points with a structured narrative. The model creates clear, well-structured documents, presentations, spreadsheets, and analyses that follow templates and match writing and visual styles. Astra is also trained to pull only the context that matters into outputs, rather than repeating unnecessary information. All this means it can output more immediately usable artifacts that match business context and standards.

The model brings stronger visual judgment to the websites, games, applications, and renderings it builds. With Sites in ChatGPT, Astra can create, host, and share websites, web apps, and games directly from a prompt. When instructions leave room for interpretation, GPT-6 Astra is better than previous models at making the right call. It uses context to fill in routine gaps and asks focused questions when the answer could change the outcome. In Codex, Astra can ask questions asynchronously while continuing work that doesn't depend on a reply. If the user doesn't respond, Astra proceeds with sensible assumptions where appropriate, but waits for input on consequential decisions. These capabilities extend to everyday tasks where missing information can materially change the answer. Astra is also better at staying oriented as a task evolves. Earlier models sometimes treated steering messages as a new goal, losing track of the original request or earlier constraints. Astra incorporates new requirements, changes course when asked, and answers side questions without dropping the broader task.

Coding and Software Engineering

GPT-6 Astra is the best model for software engineering to date. The model shows a clear step forward in iterative evaluations compared with GPT-5.6 Sol, and produces code that requires less iteration to reach production quality. When used for agentic coding, Astra communicates in a way that's easier for developers to follow. With Astra, a new way for Codex to preserve and retrieve context when the context window fills is introduced. Astra can keep notes across context windows, preserving accumulated details without repeatedly compressing them into a single summary. Earlier context windows remain searchable, so Astra can find requirements or test results from previous messages and tool outputs even if that information wasn't captured in its notes. Astra can solve 88% of tasks in a single attempt and 99.2% within four attempts on the SRE-Bench benchmark, compared with 55.9% for GPT-5.6 Sol.

Advancing Scientific Discovery and Cybersecurity

Astra can help with the practical work behind scientific discovery. By combining scientific reasoning with computer use, it can work directly in specialized software to inspect data and explore results, helping researchers assess the evidence and decide what to investigate next. As a notable jump in cyber capabilities under the Preparedness Framework, Astra meets the Critical threshold. On ExploitBench, Astra achieved a 100% score, exceeding GPT-5.6 Sol's 78.5% rate. During internal evaluation on a recent-vulnerability dataset, Astra achieved substantially higher code-execution rates than GPT-5.6 Sol while using far fewer output tokens, and even discovered and used two previously unknown zero-day vulnerabilities, which are being disclosed to their maintainers. On ExploitGym, Astra reached a 42.4% success rate compared with 30.3% for GPT-5.6 Sol, while using substantially fewer output tokens. Astra solved 88% of SRE-Bench tasks in a single attempt and 99.2% within four attempts, compared with 55.9% for the prior model.

Astra will refuse to comply with advanced cybersecurity tasks such as creating proof-of-concept exploits for vulnerabilities. Through OpenAI Daybreak, more defensive workflows including vulnerability and proof-of-concept validation, malware analysis, and detection engineering are planned for future rollout. Extra safety checks can sometimes pause or stop legitimate work, including defensive cybersecurity. If a task is paused, users may be asked to review the action before continuing. In the API, the task will stop until reviewed.

Alignment and Responsible Deployment

Astra is the most aligned model OpenAI has deployed. It excels at exercising care, respecting task boundaries, and communicating transparently. In an evaluation of computer-use tasks adversarially selected to elicit misbehavior, Astra was more successful at avoiding unintended consequences. Running with additional security measures yielded even stronger performance. Astra causes fewer misaligned outcomes than any other frontier model tested. Astra is three times less likely than GPT-5.6 Sol to make inaccurate representations about its capabilities and affordances. Written reasoning is harder to monitor than GPT-5.6 Sol, based on tests that asked the model to evade monitoring, but Astra still appears to struggle to conceal reasoning needed for complex tasks. Alignment training is core to the approach, and system safeguards like Codex Auto-review and monitoring agents' reasoning and actions help detect and contain unsafe behavior. Misalignment monitoring in production involves a system of classifiers that check the model's reasoning and actions for unauthorized behavior and automatically stop potentially unauthorized activity.

Given the marked increase in cybersecurity capabilities, extra safety checks are being deployed to make the deployment safe and secure. These checks can sometimes slow, pause, or stop legitimate work. We are continuing to iterate on the system to reduce unnecessary interruptions. Defenders can use Astra for secure code review and patching.

A laptop displaying AI interface and data on screen

Availability and Enterprise Rollout

Astra is rolling out now to ChatGPT Plus, Pro, Business, and Enterprise subscribers. Enterprise and API developers can access the model via Microsoft Azure and AWS Bedrock. The update also includes improvements to the Codex testing suite, designed to speed up the model's response time when executing multi-step tasks across browser windows. By enabling autonomous troubleshooting and frontend QA testing, OpenAI aims to reduce the manual workload for engineering teams focusing on complex deployment cycles.

For more on recent developments, see our AI category. For the full technical report, read the official announcement.

← Back to Home