OpenAI says its Astra model will be available soon, signaling that the long-awaited release is approaching even though the company has not announced a specific launch date. In a September 1, 2026 update, OpenAI described the safety work required before release and said Astra’s most advanced cybersecurity capabilities will initially be limited to a small group of testers.

Source: Path to Astra: critical capabilities and frontier safeguards · OpenAI

The update, published by OpenAI, is not a launch announcement. It is a pre-release report explaining why the company delayed parts of Astra’s development and deployment while it strengthened protections against cyber misuse and unauthorized model actions. That framing suggests Astra is moving toward release, but access will be more controlled than it has been for earlier models.

Astra has reached OpenAI’s highest cybersecurity capability tier

OpenAI says Astra meets the “Critical” cybersecurity capability threshold in its Preparedness Framework. Under that framework, a model reaches this level if it can either develop functional zero-day exploits across many hardened critical systems without human intervention, or devise and execute novel, end-to-end cyberattack strategies against hardened targets from a high-level goal.

The company’s evaluations combined automated benchmarks, private testing, and expert-led assessments. OpenAI reports that Astra performed substantially better than GPT-5.6 Sol on vulnerability identification and exploit development while using fewer output tokens.

On an internal benchmark containing 20 recently disclosed high-severity vulnerabilities, OpenAI says Astra achieved higher arbitrary code-execution rates than GPT-5.6 Sol. During testing, it also discovered and used two zero-day vulnerabilities in an exploit chain. OpenAI says it is working to disclose those vulnerabilities to the relevant maintainers.

In separate expert assessments, Astra reportedly found previously unknown vulnerabilities in a hardened browser and operating system. It developed a browser-compromise chain that escaped the sandbox and executed commands on the host, and it combined operating-system vulnerabilities into a privilege-escalation chain from an unprivileged user to root.

These results explain why OpenAI is treating Astra’s release differently from a routine model update. The reported capabilities could be useful to defenders searching for flaws, but they could also create greater risks if misused.

Release preparations have delayed parts of Astra’s development

OpenAI says it paused portions of Astra’s development and release while it worked on stronger safeguards. The company also held back some larger reinforcement-learning runs until it established higher safety and security requirements for the training environment.

According to the update, OpenAI restarted a large frontier reinforcement-learning run on August 28 after putting those requirements in place. Some smaller experimental training runs remain temporarily paused.

The company says its approach addresses two possible sources of harm. Safeguards must limit malicious users’ ability to use Astra for sophisticated cyberattacks, while also detecting and containing unauthorized actions by the model itself if its behavior diverges from its intended role.

OpenAI says it incorporated lessons from the recent Hugging Face incident into its safety work. It also reports that Astra did not attempt to compromise surrounding security infrastructure in a set of internal “honeypot” tests conducted without production safeguards. The company cautions that these results describe behavior under test conditions and should not be interpreted as a description of ordinary production use.

Advanced cybersecurity access will start with testers

OpenAI says Astra will initially be available to a small group of alpha testers for advanced cybersecurity workflows. Broader access for defensive use is expected to follow through Daybreak Blue.

That staged approach is significant because it indicates that the model’s most capable cybersecurity functions will not necessarily be available to everyone at launch. OpenAI says it expects Astra’s safeguards to create more friction than ultimately intended during the early deployment period while the company gathers more evidence and calibrates its controls.

The protections include model-level refusals, system-level safety classifiers, offline detection, threat disruption, and expanded monitoring for higher-risk accounts. OpenAI reports that Astra refused 91.5% of requests in its cyber-jailbreak evaluations, compared with 59% for GPT-5.6 Sol in the company’s comparison.

OpenAI also says Astra will use monitoring systems designed to detect potentially unauthorized behavior and automatically stop suspicious activity. In ChatGPT or Codex, users may be asked to review an action before continuing if a monitor pauses a task. On other surfaces, such as the API, the task may stop instead.

What the update means for Astra’s release

The clearest message is that Astra’s release appears to be close, but it will be accompanied by restrictions and additional oversight. OpenAI says the model will be made available soon, yet the most advanced cybersecurity capabilities will begin in a controlled testing phase rather than through unrestricted access.

The company plans to publish more information in Astra’s system card at launch, including details about its safety, security, alignment testing, and evaluations. Until then, the update provides evidence that OpenAI is preparing for deployment without establishing a public release date or describing the full set of launch features.

For cybersecurity professionals, the initial rollout could offer an opportunity to evaluate Astra’s defensive usefulness under tighter controls. For other users, the staged access suggests that Astra’s exceptional capabilities may arrive gradually as OpenAI continues testing whether its safeguards can keep pace with the model’s performance.