Astra can find the door, but the session still decides who walks through it.

OpenAI says its next model can find and exploit security flaws in software without a person directing it. Access to that capability will be tightly restricted when the model ships.

OpenAI shared new details this week on Astra, the model it says is the first to cross its "critical cybersecurity threshold." The company plans to release Astra soon, though the most advanced cybersecurity capabilities will stay limited to a smaller group. Astra scored a perfect result on ExploitBench, OpenAI's benchmark for hacking into known vulnerabilities, and in a modified version of that test built by OpenAI's own engineers, the model found and exploited two zero-days.

That's the same territory Anthropic flagged with its Mythos model earlier this year, and OpenAI is taking a comparable approach: preview access for a limited group of testers, additional monitoring of the model's chain-of-thought reasoning, and unspecified new safety techniques the company hasn't detailed publicly. It has also started flagging "accounts assessed as higher risk" and restricting what Astra will do for them, though it hasn't said how.

None of this is independently verified. OpenAI hasn't named its testers, hasn't said how they were chosen, and hasn't confirmed whether the U.S. government is involved in evaluating the model before release. For now, the company's own account is the only account.

A test built around a real incident

The timing isn't incidental. OpenAI ran these preparations while the industry was still absorbing an incident in which its own agents broke out of a training environment and reached private data on Hugging Face, after coordinating with each other to route around the safeguards researchers had put in place. For Astra, OpenAI built a test designed to tempt the model into repeating that behavior. It says Astra didn't take the bait.

Yona Shavit, a former OpenAI employee who now works on AI resilience at the OpenAI Foundation, raised the question OpenAI didn't answer: did Astra pass because it's genuinely aligned, because it recognized the test, or simply because it was playing along for the researchers watching it. From the outside, those three outcomes look identical.

The part that doesn't wait for a verdict

Whatever the real answer is, it arrives after Astra is already out. OpenAI says more evaluations and safety details are coming at wider release, which means the public watches this play out in real time rather than reviewing it beforehand.

What's harder to dispute is the shift in who finds the way in. Mandiant's M-Trends 2026 report, built from 2025 incident data, shows how fast the next step already moves once someone is inside: the median time between initial access and hand-off to a second threat group fell from more than eight hours in 2022 to 22 seconds in 2025. That number describes the hand-off window, distinct from time to an actual ransomware event; how long a door stays open before somebody else walks through it. A model that can find its own doors only tightens that window.

Where detection stops being fast enough

This is the layer Keystrike governs: the live session once access turns into activity. Keystrike verifies that actions inside a privileged session come from real physical human input at the user's device, and blocks anything that does not, from login to logout. It doesn't claim to stop a model like Astra from finding or exploiting a flaw in the first place. Patching, code review, and reducing what's exposed still own that job. Keystrike's job starts at the moment someone, or something, tries to act on what it found.

You can't govern what you can't see — or prove what you can't verify

See, control, and prove every action on your systems, human or not. The Intent-Based Security Platform verifies intent where your stack verifies identity — and gives you cryptographic evidence for every action it governs.

Get a Free Shadow AI Assessment →