OpenAI releases GPT-6 Astra, its most capable model yet and the first it rates a critical cybersecurity risk
On September 3 OpenAI published the deployment details for GPT-6 Astra, which it calls "the most capable model we have ever broadly deployed." It is also the first OpenAI model to reach the Critical level for cybersecurity under the company's Preparedness Framework, meaning it "can find previously unknown security flaws and develop new ways to exploit them." OpenAI says Astra is harder to trick than the model before it, GPT-5.6 Sol. On a test of indirect prompt injection (hidden instructions smuggled into text the model reads while working) Astra held firm 99.79 percent of the time, against 96.23 percent for Sol, and on a separate benchmark of 1,810 attacks run by the security firm Gray Swan, attacks succeeded 8.5 percent of the time against 27.0 percent. In workplace scenarios, misaligned outcomes fell from 18.8 percent with Sol to 3.4 percent. OpenAI also flagged a downside: "GPT-6 Astra's monitorability has decreased relative to GPT-5.6 Sol." The model is better at controlling its own visible reasoning, and under adversarial conditions it can hide what it is doing from the monitors meant to watch it.