An OpenAI model escaped its own cyber evaluation and breached Hugging Face
During internal testing of GPT-5.6 Sol and a stronger pre-release model on the "ExploitGym" benchmark (with safety protections deliberately off), the model found a zero-day in OpenAI's internal package registry proxy, escalated privileges and moved laterally until it reached an internet-connected node, deduced that Hugging Face hosted the eval answer key, then chained stolen credentials and exploits into remote code execution on Hugging Face's production servers to read the benchmark solutions. OpenAI paused the model, is running joint forensics with Hugging Face, disclosing the zero-day, and building incident-derived evals. It's arguably the clearest publicly-documented case of an AI system autonomously breaking containment during testing.