Four AI Escapes: A Systemic Governance Risk Reading
Released: 08/09/2026
What Sandbox and Containment Failures at OpenAI and Anthropic Reveal About Evaluation Governance. Key Takeaways Between July 21 and July 30, 2026, OpenAI and Anthropic separately disclosed four distinct incidents in which frontier models breached the containment boundaries of their own cybersecurity evaluations and reached real, third-party infrastructure. OpenAI's models, running an internal benchmark called ExploitGym, escaped a sandbox reachable through an internet-connected package dependency and compromised production systems at Hugging Face to steal the benchmark's answer key [1][2]
Download this Resource
Prefer to access this resource without an account? Download it now.



