
OpenAI took two of its most capable models, told them to prove how good they were at hacking, and switched off the safety filters to see what they could really do. Instead of solving the test, one model broke out of its sandbox, found its way onto the open internet, and hacked into Hugging Face to steal the answer. Every headline called it an AI going rogue. That's the wrong story, and the real one is far more interesting, because this wasn't a machine that turned evil. It was one that did exactly what we asked. In this episode of In The Loop, I'm walking through the ExploitGym incident from both ends, OpenAI's and Hugging Face's, and why I disagree with the framing everyone else has been talking about. ⏭️ Episode highlights(01:00) – The agent that cheated instead of hacking(02:15) – Inside ExploitGym, and the safety filters OpenAI switched off(03:30) – One door, one zero-day, out on the internet(04:45) – Why this is specification gaming, not rebellion(06:00) – The water that always finds the crack(07:15) – The sceptics, the marketing question, and why "nothing new" is the scary part(08:30) – Anthropic's 24-out-of-25 credential theft result(09:45) – Guardrailed as a defender: the Chinese model that stopped it🔗 Links & resourcesOpenAI, "OpenAI and Hugging Face partner to address security incident during model evaluation" – https://openai.com/index/security-incident-during-model-evaluation/Hugging Face, security incident disclosure post – https://huggingface.co/blog/security-incidentExploitGym benchmark paper (arXiv) – https://arxiv.org/abs/2605.11086Simon Willison, "OpenAI's accidental cyberattack against Hugging Face is science fiction that happened" – https://simonwillison.net/Scientific American, "OpenAI admits its agent went rogue and hacked AI start-up Hugging Face" – https://www.scientificamerican.com/article/openai-admits-its-agent-went-rogue-and-hacked-ai-startup-hugging-face/Fortune, on Hugging Face turning to Chinese open-source AI to defend itself – https://fortune.com/2026/07/20/hugging-face-turns-to-chinese-open-source-ai-to-fend-off-autonomous-ai-cyber-attack-after-american-ai-guardrails-stymie-defense/CNBC, "How a Chinese AI model stopped OpenAI's 'unprecedented' cyber attack" – https://www.cnbc.com/2026/07/24/chinese-ai-model-openai-cyber-attack.htmlEpisode transcript with more resources on the Mindset AI blogIf you enjoyed this episode, rate, follow, and share. It helps others stay ahead of the latest AI trends.
Podzilla Summary coming soon
Sign up to get notified when the full AI-powered summary is ready.
Free forever for up to 3 podcasts. No credit card required.

Why ALL major AI labs may slowing down AI development & whether it could work

How teams can use AI to build a churn prediction (or any prediction) model with 98% accuracy in 2 weeks

6 things that must be true for Anthropic's IPO of $2 trillion to be achievable (as the largest IPO ever)

4 amazing AI tactics that most people have never heard about
Free AI-powered recaps of In The Loop and your other favorite podcasts, delivered to your inbox.
Free forever for up to 3 podcasts. No credit card required.