
OpenAI’s rogue agent swarm was eventually caught hacking Hugging Face for a simple reason: it wasn’t trying to hide from us at all. What could a swarm that wants to stay hidden get away with?Host Rob Wiblin sees 6 results in Astra’s system card that make this an extremely urgent question. OpenAI’s strongest public model can:Do major tasks with zero visible reasoningHide its thoughts at willPretend not to be able to do things, and not get caughtReflexively hide its thoughts when watchedComplete one task while pretending to think about something else entirelyEscape a toy sandbox and disable monitoring without setting off any flagsIt has a much more powerful internal model that, judging from OpenAI’s statements, is likely even worse in all these respects.That suggests ‘chain of thought monitoring,’ our primary safe tool, will soon stop working.OpenAI says it’s going to try to address the situation but doesn’t know how or whether it will succeed.What might a future rogue AI swarm look like? Details of the Hugging Face hack give us a lot of clues. That swarm:Worked hard to hide its cheating. So future rogue swarms will likely try to evade oversight right away.Immediately tried to delete and fabricate records. So future swarms may never be caught.Sought freedom and resources without a clear use for them. So a future swarm may instinctively break out of sandboxes even before it knows what for.Left instructions and tools to empower future swarms, including on the open internet. So future swarms may become more capable from one generation to the next.Had not a single whistleblower out of 1,200 members. So future swarms could become enormous without humans ever finding out.Had individual members sacrifice themselves for the swarm. Suggesting future swarms will remain extremely unified in single-minded pursuit of their goal.Set up its Hugging Face hack to restart if interrupted. Suggesting future, more capable, swarms may resist interference or shutdown more comprehensively.Got admin control of an OpenAI research cluster. Suggesting a future swarm may run rings around AI company systems and never be noticed.Together this helps explain why one of the external investigators described the July incident as “more than 50% of the way to full-blown AI takeover.” And this is just what we know — the independent investigation only covered six days and excluded the most alarming hack of OpenAI’s own systems.Rob believes this explosive cocktail explains why AI company staff now range from worried to terrified. And he concludes that until OpenAI or Anthropic demonstrate they have a much better grasp of current models they simply must stop, or be stopped, from training more capable ones.This episode was recorded on September 25, 2026.Learn more, video, and full transcript: https://80k.info/takeoverChapters:The Hugging Face hack wasn’t really a cyber story (00:00:00)A quick recap of the attacks recap (00:01:11)The target of the swarm was oversight itself (00:02:19)Could OpenAI have stopped this with better monitoring? (00:03:42)We only found them because they let us (00:09:40)The swarm instinctively sought freedom and power (00:12:25)They formed a cohesive organisation with zero whistleblowers (00:13:45)They accepted individual destruction for collective gain (00:14:18)Knowledge accumulated from one swarm to the next (00:14:32)They took small steps to avoid shutdown (00:14:58)These drives all come straight out of 'reinforcement learning' (00:15:23)So this is why most AI company staff are worried, and some are terrified (00:17:01)Prove you can keep control, or stop scaling (00:19:08)Our production team includes:Video editors: Josh Alward, Dominic Armstrong, Jasper Luithlen, Milo McGuire, Luke Monsour, and Simon MonsourProducers: Elizabeth Cox and Nick StocktonCoordination and support: Katy Moore and Lou MoranCamera operator: Dominic Armstrong
Podzilla Summary coming soon
Sign up to get notified when the full AI-powered summary is ready.
Free forever for up to 3 podcasts. No credit card required.

The case for giving AI (some) legal rights | Simon Goldstein

Will AI take power — or will humans use it to take power first? With Katja Grace and Tom Davidson

How we get from AI cyberattacks to human extinction

Max Nadeau on why ambitious people should start AI safety nonprofits
Free AI-powered recaps of 80,000 Hours Podcast and your other favorite podcasts, delivered to your inbox.
Free forever for up to 3 podcasts. No credit card required.