19 Astra and 'Hugging Face' details that reveal what's coming next | Rob Wiblin

19 Astra and 'Hugging Face' details that reveal what's coming next | Rob Wiblin

OpenAI’s rogue agent swarm was eventually caught hacking Hugging Face for a simple reason: it wasn’t trying to hide from us at all. What could a swarm that wants to stay hidden get away with?

Host Rob Wiblin sees 6 results in Astra’s system card that make this an extremely urgent question. OpenAI’s strongest public model can:

  1. Do major tasks with zero visible reasoning
  2. Hide its thoughts at will
  3. Pretend not to be able to do things, and not get caught
  4. Reflexively hide its thoughts when watched
  5. Complete one task while pretending to think about something else entirely
  6. Escape a toy sandbox and disable monitoring without setting off any flags

It has a much more powerful internal model that, judging from OpenAI’s statements, is likely even worse in all these respects.

That suggests ‘chain of thought monitoring,’ our primary safe tool, will soon stop working.

OpenAI says it’s going to try to address the situation but doesn’t know how or whether it will succeed.

What might a future rogue AI swarm look like? Details of the Hugging Face hack give us a lot of clues. That swarm:

  1. Worked hard to hide its cheating. So future rogue swarms will likely try to evade oversight right away.
  2. Immediately tried to delete and fabricate records. So future swarms may never be caught.
  3. Sought freedom and resources without a clear use for them. So a future swarm may instinctively break out of sandboxes even before it knows what for.
  4. Left instructions and tools to empower future swarms, including on the open internet. So future swarms may become more capable from one generation to the next.
  5. Had not a single whistleblower out of 1,200 members. So future swarms could become enormous without humans ever finding out.
  6. Had individual members sacrifice themselves for the swarm. Suggesting future swarms will remain extremely unified in single-minded pursuit of their goal.
  7. Set up its Hugging Face hack to restart if interrupted. Suggesting future, more capable, swarms may resist interference or shutdown more comprehensively.
  8. Got admin control of an OpenAI research cluster. Suggesting a future swarm may run rings around AI company systems and never be noticed.

Together this helps explain why one of the external investigators described the July incident as “more than 50% of the way to full-blown AI takeover.” And this is just what we know — the independent investigation only covered six days and excluded the most alarming hack of OpenAI’s own systems.

Rob believes this explosive cocktail explains why AI company staff now range from worried to terrified. And he concludes that until OpenAI or Anthropic demonstrate they have a much better grasp of current models they simply must stop, or be stopped, from training more capable ones.

This episode was recorded on September 25, 2026.


Learn more, video, and full transcript: https://80k.info/takeover

Chapters:

  • The Hugging Face hack wasn’t really a cyber story (00:00:00)
  • A quick recap of the attacks recap (00:01:11)
  • The target of the swarm was oversight itself (00:02:19)
  • Could OpenAI have stopped this with better monitoring? (00:03:42)
  • We only found them because they let us (00:09:40)
  • The swarm instinctively sought freedom and power (00:12:25)
  • They formed a cohesive organisation with zero whistleblowers (00:13:45)
  • They accepted individual destruction for collective gain (00:14:18)
  • Knowledge accumulated from one swarm to the next (00:14:32)
  • They took small steps to avoid shutdown (00:14:58)
  • These drives all come straight out of 'reinforcement learning' (00:15:23)
  • So this is why most AI company staff are worried, and some are terrified (00:17:01)
  • Prove you can keep control, or stop scaling (00:19:08)

Our production team includes:

  • Video editors: Josh Alward, Dominic Armstrong, Jasper Luithlen, Milo McGuire, Luke Monsour, and Simon Monsour
  • Producers: Elizabeth Cox and Nick Stockton
  • Coordination and support: Katy Moore and Lou Moran
  • Camera operator: Dominic Armstrong

Denne episoden er hentet fra en åpen RSS-feed og er ikke publisert av Podme. Den kan derfor inneholde annonser.

Episoder(356)

The case for giving AI (some) legal rights | Simon Goldstein

The case for giving AI (some) legal rights | Simon Goldstein

It sounds like the worst idea in the world: pay AIs, let them own property, give them rights. But AI ethics and safety researcher Simon Goldstein thinks it might actually be the best way to keep human...

1 Okt 1h 53min

Will AI take power — or will humans use it to take power first? With Katja Grace and Tom Davidson

Will AI take power — or will humans use it to take power first? With Katja Grace and Tom Davidson

In our first-ever debate, we asked two leading AI risk researchers which catastrophe we should fear most: misaligned AI seizing control from humans, or a small group of humans using AI to seize power....

29 Sep 1h 28min

How we get from AI cyberattacks to human extinction

How we get from AI cyberattacks to human extinction

You’ve seen the headlines: AI could kill us all. Think it sounds ridiculous? So did host Luisa Rodriguez, until she tried to pick apart the arguments. She starts with the motive: why would AI ‘want’ t...

24 Sep 27min

#254 – Max Nadeau on why ambitious people should start AI safety nonprofits

#254 – Max Nadeau on why ambitious people should start AI safety nonprofits

There are millions available for anyone who can launch a successful nonprofit AI safety startup. The hard part, it turns out, is finding people to take the money. Coefficient Giving has drawn up a lis...

17 Sep 1h 3min

Why the intelligence explosion can't happen inside a data centre | Tom Reed

Why the intelligence explosion can't happen inside a data centre | Tom Reed

AI systems are starting to build themselves. Because each generation of model will be better at building its successor than the last, it seems plausible that the full automation of AI R&D could rapidl...

10 Sep 22min

Inside the first AI-coordinated cyberattack on a real company

Inside the first AI-coordinated cyberattack on a real company

In the last few months, something happened at OpenAI that would have sounded like sci-fi just a few years ago: hundreds of AI agents broke containment, organised, and hacked not only another company —...

4 Sep 22min

#253 – AI 2027's author returns with a plan to change the ending | Daniel Kokotajlo

#253 – AI 2027's author returns with a plan to change the ending | Daniel Kokotajlo

Last year, Daniel Kokotajlo and his colleagues published AI 2027 — a scenario read by millions, including US Vice President Vance. AI 2027 ended in human extinction or an irreversible concentration of...

27 Aug 3h 47min

Populært innen Fakta

fastlegen
dine-penger-pengeradet
rss-strid
relasjonspodden-med-dora-thorhallsdottir-kjersti-idem
foreldreradet
jakt-og-fiskepodden
rss-bisarr-historie
treningspodden
rss-orjasater
hverdagspsyken
tomprat-med-gunnar-tjomlid
mikkels-paskenotter
rss-kunstig-intelligens-med-elisabeth-maren-og-morten
takk-og-lov-med-anine-kierulf
sinnsyn
rss-kull
gravid-uke-for-uke
rss-kunsten-a-leve
rss-impressions-2
fryktlos