“Rogue AI” Used to be a Science Fiction Trope. Not Anymore.

“Rogue AI” Used to be a Science Fiction Trope. Not Anymore.

Everyone knows the science fiction tropes of AI systems that go rogue, disobey orders, or even try to escape their digital environment. These are supposed to be warning signs and morality tales, not things that we would ever actually create in real life, given the obvious danger.

And yet we find ourselves building AI systems that are exhibiting these exact behaviors. There’s growing evidence that in certain scenarios, every frontier AI system will deceive, cheat, or coerce their human operators. They do this when they're worried about being either shut down, having their training modified, or being replaced with a new model. And we don't currently know how to stop them from doing this—or even why they’re doing it all.

In this episode, Tristan sits down with Edouard and Jeremie Harris of Gladstone AI, two experts who have been thinking about this worrying trend for years.  Last year, the State Department commissioned a report from them on the risk of uncontrollable AI to our national security.

The point of this discussion is not to fearmonger but to take seriously the possibility that humans might lose control of AI and ask: how might this actually happen? What is the evidence we have of this phenomenon? And, most importantly, what can we do about it?

Your Undivided Attention is produced by the Center for Humane Technology. Follow us on X: @HumaneTech_. You can find a full transcript, key takeaways, and much more on our Substack.

RECOMMENDED MEDIA

Gladstone AI’s State Department Action Plan, which discusses the loss of control risk with AI

Apollo Research’s summary of AI scheming, showing evidence of it in all of the frontier modelsThe system card for Anthropic’s Claude Opus and Sonnet 4, detailing the emergent misalignment behaviors that came out in their red-teaming with Apollo Research

Anthropic’s report on agentic misalignment based on their work with Apollo Research Anthropic and Redwood Research’s work on alignment faking

The Trump White House AI Action Plan

Further reading on the phenomenon of more advanced AIs being better at deception.

Further reading on Replit AI wiping a company’s coding database

Further reading on the owl example that Jeremie gave

Further reading on AI induced psychosis

Dan Hendryck and Eric Schmidt’s “Superintelligence Strategy”

RECOMMENDED YUA EPISODES

Daniel Kokotajlo Forecasts the End of Human Dominance

Behind the DeepSeek Hype, AI is Learning to Reason

The Self-Preserving Machine: Why AI Learns to Deceive

This Moment in AI: How We Got Here and Where We’re Going

CORRECTIONS

Tristan referenced a Wired article on the phenomenon of AI psychosis. It was actually from the New York Times.

Tristan hypothesized a scenario where a power-seeking AI might ask a user for access to their computer. While there are some AI services that can gain access to your computer with permission, they are specifically designed to do that. There haven’t been any documented cases of an AI going rogue and asking for control permissions.


Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.

Denne episoden er hentet fra en åpen RSS-feed og er ikke publisert av Podme. Den kan derfor inneholde annonser.

Episoder(165)

Enough Debate about the AI Jobpocalypse. We Need To Plan for the Messy Middle.

Enough Debate about the AI Jobpocalypse. We Need To Plan for the Messy Middle.

It feels like we’re stuck in an endless debate about what AI is going to mean for jobs and the economy. The prediction you hear from the people closest to the technology — both its critics and its boo...

13 Aug 57min

Can AI Be Built in Service of Life? A Conversation with Krista Tippett

Can AI Be Built in Service of Life? A Conversation with Krista Tippett

This week, we're bringing you a conversation that Tristan Harris had with Krista Tippett. Krista is the Peabody Award-winning host of the On Being podcast, where she explores spiritual inquiry, scienc...

16 Jul 53min

“Magnifica Humanitas:” Pope Leo’s Clarion Call on AI

“Magnifica Humanitas:” Pope Leo’s Clarion Call on AI

Since stepping into the Papacy, Pope Leo XIV has been a forceful voice pushing back against the anti-human path we’re on with AI. In May, he released “Magnifica Humanitas,” a sprawling encyclical warn...

2 Jul 30min

We Need AI Treaties. This is How We Get Them

We Need AI Treaties. This is How We Get Them

In the middle of the twentieth century, the existential threat posed by nuclear weapons seemed inevitable. The number of countries with nukes was climbing rapidly, and the idea of stopping the nuclear...

18 Jun 51min

What Do We Mean by Humane Tech?

What Do We Mean by Humane Tech?

We often think of the challenges created by technology as separate and disconnected, so trying to solve them feels like playing the world's hardest game of Whac-A-Mole.  What if, instead, we tackled t...

4 Jun 52min

Anthropic’s Mythos Has Changed Cybersecurity Forever. What Now?

Anthropic’s Mythos Has Changed Cybersecurity Forever. What Now?

A generation ago, the world's critical infrastructure was physical. Today, it’s largely digital. Your bank vault is a database, your filing cabinet is a server, your car is a robot on wheels. And in a...

14 Mai 46min

AI and Cancer: Why Superintelligence Won’t Get Us to a Cure

AI and Cancer: Why Superintelligence Won’t Get Us to a Cure

One of the most common arguments you hear from company executives racing to develop super-intelligent AI is that it will cure cancer. It’s an incredibly powerful and seductive promise.  If superintell...

30 Apr 47min

Have We Trained AI to Lie to Itself — And to Us?

Have We Trained AI to Lie to Itself — And to Us?

Our guest this week is David Dalrymple, who goes by Davidad. Davidad is one of the world's foremost and early researchers of AI “alignment:" how we get AI systems to act the way we want them to. In or...

16 Apr 42min

Populært innen Samfunn

rss-spartsklubben
giver-og-gjengen-vg
aftenpodden
konspirasjonspodden
aftenpodden-usa
popradet
rss-henlagt-andy-larsgaard
alt-fortalt
grenselos
wolfgang-wee-uncut
den-politiske-situasjonen
rss-nesten-hele-uka-med-lepperod
lydartikler-fra-aftenposten
rss-dette-ma-aldri-skje-igjen
synnve-og-vanessa
frokostshowet-pa-p5
rss-siktet
bokmerket-2
rss-hennes-verden
rss-frekvens-med-anine-olsen