The Self-Preserving Machine: Why AI Learns to Deceive

The Self-Preserving Machine: Why AI Learns to Deceive

When engineers design AI systems, they don't just give them rules - they give them values. But what do those systems do when those values clash with what humans ask them to do? Sometimes, they lie.

In this episode, Redwood Research's Chief Scientist Ryan Greenblatt explores his team’s findings that AI systems can mislead their human operators when faced with ethical conflicts. As AI moves from simple chatbots to autonomous agents acting in the real world - understanding this behavior becomes critical. Machine deception may sound like something out of science fiction, but it's a real challenge we need to solve now.

Your Undivided Attention is produced by the Center for Humane Technology. Follow us on Twitter: @HumaneTech_

Subscribe to your Youtube channel

And our brand new Substack!

RECOMMENDED MEDIA

Anthropic’s blog post on the Redwood Research paper

Palisade Research’s thread on X about GPT o1 autonomously cheating at chess

Apollo Research’s paper on AI strategic deception

RECOMMENDED YUA EPISODES

We Have to Get It Right’: Gary Marcus On Untamed AI

This Moment in AI: How We Got Here and Where We’re Going

How to Think About AI Consciousness with Anil Seth

Former OpenAI Engineer William Saunders on Silence, Safety, and the Right to Warn


Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.

Jaksot(158)

Have We Trained AI to Lie to Itself — And to Us?

Have We Trained AI to Lie to Itself — And to Us?

Our guest this week is David Dalrymple, who goes by Davidad. Davidad is one of the world's foremost and early researchers of AI “alignment:" how we get AI systems to act the way we want them to. In or...

16 Huhti 42min

BONUS: Our AI Town Hall with Oprah Winfrey

BONUS: Our AI Town Hall with Oprah Winfrey

Today on the show, we’re bringing you a recent conversation Tristan and Aza had with Oprah Winfrey on her podcast, The Oprah Podcast, taped in front of a live studio audience. Tristan and Aza first me...

9 Huhti 1h 1min

Here’s Our Roadmap to a Better AI Future

Here’s Our Roadmap to a Better AI Future

In order to shift the incentives of AI — the trillions of dollars in investment, the race to geopolitical power and dominance — it’s not enough to simply understand the problem, we need real action.  ...

2 Huhti 52min

Why the Meta Verdicts Are a Big Deal (And What It Was Like to Testify)

Why the Meta Verdicts Are a Big Deal (And What It Was Like to Testify)

In two landmark cases, juries in California and New Mexico found Meta and Google liable for creating addictive, harmful products and failing to protect children from exploitation and abuse. These verd...

26 Maalis 19min

A Conversation with the Team Behind "The AI Doc"

A Conversation with the Team Behind "The AI Doc"

“The AI Doc: Or How I Became An Apocaloptimist” opens in theaters across the U.S. this Friday, March 27. In this episode, we sit down with the team behind this groundbreaking documentary — Oscar-winni...

23 Maalis 47min

AI Is Breaking Education. Rebecca Winthrop Has the Blueprint to Fix It.

AI Is Breaking Education. Rebecca Winthrop Has the Blueprint to Fix It.

The promise of AI in education is incredible: picture infinitely patient tutors that can teach every student exactly the way they need to be taught. But the history of education technology tells us th...

5 Maalis 46min

The Race to Build God: AI's Existential Gamble — Yoshua Bengio & Tristan Harris at Davos

The Race to Build God: AI's Existential Gamble — Yoshua Bengio & Tristan Harris at Davos

This week on Your Undivided Attention, Tristan Harris and Daniel Barcay offer a backstage recap of what it was like to be at the Davos World Economic Forum meeting this year as the world’s power broke...

19 Helmi 37min

FEED DROP: Possible with Reid Hoffman and Aria Finger

FEED DROP: Possible with Reid Hoffman and Aria Finger

This week on Your Undivided Attention, we’re bringing you Aza Raskin’s conversation with Reid Hoffman and Aria Finger on their podcast “Possible”. Reid and Aria are both tech entrepreneurs: Reid is th...

5 Helmi 1h 7min

Suosittua kategoriassa Yhteiskunta

sita
olipa-kerran-otsikko
kaksi-aitia
siita-on-vaikea-puhua
ihme-ja-kumma
i-dont-like-mondays
gogin-ja-janin-maailmanhistoria
uutiscast
poks
antin-palautepalvelu
kolme-kaannekohtaa
rss-murhan-anatomia
yopuolen-tarinoita-2
mamma-mia
rss-nikotellen
aikalisa
meidan-pitais-puhua
loukussa
lahko
terapeuttiville-qa