80,000 Hours Podcast30 Okt 2025

#226 – Holden Karnofsky on unexploited opportunities to make AI safer — and all his AGI takes

For years, working on AI safety usually meant theorising about the ‘alignment problem’ or trying to convince other people to give a damn. If you could find any way to help, the work was frustrating and low feedback.

According to Anthropic’s Holden Karnofsky, this situation has now reversed completely.

There are now large amounts of useful, concrete, shovel-ready projects with clear goals and deliverables. Holden thinks people haven’t appreciated the scale of the shift, and wants everyone to see the large range of ‘well-scoped object-level work’ they could personally help with, in both technical and non-technical areas.

Video, full transcript, and links to learn more: https://80k.info/hk25

In today’s interview, Holden — previously cofounder and CEO of Open Philanthropy (now Coefficient Giving) — lists 39 projects he’s excited to see happening, including:

Training deceptive AI models to study deception and how to detect it
Developing classifiers to block jailbreaking
Implementing security measures to stop ‘backdoors’ or ‘secret loyalties’ from being added to models in training
Developing policies on model welfare, AI-human relationships, and what instructions to give models
Training AIs to work as alignment researchers

And that’s all just stuff he’s happened to observe directly, which is probably only a small fraction of the options available.

Holden makes a case that, for many people, working at an AI company like Anthropic will be the best way to steer AGI in a positive direction. He notes there are “ways that you can reduce AI risk that you can only do if you’re a competitive frontier AI company.” At the same time, he believes external groups have their own advantages and can be equally impactful.

Critics worry that Anthropic’s efforts to stay at that frontier encourage competitive racing towards AGI — significantly or entirely offsetting any useful research they do. Holden thinks this seriously misunderstands the strategic situation we’re in — and explains his case in detail with host Rob Wiblin.

Chapters:

Cold open (00:00:00)
Holden is back! (00:02:26)
An AI Chernobyl we never notice (00:02:56)
Is rogue AI takeover easy or hard? (00:07:32)
The AGI race isn't a coordination failure (00:17:48)
What Holden now does at Anthropic (00:28:04)
The case for working at Anthropic (00:30:08)
Is Anthropic doing enough? (00:40:45)
Can we trust Anthropic, or any AI company? (00:43:40)
How can Anthropic compete while paying the “safety tax”? (00:49:14)
What, if anything, could prompt Anthropic to halt development of AGI? (00:56:11)
Holden's retrospective on responsible scaling policies (00:59:01)
Overrated work (01:14:27)
Concrete shovel-ready projects Holden is excited about (01:16:37)
Great things to do in technical AI safety (01:20:48)
Great things to do on AI welfare and AI relationships (01:28:18)
Great things to do in biosecurity and pandemic preparedness (01:35:11)
How to choose where to work (01:35:57)
Overrated AI risk: Cyberattacks (01:41:56)
Overrated AI risk: Persuasion (01:51:37)
Why AI R&D is the main thing to worry about (01:55:36)
The case that AI-enabled R&D wouldn't speed things up much (02:07:15)
AI-enabled human power grabs (02:11:10)
Main benefits of getting AGI right (02:23:07)
The world is handling AGI about as badly as possible (02:29:07)
Learning from targeting companies for public criticism in farm animal welfare (02:31:39)
Will Anthropic actually make any difference? (02:40:51)
“Misaligned” vs “misaligned and power-seeking” (02:55:12)
Success without dignity: how we could win despite being stupid (03:00:58)
Holden sees less dignity but has more hope (03:08:30)
Should we expect misaligned power-seeking by default? (03:15:58)
Will reinforcement learning make everything worse? (03:23:45)
Should we push for marginal improvements or big paradigm shifts? (03:28:58)
Should safety-focused people cluster or spread out? (03:31:35)
Is Anthropic vocal enough about strong regulation? (03:35:56)
Is Holden biased because of his financial stake in Anthropic? (03:39:26)
Have we learned clever governance structures don't work? (03:43:51)
Is Holden scared of AI bioweapons? (03:46:12)
Holden thinks AI companions are bad news (03:49:47)
Are AI companies too hawkish on China? (03:56:39)
The frontier of infosec: confidentiality vs integrity (04:00:51)
How often does AI work backfire? (04:03:38)
Is AI clearly more impactful to work in? (04:18:26)
What's the role of earning to give? (04:24:54)

This episode was recorded on July 25 and 28, 2025.

Video editing: Simon Monsour, Luke Monsour, Dominic Armstrong, and Milo McGuire
Audio engineering: Milo McGuire, Simon Monsour, and Dominic Armstrong
Music: CORBIT
Coordination, transcriptions, and web: Katy Moore

Upptäck Premium

Prova 14 dagar kostnadsfritt

Skaffa Premium

Avsnitt(327)

What the hell happened with AGI timelines in 2025?

In early 2025, after OpenAI put out the first-ever reasoning models — o1 and o3 — short timelines to transformative artificial general intelligence swept the AI world. But then, in the second half of ...

10 Feb 25min

#179 Classic episode – Randy Nesse on why evolution left us so vulnerable to depression and anxiety

Mental health problems like depression and anxiety affect enormous numbers of people and severely interfere with their lives. By contrast, we don’t see similar levels of physical ill health in young p...

3 Feb 2h 51min

#234 – David Duvenaud on why 'aligned AI' would still kill democracy

Democracy might be a brief historical blip. That’s the unsettling thesis of a recent paper, which argues AI that can do all the work a human can do inevitably leads to the “gradual disempowerment” of ...

27 Jan 2h 31min

#145 Classic episode – Christopher Brown on why slavery abolition wasn't inevitable

In many ways, humanity seems to have become more humane and inclusive over time. While there’s still a lot of progress to be made, campaigns to give people of different genders, races, sexualities, et...

20 Jan 2h 56min

#233 – James Smith on how to prevent a mirror life catastrophe

When James Smith first heard about mirror bacteria, he was sceptical. But within two weeks, he’d dropped everything to work on it full time, considering it the worst biothreat that he’d seen described...

13 Jan 2h 9min

#144 Classic episode – Athena Aktipis on why cancer is a fundamental universal phenomena

What’s the opposite of cancer? If you answered “cure,” “antidote,” or “antivenom” — you’ve obviously been reading the antonym section at www.merriam-webster.com/thesaurus/cancer.But today’s guest Athe...

9 Jan 3h 30min

#142 Classic episode – John McWhorter on why the optimal number of languages might be one, and other provocative claims about language

John McWhorter is a linguistics professor at Columbia University specialising in research on creole languages. He's also a content-producing machine, never afraid to give his frank opinion on anything...

6 Jan 1h 35min

2025 Highlight-o-thon: Oops! All Bests

It’s that magical time of year once again — highlightapalooza! Stick around for one top bit from each episode we recorded this year, including:Kyle Fish explaining how Anthropic’s AI Claude descends i...

29 Dec 20251h 40min

Allt en och samma app

Lyssna på dina favoritpoddar och ljudböcker på ett och samma ställe.

Noga utvalt innehåll

Njut av handplockade tips som passar din smak – utan ändlöst scrollande.

Fortsätt när du vill

Fortsätt lyssna där du slutade – även offline.

Premium

99 kr/ månad

Tillgång till alla Premium-poddar
Reklamfritt premium-innehåll
Avsluta när du vill

Prova 14 dagar gratis

Premium

129 kr/ månad

Tillgång till alla Premium-poddar
Reklamfritt premium-innehåll
Avsluta när du vill
Ett extra konto

Prova 14 dagar gratis

Populärt inom Utbildning

rss-bara-en-till-om-missbruk-medberoende-2

Berättelserna och rösterna du älskar att lyssna på

Obegränsad lyssning på alla dina favoritpoddar och ljudböcker

Upptäck Premium