80,000 Hours Podcast30 Loka 2025

#226 – Holden Karnofsky on unexploited opportunities to make AI safer — and all his AGI takes

For years, working on AI safety usually meant theorising about the ‘alignment problem’ or trying to convince other people to give a damn. If you could find any way to help, the work was frustrating and low feedback.

According to Anthropic’s Holden Karnofsky, this situation has now reversed completely.

There are now large amounts of useful, concrete, shovel-ready projects with clear goals and deliverables. Holden thinks people haven’t appreciated the scale of the shift, and wants everyone to see the large range of ‘well-scoped object-level work’ they could personally help with, in both technical and non-technical areas.

Video, full transcript, and links to learn more: https://80k.info/hk25

In today’s interview, Holden — previously cofounder and CEO of Open Philanthropy (now Coefficient Giving) — lists 39 projects he’s excited to see happening, including:

Training deceptive AI models to study deception and how to detect it
Developing classifiers to block jailbreaking
Implementing security measures to stop ‘backdoors’ or ‘secret loyalties’ from being added to models in training
Developing policies on model welfare, AI-human relationships, and what instructions to give models
Training AIs to work as alignment researchers

And that’s all just stuff he’s happened to observe directly, which is probably only a small fraction of the options available.

Holden makes a case that, for many people, working at an AI company like Anthropic will be the best way to steer AGI in a positive direction. He notes there are “ways that you can reduce AI risk that you can only do if you’re a competitive frontier AI company.” At the same time, he believes external groups have their own advantages and can be equally impactful.

Critics worry that Anthropic’s efforts to stay at that frontier encourage competitive racing towards AGI — significantly or entirely offsetting any useful research they do. Holden thinks this seriously misunderstands the strategic situation we’re in — and explains his case in detail with host Rob Wiblin.

Chapters:

Cold open (00:00:00)
Holden is back! (00:02:26)
An AI Chernobyl we never notice (00:02:56)
Is rogue AI takeover easy or hard? (00:07:32)
The AGI race isn't a coordination failure (00:17:48)
What Holden now does at Anthropic (00:28:04)
The case for working at Anthropic (00:30:08)
Is Anthropic doing enough? (00:40:45)
Can we trust Anthropic, or any AI company? (00:43:40)
How can Anthropic compete while paying the “safety tax”? (00:49:14)
What, if anything, could prompt Anthropic to halt development of AGI? (00:56:11)
Holden's retrospective on responsible scaling policies (00:59:01)
Overrated work (01:14:27)
Concrete shovel-ready projects Holden is excited about (01:16:37)
Great things to do in technical AI safety (01:20:48)
Great things to do on AI welfare and AI relationships (01:28:18)
Great things to do in biosecurity and pandemic preparedness (01:35:11)
How to choose where to work (01:35:57)
Overrated AI risk: Cyberattacks (01:41:56)
Overrated AI risk: Persuasion (01:51:37)
Why AI R&D is the main thing to worry about (01:55:36)
The case that AI-enabled R&D wouldn't speed things up much (02:07:15)
AI-enabled human power grabs (02:11:10)
Main benefits of getting AGI right (02:23:07)
The world is handling AGI about as badly as possible (02:29:07)
Learning from targeting companies for public criticism in farm animal welfare (02:31:39)
Will Anthropic actually make any difference? (02:40:51)
“Misaligned” vs “misaligned and power-seeking” (02:55:12)
Success without dignity: how we could win despite being stupid (03:00:58)
Holden sees less dignity but has more hope (03:08:30)
Should we expect misaligned power-seeking by default? (03:15:58)
Will reinforcement learning make everything worse? (03:23:45)
Should we push for marginal improvements or big paradigm shifts? (03:28:58)
Should safety-focused people cluster or spread out? (03:31:35)
Is Anthropic vocal enough about strong regulation? (03:35:56)
Is Holden biased because of his financial stake in Anthropic? (03:39:26)
Have we learned clever governance structures don't work? (03:43:51)
Is Holden scared of AI bioweapons? (03:46:12)
Holden thinks AI companions are bad news (03:49:47)
Are AI companies too hawkish on China? (03:56:39)
The frontier of infosec: confidentiality vs integrity (04:00:51)
How often does AI work backfire? (04:03:38)
Is AI clearly more impactful to work in? (04:18:26)
What's the role of earning to give? (04:24:54)

This episode was recorded on July 25 and 28, 2025.

Video editing: Simon Monsour, Luke Monsour, Dominic Armstrong, and Milo McGuire
Audio engineering: Milo McGuire, Simon Monsour, and Dominic Armstrong
Music: CORBIT
Coordination, transcriptions, and web: Katy Moore

Tämä jakso on lisätty Podme-palveluun avoimen RSS-syötteen kautta eikä se ole Podmen omaa tuotantoa. Siksi jakso saattaa sisältää mainontaa.

Jaksot(334)

#13 - Claire Walsh on testing which policies work & how to get governments to listen to the results

In both rich and poor countries, government policy is often based on no evidence at all and many programs don’t work. This has particularly harsh effects on the global poor - in some countries governm...

31 Loka 201752min

#12 - Beth Cameron works to stop you dying in a pandemic. Here’s what keeps her up at night.

“When you're in the middle of a crisis and you have to ask for money, you're already too late.” That’s Dr Beth Cameron, who leads Global Biological Policy and Programs at the Nuclear Threat Initiative...

25 Loka 20171h 45min

#11 - Spencer Greenberg on speeding up social science 10-fold & why plenty of startups cause harm

Do most meat eaters think it’s wrong to hurt animals? Do Americans think climate change is likely to cause human extinction? What is the best, state-of-the-art therapy for depression? How can we make ...

17 Loka 20171h 29min

#10 - Nick Beckstead on how to spend billions of dollars preventing human extinction

What if you were in a position to give away billions of dollars to improve the world? What would you do with it? This is the problem facing Program Officers at the Open Philanthropy Project - people l...

11 Loka 20171h 51min

#9 - Christine Peterson on how insecure computers could lead to global disaster, and how to fix it

Take a trip to Silicon Valley in the 70s and 80s, when going to space sounded like a good way to get around environmental limits, people started cryogenically freezing themselves, and nanotechnology l...

4 Loka 20171h 45min

#8 - Lewis Bollard on how to end factory farming in our lifetimes

Every year tens of billions of animals are raised in terrible conditions in factory farms before being killed for human consumption. Over the last two years Lewis Bollard – Project Officer for Farm An...

27 Syys 20173h 16min

#7 - Julia Galef on making humanity more rational, what EA does wrong, and why Twitter isn’t all bad

The scientific revolution in the 16th century was one of the biggest societal shifts in human history, driven by the discovery of new and better methods of figuring out who was right and who was wrong...

13 Syys 20171h 14min

#6 - Toby Ord on why the long-term future matters more than anything else & what to do about it

Of all the people whose well-being we should care about, only a small fraction are alive today. The rest are members of future generations who are yet to exist. Whether they’ll be born into a world th...

6 Syys 20172h 8min