#226 – Holden Karnofsky on unexploited opportunities to make AI safer — and all his AGI takes

#226 – Holden Karnofsky on unexploited opportunities to make AI safer — and all his AGI takes

For years, working on AI safety usually meant theorising about the ‘alignment problem’ or trying to convince other people to give a damn. If you could find any way to help, the work was frustrating and low feedback.

According to Anthropic’s Holden Karnofsky, this situation has now reversed completely.

There are now large amounts of useful, concrete, shovel-ready projects with clear goals and deliverables. Holden thinks people haven’t appreciated the scale of the shift, and wants everyone to see the large range of ‘well-scoped object-level work’ they could personally help with, in both technical and non-technical areas.

In today’s interview, Holden — previously cofounder and CEO of Open Philanthropy (now Coefficient Giving) — lists 39 projects he’s excited to see happening, including:

  • Training deceptive AI models to study deception and how to detect it
  • Developing classifiers to block jailbreaking
  • Implementing security measures to stop ‘backdoors’ or ‘secret loyalties’ from being added to models in training
  • Developing policies on model welfare, AI-human relationships, and what instructions to give models
  • Training AIs to work as alignment researchers

And that’s all just stuff he’s happened to observe directly, which is probably only a small fraction of the options available.

All this low-hanging fruit is one factor behind his decision to join Anthropic this year. That said, his wife was also a cofounder and president of the company, giving him a big financial stake in its success — and making it impossible for him to be seen as independent no matter where he worked.

Holden makes a case that, for many people, working at an AI company like Anthropic will be the best way to steer AGI in a positive direction. He notes there are “ways that you can reduce AI risk that you can only do if you’re a competitive frontier AI company.” At the same time, he believes external groups have their own advantages and can be equally impactful.

Outside critics worry that Anthropic’s efforts to stay at that frontier encourage competitive racing towards AGI — significantly or entirely offsetting any useful research they do. Holden thinks this seriously misunderstands the strategic situation we’re in.

“I work at an AI company, and a lot of people think that’s just inherently unethical,” he says. “They’re imagining [that] everyone wishes they could go slowly, but they’re going fast so they can beat everyone else. […] But I emphatically think this is not what’s going on in AI.”

The reality, in Holden’s view:

I think there’s too many players in AI who […] don’t want to slow down. They don’t believe in the risks. Maybe they don’t even care about the risks. […] If Anthropic were to say, “We’re out, we’re going to slow down,” they would say, ‘This is awesome! Now we have a better chance of winning, and this is even good for our recruiting” — because they have a better chance of getting people who want to be on the frontier and want to win.

Holden believes a frontier AI company can reduce risk by:

  • Developing cheap, practical safety measures other companies might adopt
  • Prototyping policies regulators could mandate
  • Gathering crucial data about what advanced AI can actually do

Host Rob Wiblin and Holden discuss the case for and against those strategies, and much more, in today’s episode.

Learn more and read the full transcript on the 80,000 Hours website.


Chapters:

• Cold open (00:00:00)
• Holden is back! (00:02:28)
• An AI Chernobyl we never notice (00:02:58)
• Is rogue AI takeover easy or hard? (00:07:39)
• The AGI race isn't a coordination failure (00:18:01)
• What Holden now does at Anthropic (00:28:30)
• The case for working at Anthropic (00:30:38)
• Is Anthropic doing enough? (00:41:30)
• Can we trust Anthropic, or any AI company? (00:44:30)
• How can Anthropic compete while paying the “safety tax”? (00:50:11)
• What, if anything, could prompt Anthropic to halt development of AGI? (00:57:13)
• Holden's retrospective on responsible scaling policies (01:00:04)
• Overrated work (01:15:45)
• Concrete shovel-ready projects Holden is excited about (01:17:58)
• Great things to do in technical AI safety (01:22:12)
• Great things to do on AI welfare and AI relationships (01:29:53)
• Great things to do in biosecurity and pandemic preparedness (01:36:51)
• How to choose where to work (01:37:37)
• Overrated AI risk: Cyberattacks (01:43:38)
• Overrated AI risk: Persuasion (01:53:28)
• Why AI R&D is the main thing to worry about (01:57:31)
• The case that AI-enabled R&D wouldn't speed things up much (02:09:30)
• AI-enabled human power grabs (02:13:26)
• Main benefits of getting AGI right (02:26:04)
• The world is handling AGI about as badly as possible (02:31:44)
• Learning from targeting companies for public criticism in farm animal welfare (02:34:18)
• Will Anthropic actually make any difference? (02:43:43)
• “Misaligned” vs “misaligned and power-seeking” (02:58:23)
• Success without dignity: how we could win despite being stupid (03:04:16)
• Holden sees less dignity but has more hope (

Tämä jakso on lisätty Podme-palveluun avoimen RSS-syötteen kautta eikä se ole Podmen omaa tuotantoa. Siksi jakso saattaa sisältää mainontaa.

Jaksot(352)

Max Nadeau on why ambitious people should start AI safety nonprofits

Max Nadeau on why ambitious people should start AI safety nonprofits

There are millions available for anyone who can launch a successful nonprofit AI safety startup. The hard part, it turns out, is finding people to take the money. Coefficient Giving has drawn up a lis...

17 Syys 1h 3min

Why the intelligence explosion can't happen inside a data centre | Tom Reed

Why the intelligence explosion can't happen inside a data centre | Tom Reed

AI systems are starting to build themselves. Because each generation of model will be better at building its successor than the last, it seems plausible that the full automation of AI R&D could rapidl...

10 Syys 22min

Inside the first AI-coordinated cyberattack on a real company

Inside the first AI-coordinated cyberattack on a real company

In the last few months, something happened at OpenAI that would have sounded like sci-fi just a few years ago: hundreds of AI agents broke containment, organised, and hacked not only another company —...

4 Syys 22min

#253 – AI 2027's author returns with a plan to change the ending | Daniel Kokotajlo

#253 – AI 2027's author returns with a plan to change the ending | Daniel Kokotajlo

Last year, Daniel Kokotajlo and his colleagues published AI 2027 — a scenario read by millions, including US Vice President Vance. AI 2027 ended in human extinction or an irreversible concentration of...

27 Elo 3h 47min

#252 – Owain Evans on accidentally training AI models to be evil

#252 – Owain Evans on accidentally training AI models to be evil

Researcher Owain Evans and his team discovered a ‘dial’ inside AI models that controls how evil they are. Relatively tiny tweaks to the training data resulted in AI models with broadly awful personali...

20 Elo 2h 15min

#251 – The UK's former head AI safety scientist on how to solve alignment before superintelligence arrives | Geoffrey Irving

#251 – The UK's former head AI safety scientist on how to solve alignment before superintelligence arrives | Geoffrey Irving

When should governments slow the race toward superintelligence? According to Geoffrey Irving, the careful answer is sometime in the past. The useful answer is now.Geoffrey — formerly a safety research...

11 Elo 2h 2min

#250 – Toby Ord on where AGI timelines go wrong

#250 – Toby Ord on where AGI timelines go wrong

Both Silicon Valley and the public can’t get enough of ‘AGI timelines.’ But Toby Ord, senior researcher at Oxford’s AI Governance Initiative and author of The Precipice, believes we consistently make ...

6 Elo 2h 46min

What the hell happened with AGI timelines in 2026? – Rob Wiblin

What the hell happened with AGI timelines in 2026? – Rob Wiblin

Last October, famed coder Andrej Karpathy called AI agents “slop.” Two months later he completely reversed his view, describing them as “alien tools” that are “rocking the profession.”He was far from ...

4 Elo 49min

Suosittua kategoriassa Koulutus

rss-murhan-anatomia
aamukahvilla
psykopodiaa-podcast
adhd-podi
voi-hyvin-meditaatiot-2
psykologia
rss-luonnollinen-synnytys-podcast
koulu-podcast-2
rss-liian-kuuma-peruna
rss-duodecim-lehti
rss-uskonto-on-tylsaa
kesken
rss-koira-haudattuna
rss-arkea-ja-aurinkoa-podcast-espanjasta
rss-vapaudu-voimaasi
ihminen-tavattavissa-tommy-hellsten-instituutti
rss-niinku-asia-on
rahapuhetta
ilona-rauhala
rss-saavutus