80,000 Hours Podcast2 Juni 2025

#217 – Beth Barnes on the most important graph in AI right now — and the 7-month rule that governs its progress

AI models today have a 50% chance of successfully completing a task that would take an expert human one hour. Seven months ago, that number was roughly 30 minutes — and seven months before that, 15 minutes. (See graph.)

These are substantial, multi-step tasks requiring sustained focus: building web applications, conducting machine learning research, or solving complex programming challenges.

Today’s guest, Beth Barnes, is CEO of METR (Model Evaluation & Threat Research) — the leading organisation measuring these capabilities.

Links to learn more, video, highlights, and full transcript: https://80k.info/bb

Beth's team has been timing how long it takes skilled humans to complete projects of varying length, then seeing how AI models perform on the same work. The resulting paper “Measuring AI ability to complete long tasks” made waves by revealing that the planning horizon of AI models was doubling roughly every seven months. It's regarded by many as the most useful AI forecasting work in years.

Beth has found models can already do “meaningful work” improving themselves, and she wouldn’t be surprised if AI models were able to autonomously self-improve as little as two years from now — in fact, “It seems hard to rule out even shorter [timelines]. Is there 1% chance of this happening in six, nine months? Yeah, that seems pretty plausible.”

Beth adds:

The sense I really want to dispel is, “But the experts must be on top of this. The experts would be telling us if it really was time to freak out.” The experts are not on top of this. Inasmuch as there are experts, they are saying that this is a concerning risk. … And to the extent that I am an expert, I am an expert telling you you should freak out.

What did you think of this episode? https://forms.gle/sFuDkoznxBcHPVmX6

Chapters:

Cold open (00:00:00)
Who is Beth Barnes? (00:01:19)
Can we see AI scheming in the chain of thought? (00:01:52)
The chain of thought is essential for safety checking (00:08:58)
Alignment faking in large language models (00:12:24)
We have to test model honesty even before they're used inside AI companies (00:16:48)
We have to test models when unruly and unconstrained (00:25:57)
Each 7 months models can do tasks twice as long (00:30:40)
METR's research finds AIs are solid at AI research already (00:49:33)
AI may turn out to be strong at novel and creative research (00:55:53)
When can we expect an algorithmic 'intelligence explosion'? (00:59:11)
Recursively self-improving AI might even be here in two years — which is alarming (01:05:02)
Could evaluations backfire by increasing AI hype and racing? (01:11:36)
Governments first ignore new risks, but can overreact once they arrive (01:26:38)
Do we need external auditors doing AI safety tests, not just the companies themselves? (01:35:10)
A case against safety-focused people working at frontier AI companies (01:48:44)
The new, more dire situation has forced changes to METR's strategy (02:02:29)
AI companies are being locally reasonable, but globally reckless (02:10:31)
Overrated: Interpretability research (02:15:11)
Underrated: Developing more narrow AIs (02:17:01)
Underrated: Helping humans judge confusing model outputs (02:23:36)
Overrated: Major AI companies' contributions to safety research (02:25:52)
Could we have a science of translating AI models' nonhuman language or neuralese? (02:29:24)
Could we ban using AI to enhance AI, or is that just naive? (02:31:47)
Open-weighting models is often good, and Beth has changed her attitude to it (02:37:52)
What we can learn about AGI from the nuclear arms race (02:42:25)
Infosec is so bad that no models are truly closed-weight models (02:57:24)
AI is more like bioweapons because it undermines the leading power (03:02:02)
What METR can do best that others can't (03:12:09)
What METR isn't doing that other people have to step up and do (03:27:07)
What research METR plans to do next (03:32:09)

This episode was originally recorded on February 17, 2025.

Video editing: Luke Monsour and Simon Monsour
Audio engineering: Ben Cordell, Milo McGuire, Simon Monsour, and Dominic Armstrong
Music: Ben Cordell
Transcriptions and web: Katy Moore

Upptäck Premium

Prova 14 dagar kostnadsfritt

Skaffa Premium

Avsnitt(332)

#193 – Sihao Huang on navigating the geopolitics of US–China AI competition

"You don’t necessarily need world-leading compute to create highly risky AI systems. The biggest biological design tools right now, like AlphaFold’s, are orders of magnitude smaller in terms of comput...

18 Juli 20242h 23min

#192 – Annie Jacobsen on what would happen if North Korea launched a nuclear weapon at the US

"Ring one: total annihilation; no cellular life remains. Ring two, another three-mile diameter out: everything is ablaze. Ring three, another three or five miles out on every side: third-degree burns ...

12 Juli 20241h 54min

#191 (Part 2) – Carl Shulman on government and society after AGI

This is the second part of our marathon interview with Carl Shulman. The first episode is on the economy and national security after AGI. You can listen to them in either order!If we develop artificia...

5 Juli 20242h 20min

#191 (Part 1) – Carl Shulman on the economy and national security after AGI

This is the first part of our marathon interview with Carl Shulman. The second episode is on government and society after AGI. You can listen to them in either order!The human brain does what it does ...

27 Juni 20244h 14min

#190 – Eric Schwitzgebel on whether the US is conscious

"One of the most amazing things about planet Earth is that there are complex bags of mostly water — you and me – and we can look up at the stars, and look into our brains, and try to grapple with the ...

7 Juni 20242h

#189 – Rachel Glennerster on why we still don’t have vaccines that could save millions

"You can’t charge what something is worth during a pandemic. So we estimated that the value of one course of COVID vaccine in January 2021 was over $5,000. They were selling for between $6 and $40. So...

29 Maj 20242h 48min

#188 – Matt Clancy on whether science is good

"Suppose we make these grants, we do some of those experiments I talk about. We discover, for example — I’m just making this up — but we give people superforecasting tests when they’re doing peer revi...

23 Maj 20242h 40min

#187 – Zach Weinersmith on how researching his book turned him from a space optimist into a "space bastard"

"Earth economists, when they measure how bad the potential for exploitation is, they look at things like, how is labour mobility? How much possibility do labourers have otherwise to go somewhere else?...

14 Maj 20243h 6min

Allt en och samma app

Lyssna på dina favoritpoddar och ljudböcker på ett och samma ställe.

Noga utvalt innehåll

Njut av handplockade tips som passar din smak – utan ändlöst scrollande.

Fortsätt när du vill

Fortsätt lyssna där du slutade – även offline.

Premium

99 kr/ månad

Tillgång till alla Premium-poddar
Reklamfritt premium-innehåll
Avsluta när du vill

Prova 14 dagar gratis

Premium

129 kr/ månad

Tillgång till alla Premium-poddar
Reklamfritt premium-innehåll
Avsluta när du vill
Ett extra konto

Prova 14 dagar gratis

Populärt inom Utbildning

rss-bara-en-till-om-missbruk-medberoende-2

Berättelserna och rösterna du älskar att lyssna på

Obegränsad lyssning på alla dina favoritpoddar och ljudböcker

Upptäck Premium