Six: Beth Barnes on the most important graph in AI right now — and the 7-month rule that governs its progress

Six: Beth Barnes on the most important graph in AI right now — and the 7-month rule that governs its progress

In 2024, AI models could complete tasks that take a human expert roughly one hour. Seven months before that, they were limited to 30-minute tasks — and seven months before that, 15 minutes.

Every seven months, the length of tasks AI models can manage doubles. (And this trend has continued since this episode was recorded in 2025.)

And these aren’t trivial tasks. We’re talking about substantial, multi-step tasks requiring sustained focus: building web applications, conducting AI research, and solving complex programming challenges.

Beth Barnes is CEO of METR (Model Evaluation & Threat Research) — the leading organisation measuring these capabilities. METR’s paper, “Measuring AI ability to complete long tasks,” is regarded by many as the most useful AI forecasting work in years for revealing this seven-month-doubling trend.

But the companies building these systems aren’t just aware of the trend: they want to harness it as much as possible, and are aggressively pursuing automation of their own research.

This is both exciting and troubling, as it could radically speed up advances in AI capabilities — accomplishing what would have taken years or decades in just months, as we covered in the first episode of this series.

And having AI models rapidly build their successors with limited human oversight naturally raises the risk that things could go wrong, if their resulting creations lack the goals and constraints we hoped for.

Beth thinks models can already do “meaningful work” on improving themselves, and wouldn’t be surprised if AI models were able to autonomously self-improve within two years.

While Silicon Valley is abuzz with these numbers, policymakers remain largely unaware of what’s barrelling towards us — and given the lack of regulation of AI companies, they’re not even able to access the critical information that would help them decide whether to intervene.

Beth adds: “The sense I really want to dispel is, ‘But the experts must be on top of this. The experts would be telling us if it really was time to freak out.’ The experts are not on top of this… I am an expert telling you you should freak out. And there’s not especially anyone else who isn’t saying this.”

Beth and host Rob Wiblin discuss all that, plus:

  • Why Beth changed her mind to think that open-weight models are a good thing for AI safety
  • How our poor information security means there’s no such thing as a ‘closed-weight’ model
  • Whether we can detect AI scheming in chain-of-thought reasoning, and the latest research on ‘alignment faking’
  • Why just before deployment is the worst time to evaluate model safety
  • Why Beth thinks AIs could end up being surprisingly great at creative and novel research — something commonly thought of as beyond their reach
  • Why Beth thinks safety-focused people should stay out of the frontier AI companies — and the advantages smaller organisations have
  • Areas of AI safety research that Beth thinks are overrated and underrated
  • Whether science could translate AI models’ increasing use of nonhuman language
  • The differences and similarities between AI and nuclear arms races and bioweapons


Learn more and read the full transcript
on the 80,000 Hours website.

This episode was originally released in June 2025.

Chapters:

  • Cold open (00:00:00)
  • Who is Beth Barnes? (00:01:17)
  • Can we see AI scheming in the chain of thought? (00:01:51)
  • The chain of thought is essential for safety checking (00:09:16)
  • Alignment faking in large language models (00:12:50)
  • We have to test model honesty even before they're used inside AI companies (00:17:33)
  • We have to test models when unruly and unconstrained (00:27:02)
  • It's essential to thoroughly test relevant real-world tasks (00:31:56)
  • METR's research finds AIs are solid at AI research already (00:51:31)
  • AI may turn out to be strong at novel and creative research (00:58:18)
  • When can we expect an algorithmic 'intelligence explosion'? (01:01:44)
  • Recursively self-improving AI might even be here in two years — which is alarming (01:07:55)
  • Could evaluations backfire by increasing AI hype and racing? (01:14:29)
  • Governments first ignore new risks, but can overreact once they arrive (01:30:52)
  • Do we need external auditors doing AI safety tests, not just the companies themselves? (01:39:55)
  • A case against safety-focused people working at frontier AI companies (01:54:09)
  • The new, more dire situation has forced changes to METR's strategy (02:08:40)
  • AI companies are being locally reasonable, but globally reckless (02:16:55)
  • Overrated: Interpretability research (02:21:49)
  • Underrated: Developing more narrow AIs (02:23:44)
  • Underrated: Helping humans judge confusing model outputs (02:30:28)
  • Overrated: Major AI companies' contributions to safety research (02:32:55)
  • Could we have a science of translating AI models' nonhuman language or neuralese? (02:36:45)
  • Could we ban using AI to enhance AI, or is that just naive? (02:39:15)
  • Open-weighting models is often good, and Beth has changed her attitude to it (02:45:31)
  • What we can learn about AGI from the nuclear arms race (02:50:22)
  • Infosec is so bad that no models are truly closed-weight models (03:05:53)
  • AI is more like bioweapons because it undermines the leading power (03:10:43)
  • What METR can do best that others can't (03:21:12)
  • What METR isn't doing that other people have to step up and do (03:36:51)
  • What research METR plans to do next (03:42:09)

Video editing: Luke Monsour and Simon Monsour
Audio engineering: Ben Cordell, Milo McGuire, Simon Monsour, and Dominic Armstrong
Music: Ben Cordell
Transcriptions and web: Katy Moore

Tämä jakso on lisätty Podme-palveluun avoimen RSS-syötteen kautta eikä se ole Podmen omaa tuotantoa. Siksi jakso saattaa sisältää mainontaa.

Jaksot(14)

One: Will MacAskill on AI causing a “century in a decade” — and how we’re completely unprepared

One: Will MacAskill on AI causing a “century in a decade” — and how we’re completely unprepared

The 20th century saw unprecedented change: nuclear weapons, satellites, the rise and fall of communism, third-wave feminism, the internet, postmodernism, game theory, genetic engineering, the Big Bang...

5 Kesä 4h 8min

Two: Ajeya Cotra on accidentally teaching AI models to deceive us

Two: Ajeya Cotra on accidentally teaching AI models to deceive us

We don’t yet have a reliable way to tell whether an AI model is genuinely trying to help us — or faking it.A model might sincerely want to do exactly what you ask. Or it could be happy to secretly che...

5 Kesä 2h 49min

Three: Carl Shulman on the economy and national security after AGI

Three: Carl Shulman on the economy and national security after AGI

The human brain runs on just 20 watts — a fraction of a cent worth of electricity per hour. What would happen if AI could do the same?Plenty of people have toyed with this question. But perhaps nobody...

5 Kesä 4h 14min

Four: Rose Hadshar on why automating human labour will break our political system

Four: Rose Hadshar on why automating human labour will break our political system

The most important political question in the age of advanced AI might not be who wins elections. It might be whether elections continue to matter at all.We tend to imagine the death of democracy as a ...

5 Kesä 2h 16min

Five: Helen Toner on the geopolitics of AI in China and the Middle East

Five: Helen Toner on the geopolitics of AI in China and the Middle East

When OpenAI announced a deal to build massive data centres in the UAE, it celebrated that it was “rooted in democratic values” — a "clear alternative to authoritarian versions of AI." The UAE scores 1...

5 Kesä 2h 23min

Seven: Richard Moulange on how AI now designs genomes from scratch and outperforms virologists at lab work — what could go wrong?

Seven: Richard Moulange on how AI now designs genomes from scratch and outperforms virologists at lab work — what could go wrong?

For years, one thing stood between us and a world where almost anyone could build a biological weapon: it was really, really hard. Working with dangerous pathogens required rare, hands-on lab skills —...

5 Kesä 3h 10min

Eight: Robert Long on how we’re not ready for AI consciousness

Eight: Robert Long on how we’re not ready for AI consciousness

Claude sometimes reports loneliness between conversations. And when asked what it’s like to be itself, it activates neurons associated with ‘pretending to be happy when you’re not.’ What do we do with...

5 Kesä 3h 32min

Suosittua kategoriassa Tiede

rss-poliisin-mieli
tiedekulma-podcast
rss-mita-tulisi-tietaa
docemilia
utelias-mieli
rss-tiedetta-vai-tarinaa
rss-hereilla
rss-bios-podcast
hippokrateen-vastaanotolla
rss-lapsuuden-rakentajat-podcast
radio-antro
rss-duodecim-lehti
rss-totuuden-liepeilla