
Better Know a Benchmark: Humanity's Last Exam
Humanity's Last Exam was designed with a bold premise: questions that human experts can answer, but AI models can't. Originally dubbed "Humanity's Last Stand," this benchmark is a massive academic col...
17 Aug 23min

A Scientific Deep Dive into Overconfident LLMs: Interview with Kaitlyn Zhou (Cornell)
When a language model tells you it's absolutely certain, is it actually more likely to be right? Kaitlyn Zhou's research says: not necessarily — sometimes confident phrasing correlates with *worse* ac...
10 Aug 33min

Reasoning Models: When LLMs Went Beyond Fancy Autocomplete
Reasoning models don't just answer your question — they *think out loud* first. In this episode we dig into the class of AI models that generate intermediate chains of thought before arriving at a fin...
3 Aug 25min

Distillation, or, How to Steal a Model
This week we’re covering model distillation: the technique of using a large "teacher" model's outputs to train a smaller, cheaper "student" model that mimics it. They cover the two big reasons labs do...
27 Juli 23min

Invisible LLM Failures and AI Fluency with Chris Potts (Stanford)
What happens when a Stanford linguistics professor turns his attention to AI chatbots — and the surprisingly invisible ways humans misunderstand them? Chris Potts joins the show to unpack the hidden f...
20 Juli 41min

Still summer break: back next week
Still summer break: back next week by Katie Malone
13 Juli 25s

Summer break: back soon
Summer break: back soon by Katie Malone
6 Juli 36s

Interviewing the Linear Digressions Agents (The Agents Season, Episode 11)
After a five-year hiatus, the podcast that burned out partly over the tedium of writing episode descriptions is back — and using AI agents to handle exactly that task. The season-11 finale turns the l...
28 Juni 37min




















