Unfaithful Chain of Thought

Unfaithful Chain of Thought

What's actually happening when an LLM "thinks out loud"? Research on human decision-making suggests that much of the reasoning we believe drives our choices is actually post hoc rationalization — we decide first, explain later. Katie and Ben get curious about whether the same might be true for large language models: when you watch a model reason through a problem in real time, is that chain of thought the genuine process, or just a plausible-sounding story told after the fact? It's a deceptively deep question with real stakes for how much we should trust model explanations. Miles Turpin et al., "Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting" (NeurIPS 2023, NYU and Anthropic): https://arxiv.org/abs/2305.04388 Anthropic, "Reasoning Models Don't Always Say What They Think" (Alignment Faking research, 2025): https://www.anthropic.com/research/reasoning-models-dont-say-think

Tämä jakso on lisätty Podme-palveluun avoimen RSS-syötteen kautta eikä se ole Podmen omaa tuotantoa. Siksi jakso saattaa sisältää mainontaa.

Jaksot(316)

Invisible LLM Failures and AI Fluency with Chris Potts (Stanford)

Invisible LLM Failures and AI Fluency with Chris Potts (Stanford)

What happens when a Stanford linguistics professor turns his attention to AI chatbots — and the surprisingly invisible ways humans misunderstand them? Chris Potts joins the show to unpack the hidden f...

20 Heinä 41min

Still summer break: back next week

Still summer break: back next week

Still summer break: back next week by Katie Malone

13 Heinä 25s

Summer break: back soon

Summer break: back soon

Summer break: back soon by Katie Malone

6 Heinä 36s

Interviewing the Linear Digressions Agents (The Agents Season, Episode 11)

Interviewing the Linear Digressions Agents (The Agents Season, Episode 11)

After a five-year hiatus, the podcast that burned out partly over the tedium of writing episode descriptions is back — and using AI agents to handle exactly that task. The season-11 finale turns the l...

28 Kesä 37min

Agent Economics (The Agents Season, Episode 10)

Agent Economics (The Agents Season, Episode 10)

What if building more highways made your commute *slower*? That's the paradox at the heart of AI agent economics: even as per-token inference costs have plummeted dramatically over the past two years,...

22 Kesä 24min

Agent Trust, Oversight and Control (The Agents Season, Episode 9)

Agent Trust, Oversight and Control (The Agents Season, Episode 9)

Capabilities get all the attention when it comes to AI agents — but what happens when a highly capable agent makes a bad decision in the real world? Trust, oversight, and control are the unglamorous b...

15 Kesä 25min

Many Agents, Many Problems (The Agents Season, Episode 8)

Many Agents, Many Problems (The Agents Season, Episode 8)

Whether you work best solo or thrive in a team, you know collaboration is complicated — and it turns out AI agents face the same tensions. This episode dives into multi-agent systems, exploring how ne...

8 Kesä 28min

How Do You Evaluate An AI Agent? (The Agents Season, Episode 7)

How Do You Evaluate An AI Agent? (The Agents Season, Episode 7)

Knowing when an AI agent has failed sounds straightforward — until it isn't. Agents have a frustrating habit of finishing confidently while quietly doing the wrong thing, or looping endlessly without ...

1 Kesä 31min