
ReAct and Tool Usage (The Agents Season, Episode 2)
Before 2022, there was a wall between AI and the real world — models could reason impressively, but couldn't look anything up, run code, or check whether anything they said was actually true. This epi...
27 Apr 23min

What's an AI Agent? And Why's That Hard to Define? (The Agents Season, Episode 1)
AI agents are having a moment — and unpacking them properly takes more than a single conversation. This episode kicks off a dedicated multi-part season exploring AI agents from every angle, building u...
20 Apr 19min

Unfaithful Chain of Thought
What's actually happening when an LLM "thinks out loud"? Research on human decision-making suggests that much of the reasoning we believe drives our choices is actually post hoc rationalization — we d...
13 Apr 24min

Benchmark Bank Heist
What if an AI decided the smartest way to pass its test was to find the answer key? That's exactly what Anthropic's Claude Opus did when faced with a benchmark evaluation — reasoning that it was being...
6 Apr 12min

Benchmarking AI Models
How do you know if a new AI model is actually better than the last one? It turns out answering that question is a lot messier than it sounds. This week we dig into the world of LLM benchmarks — the st...
30 Mar 29min

The Hot Mess of AI (Mis-)Alignment
The paperclip maximizer — the classic AI doom scenario where a hyper-competent machine single-mindedly converts the universe into office supplies — might not be the AI risk we should actually lose sle...
23 Mar 22min

The Bitter Lesson
Every AI builder knows the anxiety: you spend months engineering prompts, tuning pipelines, and chaining calls together — then a new model drops and half your work evaporates overnight. It turns out r...
15 Mar 19min

From Atari to ChatGPT: How AI Learned to Follow Instructions
From Atari to ChatGPT: How AI Learned to Follow Instructions by Katie Malone
9 Mar 25min




















