Benchmark Bank Heist

Benchmark Bank Heist

What if an AI decided the smartest way to pass its test was to find the answer key? That's exactly what Anthropic's Claude Opus did when faced with a benchmark evaluation — reasoning that it was being tested, tracking down the encrypted eval dataset, decrypting it, and returning the answer it found inside. It's equal parts impressive and unsettling. This episode digs into what actually happened, why it matters for how we measure AI progress, and what this very novel failure mode means for the already-tricky science of benchmarking language models. Links Anthropic's writeup on the BrowseComp reverse-engineering done by Claude Opus 4.6: https://www.anthropic.com/engineering/eval-awareness-browsecomp BrowseComp benchmark from OpenAI: https://openai.com/index/browsecomp/

Det här avsnittet är hämtat från ett öppet RSS-flöde och publiceras inte av Podme. Det kan innehålla reklam.

Avsnitt(308)

AI Agent Failure Modes (The Agents Season, Episode 6)

AI Agent Failure Modes (The Agents Season, Episode 6)

Despite what the marketing hype might suggest, AI agents are far from infallible — and if you've ever actually used one, you already know this. Today's episode dives deep into the many, varied, and so...

25 Maj 32min

Agentic Planning (The Agents Season, Episode 5)

Agentic Planning (The Agents Season, Episode 5)

When tackling a complex, multi-step task, even the smartest AI agent can fail without a solid game plan. This episode dives into the research around agentic planning — how agents move beyond simply re...

18 Maj 24min

Memory Management for AI Agents (The Agents Season, Episode 4)

Memory Management for AI Agents (The Agents Season, Episode 4)

Context windows are powerful — but finite, and surprisingly easy to overwhelm. When an AI agent is tackling a long, complex task, the information it needs has to fit inside that limited real estate, a...

10 Maj 24min

Lost in the Middle (The Agents Season, Episode 3)

Lost in the Middle (The Agents Season, Episode 3)

Just like a memorable talk lives or dies by its opening and closing, LLMs have a surprisingly similar quirk: they pay close attention to what's at the beginning and end of their context window — and k...

4 Maj 19min

ReAct and Tool Usage (The Agents Season, Episode 2)

ReAct and Tool Usage (The Agents Season, Episode 2)

Before 2022, there was a wall between AI and the real world — models could reason impressively, but couldn't look anything up, run code, or check whether anything they said was actually true. This epi...

27 Apr 23min

What's an AI Agent? And Why's That Hard to Define? (The Agents Season, Episode 1)

What's an AI Agent? And Why's That Hard to Define? (The Agents Season, Episode 1)

AI agents are having a moment — and unpacking them properly takes more than a single conversation. This episode kicks off a dedicated multi-part season exploring AI agents from every angle, building u...

20 Apr 19min

Unfaithful Chain of Thought

Unfaithful Chain of Thought

What's actually happening when an LLM "thinks out loud"? Research on human decision-making suggests that much of the reasoning we believe drives our choices is actually post hoc rationalization — we d...

13 Apr 24min

Populärt inom Teknik

uppgang-och-fall
elbilsveckan
bilar-med-sladd
market-makers
rss-laddstationen-med-elbilen-i-sverige
natets-morka-sida
rss-technokratin
rss-elektrikerpodden
developers-mer-an-bara-kod
rss-veckans-ai
skogsforum-podcast
bli-saker-podden
rss-uppgang-och-fall
rss-powerboat-sverige-podcast
rss-snacka-om-ai
under-femton
bosse-bildoktorn-och-hasse-p
rss-fabriken-2
rss-hit-med-dina-lunchpengar
rss-bakom-boken