Stealing Reasoning Traces from Proprietary LLM APIs — Ilia Shumailov & Alexander Panfilov

Stealing Reasoning Traces from Proprietary LLM APIs — Ilia Shumailov & Alexander Panfilov

Tim Scarfe speaks with Ilia Shumailov and Alexander Panfilov about their paper, Stealing Reasoning Traces from Proprietary LLM APIs.The core bug sounds deceptively simple: providers return encrypted reasoning state so conversations can be resumed or forked. But those blobs can be replayed across users and sibling models. A smaller model can ask the provider to decrypt the trace, then repeat the hidden reasoning in plain text. The discussion covers leaked private data, a broadly reusable jailbreak, poisoned agent traces, chain-of-thought monitoring, responsible disclosure, and possible defenses.Ilia Shumailov is an AI and security researcher, formerly at Google DeepMind, who completed his Cambridge PhD under Ross Anderson. Alexander Panfilov is a PhD researcher at the ELLIS Institute Tübingen and the Max Planck Institute for Intelligent Systems, working on AI safety, adversarial machine learning, and LLM red-teaming. They close by separating the demonstrated jailbreaking threat from ordinary benign distillation, and by arguing for controlled experiments over sweeping claims.---TIMESTAMPS:00:00:00 Intro montage00:01:33 Portable encrypted thought and decoded reasoning00:24:55 How the attack works and what it means00:39:04 Doom, defense, and scientific restraint---REFERENCES:paper:[00:00:00] Stealing Reasoning Traces from Proprietary LLM APIshttps://arxiv.org/abs/2608.09867[00:09:22] Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safetyhttps://arxiv.org/abs/2507.11473[00:11:30] Reasoning Models Don’t Always Say What They Thinkhttps://www.anthropic.com/research/reasoning-models-dont-say-think[00:37:22] PostTrainBench: Can LLM Agents Automate LLM Post-Training?https://arxiv.org/abs/2603.08640[00:41:02] Large-scale online deanonymization with LLMshttps://arxiv.org/abs/2602.16800other:[00:09:28] OpenAI and Hugging Face partner to address security incident during model evaluationhttps://openai.com/index/hugging-face-model-evaluation-security-incident/[00:10:22] Claude, GPT, and Gemini All Struggle to Evade Monitorshttps://metr.org/notes/2025-08-22-claude-gpt-gemini-struggle-evade-monitors/tool:[00:42:08] Isabelle proof assistanthttps://isabelle.in.tum.de/---RESCRIPT: https://app.rescript.info/share/07fc38276e0823dc9b8986c32e202c7f

Tämä jakso on lisätty Podme-palveluun avoimen RSS-syötteen kautta eikä se ole Podmen omaa tuotantoa. Siksi jakso saattaa sisältää mainontaa.

Jaksot(260)

Every Exponential Ends — Silicon Valley Forgot — Adam Becker

Every Exponential Ends — Silicon Valley Forgot — Adam Becker

Astrophysicist Adam Becker, author of "What Is Real?", joins Tim Scarfe to take apart the futures Silicon Valley keeps selling: the 2045 singularity, mind uploading, Mars colonies, and the AI apocalyp...

20 Elo 1h 18min

AI Is Learning at the Wrong Level of Abstraction — Matthieu Wyart

AI Is Learning at the Wrong Level of Abstraction — Matthieu Wyart

This episode is sponsored by Notion. Learn more about Notion's Developer Platform today at https://notion.com/mlstWhy can deep networks discover abstractions that shallow models miss? Statistical phys...

10 Elo 1h 18min

How Researchers Test AI for Hidden Goals — Apollo Research

How Researchers Test AI for Hidden Goals — Apollo Research

Can an AI do the right thing for the wrong reason? Tim Scarfe speaks with Apollo Research’s Alexander Meinke, Axel Højmark and Jérémy Scheurer about Measuring Reward-Seeking via Contrastive Belief Upd...

31 Heinä 1h 18min

Why a Nation Can't Outsource Its Frontier AI - Alistair Pullen (Cosine AI)

Why a Nation Can't Outsource Its Frontier AI - Alistair Pullen (Cosine AI)

This episode is sponsored by Notion. Learn more about Notion's Developer Platform today at https://notion.com/mlstBritain's most capable coding model can't be exported, and that ban is the whole reaso...

13 Heinä 55min

 The Benchmark With No Instructions — ARC-AGI-3 (winning team!)

The Benchmark With No Instructions — ARC-AGI-3 (winning team!)

Tim Scarfe travels to Zurich to sit down with the Tufa Labs ARC-AGI-3 team — founder Benjamin Crouzier, with Jeroen Cottaar, Dries Smit, Stefano Viel and Michal Tesnar — to work out what their leaderb...

1 Heinä 1h 24min

The Thermodynamic AI Computing Chip - Thomas Ahle

The Thermodynamic AI Computing Chip - Thomas Ahle

Thomas Ahle wants Normal Computing to be the Lovable for chip design: type your intent, and a swarm of agents carries it from design through optimisation, formalisation and verification to tape-out. T...

28 Kesä 1h 2min

He won a Nobel here for AlphaFold. Then he left. - John Jumper

He won a Nobel here for AlphaFold. Then he left. - John Jumper

This episode is sponsored by Notion. Learn more about Notion's Developer Platform today at https://notion.com/mlstProtein folding stalled biology for fifty years. A sequence of amino acids dictates a ...

22 Kesä 53min