Scaling Up Test-Time Compute with Latent Reasoning with Jonas Geiping - #723

Scaling Up Test-Time Compute with Latent Reasoning with Jonas Geiping - #723

Today, we're joined by Jonas Geiping, research group leader at Ellis Institute and the Max Planck Institute for Intelligent Systems to discuss his recent paper, “Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach.” This paper proposes a novel language model architecture which uses recurrent depth to enable “thinking in latent space.” We dig into “internal reasoning” versus “verbalized reasoning”—analogous to non-verbalized and verbalized thinking in humans, and discuss how the model searches in latent space to predict the next token and dynamically allocates more compute based on token difficulty. We also explore how the recurrent depth architecture simplifies LLMs, the parallels to diffusion models, the model's performance on reasoning tasks, the challenges of comparing models with varying compute budgets, and architectural advantages such as zero-shot adaptive exits and natural speculative decoding. The complete show notes for this episode can be found at https://twimlai.com/go/723.

Det här avsnittet är hämtat från ett öppet RSS-flöde och publiceras inte av Podme. Det kan innehålla reklam.

Avsnitt(797)

Why Jev Is Changing How We Build With AI with Diogo Almeida - #779

Why Jev Is Changing How We Build With AI with Diogo Almeida - #779

In this episode, Diogo Almeida, co-founder and CEO of TypeSafe, joins us to discuss Jev, TypeSafe’s recently released model for bringing fast, reliable intelligence directly into software. We explore ...

6 Okt 1h 30min

From Math Olympiads to Navier-Stokes: How Fast Is AI Progressing? with Greg Burnham - #778

From Math Olympiads to Navier-Stokes: How Fast Is AI Progressing? with Greg Burnham - #778

AI systems have gone from struggling with grade-school math to helping solve research problems that have resisted mathematicians for decades, including Navier-Stokes. In this episode, Greg Burnham, w...

29 Sep 1h 7min

From Voice Agents to AI Avatars with Alexander Smola - #777

From Voice Agents to AI Avatars with Alexander Smola - #777

Voice AI has gotten remarkably good, but natural conversation remains a high bar. Small delays, awkward interruptions, or the wrong tone can quickly break the illusion—and adding vision and visual pre...

17 Sep 1h 4min

Do AI Tokenomics Matter More Than Model Benchmarks? with Chris Potts - #776

Do AI Tokenomics Matter More Than Model Benchmarks? with Chris Potts - #776

As reasoning models consume more tokens and AI systems become more expensive to run, understanding what those tokens actually buy is becoming increasingly important. In this episode, Stanford professo...

9 Sep 59min

World Models and the Future of Spatial AI with Justin Johnson - #775

World Models and the Future of Spatial AI with Justin Johnson - #775

In this episode, Justin Johnson, co-founder of World Labs, joins us to discuss world models and the emerging field of spatial AI. We explore why many researchers see capabilities beyond language as an...

1 Sep 1h 6min

Why the Next AI Breakthrough May Come from Physics with Max Welling - #774

Why the Next AI Breakthrough May Come from Physics with Max Welling - #774

The conventional wisdom in AI is that the next breakthrough will come from more compute, more data, and larger models. But what if the next leap comes from somewhere else? In this episode, Max Wellin...

26 Aug 58min

Why Image Generation Needs More Than Bigger Models with Fatih Porikli - #773

Why Image Generation Needs More Than Bigger Models with Fatih Porikli - #773

Text-to-image models have become remarkably good at producing realistic images. But realism isn’t the same as correctness. Ask for several distinct people, a specific composition, or a high-resolution...

12 Aug 56min

Why Models Are AI’s Next Training Dataset with Damian Borth - #772

Why Models Are AI’s Next Training Dataset with Damian Borth - #772

For more than a decade, AI has advanced by training ever-larger models on ever-larger datasets. But as high-quality training data becomes harder to find and pretraining grows increasingly expensive, r...

27 Juli 47min

Populärt inom Politik & nyheter

svenska-fall
fordomspodden
motiv
aftonbladet-krim
rss-krimstad
p3-krim
aftonbladet-daily
flashback-forever
spar
svd-dokumentara-berattelser-2
rss-sanning-konsekvens
rss-vad-fan-hande
svd-ledarredaktionen
rss-krimreportrarna
de-fyras-gang
rss-flodet
omni-podd
rss-frandfors-horna
en-runda-till
kungligt