Inception Labs says its diffusion LLM is 10x faster than Claude, ChatGPT, Gemini

Inception Labs says its diffusion LLM is 10x faster than Claude, ChatGPT, Gemini

On a recent episode of the The New Stack Agents, Inception Labs CEO Stefano Ermon introduced Mercury 2, a large language model built on diffusion rather than the standard autoregressive approach. Traditional LLMs generate text token by token from left to right, which Ermon describes as “fancy autocomplete.” In contrast, diffusion models begin with a rough draft and refine it in parallel, similar to image systems like Stable Diffusion.

This parallel process allows Mercury 2 to produce over 1,000 tokens per second—five to ten times faster than optimized models from labs such as OpenAI, Anthropic, and Google, according to company tests. Ermon argues diffusion models better leverage GPUs, with support from investor Nvidia to optimize performance.

While Mercury 2 matches mid-tier models like Claude Haiku and Google Flash rather than top systems such as Claude Opus or GPT-4, Ermon believes diffusion’s speed and economic advantages will become increasingly compelling as AI applications scale.

Learn more from The New Stack about the latest developments around around large language model built on diffusion:

How Diffusion-Based LLM AI Speeds Up Reasoning

Get Ready for Faster Text Generation With Diffusion LLMs

Join our community of newsletter subscribers to stay on top of the news and at the top of your game.

Tämä jakso on lisätty Podme-palveluun avoimen RSS-syötteen kautta eikä se ole Podmen omaa tuotantoa. Siksi jakso saattaa sisältää mainontaa.

Jaksot(300)

Why CPUs still matter in the age of AI agents

Why CPUs still matter in the age of AI agents

As AI evolves from conversational chatbots to autonomous agents, CPUs are becoming an increasingly important part of the infrastructure equation. In this episode, The New Stack speaks with Bhumik Pate...

11 Elo 26min

Why Doist Says Less AI Can Deliver More

Why Doist Says Less AI Can Deliver More

Doist CTO Gonçalo Silva says AI is reshaping software development, but success depends on restraint rather than rapid feature expansion. Instead of chasing every AI capability, Doist prioritizes “subt...

31 Heinä 35min

Why your company should (try to) build its own AI SRE

Why your company should (try to) build its own AI SRE

As AI coding agents accelerate software development, they also create new challenges for site reliability engineers (SREs), who are increasingly responsible for debugging systems that no single human ...

30 Heinä 30min

Nvidia

Nvidia

In this episode with The New Stack Agents, Frederic Lardinois, NVIDIA’s Joey Conway says advances in AI over the past year have dramatically improved the capabilities of local models, making them prac...

23 Heinä 32min

Meet Brain, the AI that decides when Azure is officially down

Meet Brain, the AI that decides when Azure is officially down

In this episode, Mark Russinovich, CTO of Microsoft Azure revealed Brain, the AI-powered AIOps system that continuously monitors Azure’s health, detects incidents, identifies root causes, and increasi...

14 Heinä 19min

What comes after attention? This startup says it already knows.

What comes after attention? This startup says it already knows.

Subquadratic is beginning to back up its ambitious claims with benchmarks and third-party validation for its SubQ 1.1 Small model, which uses its proprietary Sparse Attention (SSA) architecture to dra...

7 Heinä 20min

“The harness is where the hard work is”: Harness bets on agents that enterprises can trust in production

“The harness is where the hard work is”: Harness bets on agents that enterprises can trust in production

Harness has introduced Autonomous Worker Agents, a new capability that allows enterprises to replace rigid CI/CD pipeline scripts with AI agents that can deploy applications, run tests, and perform se...

2 Heinä 19min

Public cloud vs. on-prem: Summit on where each workload belongs

Public cloud vs. on-prem: Summit on where each workload belongs

More than two decades after AWS helped usher in the public cloud era, many organizations are reassessing whether a cloud-first strategy still delivers the cost and operational benefits it once promise...

25 Kesä 36min

Suosittua kategoriassa Politiikka ja uutiset

uutiscast
aikalisa
politiikan-puskaradio
ootsa-kuullut-tasta-2
rss-ootsa-kuullut-tasta
otetaan-yhdet
rss-vaalirankkurit-podcast
rss-podme-livebox
rss-seksicast
rikosmyytit
aihe
the-ulkopolitist
rss-asiastudio
tervo-halme
rss-vain-talouselamaa
rss-kaikki-uusiksi
rss-bimbolia
linda-maria
rss-sanna-ukkola-show-verkkouutiset
rss-girls-finish-f1rst