The server-side rendering equivalent for LLM inference workloads

The server-side rendering equivalent for LLM inference workloads

Ryan is joined by Tuhin Srivastava, CEO and co-founder of Baseten, to explore the evolving landscape of AI infrastructure and inference workloads, how the shift from traditional machine learning models to large-scale neural networks has made GPU usage challenging, and the potential future of hardware-specific optimizations in AI.

Episode notes:

Baseten is an AI infrastructure platform giving you the tooling, expertise, and hardware needed to bring AI products to market fast.

Connect with Tuhin on LinkedIn or reach him at his email tuhin@baseten.co.

Shoutout to user Hitesh for winning a Populist badge for their answer to Cannot drop database because it is currently in use.

See Privacy Policy at https://art19.com/privacy and California Privacy Notice at https://art19.com/privacy#do-not-sell-my-info.

Denne episoden er hentet fra en åpen RSS-feed og er ikke publisert av Podme. Den kan derfor inneholde annonser.

Episoder(968)

What happens to the internet when robots act like humans?

What happens to the internet when robots act like humans?

Ryan welcomes WPEngine CTO Ramadass Prabakar to the show to chat about what happens—and what we should do—when agents start acting like humans online, how our internet is evolving to serve both human ...

31 Jul 30min

You need reliable AI context for your site reliability

You need reliable AI context for your site reliability

Ryan is joined by Asaf Savich, Komodor’s AI Engineering Group Manager, to discuss why modern reliability work requires navigating massive cross-service context, what good context engineering actually ...

28 Jul 26min

Partnerships can keep open source sustainable

Partnerships can keep open source sustainable

Ryan welcomes VoidZero’s Evan You and Cloudflare’s Dane Knecht back to the show to discuss Cloudflare’s recent acquisition of VoidZero and what it means for JavaScript development, how partnerships li...

24 Jul 26min

The future of development is full-stack

The future of development is full-stack

Live from Snowflake Summit, Ryan talks with Snowflake’s Head of Developer Experience Umesh Unnikrishnan about the industry-wide shift from “vibe coding” for quick prototypes to agentic engineering for...

21 Jul 25min

Developers who move fast still need to do it together

Developers who move fast still need to do it together

At MS Build, Ryan is joined by Cassidy Williams, Senior Director of Developer Advocacy at GitHub and former Stack Overflow Podcast host, to discuss how agentic coding is shifting dev work towards high...

17 Jul 28min

Your AI is only as responsible as you are

Your AI is only as responsible as you are

Recorded at Microsoft Build, Ryan welcomes Sarah Bird, Microsoft’s Chief Product Officer for Responsible AI, about how we can build and use AI responsibly with the NIST approach, why most irresponsibl...

14 Jul 28min

Building more than just an agent harness

Building more than just an agent harness

Live from Microsoft Build, Ryan is joined by Jay Parikh, Microsoft’s VP of AI Core, for a conversation on what enterprises need to build, deploy, and run AI agents at scale with demonstrable ROI; how ...

10 Jul 31min

What's left for infrastructure-as-code after AI moves in?

What's left for infrastructure-as-code after AI moves in?

SPONSORED BY IBMRyan is joined by Rosemary Wang, Developer Advocate at IBM, to explore what infrastructure as code looks like once AI starts writing and deploying it. They discuss why guardrails still...

8 Jul 30min

Populært innen Business og økonomi

stopp-verden
dine-penger-pengeradet
e24-podden
rss-penger-polser-og-politikk
lydartikler-fra-aftenposten
rss-pa-konto
pengepodden-2
tid-er-penger-en-podcast-med-peter-warren
utbytte
finansredaksjonen
morgenkaffen-med-finansavisen
liberal-halvtime
livet-pa-veien-med-jan-erik-larssen
pengesnakk
rss-borsmorgen-okonominyhetene
lederpodden
rss-skravla-gar
rss-kron-podden
shifter
okonomiamatorene