Kubernetes GPU Management Just Got a Major Upgrade

Kubernetes GPU Management Just Got a Major Upgrade

Nvidia Distinguished Engineer Kevin Klues noted that low-level systems work is invisible when done well and highly visible when it fails — a dynamic that frames current Kubernetes innovations for AI. At KubeCon + CloudNativeCon North America 2025, Klues and AWS product manager Jesse Butler discussed two emerging capabilities: dynamic resource allocation (DRA) and a new workload abstraction designed for sophisticated AI scheduling.

DRA, now generally available in Kubernetes 1.34, fixes long-standing limitations in GPU requests. Instead of simply asking for a number of GPUs, users can specify types and configurations. Modeled after persistent volumes, DRA allows any specialized hardware to be exposed through standardized interfaces, enabling vendors to deliver custom device drivers cleanly. Butler called it one of the most elegant designs in Kubernetes.

Yet complex AI workloads require more coordination. A forthcoming workload abstraction, debuting in Kubernetes 1.35, will let users define pod groups with strict scheduling and topology rules — ensuring multi-node jobs start fully or not at all. Klues emphasized that this abstraction will shape Kubernetes’ AI trajectory for the next decade and encouraged community involvement.

Learn more from The New Stack about dynamic resource allocation:

Kubernetes Primer: Dynamic Resource Allocation (DRA) for GPU Workloads

Kubernetes v1.34 Introduces Benefits but Also New Blind Spots

Join our community of newsletter subscribers to stay on top of the news and at the top of your game.

Tämä jakso on lisätty Podme-palveluun avoimen RSS-syötteen kautta eikä se ole Podmen omaa tuotantoa. Siksi jakso saattaa sisältää mainontaa.

Jaksot(300)

Bit Cloud’s next chapter starts after the AI builds your app

Bit Cloud’s next chapter starts after the AI builds your app

For many developers, turning an AI-generated prototype into maintainable software requires more than generating code—it requires infrastructure, collaboration, testing and review. In this episode of T...

1 Loka 31min

CloudBees just committed to an AI-first pivot. Here's why it matters for enterprise DevOps teams

CloudBees just committed to an AI-first pivot. Here's why it matters for enterprise DevOps teams

CloudBees CEO Mo Plassnig is leading the CI/CD company through a major transformation as generative AI reshapes software development. Returning to CloudBees eight years after joining through its acqui...

30 Syys 23min

A third option is emerging in the fight over AI and your data

A third option is emerging in the fight over AI and your data

The AI industry has faced a growing enterprise dilemma: companies want access to powerful proprietary AI models without risking sensitive data or intellectual property, while AI labs want to protect t...

23 Syys 30min

Drowning in AI pull requests: Harness's field CTO on code review and a Git repo built for agents

Drowning in AI pull requests: Harness's field CTO on code review and a Git repo built for agents

Harness Field CTO Martin Reynolds joins The New Stack to talk about what happens after coding agents start opening pull requests faster than anyone can review them. He explains how he first saw the bo...

7 Syys 26min

How to find failures without drowning in tracing data

How to find failures without drowning in tracing data

Traces provide a detailed view of a request’s journey through data, microservices and applications, helping SREs pinpoint where failures occur and resolve issues faster. But while tracing can reduce d...

3 Syys 31min

Why CPUs still matter in the age of AI agents

Why CPUs still matter in the age of AI agents

As AI evolves from conversational chatbots to autonomous agents, CPUs are becoming an increasingly important part of the infrastructure equation. In this episode, The New Stack speaks with Bhumik Pate...

11 Elo 26min

Why Doist Says Less AI Can Deliver More

Why Doist Says Less AI Can Deliver More

Doist CTO Gonçalo Silva says AI is reshaping software development, but success depends on restraint rather than rapid feature expansion. Instead of chasing every AI capability, Doist prioritizes “subt...

31 Heinä 35min

Why your company should (try to) build its own AI SRE

Why your company should (try to) build its own AI SRE

As AI coding agents accelerate software development, they also create new challenges for site reliability engineers (SREs), who are increasingly responsible for debugging systems that no single human ...

30 Heinä 30min

Suosittua kategoriassa Politiikka ja uutiset

uutiscast
vallattomat
aikalisa
rss-viihde-media
politiikan-puskaradio
rss-vaalirankkurit-podcast
rss-ootsa-kuullut-tasta
ootsa-kuullut-tasta-2
otetaan-yhdet
tervo-halme
rss-voi-venaja
rss-podme-livebox
rss-kaikki-uusiksi
rss-asiastudio
rss-seksicast
the-ulkopolitist
rss-raha-talous-ja-politiikka
rss-girls-finish-f1rst
rss-ulkopoditiikkaa
politbyroo