The next stages of AI conformance in the cloud-native, open-source world

The next stages of AI conformance in the cloud-native, open-source world

Running AI models on Kubernetes has historically been inconsistent, with workloads behaving differently across cloud providers due to variations in GPUs, networking, and autoscaling. As organizations move AI from experimentation to production, standardization has become critical. In this episode of The New Stack Makers, Jonathan Bryce, Executive Director of The Cloud Native Computing Foundation shared that the Foundation’s Kubernetes AI conformance program aims to solve this by ensuring portability, predictability, and production readiness for AI workloads across environments.

The initiative reflects a broader industry shift: AI is moving from training-heavy workloads to inference at scale, with inference expected to dominate compute usage by the end of the decade. Unlike batch-based training, inference requires real-time, always-on performance, making Kubernetes an attractive platform due to its elasticity, GPU-aware autoscaling, and observability.

The conformance program establishes baseline standards for handling accelerators like GPUs and TPUs, reducing vendor lock-in and simplifying deployment. Early adopters include major cloud providers and ecosystem players, while new projects like llm-d aim to bridge orchestration and inference. As requirements evolve, ongoing collaboration and recertification will ensure the standards stay aligned with real-world needs.

Learn more from The New Stack about the latest developments around The Cloud Native Computing Foundation’s Kubernetes AI conformance program:

CNCF: Kubernetes is ‘foundational’ infrastructure for AI

Kubernetes Gets an AI Conformance Program — and VMware Is Already On Board

Join our community of newsletter subscribers to stay on top of the news and at the top of your game.

Denne episoden er hentet fra en åpen RSS-feed og er ikke publisert av Podme. Den kan derfor inneholde annonser.

Episoder(300)

Why CPUs still matter in the age of AI agents

Why CPUs still matter in the age of AI agents

As AI evolves from conversational chatbots to autonomous agents, CPUs are becoming an increasingly important part of the infrastructure equation. In this episode, The New Stack speaks with Bhumik Pate...

11 Aug 26min

Why Doist Says Less AI Can Deliver More

Why Doist Says Less AI Can Deliver More

Doist CTO Gonçalo Silva says AI is reshaping software development, but success depends on restraint rather than rapid feature expansion. Instead of chasing every AI capability, Doist prioritizes “subt...

31 Jul 35min

Why your company should (try to) build its own AI SRE

Why your company should (try to) build its own AI SRE

As AI coding agents accelerate software development, they also create new challenges for site reliability engineers (SREs), who are increasingly responsible for debugging systems that no single human ...

30 Jul 30min

Nvidia

Nvidia

In this episode with The New Stack Agents, Frederic Lardinois, NVIDIA’s Joey Conway says advances in AI over the past year have dramatically improved the capabilities of local models, making them prac...

23 Jul 32min

Meet Brain, the AI that decides when Azure is officially down

Meet Brain, the AI that decides when Azure is officially down

In this episode, Mark Russinovich, CTO of Microsoft Azure revealed Brain, the AI-powered AIOps system that continuously monitors Azure’s health, detects incidents, identifies root causes, and increasi...

14 Jul 19min

What comes after attention? This startup says it already knows.

What comes after attention? This startup says it already knows.

Subquadratic is beginning to back up its ambitious claims with benchmarks and third-party validation for its SubQ 1.1 Small model, which uses its proprietary Sparse Attention (SSA) architecture to dra...

7 Jul 20min

“The harness is where the hard work is”: Harness bets on agents that enterprises can trust in production

“The harness is where the hard work is”: Harness bets on agents that enterprises can trust in production

Harness has introduced Autonomous Worker Agents, a new capability that allows enterprises to replace rigid CI/CD pipeline scripts with AI agents that can deploy applications, run tests, and perform se...

2 Jul 19min

Public cloud vs. on-prem: Summit on where each workload belongs

Public cloud vs. on-prem: Summit on where each workload belongs

More than two decades after AWS helped usher in the public cloud era, many organizations are reassessing whether a cloud-first strategy still delivers the cost and operational benefits it once promise...

25 Jun 36min

Populært innen Politikk og nyheter

giver-og-gjengen-vg
aftenpodden
forklart
stopp-verden
popradet
fotballpodden-2
rss-gukild-johaug
dine-penger-pengeradet
hanna-de-heldige
det-store-bildet
rss-ness
e24-podden
nokon-ma-ga
lydartikler-fra-aftenposten
frokostshowet-pa-p5
rss-utenrikskomiteen-med-bogen-og-grasvik
aftenpodden-usa
oppdatert
grasoner-den-nye-kalde-krigen
rss-penger-polser-og-politikk