Why Vision Language Models Ignore What They See with Munawar Hayat - #758

Why Vision Language Models Ignore What They See with Munawar Hayat - #758

In this episode, we’re joined by Munawar Hayat, researcher at Qualcomm AI Research, to discuss a series of papers presented at NeurIPS 2025 focusing on multimodal and generative AI. We dive into the persistent challenge of object hallucination in Vision-Language Models (VLMs), why models often discard visual information in favor of pre-trained language priors, and how his team used attention-guided alignment to enforce better visual grounding. We also explore a novel approach to generalized contrastive learning designed to solve complex, composed retrieval tasks—such as searching via combined text and image queries—without increasing inference costs. Finally, we cover the difficulties generative models face when rendering multiple human subjects, and the new "MultiHuman Testbench" his team created to measure and mitigate issues like identity leakage and attribute blending. Throughout the discussion, we examine how these innovations align with the need for efficient, on-device AI deployment. The complete show notes for this episode can be found at https://twimlai.com/go/758.

Det här avsnittet är hämtat från ett öppet RSS-flöde och publiceras inte av Podme. Det kan innehålla reklam.

Avsnitt(793)

World Models and the Future of Spatial AI with Justin Johnson - #775

World Models and the Future of Spatial AI with Justin Johnson - #775

In this episode, Justin Johnson, co-founder of World Labs, joins us to discuss world models and the emerging field of spatial AI. We explore why many researchers see capabilities beyond language as an...

1 Sep 1h 6min

Why the Next AI Breakthrough May Come from Physics with Max Welling - #774

Why the Next AI Breakthrough May Come from Physics with Max Welling - #774

The conventional wisdom in AI is that the next breakthrough will come from more compute, more data, and larger models. But what if the next leap comes from somewhere else? In this episode, Max Wellin...

26 Aug 58min

Why Image Generation Needs More Than Bigger Models with Fatih Porikli - #773

Why Image Generation Needs More Than Bigger Models with Fatih Porikli - #773

Text-to-image models have become remarkably good at producing realistic images. But realism isn’t the same as correctness. Ask for several distinct people, a specific composition, or a high-resolution...

12 Aug 56min

Why Models Are AI’s Next Training Dataset with Damian Borth - #772

Why Models Are AI’s Next Training Dataset with Damian Borth - #772

For more than a decade, AI has advanced by training ever-larger models on ever-larger datasets. But as high-quality training data becomes harder to find and pretraining grows increasingly expensive, r...

27 Juli 47min

How AI Learns to Smell with Alex Wiltschko - #771

How AI Learns to Smell with Alex Wiltschko - #771

In this episode, Alex Wiltschko, founder and CEO of Osmo, joins the show to discuss his goal of giving computers a sense of smell and what it takes to build olfactory intelligence. We explore the sci...

8 Juli 59min

Why AI Agents Break the GenAI Security Model with Devvret Rishi - #770

Why AI Agents Break the GenAI Security Model with Devvret Rishi - #770

In this episode, Sam talks with Dev Rishi, GM of AI at Rubrik, about what happens when agents move beyond answering questions and start taking action across tools, systems, and business processes. We...

16 Juni 56min

Is RAG Dead? Lessons from Building AI for Tax Law with Alex Bowcut - #769

Is RAG Dead? Lessons from Building AI for Tax Law with Alex Bowcut - #769

As context windows grow into the millions of tokens, many AI practitioners are questioning whether retrieval-augmented generation (RAG) is still necessary. If modern models can ingest entire libraries...

9 Juni 51min

Relational Foundation Models for Enterprise Data with Jure Leskovec - #768

Relational Foundation Models for Enterprise Data with Jure Leskovec - #768

In this episode, Jure Leskovec, co-founder and chief scientist at Kumo and professor of computer science at Stanford, joins us to explore two fronts of his work: AI for science and relational deep lea...

21 Maj 1h 6min

Populärt inom Politik & nyheter

aftonbladet-krim
svenska-fall
p3-krim
en-runda-till
rss-krimstad
fordomspodden
aftonbladet-daily
flashback-forever
politiken
rss-vad-fan-hande
rss-krimreportrarna
kungligt
rss-sanning-konsekvens
rss-expressen-dok
motiv
svd-ledarredaktionen
rss-flodet
rss-frandfors-horna
spar
krimmagasinet