Generative Benchmarking with Kelly Hong - #728
In this episode, Kelly Hong, a researcher at Chroma, joins us to discuss "Generative Benchmarking," a novel approach to evaluating retrieval systems, like RAG applications, using synthetic data. Kelly explains how traditional benchmarks like MTEB fail to represent real-world query patterns and how embedding models that perform well on public benchmarks often underperform in production. The conversation explores the two-step process of Generative Benchmarking: filtering documents to focus on relevant content and generating queries that mimic actual user behavior. Kelly shares insights from applying this approach to Weights & Biases' technical support bot, revealing how domain-specific evaluation provides more accurate assessments of embedding model performance. We also discuss the importance of aligning LLM judges with human preferences, the impact of chunking strategies on retrieval effectiveness, and how production queries differ from benchmark queries in ambiguity and style. Throughout the episode, Kelly emphasizes the need for systematic evaluation approaches that go beyond "vibe checks" to help developers build more effective RAG applications. The complete show notes for this episode can be found at https://twimlai.com/go/728.

Episoder(779)

Robotics at OpenAI with Jonas Schneider - TWiML Talk #76

Robotics at OpenAI with Jonas Schneider - TWiML Talk #76

The show is part of a series that I’m really excited about, in part because I’ve been working to bring them to you for quite a while now. The focus of the series is a sampling of the interesting work ...

1 Des 201745min

AI Robustness and Safety with Dario Amodei - TWiML Talk #75

AI Robustness and Safety with Dario Amodei - TWiML Talk #75

The show is part of a series that I’m really excited about, in part because I’ve been working to bring them to you for quite a while now. The focus of the series is a sampling of the interesting work ...

30 Nov 201736min

Towards Artificial General Intelligence with Greg Brockman - TWiML Talk #74

Towards Artificial General Intelligence with Greg Brockman - TWiML Talk #74

The show is part of a series that I’m really excited about, in part because I’ve been working to bring them to you for quite a while now. The focus of the series is a sampling of the interesting work ...

28 Nov 201755min

Explaining Black Box Predictions with Sam Ritchie - TWiML Talk #73

Explaining Black Box Predictions with Sam Ritchie - TWiML Talk #73

This week, we’ll be featuring a series of shows recorded from Strange Loop, a great developer-focused conference that takes place every year right in my backyard! The conference is a multi-disciplinar...

25 Nov 201738min

Experimental Creative Writing with the Vectorized Word - Allison Parish - TWIML Talk #72

Experimental Creative Writing with the Vectorized Word - Allison Parish - TWIML Talk #72

This week, we’ll be featuring a series of shows recorded from Strange Loop, a great developer-focused conference that takes place every year right in my backyard! The conference is a multi-disciplinar...

24 Nov 201728min

The Biological Path Towards Strong AI - Matthew Taylor - TWiML Talk #71

The Biological Path Towards Strong AI - Matthew Taylor - TWiML Talk #71

This week, we’ll be featuring a series of shows recorded from Strange Loop, a great developer-focused conference that takes place every year right in my backyard! The conference is a multi-disciplinar...

22 Nov 201737min

Pytorch: Fast Differentiable Dynamic Graphs in Python with Soumith Chintala - TWiML Talk #70

Pytorch: Fast Differentiable Dynamic Graphs in Python with Soumith Chintala - TWiML Talk #70

This week, we’ll be featuring a series of shows recorded from Strange Loop, a great developer-focused conference that takes place every year right in my backyard! The conference is a multi-disciplinar...

21 Nov 201742min

Accessible Machine Learning for the Enterprise Developer with Ryan Sevey & Jason Montgomery

Accessible Machine Learning for the Enterprise Developer with Ryan Sevey & Jason Montgomery

This week, we’ll be featuring a series of shows recorded from Strange Loop, a great developer-focused conference that takes place every year right in my backyard! The conference is a multi-disciplinar...

20 Nov 201745min

Populært innen Politikk og nyheter

giver-og-gjengen-vg
aftenpodden-usa
aftenpodden
i-retten
forklart
stopp-verden
popradet
fotballpodden-2
rss-gukild-johaug
nokon-ma-ga
det-store-bildet
dine-penger-pengeradet
bt-dokumentar-2
aftenbla-bla
hanna-de-heldige
rss-penger-polser-og-politikk
rss-dannet-uten-piano
frokostshowet-pa-p5
rss-ness
e24-podden