Generative Benchmarking with Kelly Hong - #728
In this episode, Kelly Hong, a researcher at Chroma, joins us to discuss "Generative Benchmarking," a novel approach to evaluating retrieval systems, like RAG applications, using synthetic data. Kelly explains how traditional benchmarks like MTEB fail to represent real-world query patterns and how embedding models that perform well on public benchmarks often underperform in production. The conversation explores the two-step process of Generative Benchmarking: filtering documents to focus on relevant content and generating queries that mimic actual user behavior. Kelly shares insights from applying this approach to Weights & Biases' technical support bot, revealing how domain-specific evaluation provides more accurate assessments of embedding model performance. We also discuss the importance of aligning LLM judges with human preferences, the impact of chunking strategies on retrieval effectiveness, and how production queries differ from benchmark queries in ambiguity and style. Throughout the episode, Kelly emphasizes the need for systematic evaluation approaches that go beyond "vibe checks" to help developers build more effective RAG applications. The complete show notes for this episode can be found at https://twimlai.com/go/728.

Avsnitt(781)

Growth Hacking Sports w/ Machine Learning with Noah Gift - TWiML Talk #158

Growth Hacking Sports w/ Machine Learning with Noah Gift - TWiML Talk #158

In this episode of our AI in Sports series I'm joined by Noah Gift, Founder and Consulting CTO at Pragmatic Labs and professor at UC Davis. Noah and I discuss some of his recent work in using social m...

28 Juni 201850min

Fine-Grained Player Prediction in Sports with Jennifer Hobbs - TWiML Talk #157

Fine-Grained Player Prediction in Sports with Jennifer Hobbs - TWiML Talk #157

In this episode of our AI in Sports series, I'm joined by Jennifer Hobbs, Senior Data Scientist at STATS, a collector and distributor of sports data, to discuss the STATS data pipeline and how they co...

27 Juni 201842min

Targeted Ticket Sales Using Azure ML with the Trail Blazers w/ Mike Schumacher & Chenhui Hu - TWiML Talk #156

Targeted Ticket Sales Using Azure ML with the Trail Blazers w/ Mike Schumacher & Chenhui Hu - TWiML Talk #156

In today’s episode of our AI in Sports series I'm joined by Mike Schumacher, director of business analytics for the Portland Trail Blazers, and Chenhui Hu, a data scientist at Microsoft to discuss how...

26 Juni 201837min

AI for Athlete Optimization with Sinead Flahive - TWiML Talk #155

AI for Athlete Optimization with Sinead Flahive - TWiML Talk #155

This week we’re excited to kick off a series of shows on AI in sports. In this episode I'm joined by Sinead Flahive, data scientist at Dublin, Ireland based Kitman Labs to discuss Kitman’s Athlete Opt...

25 Juni 201840min

Omni-Channel Customer Experiences with Vince Jeffs - TWiML Talk #154

Omni-Channel Customer Experiences with Vince Jeffs - TWiML Talk #154

In this, the final episode of our PegaWorld series I’m joined by Vince Jeffs, Senior Director of Product Strategy for AI and Decisioning at Pegasystems. Vince and I had a great talk about the role AI ...

21 Juni 201843min

Workforce Intelligence for Automation & Productivity with Michael Kempe - TWiML Talk #153

Workforce Intelligence for Automation & Productivity with Michael Kempe - TWiML Talk #153

In this episode of our PegaWorld series, I’m joined by Michael Kempe, chief operating officer at global share registry and financial services provider Link Market Services. In the interview, Michael a...

20 Juni 201836min

Data Platforms for Decision Automation at Scotiabank with Jim Saleh - TWiML Talk #152

Data Platforms for Decision Automation at Scotiabank with Jim Saleh - TWiML Talk #152

In this show, part of our PegaWorld 18 series, I'm joined by Jim Saleh, Senior Director of process and decision automation at Scotiabank. Jim is tasked with helping the bank transition from a world wh...

19 Juni 201832min

Towards the Self-Driving Enterprise with Kirk Borne - TWiML Talk #151

Towards the Self-Driving Enterprise with Kirk Borne - TWiML Talk #151

In this show, the first of our PegaWorld 18 series, I'm joined by Kirk Borne, Principal Data Scientist at management consulting firm Booz Allen Hamilton. In our conversation, Kirk shares his views on ...

18 Juni 201841min

Populärt inom Politik & nyheter

aftonbladet-krim
svenska-fall
p3-krim
rss-krimstad
fordomspodden
rss-expressen-dok
flashback-forever
rss-sanning-konsekvens
motiv
aftonbladet-daily
spar
rss-vad-fan-hande
blenda-2
olyckan-inifran
rss-krimreportrarna
rss-frandfors-horna
rss-flodet
dagens-eko
svd-ledarredaktionen
grans