Interleaving
Linear Digressions22 Juli 2019

Interleaving

If you’re Google or Netflix, and you have a recommendation or search system as part of your bread and butter, what’s the best way to test improvements to your algorithm? A/B testing is the canonical answer for testing how users respond to software changes, but it gets tricky really fast to think about what an A/B test means in the context of an algorithm that returns a ranked list. That’s why we’re talking about interleaving this week—it’s a simple modification to A/B testing that makes it much easier to race two algorithms against each other and find the winner, and it allows you to do it with much less data than a traditional A/B test. Relevant links: https://medium.com/netflix-techblog/interleaving-in-online-experiments-at-netflix-a04ee392ec55 https://www.microsoft.com/en-us/research/publication/predicting-search-satisfaction-metrics-with-interleaved-comparisons/ https://www.cs.cornell.edu/people/tj/publications/joachims_02b.pdf

Det här avsnittet är hämtat från ett öppet RSS-flöde och publiceras inte av Podme. Det kan innehålla reklam.

Avsnitt(321)

Understanding AI Text Watermarking

Understanding AI Text Watermarking

Anthropic just announced they're baking invisible watermarks directly into Claude's generated text — and while everyone else was busy having opinions about it, we were busy asking the more interesting...

24 Aug 29min

Better Know a Benchmark: Humanity's Last Exam

Better Know a Benchmark: Humanity's Last Exam

Humanity's Last Exam was designed with a bold premise: questions that human experts can answer, but AI models can't. Originally dubbed "Humanity's Last Stand," this benchmark is a massive academic col...

17 Aug 23min

A Scientific Deep Dive into Overconfident LLMs: Interview with Kaitlyn Zhou (Cornell)

A Scientific Deep Dive into Overconfident LLMs: Interview with Kaitlyn Zhou (Cornell)

When a language model tells you it's absolutely certain, is it actually more likely to be right? Kaitlyn Zhou's research says: not necessarily — sometimes confident phrasing correlates with *worse* ac...

10 Aug 33min

Reasoning Models: When LLMs Went Beyond Fancy Autocomplete

Reasoning Models: When LLMs Went Beyond Fancy Autocomplete

Reasoning models don't just answer your question — they *think out loud* first. In this episode we dig into the class of AI models that generate intermediate chains of thought before arriving at a fin...

3 Aug 25min

Distillation, or, How to Steal a Model

Distillation, or, How to Steal a Model

This week we’re covering model distillation: the technique of using a large "teacher" model's outputs to train a smaller, cheaper "student" model that mimics it. They cover the two big reasons labs do...

27 Juli 23min

Invisible LLM Failures and AI Fluency with Chris Potts (Stanford)

Invisible LLM Failures and AI Fluency with Chris Potts (Stanford)

What happens when a Stanford linguistics professor turns his attention to AI chatbots — and the surprisingly invisible ways humans misunderstand them? Chris Potts joins the show to unpack the hidden f...

20 Juli 41min

Still summer break: back next week

Still summer break: back next week

Still summer break: back next week by Katie Malone

13 Juli 25s

Summer break: back soon

Summer break: back soon

Summer break: back soon by Katie Malone

6 Juli 36s

Populärt inom Teknik

uppgang-och-fall
market-makers
rss-elektrikerpodden
 och-bilen-gar-bra
rss-laddstationen-med-elbilen-i-sverige
bli-saker-podden
rss-technokratin
rss-en-ai-till-kaffet
rss-uppgang-och-fall
skogsforum-podcast
hej-bruksbil
gubbar-som-tjotar-om-bilar
bilar-med-sladd
rss-milpodden
natets-morka-sida
klocksnack-tillsammans-med-nymans-ur-1851
rss-fabriken-2
rss-veckans-ai
elbilsveckan
developers-mer-an-bara-kod