What makes a machine learning algorithm "superhuman"?

What makes a machine learning algorithm "superhuman"?

A few weeks ago, we podcasted about a neural network that was being touted as "better than doctors" in diagnosing pneumonia from chest x-rays, and how the underlying dataset used to train the algorithm raised some serious questions. We're back again this week with further developments, as the author of the original blog post pointed us toward more developments. All in all, there's a lot more clarity now around how the authors arrived at their original "better than doctors" claim, and a number of adjustments and improvements as the original result was de/re-constructed. Anyway, there are a few things that are cool about this. First, it's a worthwhile follow-up to a popular recent episode. Second, it goes *inside* an analysis to see what things like imbalanced classes, outliers, and (possible) signal leakage can do to real science. And last, it raises a really interesting question in an age when computers are often claimed to be better than humans: what do those claims really mean? Relevant links: https://lukeoakdenrayner.wordpress.com/2018/01/24/chexnet-an-in-depth-review/

Det här avsnittet är hämtat från ett öppet RSS-flöde och publiceras inte av Podme. Det kan innehålla reklam.

Avsnitt(321)

Understanding AI Text Watermarking

Understanding AI Text Watermarking

Anthropic just announced they're baking invisible watermarks directly into Claude's generated text — and while everyone else was busy having opinions about it, we were busy asking the more interesting...

24 Aug 29min

Better Know a Benchmark: Humanity's Last Exam

Better Know a Benchmark: Humanity's Last Exam

Humanity's Last Exam was designed with a bold premise: questions that human experts can answer, but AI models can't. Originally dubbed "Humanity's Last Stand," this benchmark is a massive academic col...

17 Aug 23min

A Scientific Deep Dive into Overconfident LLMs: Interview with Kaitlyn Zhou (Cornell)

A Scientific Deep Dive into Overconfident LLMs: Interview with Kaitlyn Zhou (Cornell)

When a language model tells you it's absolutely certain, is it actually more likely to be right? Kaitlyn Zhou's research says: not necessarily — sometimes confident phrasing correlates with *worse* ac...

10 Aug 33min

Reasoning Models: When LLMs Went Beyond Fancy Autocomplete

Reasoning Models: When LLMs Went Beyond Fancy Autocomplete

Reasoning models don't just answer your question — they *think out loud* first. In this episode we dig into the class of AI models that generate intermediate chains of thought before arriving at a fin...

3 Aug 25min

Distillation, or, How to Steal a Model

Distillation, or, How to Steal a Model

This week we’re covering model distillation: the technique of using a large "teacher" model's outputs to train a smaller, cheaper "student" model that mimics it. They cover the two big reasons labs do...

27 Juli 23min

Invisible LLM Failures and AI Fluency with Chris Potts (Stanford)

Invisible LLM Failures and AI Fluency with Chris Potts (Stanford)

What happens when a Stanford linguistics professor turns his attention to AI chatbots — and the surprisingly invisible ways humans misunderstand them? Chris Potts joins the show to unpack the hidden f...

20 Juli 41min

Still summer break: back next week

Still summer break: back next week

Still summer break: back next week by Katie Malone

13 Juli 25s

Summer break: back soon

Summer break: back soon

Summer break: back soon by Katie Malone

6 Juli 36s

Populärt inom Teknik

uppgang-och-fall
market-makers
rss-elektrikerpodden
rss-laddstationen-med-elbilen-i-sverige
 och-bilen-gar-bra
rss-technokratin
bilar-med-sladd
bli-saker-podden
skogsforum-podcast
gubbar-som-tjotar-om-bilar
natets-morka-sida
rss-en-ai-till-kaffet
rss-uppgang-och-fall
bosse-bildoktorn-och-hasse-p
rss-veckans-ai
hej-bruksbil
rss-fabriken-2
klocksnack-tillsammans-med-nymans-ur-1851
elbilsveckan
developers-mer-an-bara-kod