Keeping ourselves honest when we work with observational healthcare data

Keeping ourselves honest when we work with observational healthcare data

The abundance of data in healthcare, and the value we could capture from structuring and analyzing that data, is a huge opportunity. It also presents huge challenges. One of the biggest challenges is how, exactly, to do that structuring and analysis—data scientists working with this data have hundreds or thousands of small, and sometimes large, decisions to make in their day-to-day analysis work. What data should they include in their studies? What method should they use to analyze it? What hyperparameter settings should they explore, and how should they pick a value for their hyperparameters? The thing that’s really difficult here is that, depending on which path they choose among many reasonable options, a data scientist can get really different answers to the underlying question, which makes you wonder how to conclude anything with certainty at all. The paper for this week’s episode performs a systematic study of many, many different permutations of the questions above on a set of benchmark datasets where the “right” answers are known. Which strategies are most likely to yield the “right” answers? That’s the whole topic of discussion. Relevant links: https://hdsr.mitpress.mit.edu/pub/fxz7kr65

Denne episoden er hentet fra en åpen RSS-feed og er ikke publisert av Podme. Den kan derfor inneholde annonser.

Episoder(320)

Better Know a Benchmark: Humanity's Last Exam

Better Know a Benchmark: Humanity's Last Exam

Humanity's Last Exam was designed with a bold premise: questions that human experts can answer, but AI models can't. Originally dubbed "Humanity's Last Stand," this benchmark is a massive academic col...

17 Aug 23min

A Scientific Deep Dive into Overconfident LLMs: Interview with Kaitlyn Zhou (Cornell)

A Scientific Deep Dive into Overconfident LLMs: Interview with Kaitlyn Zhou (Cornell)

When a language model tells you it's absolutely certain, is it actually more likely to be right? Kaitlyn Zhou's research says: not necessarily — sometimes confident phrasing correlates with *worse* ac...

10 Aug 33min

Reasoning Models: When LLMs Went Beyond Fancy Autocomplete

Reasoning Models: When LLMs Went Beyond Fancy Autocomplete

Reasoning models don't just answer your question — they *think out loud* first. In this episode we dig into the class of AI models that generate intermediate chains of thought before arriving at a fin...

3 Aug 25min

Distillation, or, How to Steal a Model

Distillation, or, How to Steal a Model

This week we’re covering model distillation: the technique of using a large "teacher" model's outputs to train a smaller, cheaper "student" model that mimics it. They cover the two big reasons labs do...

27 Jul 23min

Invisible LLM Failures and AI Fluency with Chris Potts (Stanford)

Invisible LLM Failures and AI Fluency with Chris Potts (Stanford)

What happens when a Stanford linguistics professor turns his attention to AI chatbots — and the surprisingly invisible ways humans misunderstand them? Chris Potts joins the show to unpack the hidden f...

20 Jul 41min

Still summer break: back next week

Still summer break: back next week

Still summer break: back next week by Katie Malone

13 Jul 25s

Summer break: back soon

Summer break: back soon

Summer break: back soon by Katie Malone

6 Jul 36s

Interviewing the Linear Digressions Agents (The Agents Season, Episode 11)

Interviewing the Linear Digressions Agents (The Agents Season, Episode 11)

After a five-year hiatus, the podcast that burned out partly over the tedium of writing episode descriptions is back — and using AI agents to handle exactly that task. The season-11 finale turns the l...

28 Jun 37min

Populært innen Teknologi

lydartikler-fra-aftenposten
tomprat-med-gunnar-tjomlid
teknisk-sett
smart-forklart
elektropodden
nasjonal-sikkerhetsmyndighet-nsm
rss-ki-praten
fornybaren
shifter
rss-kunstig-intelligens-med-elisabeth-maren-og-morten
teknologi-og-mennesker
rss-alt-som-gar-pa-strom
rss-snakk-om-sikkerhet
rss-ai-forklart
rss-alt-vi-kan
rss-fisketimen
digital-forretningsforstaelse
rss-bouvet-bobler
pedagogisk-intelligens
rss-polypod