A Key Concept in AI Alignment: Deep Reinforcement Learning from Human Preferences

A Key Concept in AI Alignment: Deep Reinforcement Learning from Human Preferences

Modern AI chatbots have a few different things that go into creating them. Today we're going to talk about a really important part of the process: the alignment training, where the chatbot goes from being just a pre-trained model—something that's kind of a fancy autocomplete—to something that really gives responses to human prompts that are more conversational, that are closer to the ones that we experience when we actually use a model like ChatGPT or Gemini or Claude. To go from the pre-trained model to one that's aligned, that's ready for a human to talk with, it uses reinforcement learning. And a really important step in figuring out the right way to frame the reinforcement learning problem happened in 2017 with a paper that we're going to talk about today: Deep Reinforcement Learning from Human Preferences. You are listening to Linear Digressions. The paper discussed in this episode is Deep Reinforcement Learning from Human Preferences https://arxiv.org/abs/1706.03741

Det här avsnittet är hämtat från ett öppet RSS-flöde och publiceras inte av Podme. Det kan innehålla reklam.

Avsnitt(320)

Better Know a Benchmark: Humanity's Last Exam

Better Know a Benchmark: Humanity's Last Exam

Humanity's Last Exam was designed with a bold premise: questions that human experts can answer, but AI models can't. Originally dubbed "Humanity's Last Stand," this benchmark is a massive academic col...

17 Aug 23min

A Scientific Deep Dive into Overconfident LLMs: Interview with Kaitlyn Zhou (Cornell)

A Scientific Deep Dive into Overconfident LLMs: Interview with Kaitlyn Zhou (Cornell)

When a language model tells you it's absolutely certain, is it actually more likely to be right? Kaitlyn Zhou's research says: not necessarily — sometimes confident phrasing correlates with *worse* ac...

10 Aug 33min

Reasoning Models: When LLMs Went Beyond Fancy Autocomplete

Reasoning Models: When LLMs Went Beyond Fancy Autocomplete

Reasoning models don't just answer your question — they *think out loud* first. In this episode we dig into the class of AI models that generate intermediate chains of thought before arriving at a fin...

3 Aug 25min

Distillation, or, How to Steal a Model

Distillation, or, How to Steal a Model

This week we’re covering model distillation: the technique of using a large "teacher" model's outputs to train a smaller, cheaper "student" model that mimics it. They cover the two big reasons labs do...

27 Juli 23min

Invisible LLM Failures and AI Fluency with Chris Potts (Stanford)

Invisible LLM Failures and AI Fluency with Chris Potts (Stanford)

What happens when a Stanford linguistics professor turns his attention to AI chatbots — and the surprisingly invisible ways humans misunderstand them? Chris Potts joins the show to unpack the hidden f...

20 Juli 41min

Still summer break: back next week

Still summer break: back next week

Still summer break: back next week by Katie Malone

13 Juli 25s

Summer break: back soon

Summer break: back soon

Summer break: back soon by Katie Malone

6 Juli 36s

Interviewing the Linear Digressions Agents (The Agents Season, Episode 11)

Interviewing the Linear Digressions Agents (The Agents Season, Episode 11)

After a five-year hiatus, the podcast that burned out partly over the tedium of writing episode descriptions is back — and using AI agents to handle exactly that task. The season-11 finale turns the l...

28 Juni 37min

Populärt inom Teknik

uppgang-och-fall
market-makers
rss-elektrikerpodden
rss-laddstationen-med-elbilen-i-sverige
skogsforum-podcast
 och-bilen-gar-bra
rss-en-ai-till-kaffet
rss-uppgang-och-fall
bli-saker-podden
rss-technokratin
hej-bruksbil
bilar-med-sladd
gubbar-som-tjotar-om-bilar
rss-veckans-ai
under-femton
klocksnack-tillsammans-med-nymans-ur-1851
rss-milpodden
bosse-bildoktorn-och-hasse-p
elbilsveckan
natets-morka-sida