OpenAI's Noam Brown, Ilge Akkaya and Hunter Lightman on o1 and Teaching LLMs to Reason Better
Training Data2 Loka 2024

OpenAI's Noam Brown, Ilge Akkaya and Hunter Lightman on o1 and Teaching LLMs to Reason Better

Combining LLMs with AlphaGo-style deep reinforcement learning has been a holy grail for many leading AI labs, and with o1 (aka Strawberry) we are seeing the most general merging of the two modes to date. o1 is admittedly better at math than essay writing, but it has already achieved SOTA on a number of math, coding and reasoning benchmarks. Deep RL legend and now OpenAI researcher Noam Brown and teammates Ilge Akkaya and Hunter Lightman discuss the ah-ha moments on the way to the release of o1, how it uses chains of thought and backtracking to think through problems, the discovery of strong test-time compute scaling laws and what to expect as the model gets better. Hosted by: Sonya Huang and Pat Grady, Sequoia Capital Mentioned in this episode: Learning to Reason with LLMs: Technical report accompanying the launch of OpenAI o1. Generator verifier gap: Concept Noam explains in terms of what kinds of problems benefit from more inference-time compute. Agent57: Outperforming the human Atari benchmark, 2020 paper where DeepMind demonstrated “the first deep reinforcement learning agent to obtain a score that is above the human baseline on all 57 Atari 2600 games.” Move 37: Pivotal move in AlphaGo’s second game against Lee Sedol where it made a move so surprising that Sedol thought it must be a mistake, and only later discovered he had lost the game to a superhuman move. IOI competition: OpenAI entered o1 into the International Olympiad in Informatics and received a Silver Medal. System 1, System 2: The thesis if Danial Khaneman’s pivotal book of behavioral economics, Thinking, Fast and Slow, that positied two distinct modes of thought, with System 1 being fast and instinctive and System 2 being slow and rational. AlphaZero: The predecessor to AlphaGo which learned a variety of games completely from scratch through self-play. Interestingly, self-play doesn’t seem to have a role in o1. Solving Rubik’s Cube with a robot hand: Early OpenAI robotics paper that Ilge Akkaya worked on. The Last Question: Science fiction story by Isaac Asimov with interesting parallels to scaling inference-time compute. Strawberry: Why? O1-mini: A smaller, more efficient version of 1 for applications that require reasoning without broad world knowledge. 00:00 - Introduction 01:33 - Conviction in o1 04:24 - How o1 works 05:04 - What is reasoning? 07:02 - Lessons from gameplay 09:14 - Generation vs verification 10:31 - What is surprising about o1 so far 11:37 - The trough of disillusionment 14:03 - Applying deep RL 14:45 - o1’s AlphaGo moment? 17:38 - A-ha moments 21:10 - Why is o1 good at STEM? 24:10 - Capabilities vs usefulness 25:29 - Defining AGI 26:13 - The importance of reasoning 28:39 - Chain of thought 30:41 - Implication of inference-time scaling laws 35:10 - Bottlenecks to scaling test-time compute 38:46 - Biggest misunderstanding about o1? 41:13 - o1-mini 42:15 - How should founders think about o1?

Tämä jakso on lisätty Podme-palveluun avoimen RSS-syötteen kautta eikä se ole Podmen omaa tuotantoa. Siksi jakso saattaa sisältää mainontaa.

Jaksot(105)

Chai Discovery's Bitter Lesson: Drug Design Is Another Scaling Problem

Chai Discovery's Bitter Lesson: Drug Design Is Another Scaling Problem

Most people treat biology as a bespoke, messy science. Josh Meier and Matt McPartlon, co-founders of Chai Discovery, treat it as an engineering problem. They make the case that drug design obeys the b...

4 Elo 47min

Building the Automated AGI Lab: Core Automation's Jerry Tworek and Rohan Anil

Building the Automated AGI Lab: Core Automation's Jerry Tworek and Rohan Anil

Jerry Tworek led reasoning at OpenAI, convinced that scaling reinforcement learning was the path to AGI. Rohan Anil co-led Gemini pre-training and built the Shampoo optimizer. Now they've teamed up at...

29 Heinä 49min

Factory's Matan Grinberg: The Coming ‘Dark Factory’ Where Software Builds Itself

Factory's Matan Grinberg: The Coming ‘Dark Factory’ Where Software Builds Itself

Factory started building fully autonomous coding agents in April 2023, two years before enterprises were ready. Matan Grinberg now says this is indistinguishable from being wrong. The Factory co-found...

21 Heinä 51min

Anthropic's Katelyn Lesse & Angela Jiang: Building an Ecosystem, not a Walled Garden

Anthropic's Katelyn Lesse & Angela Jiang: Building an Ecosystem, not a Walled Garden

Katelyn Lesse and Angela Jiang lead the team building Anthropic's developer platform - the layer that both outside builders and Anthropic's own products run on top of. Angela frames the platform as a ...

14 Heinä 48min

Inside Zipline's Autonomous System: 140M Miles, Zero Incidents

Inside Zipline's Autonomous System: 140M Miles, Zero Incidents

The largest commercial autonomous system on earth isn't a robotaxi fleet — it's Zipline, which has flown 140 million autonomous miles with zero safety incidents. Co-founder Keller Rinaudo Cliffton and...

7 Heinä 55min

Why Hardware-Software Co-Design Is AI's Real 100x: Dylan Patel of SemiAnalysis

Why Hardware-Software Co-Design Is AI's Real 100x: Dylan Patel of SemiAnalysis

Dylan Patel, founder of SemiAnalysis, argues the biggest gains in AI don't come from faster chips, they come from software-hardware co-design. Optimizing the model, the kernels, and the silicon togeth...

30 Kesä 1h 10min

Memory and Continual Learning: Engram's Dan Biderman and Jessy Lin

Memory and Continual Learning: Engram's Dan Biderman and Jessy Lin

Dan Biderman and Jessy Lin, co-founders of Engram, are building a neolab around memory and continual learning, which they call two sides of the same coin. Their contrarian premise: instead of stuffing...

24 Kesä 44min

Simulating Humans at Scale: Simile's Joon Sung Park

Simulating Humans at Scale: Simile's Joon Sung Park

The race to build superintelligence is producing models that keep getting better at objective problems, but not at behaving like actual people. Joon Sung Park, founder and CEO of Simile and creator of...

16 Kesä 38min

Suosittua kategoriassa Liike-elämä ja talous

sijotuskasti
psykopodiaa-podcast
mimmit-sijoittaa
ostan-asuntoja-podcast
rss-rahapodi
raharesepti
pomojen-suusta
pari-sanaa-lastensuojelusta
rss-sami-miettinen-neuvottelija
rss-porssipodi
rss-oivalluksia-rahasta-elamasta
rss-h-asselmoilanen
rss-laiskan-yrittajan-markkinointi-podcast
rss-myyntiradio
rss-deloitten-podcastit
the-mel-schmidt-show-yrittajyys-lifestyle-kasvu
pelkkana-korvana
rss-regenerative-business-uudistava-liiketoiminta
juristipodi
rss-lahtijat