Why Tejal Patwardhan stopped underestimating the models - Episode 21

Why Tejal Patwardhan stopped underestimating the models - Episode 21

The old tests are getting too easy. Tejal Patwardhan leads OpenAI’s frontier evals team, which is finding new ways to measure and forecast progress as models become more capable. She and host Andrew Mayne discuss why evals matter for research, how benchmarks can break or get gamed, and what models need to be judged on next.


Chapters


00:00:24 Growing up at OpenAI

00:03:10 Why reasoning changed everything

00:06:28 What made o1 surprising

00:11:20 Why old benchmarks stopped working

00:14:45 What makes a good benchmark

00:17:35 Why evals are getting harder

00:22:09 Measuring voice and vision models

00:24:48 Testing models on real science

00:33:23 How OpenAI tracks frontier progress

00:40:47 What AI means for work



Hosted on Acast. See acast.com/privacy for more information.

Tämä jakso on lisätty Podme-palveluun avoimen RSS-syötteen kautta eikä se ole Podmen omaa tuotantoa. Siksi jakso saattaa sisältää mainontaa.

Jaksot(23)

What racing reveals about working with AI - Episode 22

What racing reveals about working with AI - Episode 22

Racing is a sport of tiny margins and mountains of data. OpenAI researcher Joyce Ruffell and RaceTek Systems co-founder Chase Holden are each using AI to help teams make better use of information from...

16 Heinä 41min

How a reasoning model cracked an 80-year-old math problem - Episode 20

How a reasoning model cracked an 80-year-old math problem - Episode 20

Last month AI found something mathematicians had missed for decades. Reasoning researchers Alexander Wei, Hongxun Wu, and Lijie Chen join the podcast to discuss how a general-purpose model helped disp...

4 Kesä 41min

Inside image generation’s Renaissance moment - Episode 19

Inside image generation’s Renaissance moment - Episode 19

People are generating over 1.5 billion images a week in ChatGPT. In this episode, Product lead Adele Li and researcher Kenji Hata share some of the new use cases and trends since the launch of Images ...

14 Touko 29min

Why AI needs a new kind of supercomputer network - Episode 18

Why AI needs a new kind of supercomputer network - Episode 18

Training frontier models isn’t as simple as adding more GPUs—one small problem and the whole coordinated dance falls apart. OpenAI’s Mark Handley and Greg Steinbrecher discuss how a new supercomputer ...

6 Touko 37min

What happens now that AI is good at math? - Episode 17

What happens now that AI is good at math? - Episode 17

Math is one of the clearest ways to see how far AI has come in a short span. OpenAI researchers Sébastien Bubeck and Ernest Ryu join host Andrew Mayne to explain what changed and what it could mean fo...

28 Huhti 43min

Building AI for Life Sciences - Episode 16

Building AI for Life Sciences - Episode 16

What does it take to build AI systems that can actually help scientists? Research lead Joy Jiao and product lead Yunyun Wang discuss how OpenAI is developing models for life sciences and what responsi...

16 Huhti 44min

Inside the Model Spec - Episode 15

Inside the Model Spec - Episode 15

The more AI can do, the more we need to ask what it should and shouldn’t do. In this episode, OpenAI researcher Jason Wolfe joins host Andrew Mayne to talk about the Model Spec, the public framework t...

25 Maalis 37min

Suosittua kategoriassa Tiede

rss-poliisin-mieli
tiedekulma-podcast
rss-mita-tulisi-tietaa
rss-tiedetta-vai-tarinaa
docemilia
rss-duodecim-lehti
rss-radplus
utelias-mieli
hippokrateen-vastaanotolla
rss-hereilla
rss-ranskaa-raakana
rss-bios-podcast
filocast-filosofian-perusteet
rss-tervetta-skeptisyytta
mielipaivakirja
radio-antro