Why Tejal Patwardhan stopped underestimating the models - Episode 21

Why Tejal Patwardhan stopped underestimating the models - Episode 21

The old tests are getting too easy. Tejal Patwardhan leads OpenAI’s frontier evals team, which is finding new ways to measure and forecast progress as models become more capable. She and host Andrew Mayne discuss why evals matter for research, how benchmarks can break or get gamed, and what models need to be judged on next.


Chapters


00:00:24 Growing up at OpenAI

00:03:10 Why reasoning changed everything

00:06:28 What made o1 surprising

00:11:20 Why old benchmarks stopped working

00:14:45 What makes a good benchmark

00:17:35 Why evals are getting harder

00:22:09 Measuring voice and vision models

00:24:48 Testing models on real science

00:33:23 How OpenAI tracks frontier progress

00:40:47 What AI means for work



Hosted on Acast. See acast.com/privacy for more information.

Denne episoden er hentet fra en åpen RSS-feed og er ikke publisert av Podme. Den kan derfor inneholde annonser.

Episoder(23)

What racing reveals about working with AI - Episode 22

What racing reveals about working with AI - Episode 22

Racing is a sport of tiny margins and mountains of data. OpenAI researcher Joyce Ruffell and RaceTek Systems co-founder Chase Holden are each using AI to help teams make better use of information from...

16 Jul 41min

How a reasoning model cracked an 80-year-old math problem - Episode 20

How a reasoning model cracked an 80-year-old math problem - Episode 20

Last month AI found something mathematicians had missed for decades. Reasoning researchers Alexander Wei, Hongxun Wu, and Lijie Chen join the podcast to discuss how a general-purpose model helped disp...

4 Jun 41min

Inside image generation’s Renaissance moment - Episode 19

Inside image generation’s Renaissance moment - Episode 19

People are generating over 1.5 billion images a week in ChatGPT. In this episode, Product lead Adele Li and researcher Kenji Hata share some of the new use cases and trends since the launch of Images ...

14 Mai 29min

Why AI needs a new kind of supercomputer network - Episode 18

Why AI needs a new kind of supercomputer network - Episode 18

Training frontier models isn’t as simple as adding more GPUs—one small problem and the whole coordinated dance falls apart. OpenAI’s Mark Handley and Greg Steinbrecher discuss how a new supercomputer ...

6 Mai 37min

What happens now that AI is good at math? - Episode 17

What happens now that AI is good at math? - Episode 17

Math is one of the clearest ways to see how far AI has come in a short span. OpenAI researchers Sébastien Bubeck and Ernest Ryu join host Andrew Mayne to explain what changed and what it could mean fo...

28 Apr 43min

Building AI for Life Sciences - Episode 16

Building AI for Life Sciences - Episode 16

What does it take to build AI systems that can actually help scientists? Research lead Joy Jiao and product lead Yunyun Wang discuss how OpenAI is developing models for life sciences and what responsi...

16 Apr 44min

Inside the Model Spec - Episode 15

Inside the Model Spec - Episode 15

The more AI can do, the more we need to ask what it should and shouldn’t do. In this episode, OpenAI researcher Jason Wolfe joins host Andrew Mayne to talk about the Model Spec, the public framework t...

25 Mar 37min

Populært innen Vitenskap

fastlegen
tingenes-tilstand
romkapsel
jss
liberal-halvtime
rekommandert
villmarksliv
abels-tarn
dekodet-2
vett-og-vitenskap-med-gaute-einevoll
sinnsyn
fjellsportpodden
rss-overskuddsliv
rss-rekommandert
tomprat-med-gunnar-tjomlid
rss-inn-til-kjernen-med-sunniva-rose
hva-er-greia-med
rss-nysgjerrige-norge
diagnose
kvinnehelsepodden