AI Agents: Substance or Snake Oil with Arvind Narayanan - #704

AI Agents: Substance or Snake Oil with Arvind Narayanan - #704

Today, we're joined by Arvind Narayanan, professor of Computer Science at Princeton University to discuss his recent works, AI Agents That Matter and AI Snake Oil. In “AI Agents That Matter”, we explore the range of agentic behaviors, the challenges in benchmarking agents, and the ‘capability and reliability gap’, which creates risks when deploying AI agents in real-world applications. We also discuss the importance of verifiers as a technique for safeguarding agent behavior. We then dig into the AI Snake Oil book, which uncovers examples of problematic and overhyped claims in AI. Arvind shares various use cases of failed applications of AI, outlines a taxonomy of AI risks, and shares his insights on AI’s catastrophic risks. Additionally, we also touched on different approaches to LLM-based reasoning, his views on tech policy and regulation, and his work on CORE-Bench, a benchmark designed to measure AI agents' accuracy in computational reproducibility tasks. The complete show notes for this episode can be found at https://twimlai.com/go/704.

Episoder(781)

Robotic Perception and Control with Chelsea Finn - TWiML Talk #29

Robotic Perception and Control with Chelsea Finn - TWiML Talk #29

This week we continue our series on industrial applications of machine learning and AI with a conversation with Chelsea Finn, a PhD student at UC Berkeley. Chelsea’s research is focused on machine lea...

23 Jun 201754min

Reinforcement Learning Deep Dive with Pieter Abbeel - TWiML Talk #28

Reinforcement Learning Deep Dive with Pieter Abbeel - TWiML Talk #28

This week our guest is Pieter Abbeel, Assistant Professor at UC Berkeley, Research Scientist at OpenAI, and Cofounder of Gradescope. Pieter has an extensive background in AI research, going way back t...

17 Jun 201752min

Intelligent Autonomous Robots with Ilia Baranov - TWiML Talk #27

Intelligent Autonomous Robots with Ilia Baranov - TWiML Talk #27

Our first guest in the Industrial AI series is Ilia Baranov, engineering manager at Clearpath Robotics. Ilia is responsible for setting the engineering direction for all of Clearpath’s research platfo...

9 Jun 201753min

Global AI Trends with Ben Lorica - TWiML Talk #26

Global AI Trends with Ben Lorica - TWiML Talk #26

This week I’ve invited my friend Ben Lorica onto the show. Ben is Chief Data Scientist for O’Reilly Media, and Program Director of Strata Data & the O'Reilly A.I. conference. Ben has worked on analyti...

2 Jun 201754min

Offensive vs Defensive Data Science with Deep Varma - TWiML Talk #25

Offensive vs Defensive Data Science with Deep Varma - TWiML Talk #25

This week on the show my guest is Deep Varma, Vice President of Data Engineering at real estate startup Trulia. Deep has run data engineering teams in silicon valley for well over a decade, and is now...

26 Mai 201753min

Reinforcement Learning: The Next Frontier of Gaming with Danny Lange - TWiML Talk #24

Reinforcement Learning: The Next Frontier of Gaming with Danny Lange - TWiML Talk #24

My guest on the show this week is Danny Lange, VP for Machine Learning & AI at video game technology developer Unity Technologies. Danny is well traveled in the world of ML and AI, and has had a hand ...

20 Mai 201754min

Integrating Psycholinguistics into AI with Dominique Simmons - TWiML Talk #23

Integrating Psycholinguistics into AI with Dominique Simmons - TWiML Talk #23

I think you’re really going to enjoy today’s show. Our guest this week is Dominique Simmons, Applied research Scientist at AI tools vendor Dimensional Mechanics. Dominique brings an interesting backgr...

12 Mai 20171h

Deep Neural Nets for Visual Recognition with Matt Zeiler - TWiML Talk #22

Deep Neural Nets for Visual Recognition with Matt Zeiler - TWiML Talk #22

Today we bring you our final interview from backstage at the NYU FutureLabs AI Summit. Our guest this week is Matt Zeiler. Matt graduated from the University of Toronto where he worked with deep learn...

5 Mai 201722min

Populært innen Politikk og nyheter

giver-og-gjengen-vg
aftenpodden
forklart
aftenpodden-usa
popradet
stopp-verden
lydartikler-fra-aftenposten
i-retten
rss-gukild-johaug
fotballpodden-2
det-store-bildet
dine-penger-pengeradet
nokon-ma-ga
rss-ness
hanna-de-heldige
aftenbla-bla
bt-dokumentar-2
e24-podden
frokostshowet-pa-p5
rss-dannet-uten-piano