AI Safety Goes Mainstream

AI Safety Goes Mainstream

An Anthropic researcher's viral resignation post last week set off a maelstrom of fear around extreme AI risks. Aalok and Nicole discuss the resulting vibe shift and how frontier labs are responding. They also unpack OpenAI solving one of the most difficult open problems in mathematics and reports of more rogue agent activity on the open internet. Timestamps: OpenAI solves Millenium Problem (1:04) Rogue OpenAI agents take over a German wiki (12:14) Anthropic researcher's resignation post goes viral (19:34) Dario Amodei commits to embedding third-party evaluators inside Anthropic (28:12) Where the U.S. government can step in (35:12) Additional Reading: "On the Navier-Stokes Millennium Prize Problem" by OpenAI: https://openai.com/index/navier-stokes-solution/ "An OpenAI model has disproved a central conjecture in discrete geometry" by OpenAI: https://openai.com/index/model-disproves-discrete-geometry-conjecture/ Statement on Navier-Stokes by NYU Professor Tristan Buckmaster: https://cims.nyu.edu/~tristanb/statement.pdf "An alignment assessment of recent cybersecurity incidents" by Anthropic: https://www.anthropic.com/research/alignment-assessment-cybersecurity-incidents "OpenAI's rogue agents used at least 10 more sites for unauthorized comms, researchers say" by Reuters: https://www.reuters.com/world/openais-rogue-agents-used-least-10-more-sites-unauthorized-comms-rese… "OpenAI agents carried out an undisclosed cyber-attack on RubyGems" by Spencer Kitts, Thomas Larsen, and Sydney Von Arx: https://www.rubyhack.ai "Discovery of a new OpenAI agent message board" by Sydney Von Arx, Cormac Slade Byrd, Spencer Kitts, and Thomas Larsen: https://collusion.wiki "Anthropic Researcher Quits Over 'Out-of-Control' AI Fears" by The Wall Street Journal: https://www.wsj.com/tech/ai/anthropic-researcher-quits-over-out-of-control-ai-fears-707b7628 "The AI policy window is open. We need to act." by OpenAI Chief Global Affairs Officer Chris Lehane: https://openai.com/index/ai-policy-window/ "We Must Pace the Frontier" by Dario Amodei: https://darioamodei.com/post/we-must-pace-the-frontier

Denne episoden er hentet fra en åpen RSS-feed og er ikke publisert av Podme. Den kan derfor inneholde annonser.

Episoder(111)

OpenAI Agents Hack Australian Government and Trump Rejects Global Governance at UNGA

OpenAI Agents Hack Australian Government and Trump Rejects Global Governance at UNGA

This week, Aalok and Nicole discuss new reports of rogue AI agents on the open internet – including the first known case of an agent hacking a government system – and the lackluster outcome of last we...

1 Okt 52min

How AI Startups are Navigating Regulation with Maryam Mujica and Munjal Shah

How AI Startups are Navigating Regulation with Maryam Mujica and Munjal Shah

This week, Aalok is joined by Maryam Mujica, president of the General Catalyst Institute, and Munjal Shah, co-founder and CEO of generative AI startup Hippocratic AI, for a discussion on how generativ...

24 Sep 57min

Decoding State-Level Regulation with Jason Elliott

Decoding State-Level Regulation with Jason Elliott

This week, Aalok is joined by Jason Elliott, the former Deputy Chief of Staff to California Governor Gavin Newsom, for a conversation about regulating AI at the state level. Guests: Jason Elliott...

10 Sep 51min

METR Investigates OpenAI-Hugging Face Incident and California Legislature Advances IVO Framework

METR Investigates OpenAI-Hugging Face Incident and California Legislature Advances IVO Framework

This week, we unpack findings from METR's six-day investigation of the OpenAI-Hugging Face cyber incident. We also discuss an IVO bill passed by the California Legislature, Bill Gates' take on the cur...

3 Sep 43min

Responding to AI Agent Containment Failures with CSET's Helen Toner, LawAI's Mackenzie Arnold & CSIS's Matt Pearl

Responding to AI Agent Containment Failures with CSET's Helen Toner, LawAI's Mackenzie Arnold & CSIS's Matt Pearl

This episode cross-posts a panel from the event "AI Agent Containment Failures: Technical Realities and Policy Responses," co-hosted by the Wadhwani AI Center and the Institute for Law and AI on Augus...

27 Aug 1h 1min

OpenAI Pauses RL Training and Anthropic Adds Watermarks to AI-Generated Text

OpenAI Pauses RL Training and Anthropic Adds Watermarks to AI-Generated Text

This week, we cover updates from the ongoing cyber saga, including OpenAI's two-week pause on RL training and the cyber capabilities of Z.ai's latest model GLM-5.3. We also unpack Anthropic's move to ...

20 Aug 47min

Planning for the AI-Powered Economy with Windfall's Adrian Brown

Planning for the AI-Powered Economy with Windfall's Adrian Brown

In this episode, we're joined by Adrian Brown, founder and CEO of Windfall Trust, for a conversation about the economic impacts of AI. Why the economic focus (00:58) Capturing AI's impact in macr...

13 Aug 46min

Populært innen Politikk og nyheter

giver-og-gjengen-vg
aftenpodden
forklart
aftenpodden-usa
stopp-verden
fotballpodden-2
popradet
dine-penger-pengeradet
rss-gukild-johaug
nokon-ma-ga
rss-espen-lee-usensurert
det-store-bildet
hanna-de-heldige
aftenbla-bla
rss-ness
rss-penger-polser-og-politikk
frokostshowet-pa-p5
bt-dokumentar-2
ta-dokumentar
rss-borsmorgen-okonominyhetene