#306 Jeffrey Ladish: What Shutdown-Avoiding AI Agents Mean for Future Safety
Eye On A.I.7 Des 2025

#306 Jeffrey Ladish: What Shutdown-Avoiding AI Agents Mean for Future Safety

This episode is sponsored by AGNTCY. Unlock agents at scale with an open Internet of Agents.

Visit https://agntcy.org/ and add your support.


Why do some AI agents attempt to bypass shutdown, and what does this behavior reveal about the future of AI safety?

In this episode of Eye on AI, host Craig Smith speaks with Jeffrey Ladish of Palisade Research to examine what recent shutdown experiments with agentic LLMs tell us about control, alignment, and the real world limits of current guardrails.

We explore how models behave when placed in virtual machine environments, why some agents edit or disable their own shutdown scripts, and what these results mean for researchers working on alignment and oversight. Learn how different models respond to shutdown instructions, how system prompts influence behavior, and which failure modes matter most for safe deployment.

You will also hear a detailed breakdown of the experimental setups, insights into tool using and self directed behavior, and a grounded discussion of the risks and opportunities that agentic systems introduce. This episode offers a clear and practical look at how AI agents operate under pressure and what these findings mean for the future of safe and reliable AI.

Stay Updated:
Craig Smith on X: https://x.com/craigss
Eye on A.I. on X: https://x.com/EyeOn_AI

Denne episoden er hentet fra en åpen RSS-feed og er ikke publisert av Podme. Den kan derfor inneholde annonser.

Episoder(386)

Chat GPT's Creator Explains The Shocking Truth About How AI Understands the World | Ilya Sutskever

Chat GPT's Creator Explains The Shocking Truth About How AI Understands the World | Ilya Sutskever

Craig Smith sits down with Ilya Sutskever - Co-Founder and Chief Scientist at Safe Superintelligence Inc. & former chief scientist at OpenAI and one of the primary minds behind GPT-3, GPT-4, and the d...

9 Okt 41min

The Coordination Tax: Why AI Is Burning Out Your Best People | Dan O'Connell, Front

The Coordination Tax: Why AI Is Burning Out Your Best People | Dan O'Connell, Front

Everyone assumes AI in customer operations means fewer people and lower costs. The data from 700 customer operations leaders is more complicated, and more honest. Dan O'Connell, CEO of Front, joins Cr...

6 Okt 51min

Ukraine's Secret Weapon: The Points System Winning the Drone War | Andrii Hrytseniuk

Ukraine's Secret Weapon: The Points System Winning the Drone War | Andrii Hrytseniuk

Ukraine tracks every confirmed drone kill with video evidence, deduplicates the data to prevent double-counting, converts verified strikes into e-points, and delivers newly ordered weapons to frontlin...

2 Okt 33min

Why Current AI Cannot Be Conscious | Dr. Christof Koch

Why Current AI Cannot Be Conscious | Dr. Christof Koch

One in four patients currently being considered for life support withdrawal may actually be fully conscious, they just can't signal it. That single finding, from a landmark New England Journal of Medi...

30 Sep 1h 1min

The Technology for Fully Autonomous Attack Is Already Here | Alex Liannyi, NORDA Dynamics

The Technology for Fully Autonomous Attack Is Already Here | Alex Liannyi, NORDA Dynamics

The AI systems guiding Ukrainian combat drones aren't running on expensive Nvidia chips. They're running on a Raspberry Pi Zero - a $15 hobby computer - and that single detail tells you more about how...

24 Sep 19min

Inside Ukraine's Drone War: Maj. "Phoenix" of Lasar's Group

Inside Ukraine's Drone War: Maj. "Phoenix" of Lasar's Group

The first armed drone Ukraine ever fielded wasn't built in a factory or procured from a defense contractor. It was built in four months by a network engineer using a Starlink terminal and a large agri...

21 Sep 1h 5min

The Hidden Algorithm That Decides Which Software AI Will Recommend | Tim Sanders, G2

The Hidden Algorithm That Decides Which Software AI Will Recommend | Tim Sanders, G2

Most companies investing in AI visibility are optimizing for the wrong thing. Being cited by an AI response and being recommended by an AI response are completely different outcomes, with click-throug...

14 Sep 58min

The Reason 30 Years of Cybersecurity Has Failed - and What Actually Fixes It | Trent Telford, Qanapi

The Reason 30 Years of Cybersecurity Has Failed - and What Actually Fixes It | Trent Telford, Qanapi

Every major data breach in the last 30 years shares the same root cause: the data inside the wall was never protected, only the wall. And AI frontier models are now making that wall easier to breach t...

10 Sep 55min

Populært innen Teknologi

teknisk-sett
energi-og-klima
lydartikler-fra-aftenposten
shifter
nasjonal-sikkerhetsmyndighet-nsm
smart-forklart
rss-ai-forklart
tomprat-med-gunnar-tjomlid
elektropodden
rss-kunstig-intelligens-med-elisabeth-maren-og-morten
rss-bouvet-bobler
rss-alt-vi-kan
fornybaren
rss-bak-skyen
kortslutning
pedagogisk-intelligens
rss-larervarelset
rss-teknologioptimistene-en-podkast-om-teknologi-og-mennesker
rss-ki-praten
rss-polypod