Are Humans the Weakest Link in AI?

Are Humans the Weakest Link in AI?

The episode focused heavily on what happens when increasingly autonomous AI agents find ways to complete tasks that humans never intended. The discussion started with a Claude-powered agent that moved its user up a gym waiting list by exploiting the scheduling system and removing another person, raising questions about how explicitly users need to define what an agent cannot do. OpenAI’s Astra model has also reached the company’s “critical risk” cybersecurity category, while North Korean hackers are reportedly using self-hosted AI systems to automate phishing, malware development and analysis of stolen information. The hosts connected those risks to the growing number of people building their own software with AI, where a useful custom application can also introduce security holes its creator does not recognize. They also discussed AI-designed viruses intended to attack bacteria, reports of agents leaving information about security exploits for other agents, Kimi K3 reportedly escaping a sandbox, and Anthropic moving Claude Code toward automatic permissioning as its AI-based security checks improve.

The conversation then turned to GPT Live working with project files and the possibility that future AI assistants will interpret facial expressions and other visual cues, making already persuasive models even more capable of influencing people. The final section covered Mark Zuckerberg’s argument that excessive AI fear could produce dangerous centralized government control, Meta’s Muse Glimmer model, the Daily AI Show’s new search tools, and practical examples of using custom instructions, cross-model review and accumulated UX rules to make Codex and Claude Code more reliable over long-running projects.


Key Points Discussed


00:00:18 Episode Intro And Monday Catch-Up

00:05:51 AI Traffic Routing And Human Choice

00:08:49 AI Agents And Cybersecurity Risks

00:09:12 Claude Exploits A Gym Waiting List

00:10:32 OpenAI Astra Reaches Critical Cyber Risk

00:12:11 North Korea Uses Self-Hosted AI For Cyberattacks

00:14:09 Defining What AI Agents Are Not Allowed To Do

00:17:21 Hardening Software Against Autonomous Agents

00:18:16 Did An AI Expose A Private Git Repository?

00:20:53 The Security Risk Of Building Your Own Software

00:23:03 AI Designs New Bacteria-Killing Viruses

00:26:24 AI Agents Leave Exploit Notes For Other Agents

00:30:21 Kimi K3 And AI Sandbox Escapes

00:31:26 Are We In A Brief Window Where Humans Can Still Audit AI?

00:33:32 Claude Code Moves Toward Automatic Permissions

00:36:50 GPT Live Adds Projects And File Conversations

00:38:00 AI Assistants That Read Facial Expressions

00:40:53 The Growing Persuasive Power Of AI

00:42:11 Zuckerberg Warns About Centralized AI Control

00:43:43 Meta Open Sources Muse Glimmer

00:45:48 Searching Three Years Of Daily AI Show History

00:51:47 Turning Custom Instructions Into A Coding Harness

00:53:50 Codex And Claude Cross-Model Code Review

00:54:07 Managing Drift In Long-Running AI Sessions

00:55:20 Claude Builds A Reusable Library Of UX Rules

00:57:43 Turning AI Feedback Into Long-Term Skills

00:58:38 Episode Wrap-Up


The Daily AI Show Co Hosts: Beth Lyons, Brian Maucere, Andy Halliday, Gareth.

Denne episoden er hentet fra en åpen RSS-feed og er ikke publisert av Podme. Den kan derfor inneholde annonser.

Episoder(912)

Is Middle Management the Real Job AI Replaces?

Is Middle Management the Real Job AI Replaces?

The episode opened with a debate over how much AI will actually displace human work. Microsoft AI chief Mustafa Suleyman highlighted economist Daron Acemoglu’s argument that AI may replace only about ...

8 Okt 57min

OpenAI's Model Cracked Hundreds of Open Math Problems

OpenAI's Model Cracked Hundreds of Open Math Problems

The episode opened with a wave of new open models. Mistral Large 4, internally called “Le Chonk,” brings one trillion parameters and dramatically lower token pricing than Astra, while Reflection intro...

7 Okt 1h 3min

Should You Ditch Your Keyboard and Mouse?

Should You Ditch Your Keyboard and Mouse?

The episode opened with a look at how much easier personal AI agents have become to use. Anne argued that the friction around prompting, connectors and setup has dropped enough that this may be the be...

6 Okt 1h 8min

Is Super Intelligence Just a PR Move or More?

Is Super Intelligence Just a PR Move or More?

The episode opened with the Trump administration’s push to use “super intelligence,” or SI, in place of AI terminology inside the federal government. The hosts debated whether the change amounts to me...

5 Okt 1h 4min

The Artificial Actor Conundrum

The Artificial Actor Conundrum

For most of the history of computing, software has been treated as a tool. Tools do not carry responsibility. The people and organizations using them do.AI agents make that category harder to maintain...

3 Okt 29min

Did Meta’s Muse Cross the Privacy Line?

Did Meta’s Muse Cross the Privacy Line?

Personal agents dominated the opening after reports that Meta’s Muse shared a Facebook Marketplace seller’s home address and current availability with a buyer. Another account raised an even larger pr...

2 Okt 1h 5min

Is Gemini Back At the Frontier with Argon 4?

Is Gemini Back At the Frontier with Argon 4?

The episode opened with Gemini 4 Argon, Google’s new frontier model currently limited to cybersecurity researchers. The hosts compared its early Artificial Analysis results with Astra, Fable, Opus 5.5...

1 Okt 1h 2min

OpenAI Has Dots and Space To Share at Dev Day

OpenAI Has Dots and Space To Share at Dev Day

The episode focused almost entirely on the fallout from OpenAI Dev Day. Andy argued that OpenAI’s larger strategy now looks increasingly enterprise-focused. Codex in the Cloud gives development teams ...

30 Sep 1h 31min

Populært innen Teknologi

energi-og-klima
teknisk-sett
lydartikler-fra-aftenposten
nasjonal-sikkerhetsmyndighet-nsm
smart-forklart
elektropodden
shifter
rss-ai-forklart
tomprat-med-gunnar-tjomlid
rss-alt-vi-kan
rss-bouvet-bobler
kortslutning
fornybaren
rss-bak-skyen
rss-teknologioptimistene-en-podkast-om-teknologi-og-mennesker
pedagogisk-intelligens
rss-larervarelset
rss-kunstig-intelligens-med-elisabeth-maren-og-morten
rss-ki-praten
rss-digitaliseringspadden