How Claude Mythos found a 15-year-old bug in Mozilla Firefox | Brian Grinstead
How I AI22 Kesä

How Claude Mythos found a 15-year-old bug in Mozilla Firefox | Brian Grinstead

Brian Grinstead is a distinguished engineer at Mozilla, where he’s worked on Firefox and the web platform since 2013 (he joined to help launch Firefox DevTools). Recently he and his team pointed an agentic bug-finding pipeline at Firefox—a codebase with tens of thousands of files and tens of millions of lines of code—and shipped a record month of security fixes. The viral chart everyone saw gave the credit to Anthropic’s new Mythos model. Brian’s take is that the harness and pipeline did just as much of the work, and he walks through exactly how it runs and how anyone can build a starter version.


What you’ll learn:

  1. How to build a basic bug-finding harness by running Claude Code or Codex with one prompt and the -p flag, no SDK required
  2. Why pointing an agent at a whole codebase fails, and how an LLM judge can score and rank files before you spend any compute
  3. How a verifier subagent kills false positives by catching the agent when it cheats
  4. The goal-loop pattern: give an agent a tightly scoped problem, a clear pass/fail signal, and let it retry far past the point a human would quit
  5. Why teams that already invested in fuzzing, CI, and dev tooling are so far ahead
  6. How to weigh model versus harness, and why Brian splits the credit close to 50-50
  7. How a non-engineer can reuse the same score, verify, and fix the loop for design quality, conversion rate, or tech debt
  8. Why AI-generated patches still can’t ship on their own, and where humans stay in the loop

Brought to you by:

WorkOS—Make your app enterprise-ready today

Metaview—The agentic recruiting platform for winning teams

In this episode, we cover:

(00:00) Introduction to Brian Grinstead

(02:43) The viral chart: Firefox Security Bug Fixes by Month

(05:32) How the custom harness works

(10:22) Goal loops and guardrails

(14:45) How they built it

(16:55) Real bugs, including a 15-year-old one

(23:00) Open-sourcing it

(26:26) Why humans still review every fix

(32:30) Live demo and prioritizing files

(40:18) Mobilizing the team and recap

(42:33) Lightning round

Tools referenced:

• Claude Code: https://claude.ai/code

• Claude Agent SDK: https://code.claude.com/docs/en/agent-sdk/overview

• Codex: https://openai.com/index/openai-codex/

• OpenAI Agent SDK: https://developers.openai.com/api/docs/guides/agents

• VS Code: https://code.visualstudio.com/

• Docker: https://www.docker.com/

• Firefox: https://www.mozilla.org/firefox/

• Address Sanitizer: https://github.com/google/sanitizers

• RLBox: https://rlbox.dev/

Other references:

• Mozilla Bug Bounty Program: https://www.mozilla.org/security/bug-bounty/

• Mozilla GitHub: https://github.com/mozilla

Where to find Brian Grinstead:

LinkedIn: https://www.linkedin.com/in/bgrins/

GitHub: https://github.com/bgrins

Where to find Claire Vo:

ChatPRD: https://www.chatprd.ai/

Website: https://clairevo.com/

LinkedIn: https://www.linkedin.com/in/clairevo/

X: https://x.com/clairevo

Production and marketing by https://penname.co/. For inquiries about sponsoring the podcast, email jordan@penname.co.

Tämä jakso on lisätty Podme-palveluun avoimen RSS-syötteen kautta eikä se ole Podmen omaa tuotantoa. Siksi jakso saattaa sisältää mainontaa.

Jaksot(103)

I spent $20,000 on Devin in a month. Here’s what I learned | Ryan Carson (solo founder)

I spent $20,000 on Devin in a month. Here’s what I learned | Ryan Carson (solo founder)

Ryan Carson is a five-time founder and the current solo founder of Untangle, a B2B SaaS platform for family law firms. Before Untangle, he co-founded Treehouse, an online coding education platform, an...

24 Elo 44min

I tested Grok Bot, Grok 4.6, and Cursor Origin - here’s my honest take

I tested Grok Bot, Grok 4.6, and Cursor Origin - here’s my honest take

This week I’m doing a solo breakdown of everything xAI and Cursor have shipped recently, including Grok Bot, Cursor Origin, and the Grok 4.6 model. I set up five Grok Bots, ran Grok 4.6 through my Cla...

18 Elo 27min

How a solo founder used Codex and ChatGPT to launch a fashion brand without engineers | Yana Welinder

How a solo founder used Codex and ChatGPT to launch a fashion brand without engineers | Yana Welinder

Yana Welinder is the solo founder of Yana Bana, an AI-native fashion brand built with AI as her technical co-founder, starting from hand-drawn sketches and ending with runway photos, CAD files for 3D ...

17 Elo 32min

Claude Code for normal people: skills, voice mode, and how to collaborate with AI

Claude Code for normal people: skills, voice mode, and how to collaborate with AI

Grace Clarke is an AI educator and former marketing consultant who taught herself Claude Code earlier this year and built a curriculum out of the process. She now runs her entire service business on t...

10 Elo 43min

Build an AI code review bot in 30 minutes with Vercel Eve

Build an AI code review bot in 30 minutes with Vercel Eve

AI writes most of my code now, and that created a new problem: a PR queue I couldn’t keep up with. In this episode, I walk through how I built Merge Mommy, a Vercel Eve agent that reads every PR after...

5 Elo 24min

ChatGPT Codex Voice + browser + Sites: an expert’s AI workflow | Nick Baumann (OpenAI)

ChatGPT Codex Voice + browser + Sites: an expert’s AI workflow | Nick Baumann (OpenAI)

Nick Baumann is on the Developer Experience team at OpenAI, where he spends his days building with, testing, and communicating the capabilities of ChatGPT Codex and ChatGPT Work. In this episode, Nick...

3 Elo 41min

From zero coding background to hardware hacker: How Cursor + a Raspberry Pi makes AI fun

From zero coding background to hardware hacker: How Cursor + a Raspberry Pi makes AI fun

Maddie Reese is a vibe coder, hardware tinkerer, and builder. She builds things at the intersection of software and hardware, including a thermal receipt printer that people around the world can messa...

27 Heinä 28min

Claude Opus 5 review: this model is brilliant (but annoying)

Claude Opus 5 review: this model is brilliant (but annoying)

I’m tired of new models. Every week there’s a new benchmark, a new frontier intelligence claim, a new thing to test. But here we are, because Opus 5 just dropped and I’ve had real hands-on time with i...

24 Heinä 24min