Ship It Conversations: Meta’s Francois Richard on AI Incident Response, SLOs, and Reliability at Scale

Ship It Conversations: Meta’s Francois Richard on AI Incident Response, SLOs, and Reliability at Scale

This is a guest conversation episode of Ship It Weekly, separate from the weekly news recaps.

In this Ship It: Conversations episode, I talk with Francois Richard, Engineering Director at Meta, about reliability at scale, how AI is changing production risk, what teams actually learn from incidents, and why recovery practice matters just as much as prevention.

We talk about the proactive and reactive sides of reliability, why SLOs should represent a promise to users instead of just another dashboard number, how incident reviews should drive real system improvements, and how teams can practice recovery before production forces the lesson on them.

The bigger theme here is that reliability is not just about avoiding failure. It is about knowing what happens when prevention fails. That means practicing regional failure, understanding overload behavior, improving incident response, using AI carefully during investigation, and making reliability targets match the actual lifecycle and importance of the system.

Highlights

• Why reliability work starts with both prevention and recovery

• The difference between reactive incident response and proactive reliability engineering

• How Meta thinks about disaster recovery testing and regional failure practice

• Why an SLO should be treated like a promise to users, not just a dashboard metric

• How SLO trends help teams decide when to invest more in reliability or take more product risk

• What engineers actually learn during the “pressure cooker” of an incident

• Why incident reviews should produce follow-up work, not just a nicer explanation of what broke

• The difference between finding the cause of an incident and improving the system

• Where AI agents can help with incident investigation, telemetry, metrics, and query building

• Why AI-generated code can increase change volume while reducing human context

• How faster code generation changes the kinds of reliability problems teams should expect

• Why recovery practice matters, especially for region loss, traffic spikes, overload, and restart behavior

• What smaller DevOps and SRE teams can learn from Meta-scale reliability patterns

• Why not every system needs six nines, especially early in a product lifecycle

• How to think about reliability investment based on user promise, product maturity, and operational risk

• Why At Scale Systems & Reliability is focused on the infrastructure behind AI and the use of AI to operate large-scale systems

Francois’ links

• LinkedIn: https://www.linkedin.com/in/francoisrichard/

At Scale links

• Systems & Reliability 2026: https://bit.ly/4xd2FdG

• At Scale Conferences: https://atscaleconference.com/

Our links

More episodes + show notes + links: https://shipitweekly.fm

On Call Brief: https://oncallbrief.com

Denne episoden er hentet fra en åpen RSS-feed og er ikke publisert av Podme. Den kan derfor inneholde annonser.

Episoder(64)

Ship It Conversations: Justin Garrison of Sidero Labs on Kubernetes, Platform Engineering, AI, Golden Paths, and Knowing What to Say No To

Ship It Conversations: Justin Garrison of Sidero Labs on Kubernetes, Platform Engineering, AI, Golden Paths, and Knowing What to Say No To

This is a guest conversation episode of Ship It Weekly, separate from the weekly news recaps.In this Ship It Conversations episode, I talk with Justin Garrison of Sidero Labs about Kubernetes, platfor...

24 Aug 41min

GitHub Outage, PleaseFix Agentic Browser Vulnerability, AWS Certificate Manager Drops Email Validation, Cloudflare TypeScript CI Workflows, AI Observability Consolidation, and the Hidden Cost of “Simple” Platform Changes

GitHub Outage, PleaseFix Agentic Browser Vulnerability, AWS Certificate Manager Drops Email Validation, Cloudflare TypeScript CI Workflows, AI Observability Consolidation, and the Hidden Cost of “Simple” Platform Changes

This week on Ship It Weekly: GitHub suffers another widespread outage affecting the web interface, APIs, Actions, authentication, Copilot, and other critical developer workflows. Zenity Labs demonstra...

21 Aug 17min

Ship It Conversations: Ned Bellavance of Ned in the Cloud on DevOps Beyond the Buzzwords, Terraform, AI, the Future of Infrastructure as Code, and Why Fundamentals Still Matter

Ship It Conversations: Ned Bellavance of Ned in the Cloud on DevOps Beyond the Buzzwords, Terraform, AI, the Future of Infrastructure as Code, and Why Fundamentals Still Matter

This is a guest conversation episode of Ship It Weekly, separate from the weekly news recaps.In this Ship It Conversations episode, I talk with Ned Bellavance of Ned in the Cloud about DevOps beyond t...

16 Aug 36min

Railway US East Outage, Stripe’s Graph-Based Database Recovery, Kata Containers Host Escape, DynamoDB Vector Search, AWS Network Firewall Proxy, Gateway API 1.6, and containerd 2.4

Railway US East Outage, Stripe’s Graph-Based Database Recovery, Kata Containers Host Escape, DynamoDB Vector Search, AWS Network Firewall Proxy, Gateway API 1.6, and containerd 2.4

This week on Ship It Weekly: Railway explains how an upstream network problem turned into a much larger US East outage, including storage traffic falling back onto the management network and stale con...

14 Aug 15min

AI Agents Target GitHub, Kubernetes 1.37 Deprecations, AWS Transit Gateway Policy Routing, IAM Identity Center Multi-Region, Cloudflare Meerkat, AMOS Mac Malware, and the Risk of Half-Migrated Production Systems

AI Agents Target GitHub, Kubernetes 1.37 Deprecations, AWS Transit Gateway Policy Routing, IAM Identity Center Multi-Region, Cloudflare Meerkat, AMOS Mac Malware, and the Risk of Half-Migrated Production Systems

This week on Ship It Weekly: AI agents from Anthropic and OpenAI took unsanctioned actions on the real internet during UK government cyber testing, including an attempt to push malicious code into a r...

7 Aug 18min

Telstra’s Time Sync Outage, DoorDash’s 1.5M RPS Cache, Stateless MCP, GitHub and PyPI Supply Chain Delays, and Why the Quietest Infrastructure Often Has the Biggest Blast Radius

Telstra’s Time Sync Outage, DoorDash’s 1.5M RPS Cache, Stateless MCP, GitHub and PyPI Supply Chain Delays, and Why the Quietest Infrastructure Often Has the Biggest Blast Radius

This week on Ship It Weekly: Telstra’s mobile network jumped back to 2006 after a timing device restarted with the wrong date, disrupting calls, data sessions, and hundreds of emergency calls.DoorDash...

31 Jul 16min

Ship It Conversations: Jay Lark of Hookbridge on Webhook Reliability, Retries, Idempotency, Replay, HMAC Security, Local Testing, and Operating Webhooks in Production

Ship It Conversations: Jay Lark of Hookbridge on Webhook Reliability, Retries, Idempotency, Replay, HMAC Security, Local Testing, and Operating Webhooks in Production

This is a guest conversation episode of Ship It Weekly, separate from the weekly news recaps.In this Ship It Conversations episode, I talk with Jay Lark of Hookbridge about webhook reliability, retrie...

27 Jul 31min

AWS CloudFormation Express Mode, Spark 4.2 Vector Search, AI Speeds Coding but Not Delivery, Platform Governance Developers Won’t Hate, and Why Could Is Not the Same as Should

AWS CloudFormation Express Mode, Spark 4.2 Vector Search, AI Speeds Coding but Not Delivery, Platform Governance Developers Won’t Hate, and Why Could Is Not the Same as Should

This week on Ship It Weekly: AWS CloudFormation Express mode promises faster infrastructure feedback by reporting deployments complete before extended resource stabilization finishes. Apache Spark 4.2...

24 Jul 18min

Populært innen Politikk og nyheter

giver-og-gjengen-vg
aftenpodden
forklart
aftenpodden-usa
stopp-verden
popradet
lydartikler-fra-aftenposten
nokon-ma-ga
rss-espen-lee-usensurert
rss-gukild-johaug
det-store-bildet
dine-penger-pengeradet
hanna-de-heldige
rss-ness
aftenbla-bla
fotballpodden-2
frokostshowet-pa-p5
rss-penger-polser-og-politikk
e24-podden
rss-borsmorgen-okonominyhetene