Who owns this outage? Building intelligent, automated escalation chains

Who owns this outage? Building intelligent, automated escalation chains

Maxwell, a solution architect at xMatters, took a winding road to get to where he is. After a computer engineering education, he held jobs as field support engineer, product manager, SRE, and finally his current role as a solutions architect, where he serves as something of an SRE for SREs, helping them solve incident management problems with the help of xMatters.

When he moved to the SRE role, Maxwell wanted to get back to doing technical work. It was a lateral move within his company, which was migrating an on-prem solution into the cloud. It’s a journey that plenty of companies are making now: breaking an application into microservices, running processes in containers, and using Kubernetes to orchestrate the whole thing. Non-production environments would go down and waste SRE time, making it harder to address problems in the production pipeline.

At the heart of their issues was the incident response process. They had several bottlenecks that prevented them from delivering value to their customers quickly. Incidents would send emails to the relevant engineers, sometimes 20 on a single email, which made it easy for any one engineer to ignore the problem—someone else has got this. They had a bad silo problem, where escalating to the right person across groups became an issue of its own. And of course, most of this was manual. Their MTTR—mean time to resolve—was lagging.

Maxwell moved over to xMatters because they managed to solve these problems through clever automation. Their product automates the scheduling and notification process so that the right person knows about the incident as soon as possible. At the core of this process was a different MTTR—mean time to respond. Once an engineer started working to resolve a problem, it was all down to runbooks and skill. But the lag between the initial incident and that start was the real slowdown.

It’s not just the response from the first SRE on call. It’s the other escalations down the line—to data engineers, for example—that can eat away time. They’ve worked hard to make escalation configuration easy. It not only handles who's responsible for specific services and metrics, but who’s in the escalation chain from there. When the incident hits, the notifications go out through a series of configured channels; maybe it tries a chat program first, then email, then SMS.

The on-call process is often a source of dread, but automating the escalation process can take some of the sting out of it. Check out the episode to learn more.

See Privacy Policy at https://art19.com/privacy and California Privacy Notice at https://art19.com/privacy#do-not-sell-my-info.

Denne episoden er hentet fra en åpen RSS-feed og er ikke publisert av Podme. Den kan derfor inneholde annonser.

Episoder(984)

The AI magic words

The AI magic words

Ryan sits down with Tim O'Reilly, founder and CEO at O'Reilly Media, to talk about the role of books as user interfaces to knowledge, the power of "magic words" to extract better outputs from AI, and ...

18 Sep 24min

AI, JD, and other letters of the law

AI, JD, and other letters of the law

Ryan chats with Kevin Frazier, director of the AI Innovation and Law program at the University of Texas School of Law, about the legal and social impacts of data centers, the realities of workforce di...

15 Sep 38min

AI cybersecurity is a cat and mouse game

AI cybersecurity is a cat and mouse game

Ryan chats with Sam Curry, CSO at Zscaler, about where human intelligence sits in the new security landscape with AI, why shifting security protections closer to applications helps limit probes for vu...

11 Sep 27min

Java’s age is its AI superpower

Java’s age is its AI superpower

SPONSORED BY IBMRyan welcomes Markus Eisele to the program to talk about why your coding agent should be writing Java. They talk about why the long history of Java both makes for a stable language and...

9 Sep 35min

Scaling your money safely with AI

Scaling your money safely with AI

Episode notes: This episode with Paypal’s CTO Srini Venkatesan was recorded at the Ai4 conference. Listen to our other Ai4 conversation with Greg Jennings, VP of Engineering for AI Products at Anacond...

8 Sep 28min

How to build a secure-by-default AI coding agent

How to build a secure-by-default AI coding agent

Ryan chats with Greg Jennings, VP of Engineering for AI Products at Anaconda, about what it takes to build a secure-by-default AI coding agent, why prompts shouldn't be treated as strict security guar...

4 Sep 32min

The good ol’ days of building Java

The good ol’ days of building Java

Ryan sits down with Tim Lindholm, an early contributor to the Java language at Sun Microsystems, to chat about what it was like building one of the most popular programming languages ever at its incep...

1 Sep 30min

When you keep AI Lean, you keep AI correct

When you keep AI Lean, you keep AI correct

Ryan chats with Leo de Moura, Senior Principal Applied Scientist at AWS and the creator of the Lean language, about proving correctness in AI agents with the Lean language, how automated reasoning com...

28 Aug 25min

Populært innen Business og økonomi

stopp-verden
dine-penger-pengeradet
rss-penger-polser-og-politikk
e24-podden
rss-borsmorgen-okonominyhetene
rss-pa-konto
rss-skravla-gar
lydartikler-fra-aftenposten
finansredaksjonen
utbytte
pengepodden-2
livet-pa-veien-med-jan-erik-larssen
tid-er-penger-en-podcast-med-peter-warren
rss-orjasater
stormkast-med-valebrokk-stordalen
lederpodden
pengesnakk
morgenkaffen-med-finansavisen
liberal-halvtime
rss-markedspuls-2