Site Reliability Engineering

Betsy Beyer, Chris Jones, Jennifer Petoff and Niall Richard Murphy

About this book

Google's account of how it runs production systems, written by the engineers who do it. The central ideas are the service level objective as an explicit reliability target and the error budget that follows from it, which turns arguments about release pace into an arithmetic question anyone can check. Chapters cover monitoring and alerting philosophy, on call practice, incident response and postmortems, release engineering, and the effort to keep operational toil below a fixed share of an engineer's time. Later sections deal with distributed systems concerns including load balancing, cascading failures and data integrity. Worth reading for the error budget alone, which gives product and engineering a shared way to trade reliability against speed.

Description via Product Digest.

Topics

Related books