The Site Reliability Workbook
About this book
The companion volume to Site Reliability Engineering, written to answer the question of how to put the ideas into practice. Where the first book describes how Google runs production, this one gives worked examples of service level objectives, error budgets, monitoring, incident response and postmortem culture, with case studies from companies outside Google. Several chapters walk through the arithmetic of setting an objective and deciding what an error budget policy should trigger. There is also material on toil, on simplicity, and on the organisational side of introducing site reliability engineering where an operations team already exists. Read it alongside the first book rather than instead of it.
Description via Product Digest.