Skip to content

Infrastructure Management

Monitoring tells you something broke. A runbook tells you what to do

Monitoring is easy to buy and satisfying to install. Within a week you have dashboards, and every service is green. What is usually missing is the other half: when one of those turns red at two in the morning, what does the person on call actually do?

Why the gap exists

Monitoring is a product; response is a habit. You can procure the first. The second has to be written down by the people who know, at a time when nothing is broken, which is exactly when it feels least urgent.

So the knowledge stays in three or four heads, and the incident process becomes: alert fires, first responder recognizes it or does not, and if not, someone senior gets woken up. That works until the people in whose heads it lives are unavailable, on holiday, or gone.

What a runbook actually needs

A runbook that gets used at three in the morning is short, specific and ordered. It answers, in this sequence:

  1. What this alert means in one sentence, in terms of impact. “The queue consumer has stopped, so orders are not being processed” — not “consumer lag exceeded threshold”.
  2. How to confirm it is real. The one check that distinguishes a genuine fault from a monitoring blip. This prevents the most common wasted hour.
  3. Who is affected and how badly, so the responder can judge whether to escalate now or continue alone.
  4. The first thing to try, with the exact command or the exact screen. Not “restart the service” but the command, on the right host, with the right flags.
  5. What to do if that does not work — the next two or three steps, then a stop point.
  6. Who to escalate to, by role and by actual contact method, and at what point.
  7. What to check afterwards to confirm recovery, since a service that starts is not the same as a service that is working.

Things that make runbooks useless

  • Written for someone who already knows.“Clear the backlog as usual” is a note to self, not a runbook.
  • Out of date.A runbook naming a decommissioned server destroys trust in every other runbook.
  • Not linked from the alert.If finding the runbook means searching a wiki at two in the morning, it will not be found. The link belongs in the alert payload.
  • Written as prose.Nobody reads paragraphs during an incident. Numbered steps, commands in a form that can be copied.
  • Covering everything.A forty-page document for one alert is a document nobody opens.

Where to start when there are none

Do not try to write a runbook for every alert. Start from what actually happens:

  1. List the alerts that fired in the last three months, sorted by frequency.
  2. Take the top five. They are the majority of the volume.
  3. Write the runbook for each one the next time it fires, by having the responder record what they did as they did it.
  4. Have someone else use it at the next occurrence, and fix what was unclear.

The recording-as-you-go step is what makes this achievable. Writing a runbook from memory produces the version where the responder already knows; writing it during the incident produces the version that works.

Then use the runbook to remove the alert

The best outcome for a frequently-used runbook is that it stops being needed. If the response to an alert is the same three commands every time, that is a candidate for automation — either the fix, or the condition that causes it.

Reviewing the most-used runbooks quarterly and asking “why does this still happen?” is how the alert volume comes down. A team whose runbook collection keeps growing without anything ever leaving it is documenting its problems rather than solving them.

Keep them with the system

Runbooks in version control, beside the code or the infrastructure definitions, change when the system changes and get reviewed when the change is reviewed. Runbooks in a separate wiki drift, because nothing connects them to the thing they describe.

Our infrastructure managed services treat the runbook as part of taking on a system rather than something written later — an alert we monitor without a documented response is an alert we have not finished onboarding.

What can we help you achieve?

We empower your vision with innovative and effective strategies.

Aura
AI Agent

Hi there 👋

AI-powered assistant for services, careers & support

👋

Quick intro

So we can assist you better

Please enter your name
Please enter a valid email
Please enter a valid phone number
Your data is secure
Powered by AcmaCorp Solutions