The On-Call Survival Guide for Humans
Laugh first, learn second.
Nobody ever said "great retro" about a 3 a.m. incident. But they should. Here is how teams make on-call less awful.
The humane version
- Write the runbook before the incident, not during.
- Practise failure in a safe place so the real one feels familiar.
- Rest is part of the process, not a reward for finishing it.
Rule one: the person on call is a person
Sleep-deprived engineers make worse decisions. Rotate fairly, give time back after a rough night, and never treat a pager as a personality trait.
Take a break. The bug will still be there.
Rule two: rehearse the bad day
Teams handle real incidents better when they have seen fake ones. Use a browser interceptor to force a 500 or a timeout on a staging build and let the on-call person walk the runbook. Nobody is harmed, and the runbook gets fixed where it was wrong.
Rule three: kind feedback, blameless retros
- Describe what happened, not who did it.
- Ask "what made this easy to get wrong?"
- End every retro with one change, one owner, one date.
- Say thank you to the person who got paged.
Hiring juniors well and supporting on-call are the same skill: making it safe to ask a question.
Read next
Intercept your first request in under a minute
Create a free ProxyCeptor account to mock, delay, block and rewrite API traffic, then share the same rules with your team.