On-call rotations are often the invisible burden in engineering teams. When managed well, they build resilience and shared ownership. When managed poorly, they drive burnout and attrition faster than almost any other management failure. The difference is rarely the technical setup. It is whether the manager treats on-call as a system worth actively managing, not just a rota to fill.
A rotation that nobody respects is not a safety net for your system. It is a slow drain on your team. The manager's job is to make on-call fair, well-supported, and genuinely worth showing up for.
Setting up the rotation fairly
The most common mistake managers make with on-call is treating it as a scheduling problem rather than a people problem. A rota that looks balanced on a spreadsheet can still feel deeply unfair if some weeks are consistently worse than others, if certain engineers always get the difficult incidents, or if recovery time is not protected. Fairness is perceived as much as it is measured.
The minimum viable on-call rotation has four components: a clear schedule shared well in advance, named primary and secondary engineers for each shift, a defined escalation path when the primary cannot respond, and protected time for the engineer coming off call. Without these, incidents route to whoever is most visible rather than whoever is on duty, and on-call fatigue accumulates invisibly until someone burns out.
- Minimum team sizeYou need at least four engineers to run a sustainable weekly rotation without someone being on call more than once a month. Below that, consider pooling with another team or adjusting shift length.
- Primary and secondaryAlways name a secondary engineer for each shift. The secondary is the failsafe when the primary is unreachable or overwhelmed. Without this, an unavailable primary means an unanswered incident.
- Recovery timeAn engineer who handles a major incident at 3am should not be expected to work a full day. Protect at least half a day of recovery time after any disrupted night. This is a policy, not a courtesy.
- Handover ritualA short handover note at the start of each shift sets the incoming engineer up with context. What is currently fragile? What is being monitored? What changed in the last 48 hours?
What your on-call runbook needs to include
An on-call runbook is the single most valuable thing you can give an engineer going on call for the first time. Without it, they are guessing what to do when something breaks at midnight. With it, they have a clear path from alert to action. The runbook does not need to be exhaustive, but it does need to answer the questions an anxious engineer will have at 2am.
Runbook checklist
Alert definitions
What each alert means, its severity, and the expected response time.
First-response steps
The immediate actions to take for the most common incident types.
Escalation contacts
Who to call if you cannot resolve it, and how to reach them out of hours.
Communication template
A standard format for updating stakeholders during an active incident.
Post-incident process
What to file, when to file it, and who needs to be informed once resolved.
Review the runbook after every incident and update anything that turned out to be wrong or missing.
The runbook is a living document. One of the best signals that on-call is working well is when the runbook improves after every incident, because engineers are adding what they learnt. If your runbook has not changed in six months, it is either perfect or nobody is reading it.
Supporting your team during on-call
The support an engineer receives during their on-call week signals more about your team culture than almost anything else. An engineer who gets no acknowledgement of a difficult week, no compensation for disrupted sleep, and no help reducing the noise will not stay on the rotation. They will find a team where on-call is taken seriously, or they will leave.
Check in at the start and middle of each on-call shift. A two-minute catchup asking how it is going so far is not micromanagement. It is how you catch a rough week before it becomes a crisis. If someone is being paged five times a night, that is urgent management information, not just an engineering problem. Log it, treat it as a Target to reduce noise, and make progress visible.
- CompensationBe explicit about how on-call is compensated, whether that is additional pay, time off in lieu, or a reduced workload the following week. Ambiguity breeds resentment.
- Noise reductionAlert fatigue is a management problem, not just a technical one. If your team is being paged for things that do not require immediate action, fix the alerts. Every unnecessary page erodes trust in the rotation.
- Toil trackingAsk engineers to log the toil they handle during on-call: the manual tasks, repetitive fixes, and chores that do not improve the system. Visible toil gets prioritised. Hidden toil accumulates.
- Check-insA brief catchup at the start and midpoint of each on-call shift lets you catch problems early. Do not wait for an engineer to tell you it was a bad week. Ask.
Using incidents to improve the system
Every incident is a signal. Whether you treat it as noise or information is a management choice. Teams that run blameless post-mortems after meaningful incidents, review findings in the next retrospective, and track improvement actions as first-class work consistently reduce their incident rate over time. Teams that do not tend to fight the same fires indefinitely.
You do not need a post-mortem for every page. Save the full process for incidents that meet a threshold: significant customer impact, more than thirty minutes to resolve, or a repeat occurrence of the same failure. For smaller events, a short note in your incident log is enough. What matters is that the pattern is visible. See our guide on how to run a post-mortem for a full walkthrough.
After each incident
In the next retro
In Manager Toolkit, actions from post-mortems connect directly to the relevant retrospective or meeting note, so the thread of what happened, what was decided, and what changed is always traceable. When the same issue surfaces six months later, you can see exactly what was tried before.
Common on-call failure modes and how to fix them
Most on-call problems fall into a small number of recurring patterns. Recognising them early is the fastest path to a healthier rotation. The patterns below are not engineering problems. They are management problems, and they require management solutions.
- One person carries itWhen incidents always route to the same engineer, it means the runbook is missing, the rotation is not trusted, or knowledge is too concentrated. Fix the runbook, invest in spreading expertise, and enforce the rotation even when it is inconvenient.
- Alert noise is too highIf engineers are paged for things that do not require immediate action, the alerts are wrong. Work with the team to categorise alerts by urgency and remove anything that is not genuinely actionable. A noisy pager makes the real signal impossible to hear.
- No time to fix toilOn-call toil stays invisible because it is never prioritised over feature work. Block dedicated time each sprint for reliability improvements. Even two days per sprint compounds significantly over a quarter.
- The rotation feels punishingIf nobody wants to be on call, ask why. The answer is usually one of four things: the alerts are too noisy, the runbook is inadequate, recovery time is not protected, or the compensation does not reflect the burden. Fix the cause, not the symptom.
Frequently asked questions
Make on-call work for your team
Track incident actions, run retrospectives, and reduce toil with tools built for engineering managers.
