A production incident arrives without warning and immediately tests everything you have built as a manager. The engineers are under pressure, stakeholders are asking questions, and the temptation to dive in and help is strong. But your job during an incident is not to fix the problem. It is to create the conditions in which your team can fix it well, fast, and without burning out. The managers who handle incidents best are not always the most technically capable. They are the ones who stay calm, manage the noise, and let the engineers do what they are best at.
Your role during an incident is not to fix the problem. It is to handle everything else so the engineers can fix it without distraction.
Your role and what it is not
The technical lead on the incident owns the fix. Your role is everything else: keeping stakeholders informed, protecting the team's focus, making decisions when escalation is needed, and ensuring the response is coordinated and visible. These are not glamorous jobs. But without them, incidents drag on longer, communication breaks down, and the team emerges demoralised even when they resolve the problem.
- Stay out of the fixYour presence in the debugging conversation adds noise, not signal. Trust the engineers to own the technical response. You can monitor the incident channel without contributing to it. If you have technical context that changes the diagnosis, share it briefly and step back.
- Own communicationsWrite the status updates. Field questions from senior leadership. Shield the engineers from the constant interruption of stakeholders wanting live updates. Every question you handle is time the engineers can spend on the fix.
- Clear decisions above the teamIf the incident requires a call about rollback, public disclosure, or engaging a third-party vendor, that decision is yours to own and make quickly. Waiting for a decision is one of the most common reasons incidents run longer than they need to.
- Track the timelineNote when the incident started, what was tried, and when changes were made. This information is invaluable for the post-mortem and for accurate stakeholder communication. It is hard to reconstruct accurately after the fact.
The first fifteen minutes
The first fifteen minutes often determine how well an incident goes. Clear action under pressure is hard, but having a mental model of what to do first makes it possible. Speed matters, but accuracy matters more - a rushed, incorrect first update creates confusion that takes longer to undo than a slightly delayed accurate one. Take two minutes to understand what is actually happening before you communicate it externally.
First fifteen minutes checklist
An acknowledgement that says "We are aware of an issue affecting checkout. Engineers are investigating. Next update in fifteen minutes" is enough. It gives stakeholders a point of contact and commits you to a cadence without over-promising on cause or timeline.
Communicating with stakeholders
Stakeholder communication during an incident is one of the most draining and most important parts of your role. Done well, it keeps trust intact even when things go wrong. Done badly, it turns a technical incident into a leadership crisis. The goal is not to have all the answers. It is to show that the situation is being managed and that people will hear from you regularly.
- Set a cadenceCommit to a regular update interval, typically every fifteen or thirty minutes, and stick to it even when you have no new information. The uncertainty of silence is worse than an honest "still investigating." Once the acute phase passes, you can reduce the frequency.
- Share facts, not theoriesAvoid speculating about root cause or recovery time unless you are confident. "We believe the issue is in the payment service" is dangerous if it turns out to be wrong. Tell stakeholders what you know and what you are doing, not what you think might be the case.
- Know who needs whatA VP needs a business impact summary. An on-call engineer needs technical context. A customer success team needs language they can use with affected clients. Write different updates for different audiences rather than one long message that serves nobody well.
- Separate cause from resolutionIn the first update, focus on what is affected and what you are doing to fix it. Root cause analysis belongs in the post-incident review, not the live incident thread. Forcing a premature explanation under pressure leads to incorrect or misleading information.
Supporting the team under pressure
Engineers solving production incidents are under significant stress. The human side of your role matters as much as the coordination. How you manage the team during an incident shapes how they feel about responding to the next one. A well-managed incident builds psychological safety. A poorly managed one creates dread around being on-call.
- Control who is involvedExtra people in the incident channel slow things down. Limit active involvement to engineers who are directly contributing. Observers, the curious, and managers who want to help without a clear role should wait for the post-incident summary. A focused channel resolves incidents faster.
- Rotate for long incidentsIf the incident extends beyond two or three hours, bring in fresh engineers and let tired ones step back. Exhausted engineers make mistakes and miss things. A handover costs twenty minutes and is almost always worth it. Protect your strongest engineers for the most critical phases.
- Acknowledge the pressureA brief "I know this is tough, you are doing well" during a long incident matters more than most managers expect. It does not need to be elaborate. People working under stress need to know their effort is seen and that the difficulty is acknowledged, not just expected.
- Protect their focusEvery question you route through yourself instead of directly to an engineer is time they can spend on the fix. Batch non-urgent queries. Push back on stakeholders who want live briefings from the technical team during the incident. Your job is to be the buffer.
After the incident resolves
The moment an incident is resolved is not the moment to relax. There is a short window in which you can support the team well and set up the learning process before the adrenaline fades and memory becomes less reliable. Schedule the post-mortem within forty-eight hours while the detail is fresh. What you do in the first hour after resolution shapes whether the incident produces lasting improvement or just relief.
- Communicate resolutionSend a final update to all stakeholders confirming the issue is resolved, what the customer impact was, and that a full post-incident review will follow with root cause and next steps. This closes the loop and prevents people continuing to wonder if things are actually fixed.
- Let the team breatheDo not immediately pile in with questions or analysis. Give the engineers who fixed the incident fifteen minutes to decompress before any debrief conversation starts. The instinct to debrief immediately is understandable but rarely produces the best outcomes.
- Schedule the post-mortemThe post-mortem should happen within forty-eight hours while memory is fresh. Delaying it risks losing the detail and sends the message that learning is optional. Blameless retrospective practice applies directly here - the goal is improving the system, not assigning fault.
- Capture actions nowThings that will prevent recurrence or reduce impact in future should be captured before the team moves on. The first half hour after resolution is when the team knows most clearly what needs to change. An hour later, other priorities will start to compete for attention.
Building a team that handles incidents well
How a team handles incidents is a lagging indicator of how healthy it is day to day. The patterns that make incidents go badly - unclear ownership, poor communication, blame - exist in how the team operates ordinarily. Improving incident response is not just about process. It is about the culture you build through your regular management habits.
- Practise before it is realTeams that review runbooks, test alerting, and practise handovers handle real incidents better. You do not need a formal chaos engineering programme. A thirty-minute session reviewing your monitoring setup or on-call runbook is more useful than most teams give it credit for.
- Make post-mortems blamelessIf people dread post-mortems or feel judged, the process is not working. Approach every post-mortem as an investigation of the system, not of the individuals involved. When people feel safe to be honest, you get far better information about what actually needs to change.
- Rotate the on-call load fairlyChronic on-call burden on the same engineers creates burnout and knowledge silos. If the same two people fix every incident, that is a risk to the team and to those individuals. Broad rotation builds resilience and surfaces systemic issues faster because more engineers encounter them.
- Acknowledge good incident responseWhen your team handles an incident well, say so publicly. In your next all-hands, in a message to leadership, in the next retrospective. Incident response is hard and often invisible. Making it visible and valued is what keeps people from quietly resenting the on-call rotation.
Frequently asked questions
Stay across what matters after the incident
Capture follow-up actions, connect them to context, and make sure nothing falls through once the adrenaline fades.
