Incident Management
Incident management is the structured coordination of people, technical work, and communication during a service disruption. It helps a team reduce user impact, restore service, and learn from the event.
itPlatform engineering and SRE | OpenSkills.info
Course pathWalk it in order
Look it upDip in anytime
Go furtherLeaves this page
Don't Panic
Don't Panic - Incident Management
Incident management is how a team coordinates an urgent response to a service disruption. Technical diagnosis matters, but it is only one part of the work. Someone must keep the response organized, communicate what is known, and prevent uncontrolled changes from making the situation worse.
Declare when impact is visible, more than one team is needed, or focused investigation has not resolved the issue. Declare early. A small, quickly closed incident is cheaper than trying to build a response structure while impact grows. The immediate objective is mitigation: stop or reduce user impact and restore an acceptable service state. Root-cause analysis follows stabilization unless the cause is already clear.
A useful response framework separates command, operations, communication, and planning. The incident commander owns overall state, priorities, roles, and handoffs. The operations lead directs technical investigation and approved production changes. The communications lead gives responders and stakeholders regular, accurate updates. Planning tracks follow-up work, staffing, and changing system state. In a small incident one person can hold several roles. Split them as load grows so the person changing the system is not also answering every stakeholder question.
Keep one shared incident record with current impact, timeline, hypotheses, decisions, actions, owners, and the next update time near the top. Use a collaboration system that remains available when the affected service fails. Only the designated operations group should make production changes. Everyone else contributes through the agreed channel.
The response loop is declare, assign command, reduce harm, observe, update the record, and reassess. Hand off explicitly across shifts. After service is stable, close deliberately, preserve the record, and schedule the review.
A framework does not diagnose the fault for you. It makes observability, on-call coverage, and change control usable under pressure. Read the Intro for the role model and loop. Use the Cheatsheet when you need the declaration and handoff checklist.
Where this skill leads
Relevant careers
See how this topic contributes to broader role-level skill maps.
Sources
- https://sre.google/sre-book/managing-incidents/
Supports
- Incident management limits disruption and restores normal operations through prepared, structured coordination
- Role separation for incident command, operations, communication, and planning
- The operations team is the only group modifying the system during an incident
- Live incident documents, clear handoffs, early declaration, mitigation before root cause, and practice
- https://sre.google/workbook/incident-response/
Supports
- Incident response coordinates responding teams and communication while incident resolution mitigates impact or restores service
- Clear command, defined roles, working records, and early declaration as incident-response principles
- Incident commander, communications lead, and operations lead responsibilities
- https://sre.google/sre-book/postmortem-culture/
Supports
- Postmortem culture as follow-up learning after incidents
- https://response.pagerduty.com/
Supports
- Open documentation for before, during, and after incident guidance plus role-specific training
- https://github.com/dastergon/awesome-sre
Supports
- Curated discovery of Atlassian Incident Handbook, PagerDuty Incident Response Handbook, and incident.io incident-management resources
- https://www.atlassian.com/incident-management/handbook
Supports
- Incident-management handbook sections for roles, lifecycle, playbooks, communication, on-call, and metrics
- https://www.incident.io/guide
Supports
- Incident-management guides and advice as an ecosystem learner resource
