openskills.info
Course Preview

Incident Management

Incident management is the structured coordination of people, technical work, and communication during a service disruption. It helps a team reduce user impact, restore service, and learn from the event.

itPlatform engineering and SRE

Don't Panic - Incident Management

Incident management is how a team coordinates an urgent response to a service disruption. Technical diagnosis matters, but it is only one part of the work. Someone must keep the response organized, communicate what is known, and prevent uncontrolled changes from making the situation worse.

Declare when impact is visible, more than one team is needed, or focused investigation has not resolved the issue. Declare early. A small, quickly closed incident is cheaper than trying to build a response structure while impact grows. The immediate objective is mitigation: stop or reduce user impact and restore an acceptable service state. Root-cause analysis follows stabilization unless the cause is already clear.

A useful response framework separates command, operations, communication, and planning. The incident commander owns overall state, priorities, roles, and handoffs. The operations lead directs technical investigation and approved production changes. The communications lead gives responders and stakeholders regular, accurate updates. Planning tracks follow-up work, staffing, and changing system state. In a small incident one person can hold several roles. Split them as load grows so the person changing the system is not also answering every stakeholder question.

Keep one shared incident record with current impact, timeline, hypotheses, decisions, actions, owners, and the next update time near the top. Use a collaboration system that remains available when the affected service fails. Only the designated operations group should make production changes. Everyone else contributes through the agreed channel.

The response loop is declare, assign command, reduce harm, observe, update the record, and reassess. Hand off explicitly across shifts. After service is stable, close deliberately, preserve the record, and schedule the review.

A framework does not diagnose the fault for you. It makes observability, on-call coverage, and change control usable under pressure. Read the Intro for the role model and loop. Use the Cheatsheet when you need the declaration and handoff checklist.

Where this skill leads

Relevant careers

See how this topic contributes to broader role-level skill maps.

Sources

  • https://sre.google/sre-book/managing-incidents/
  • https://sre.google/workbook/incident-response/
  • https://sre.google/sre-book/postmortem-culture/
  • https://response.pagerduty.com/
  • https://github.com/dastergon/awesome-sre
  • https://www.atlassian.com/incident-management/handbook
  • https://www.incident.io/guide