openskills.info
Course Preview

On-Call Engineering

On-call engineering is the practice of assigning engineers to respond when production services need urgent attention. It combines rotation design, actionable alerts, incident response, and follow-up work so coverage stays reliable and sustainable.

itPlatform engineering and SRE

Don't Panic: On-Call Engineering

On-call engineering gives a production service a human response path for the moments when automation cannot safely take the wheel. Production has a habit of choosing inconvenient moments for this, because systems do not consult calendars. The job is not to sit beside a pager and hope it remains quiet. The job is to make sure an urgent signal reaches someone who can understand it, act safely, and bring in help when the problem grows.

A page is the expensive form of notification. It interrupts the assigned responder because the condition is urgent, actionable, and needs a human now. A ticket is for work that can wait. Automation is for a response that is safe, understood, and repeatable. Mixing these routes turns every warning into an alarm, and eventually teaches people that alarms are not evidence. That is not a feature. It is an attention leak with a siren attached.

Acknowledgment means that somebody owns the response. It does not mean the service is fixed, the cause is known, or the coffee has been located. Triage starts with visible impact, whether it is growing, recent changes, and service signals. Mitigation reduces current impact. Diagnosis can continue once the system is stable, which is fortunate because outages are poor venues for speculative archaeology.

An escalation policy decides who comes next and when. It is a safety mechanism, not a report card on the first responder. Use it when a page is not acknowledged, the runbook ends, access is missing, another service owner is needed, or one person cannot both investigate and coordinate. A primary and secondary rotation help, but both need context and access to respond. Otherwise the schedule is a carefully formatted way to move uncertainty around.

Handoff transfers incident state, not only the pager. Active pages, temporary mitigations, elevated risks, recent changes, muted alerts, and unfinished follow-up all matter to the next responder. The loop closes when the team removes non-actionable alerts, repairs recurring causes, improves runbooks, clarifies ownership, or automates a safe recovery. Fewer unnecessary pages are evidence that the service improved.

Read the Course tab for the full response loop and its limits. Keep the Cheatsheet open during a tabletop for triage, escalation, and handoff prompts. Field Notes names operational traps that a polished schedule can hide. Then use the exercise to rehearse one fictional page before a real system decides to provide the rehearsal itself.

Where this skill leads

Relevant careers

See how this topic contributes to broader role-level skill maps.

Sources