On-Call Engineering
On-call engineering is the practice of assigning engineers to respond when production services need urgent attention. It combines rotation design, actionable alerts, incident response, and follow-up work so coverage stays reliable and sustainable.
itPlatform engineering and SRE | OpenSkills.info
Course pathWalk it in order
Look it upDip in anytime
Go furtherLeaves this page
Don't Panic
Don't Panic: On-Call Engineering
On-call engineering gives a production service a human response path for the moments when automation cannot safely take the wheel. Production has a habit of choosing inconvenient moments for this, because systems do not consult calendars. The job is not to sit beside a pager and hope it remains quiet. The job is to make sure an urgent signal reaches someone who can understand it, act safely, and bring in help when the problem grows.
A page is the expensive form of notification. It interrupts the assigned responder because the condition is urgent, actionable, and needs a human now. A ticket is for work that can wait. Automation is for a response that is safe, understood, and repeatable. Mixing these routes turns every warning into an alarm, and eventually teaches people that alarms are not evidence. That is not a feature. It is an attention leak with a siren attached.
Acknowledgment means that somebody owns the response. It does not mean the service is fixed, the cause is known, or the coffee has been located. Triage starts with visible impact, whether it is growing, recent changes, and service signals. Mitigation reduces current impact. Diagnosis can continue once the system is stable, which is fortunate because outages are poor venues for speculative archaeology.
An escalation policy decides who comes next and when. It is a safety mechanism, not a report card on the first responder. Use it when a page is not acknowledged, the runbook ends, access is missing, another service owner is needed, or one person cannot both investigate and coordinate. A primary and secondary rotation help, but both need context and access to respond. Otherwise the schedule is a carefully formatted way to move uncertainty around.
Handoff transfers incident state, not only the pager. Active pages, temporary mitigations, elevated risks, recent changes, muted alerts, and unfinished follow-up all matter to the next responder. The loop closes when the team removes non-actionable alerts, repairs recurring causes, improves runbooks, clarifies ownership, or automates a safe recovery. Fewer unnecessary pages are evidence that the service improved.
Read the Course tab for the full response loop and its limits. Keep the Cheatsheet open during a tabletop for triage, escalation, and handoff prompts. Field Notes names operational traps that a polished schedule can hide. Then use the exercise to rehearse one fictional page before a real system decides to provide the rehearsal itself.
Where this skill leads
Relevant careers
See how this topic contributes to broader role-level skill maps.
Sources
- https://sre.google/sre-book/being-on-call/
Supports
- On-call duties, acknowledgment, triage, mitigation, and escalation
- Primary and secondary rotation patterns
- Balanced workload, operational overload, and operational underload
- Engineering work as the means to reduce and scale operational work
- Controlled exercises for maintaining troubleshooting readiness
- https://sre.google/workbook/on-call/
Supports
- Pager-load definition and reduction
- Choosing pages, tickets, or automated repair according to urgency
- Scheduling flexibility, team health, incentives, and on-call dynamics
- Sustainable coverage that protects project work and responder health
- https://sre.google/sre-book/managing-incidents/
Supports
- Risks of unmanaged parallel response work
- Explicit coordination roles, shared incident state, and communication
- Separation of technical investigation from broader incident management
- https://sre.google/sre-book/accelerating-sre-on-call/
Supports
- Service knowledge and structured on-call learning checklists
- Shadowing, reverse shadowing, supervised response, and exercises
- Readiness as demonstrated operational competence
- https://response.pagerduty.com/
Supports
- Practitioner guidance spanning on-call preparation, incident response, roles, communication, and follow-up
- Scope and learning sequence described in the reference-link rationale
- https://response.pagerduty.com/oncall/being_oncall/
Supports
- Responder responsibilities and the purpose of paging
- Acknowledgment and escalation within an on-call response
- https://response.pagerduty.com/oncall/alerting_principles/
Supports
- Actionable paging and separation of urgent pages from lower-priority work
- https://response.pagerduty.com/during/during_an_incident/
Supports
- Constructive participation, escalation, and coordinated work during incidents
- https://response.pagerduty.com/before/different_roles/
Supports
- Incident coordinator, scribe, liaison, and subject-matter expert responsibilities
- Separation of coordination, communication, and technical work
- https://github.com/sindresorhus/awesome
Supports
- Discovery of the curated Awesome Site Reliability Engineering list
- https://github.com/dastergon/awesome-sre
Supports
- Curated On-Call section
- Discovery of the Awesome SRE Tools list
- https://github.com/SquadcastHub/awesome-sre-tools
Supports
- Discovery of FireHydrant, Rootly, PagerTree, and Cabot in incident-management, alerting, on-call, or monitoring categories
- https://docs.firehydrant.com/docs/signals-introduction
Supports
- Events, alerts, on-call schedules, escalation policies, notification methods, and incident promotion
- FireHydrant Awesome Links rationale
- https://docs.rootly.com/on-call/escalation-policies
Supports
- Escalation targets, delays, repeat behavior, dynamic paths, and assignment to services or teams
- Rootly Awesome Links rationale
- https://pagertree.com/docs/teams
Supports
- Teams grouping users, schedules, escalation policies, and alerts
- On-call calendar export and PagerTree Awesome Links rationale
- https://cabotapp.com/
Supports
- Self-hosted monitoring of metrics, jobs, and web endpoints
- Alert delivery to support staff and Cabot Awesome Links rationale
- https://www.pagerduty.com/platform/incident-management/on-call-management/schedules-and-escalations/
Supports
- PagerDuty landscape placement for rotations, overrides, and automated escalation
- https://docs.incident.io/on-call/getting-started
Supports
- incident.io landscape placement for alerts, escalation paths, schedules, and team context
- https://docs.firehydrant.com/docs/signals-on-call-schedules
Supports
- FireHydrant landscape placement for schedules, primary and secondary rotations, and coverage gaps
- Field Notes card about coverage gaps as a pre-incident signal
- https://www.xmatters.com/features/on-call-management/
Supports
- xMatters landscape placement for rotations, coverage gaps, escalation, and workflow actions
- https://www.splunk.com/en_us/products/on-call.html
Supports
- Splunk On-Call landscape placement for rotations, overrides, and incident context
- https://betterstack.com/docs/uptime/getting-started-with-oncall-v2/
Supports
- Better Stack landscape placement for primary schedules, acknowledgment, and escalation
- https://www.nagios.org/about/history/
Supports
- 1996 numeric paging, 1999 NetSaint release, and 2002 Nagios rename in the timeline
- https://www.pagerduty.com/blog/company/decade-of-duty/
Supports
- PagerDuty's August 2009 beta launch in the timeline
- https://www.oreilly.com/library/view/site-reliability-engineering/9781491929117/copyright-page01.html
Supports
- 2016 publication of the first Site Reliability Engineering edition in the timeline
- https://www.oreilly.com/library/view/the-site-reliability/9781492029496/
Supports
- The Site Reliability Workbook practical SRE focus in the timeline
- https://www.splunk.com/en_us/about-splunk/acquisitions.html
Supports
- Splunk's June 2018 VictorOps acquisition in the timeline
- https://investors.atlassian.com/files/doc_financials/2019/ar/TEAM-2019_Annual_Report.pdf
Supports
- Atlassian's October 2018 Opsgenie acquisition in the timeline
- https://incident.io/hubs/on-call/announcing-on-call
Supports
- incident.io's March 2024 On-call launch in the timeline
- Field Notes card about preserving alert context through response
