On-Call Engineering
On-call engineering is the practice of assigning engineers to respond when production services need urgent attention. It combines rotation design, actionable alerts, incident response, and follow-up work so coverage stays reliable and sustainable.
itPlatform engineering and SRE | OpenSkills.info
Intro
On-Call Engineering
On-call engineering gives a production service a clear human response path when automation cannot safely resolve a problem. During an assigned shift, an engineer accepts urgent notifications, assesses their impact, and coordinates the work needed to restore service.
The pager is only the visible edge of the system. A healthy on-call program also needs service ownership, useful alerts, schedules, escalation paths, response guidance, incident roles, and follow-up. If any link is missing, a page can reach someone who lacks the context or authority to act.
The operating loop
Think of on-call as a feedback loop:
- A monitoring system detects a condition that may need urgent human action.
- Alerting routes a page to the current primary responder.
- The responder acknowledges, triages, and chooses an immediate action.
- The responder escalates when the issue exceeds their knowledge, authority, or available capacity.
- The team mitigates the impact, communicates, and restores normal operation.
- Follow-up work improves alerts, runbooks, automation, or the service itself.
Continue the course
This section is part of the paid course.
See pricing to subscribe, or log in if you already have access.
Where this skill leads
Relevant careers
See how this topic contributes to broader role-level skill maps.
Sources
- https://sre.google/sre-book/being-on-call/
Supports
- On-call duties, acknowledgment, triage, mitigation, and escalation
- Primary and secondary rotation patterns
- Balanced workload, operational overload, and operational underload
- Engineering work as the means to reduce and scale operational work
- Controlled exercises for maintaining troubleshooting readiness
- https://sre.google/workbook/on-call/
Supports
- Pager-load definition and reduction
- Choosing pages, tickets, or automated repair according to urgency
- Scheduling flexibility, team health, incentives, and on-call dynamics
- Sustainable coverage that protects project work and responder health
- https://sre.google/sre-book/managing-incidents/
Supports
- Risks of unmanaged parallel response work
- Explicit coordination roles, shared incident state, and communication
- Separation of technical investigation from broader incident management
- https://sre.google/sre-book/accelerating-sre-on-call/
Supports
- Service knowledge and structured on-call learning checklists
- Shadowing, reverse shadowing, supervised response, and exercises
- Readiness as demonstrated operational competence
- https://response.pagerduty.com/
Supports
- Practitioner guidance spanning on-call preparation, incident response, roles, communication, and follow-up
- Scope and learning sequence described in the reference-link rationale
- https://response.pagerduty.com/oncall/being_oncall/
Supports
- Responder responsibilities and the purpose of paging
- Acknowledgment and escalation within an on-call response
- https://response.pagerduty.com/oncall/alerting_principles/
Supports
- Actionable paging and separation of urgent pages from lower-priority work
- https://response.pagerduty.com/during/during_an_incident/
Supports
- Constructive participation, escalation, and coordinated work during incidents
- https://response.pagerduty.com/before/different_roles/
Supports
- Incident coordinator, scribe, liaison, and subject-matter expert responsibilities
- Separation of coordination, communication, and technical work
- https://github.com/sindresorhus/awesome
Supports
- Discovery of the curated Awesome Site Reliability Engineering list
- https://github.com/dastergon/awesome-sre
Supports
- Curated On-Call section
- Discovery of the Awesome SRE Tools list
- https://github.com/SquadcastHub/awesome-sre-tools
Supports
- Discovery of FireHydrant, Rootly, PagerTree, and Cabot in incident-management, alerting, on-call, or monitoring categories
- https://docs.firehydrant.com/docs/signals-introduction
Supports
- Events, alerts, on-call schedules, escalation policies, notification methods, and incident promotion
- FireHydrant Awesome Links rationale
- https://docs.rootly.com/on-call/escalation-policies
Supports
- Escalation targets, delays, repeat behavior, dynamic paths, and assignment to services or teams
- Rootly Awesome Links rationale
- https://pagertree.com/docs/teams
Supports
- Teams grouping users, schedules, escalation policies, and alerts
- On-call calendar export and PagerTree Awesome Links rationale
- https://cabotapp.com/
Supports
- Self-hosted monitoring of metrics, jobs, and web endpoints
- Alert delivery to support staff and Cabot Awesome Links rationale
