Reliability engineering
Importance: Essential (5 of 5)
Designs services and operating practices around explicit reliability risks.
Infrastructure and operations
Applies software engineering to production reliability, scalable operations, and evidence-based service objectives.
Skill coverage
100%
Course coverage
100%
Skill coverage shows how much of the role's capability map already has a published course. Course coverage shows how much of the mapped course pipeline for this role has shipped.
Capability map
Importance describes how central each capability is to the role. Course coverage shows where you can build the skill in the current OpenSkills catalog.
Site Reliability Engineer capability map
3 capabilities
Importance: Essential (5 of 5)
Designs services and operating practices around explicit reliability risks.
Importance: Essential (5 of 5)
Defines service indicators, objectives, and error-budget policies.
Course coverage
Importance: Essential (5 of 5)
Coordinates detection, mitigation, communication, and learning during incidents.
4 capabilities
Importance: Essential (5 of 5)
Uses telemetry to investigate system behavior and protect reliability.
Course coverage
Importance: Essential (5 of 5)
Operates distributed services safely through change and failure.
Importance: Very important (4 of 5)
Eliminates repetitive operational work with maintainable software.
Course coverage
Importance: Very important (4 of 5)
Diagnoses and operates orchestrated production workloads.
Learning path
The concepts this role is built on.
Adjacent knowledge that improves day-to-day judgment.
Deeper paths for particular environments or directions.
Concrete operational procedures from the courses on this page.
Moving a self-managed kubeadm cluster forward one Kubernetes minor version, including the version-skew rules and pre-flight checks that decide whether the upgrade can proceed at all.
Kubernetes Operations
Taking a worker node out of service for disruptive work or retirement without breaching any workload's availability guarantee, and returning it or removing it cleanly afterward.
Kubernetes Operations
Renewing the one-year kubeadm control-plane certificates before they expire, and the offline recovery path for a cluster whose certificates have already lapsed and whose API server is down.
Kubernetes Operations
The last-resort recovery when etcd data is lost or corrupt beyond quorum: replacing entire cluster state from a snapshot, with the data-loss and reconciliation consequences that come with it.
Kubernetes Operations
Clearing an upgrade blocker by locating every client, manifest, and controller still calling a Kubernetes API version the next minor removes, and proving the usage is gone before the upgrade.
Kubernetes Operations
Put the map to work
Search multiple job sources at once with queries tailored to this career, then use the skill map above to evaluate what each role actually asks for.