Cloud architecture
Importance: Essential (5 of 5)
Selects cloud service patterns that satisfy availability, scalability, and cost constraints.
Infrastructure and operations
Designs, provisions, secures, and operates cloud infrastructure and managed services.
Skill coverage
100%
Course coverage
100%
Skill coverage shows how much of the role's capability map already has a published course. Course coverage shows how much of the mapped course pipeline for this role has shipped.
Capability map
Importance describes how central each capability is to the role. Course coverage shows where you can build the skill in the current OpenSkills catalog.
Cloud Engineer capability map
5 capabilities
Importance: Essential (5 of 5)
Selects cloud service patterns that satisfy availability, scalability, and cost constraints.
Importance: Essential (5 of 5)
Designs connectivity, traffic flow, and isolation for cloud workloads.
Importance: Essential (5 of 5)
Applies identity, least privilege, workload protection, and policy controls.
Importance: Very important (4 of 5)
Selects and operates durable storage for workload requirements.
Importance: Essential (5 of 5)
Selects and operates virtual-machine, container, and serverless compute services.
4 capabilities
Importance: Very important (4 of 5)
Deploys and operates cloud-native workloads on orchestrated platforms.
Course coverage
Importance: Essential (5 of 5)
Maintains service health, capacity, upgrades, and recovery procedures.
Importance: Essential (5 of 5)
Provisions cloud resources through reviewed and reproducible definitions.
Importance: Very important (4 of 5)
Uses cloud and workload telemetry to diagnose behavior and failures.
Learning path
The concepts this role is built on.
Adjacent knowledge that improves day-to-day judgment.
Deeper paths for particular environments or directions.
Concrete operational procedures from the courses on this page.
Moving a live relational database onto RDS or Aurora with an AWS DMS full-load-and-CDC task: the schema and index work around the load, the CDC latency gate that decides when a cutover is safe, and the reverse replication path that keeps rollback open.
AWS Databases
Recovering from a bad write by restoring RDS or Aurora point-in-time recovery into a new instance beside the untouched source, validating the restore time, and choosing between a full cutover and a selective repair.
AWS Databases
Disabling the credential, revoking the temporary sessions it already started, and scoping the principals, identity providers, and resources an attacker may have created for persistence, across AWS, Microsoft Entra, and Google Cloud.
Cloud Security
Isolating one compromised virtual machine without destroying evidence: the order of memory capture, disk snapshots, network isolation, and revoking the instance's own cloud credentials, then rebuilding from a known-good image rather than cleaning in place.
Cloud Security
Replacing the shared Serf gossip key on a running datacenter without partitioning any agent, and what changes when the datacenters are WAN-federated.
HashiCorp Consul
Recovering a Consul server cluster that has lost quorum, and where the peers.json Raft rebuild ends and a consul snapshot restore begins.
HashiCorp Consul
Turning on the Consul ACL system on a datacenter that has none without a service-discovery outage, and why the permissive-first, deny-last order matters.
HashiCorp Consul
Moving a self-managed kubeadm cluster forward one Kubernetes minor version, including the version-skew rules and pre-flight checks that decide whether the upgrade can proceed at all.
Kubernetes Operations
Taking a worker node out of service for disruptive work or retirement without breaching any workload's availability guarantee, and returning it or removing it cleanly afterward.
Kubernetes Operations
Renewing the one-year kubeadm control-plane certificates before they expire, and the offline recovery path for a cluster whose certificates have already lapsed and whose API server is down.
Kubernetes Operations
The last-resort recovery when etcd data is lost or corrupt beyond quorum: replacing entire cluster state from a snapshot, with the data-loss and reconciliation consequences that come with it.
Kubernetes Operations
Clearing an upgrade blocker by locating every client, manifest, and controller still calling a Kubernetes API version the next minor removes, and proving the usage is gone before the upgrade.
Kubernetes Operations
Put the map to work
Search multiple job sources at once with queries tailored to this career, then use the skill map above to evaluate what each role actually asks for.