Infrastructure Performance Engineering
Infrastructure performance engineering is the disciplined work of measuring how compute, memory, storage, and networks serve demand, then removing the constraint that limits a service. It turns latency and capacity symptoms into evidence-based engineering decisions.
itInfrastructure and operations | OpenSkills.info
Course pathWalk it in order
Look it upDip in anytime
Go furtherLeaves this page
Don't Panic
Don't Panic - Infrastructure Performance Engineering
Infrastructure performance engineering connects a service's user-visible behavior to the finite resources that run it. You measure demand, latency, errors, saturation, and resource pressure together. You then form a testable explanation before changing a limit, instance type, query, or deployment.
The goal is not to keep every graph low. A resource can be busy while a service meets its objectives. The concern is sustained contention, queues, throttling, or stalls that increase latency, reduce throughput, or cause errors.
Start with the service contract: the request, job, stream, or batch window that matters. Define success measures and the workload dimensions that change them. Compare a healthy period and an unhealthy period with the same window and scope. This prevents treating a high host CPU percentage as the incident when the real question is whether user-facing latency or errors moved.
A useful first pass is the USE method: for each resource, inspect utilization, saturation, and errors across CPU, memory, storage, network, and container or cgroup limits. On Linux, Pressure Stall Information (PSI) reports time tasks were delayed by CPU, memory, or I/O pressure, revealing harm even when utilization looks ordinary. Cgroup version two can impose limits, so inspect the workload's effective controls as well as host capacity.
Instrument work units at boundaries you can explain. Prefer counters for events your service performs. Keep labels bounded: request IDs and unbounded URL paths explode cardinality and can make monitoring part of the problem.
Investigate with controlled comparisons. Establish a baseline, state a hypothesis, gather evidence, and make one reversible change. A capacity increase may drain a queue without proving efficiency. Read the Intro for the resource model. Use the Cheatsheet when you need the USE and PSI checklist.
Where this skill leads
Relevant careers
See how this topic contributes to broader role-level skill maps.
