Kubernetes Observability
Kubernetes observability instruments clusters and workloads to surface metrics, logs, and traces that reveal what is happening inside a dynamic container environment. It covers monitoring infrastructure health, application performance, and debugging failed deployments.
itCloud native tools and technologies | OpenSkills.info
Recommended first:kubernetes-fundamentals
Course pathWalk it in order
Look it upDip in anytime
Go furtherLeaves this page
Don't Panic
Don't Panic - Kubernetes Observability
Kubernetes observability is about knowing where the platform's signals live (health probes, events, logs, metrics) and how they compose into an investigation path. A cluster is a distributed system running your distributed systems, so "look at the server" is rarely enough.
Work at two levels that must not be confused. The platform level asks whether a Pod is scheduled, a container is restarting, or a node is under pressure. The application level asks whether your code is doing the right thing. Kubernetes gives you the first natively and gives applications plumbing for the second.
Probes declare what healthy means. A liveness probe answers whether the kubelet should restart a container. A readiness probe answers whether the Pod should receive traffic. A startup probe protects slow starters by holding off the others until first success. Misconfigured liveness that checks a downstream dependency can turn one slow database into a restart storm.
Events record scheduling, image pulls, probe failures, and evictions. They answer many "why is my Pod not running" questions, and they expire quickly. Logs follow stdout/stderr: kubectl logs reads them, including --previous after a crash. Durable history needs cluster-level logging, usually a node agent shipping streams to a backend. Kubernetes does not store that history for you.
Resource metrics for kubectl top and HPA flow through metrics-server with no history. Rich monitoring scrapes Prometheus-format /metrics endpoints plus kube-state-metrics and application metrics into a compatible backend.
Investigate in order: describe (events), logs --previous, then widen to node conditions and control plane. CrashLoopBackOff is a symptom, not a cause.
Read the Intro for the two-level model. Use the Cheatsheet when you need the probe and command maps. Updates tracks Kubernetes releases that change these interfaces.
Where this skill leads
Relevant careers
See how this topic contributes to broader role-level skill maps.
Sources
- https://kubernetes.io/docs/tasks/configure-pod-container/configure-liveness-readiness-startup-probes/
Supports
- Liveness/readiness/startup probe semantics and failure consequences
- Probe mechanisms (httpGet, tcpSocket, grpc, exec) and tuning parameters
- Guidance against liveness probes with external dependencies; startup probes for slow-starting apps
- https://kubernetes.io/docs/concepts/workloads/pods/pod-lifecycle/
Supports
- Probe definitions within the Pod lifecycle; readiness gating Service endpoints
- https://kubernetes.io/docs/concepts/cluster-administration/logging/
Supports
- stdout/stderr convention, runtime capture, kubectl logs including --previous
- Node-level rotation; log loss on node failure
- Cluster-level logging architectures — node agent DaemonSet, sidecars, direct shipping
- Kubernetes providing no native cluster-level log storage
- https://kubernetes.io/docs/tasks/debug/debug-cluster/resource-usage-monitoring/
Supports
- Resource metrics pipeline — kubelet/cAdvisor, metrics-server, Metrics API, kubectl top, HPA
- Full monitoring pipeline distinction; metrics-server not a full monitoring solution
- https://kubernetes.io/docs/concepts/cluster-administration/system-metrics/
Supports
- Prometheus-format /metrics endpoints on Kubernetes components
- https://kubernetes.io/docs/concepts/cluster-administration/kube-state-metrics/
Supports
- Object-state metrics exported from the API (replicas, phases, conditions)
- https://kubernetes.io/docs/tasks/debug/debug-application/debug-running-pod/
Supports
- Debugging workflow — describe, logs, exec, ephemeral debug containers (kubectl debug), node debugging
- CrashLoopBackOff investigation via previous logs and events
- https://kubernetes.io/docs/tasks/debug/
Supports
- Application-level vs cluster-level troubleshooting split
- https://kubernetes.io/docs/reference/kubernetes-api/cluster-resources/event-v1/
Supports
- Events as API objects with limited retention (default ~1h TTL)
- https://kubernetes.io/docs/reference/using-api/health-checks/
Supports
- API server livez/readyz/healthz endpoints and verbose checks
