ML Observability Tooling
ML observability tooling collects production signals about models, data, and prediction services so teams can detect changes and investigate failures. It connects model-quality evidence with the logs, metrics, traces, and alerts used to operate software.
itArtificial intelligence and machine learning | OpenSkills.info
Course pathWalk it in order
Look it upDip in anytime
Go furtherLeaves this page
Don't Panic
Don't Panic - ML Observability Tooling
ML observability tooling turns production model behavior into evidence engineers can inspect and act on. It sits between model serving and operational response. Instrumentation records events at inference time. A collector or batch job carries those events to storage. Analysis compares current behavior with expectations. Dashboards, alerts, and investigation views present the results so teams can roll back, retrain, fix data quality, or change a prompt.
No single signal describes model health. Service metrics show latency, errors, saturation, and traffic. Prediction records connect inputs and outputs to a model version. Data profiles summarize distributions and schema. Ground-truth labels make direct performance measurement possible when they arrive. Traces show multi-step paths through retrieval and model calls. Business outcomes show whether technically valid predictions still serve the intended purpose.
An inference event needs stable identity before sophisticated statistics: timestamp, request or trace id, model and pipeline versions, prediction, and privacy-permitted fields. Labels often arrive late. Until they join back to the original prediction, drift and quality checks are warnings, not proof that accuracy changed. Sampling reduces volume and can miss rare slices. Many systems combine raw events, profiles, and aggregates for that reason.
Tooling patterns differ. Managed platforms reduce assembly work and impose a vendor data model. Open-source libraries keep data placement flexible and leave scheduling, storage, and alerting to the team. Cloud-native monitors fit models already served in one provider. Composable stacks send ML signals into general OpenTelemetry and metric systems, which still need model identity and label joins defined explicitly.
Start from failure questions, not feature counts. Confirm delayed labels, schema evolution, high-cardinality features, and ingestion freshness before trusting a green dashboard. Read the Intro for the evidence path. Use the Cheatsheet when you need the signal map. Landscape places the tools beside each other; Updates tracks MLflow and Phoenix releases that change concrete workflows this course uses as examples.
Where this skill leads
Relevant careers
See how this topic contributes to broader role-level skill maps.
Sources
- https://developers.google.com/machine-learning/guides/rules-of-ml
Supports
- Production ML systems require monitoring of serving behavior, features, and data pipelines
- Reference-path placement for operational ML foundations
- https://cloud.google.com/vertex-ai/docs/model-monitoring/overview
Supports
- Training-serving skew and inference drift baseline definitions
- Scheduled monitoring jobs, thresholds, feature distributions, and attribution monitoring
- Google Cloud Model Monitoring placement in Landscape
- https://opentelemetry.io/docs/specs/semconv/
Supports
- Semantic conventions provide common names and meaning across telemetry producers and consumers
- OpenTelemetry placement as an interoperability layer
- https://opentelemetry.io/docs/specs/semconv/registry/attributes/gen-ai/
Supports
- Generative AI telemetry attributes for provider, model, operations, and usage
- https://docs.evidentlyai.com/introduction
Supports
- Evidently evaluation and monitoring workflow
- Evidently placement in Awesome Links and Landscape
- https://docs.evidentlyai.com/quickstart_ml
Supports
- Prediction quality, input quality, data drift, reports, tests, and continuous monitoring
- Direct performance and drift distinction used in the course and quiz
- https://docs.whylabs.ai/docs/overview-profiles/
Supports
- Statistical profiles as efficient, customizable, mergeable dataset summaries
- Reference profiles as baselines for data drift monitoring
- https://docs.whylabs.ai/docs/whylogs-overview/
Supports
- Local creation of statistical summaries for data and model health analysis
- whylogs placement in Awesome Links
- https://docs.whylabs.ai/docs/performance-metrics/
Supports
- Performance metrics from predictions, targets, and optional scores
- Partial and delayed ground truth handling
- https://docs.whylabs.ai/docs/monitor-manager/
Supports
- Reference, trailing-window, and date-range baselines
- Data quality, drift, performance, and integration-health monitors
- https://docs.whylabs.ai/docs/whylabs-alerts/
Supports
- Drift, quality, performance, and missing-profile anomalies
- Independent detection of monitoring integration failure
- https://docs.whylabs.ai/docs/whylabs-overview-observe/
Supports
- Model and dataset dashboards, slices, alerts, and production investigation
- WhyLabs placement in Landscape
- https://docs.nannyml.com/cloud/
Supports
- Post-deployment monitoring workflow and project organization
- NannyML placement in Reference, Awesome Links, and Landscape
- https://docs.fiddler.ai/observability/monitoring
Supports
- Managed monitoring dashboards, alerting, performance, integrity, and drift capabilities
- Fiddler placement in Reference and Landscape
- https://docs.fiddler.ai/observability/platform/monitoring-charts-platform
Supports
- Multi-model and multi-feature charts, root-cause analysis, and performance views
- Slice and cohort investigation guidance
- https://docs.fiddler.ai/reference/ml-metrics-reference
Supports
- Built-in measures spanning performance, drift, integrity, traffic, and statistics
- https://docs.aporia.com/
Supports
- Production model visualization, drift, performance, and integrity monitoring
- Aporia placement in Landscape
- https://docs.aporia.com/monitors-and-alerts/data-drift
Supports
- Baseline windows, drift measures, thresholds, and field-level monitoring
- https://arize.com/docs/ax
Supports
- Tracing, production evaluations, dashboards, alerts, datasets, and experiments
- Arize AX placement in Landscape
- https://arize.com/docs/phoenix
Supports
- OpenTelemetry-based tracing, evaluations, datasets, experiments, and self-hosting
- Arize Phoenix placement in Awesome Links and Landscape
- https://www.mlflow.org/docs/latest/genai/eval-monitor/running-evaluation/traces/
Supports
- Search, annotation, evaluation, feedback, and monitoring of stored production traces
- MLflow placement in Reference and Landscape
- https://docs.deepchecks.com/stable/getting-started/welcome.html
Supports
- Reusable data integrity, distribution, comparison, performance, integration, and production checks
- Deepchecks placement in Awesome Links
- https://www.comet.com/docs/opik/
Supports
- Trace logging, online evaluation, feedback, latency, cost, errors, and production-derived test cases
- Opik placement in Awesome Links
- https://github.com/sindresorhus/awesome
Supports
- Discovery path to the Awesome MLOps list
- https://github.com/kelvins/awesome-mlops
Supports
- Curated discovery of Deepchecks, NannyML, Phoenix, whylogs, Opik, and Evidently
- Ecosystem and Landscape research decision
