KServe and Kubernetes Model Serving
KServe is a Kubernetes platform for running machine-learning models behind network APIs. It adds resources and controllers that turn model, runtime, scaling, and routing declarations into managed inference workloads.
itArtificial intelligence and machine learning | OpenSkills.info
Course pathWalk it in order
Look it upDip in anytime
Go furtherLeaves this page
Don't Panic
Don't Panic - KServe and Kubernetes Model Serving
KServe is a model-serving control plane on Kubernetes. You declare an inference workload, and KServe reconciles that declaration into compute, storage, networking, and scaling resources. The model server then exposes a data-plane API that applications call for predictions or generated output. KServe does not train models, replace a model registry, or make a model accurate. It standardizes how trained artifacts become operated services.
The primary resource is InferenceService. Its predictor names the serving
runtime, model format, model location, and compute needs. Optional transformer
and explainer components add pre-processing, post-processing, or explanations.
ServingRuntime and ClusterServingRuntime separate runtime images and
defaults from a particular model so platform teams and model teams can own
different layers.
Two deployment modes change the operating story. Standard mode uses ordinary Deployments, Services, and HPA or KEDA. Knative mode can scale to zero and reactivate on traffic, which saves idle capacity and can put cold start on the first request. Choose the mode from latency and utilization behavior, not from feature count. Capacity planning must include model memory, accelerator memory, concurrency, batching, and load time. CPU alone rarely describes inference saturation.
The Open Inference Protocol (V2) gives runtimes a common HTTP or gRPC shape for readiness, metadata, and inference. Protocol compatibility does not imply equal performance or framework features. InferenceGraph can sequence, switch, ensemble, or split calls when composition is required, at the cost of more network hops and failure modes.
Read the Intro for control-plane versus data-plane separation. Use the Cheatsheet when you need the resource and mode map. Landscape places KServe among related serving tools; Updates and Upstream track the kserve/kserve release line that changes these APIs.
Where this skill leads
Relevant careers
See how this topic contributes to broader role-level skill maps.
Sources
- https://kserve.github.io/website/docs/intro
Supports
- KServe purpose, control and data planes, resources, use cases, protocols, and model-serving boundary
- https://kserve.github.io/website/docs/concepts/architecture
Supports
- Control-plane and data-plane separation
- Standard and Knative deployment-mode comparison
- https://kserve.github.io/website/docs/concepts/architecture/control-plane
Supports
- Controller reconciliation, generated resources, Gateway API, autoscaling, scale-to-zero, and operational tradeoffs
- https://kserve.github.io/website/docs/concepts
Supports
- InferenceService, ServingRuntime, ClusterServingRuntime, InferenceGraph, storage and cache resource roles
- https://kserve.github.io/website/docs/reference/crd-api
Supports
- Predictor requirement, optional transformer and explainer, status conditions, and runtime fields
- https://kserve.github.io/website/docs/concepts/architecture/data-plane/v2-protocol
Supports
- Open Inference Protocol V2 operations and HTTP or gRPC implementation choice
- https://kserve.github.io/website/docs/concepts/resources/inferencegraph
Supports
- Sequence, Switch, Ensemble and Splitter semantics and independently scalable targets
- https://kserve.github.io/website/docs/admin-guide/configurations
Supports
- Deployment modes, HPA and KEDA configuration, resource defaults, and per-service overrides
- https://kserve.github.io/website/docs/model-serving/predictive-inference/observability/prometheus-metrics
Supports
- Prometheus exposure and lack of one uniform runtime metric set
- https://kserve.github.io/website/docs/model-serving/node-scheduling/isvc-node-scheduling
Supports
- Node selectors, affinity and tolerations for inference workloads
- https://kserve.github.io/website/docs/admin-guide/serverless/servicemesh
Supports
- Service-mesh TLS, authentication and authorization for inference traffic
- https://github.com/sindresorhus/awesome
Supports
- Discovery route to the curated Awesome Kubernetes list
- https://github.com/ramitsurana/awesome-kubernetes
Supports
- Discovery of Kubeflow, Seldon Core and Polyaxon in the Kubernetes machine-learning ecosystem
- https://www.kubeflow.org/docs/components/kserve/
Supports
- KServe placement within Kubeflow
- https://docs.seldon.ai/seldon-core-2
Supports
- Seldon Core as a Kubernetes inference platform comparison
- https://polyaxon.com/docs/
Supports
- Polyaxon experiment and model workflows on Kubernetes
- https://www.kubeflow.org/docs/components/model-registry/
Supports
- Registry metadata and lifecycle role adjacent to serving
- https://kserve.github.io/website/
Supports
- KServe product homepage and Landscape placement
- https://www.seldon.io/solutions/open-source-projects/core
Supports
- Seldon Core product and licensing placement
- https://www.bentoml.com/
Supports
- BentoML model-serving platform placement
- https://docs.ray.io/en/latest/serve/
Supports
- Ray Serve scalable online inference placement
- https://developer.nvidia.com/triton-inference-server
Supports
- NVIDIA Triton runtime and inference-server placement
- https://mlserver.readthedocs.io/
Supports
- MLServer V2-compatible inference runtime placement
- https://www.tensorflow.org/tfx/guide/serving
Supports
- TensorFlow Serving runtime placement
- https://docs.openvino.ai/2025/model-server/ovms_what_is_openvino_model_server.html
Supports
- OpenVINO Model Server runtime placement
