The inference serving control plane for vLLM + Kubernetes
Kinference [kay-inference] turns Kubernetes and vLLM into a fully-featured distributed inference platform, with token rate limits, routing policies, cache-aware load balancing, and everything else needed to serve multiple models, tenants, and SLOs.
Solve routing, limiting, and load balancing in one place
Optimize performance, increase GPU utilization, and protect overall system health with Kinference’s intelligent inference routing, token rate limiting, and KV-cache-aware load balancing—all in one unified layer.
Build an inference platform, not a science project
Provide a single endpoint for every inference use case and every model, while tracking capacity, per-tenant costs, and inference-specific metrics.
KV-cache-aware balancing
Cache-aware balancing optimizes inference response times and maximizes GPU utilization using each replica’s live queue and cache state.
Token rate limiting
Prevent tenants from consuming more capacity than they should.
Per-tenant cost attribution
Measure and report actual GPU spend per tenant.
Inference metrics and SLOs
Measure per-tenant TTFT, ITL, TPOT, and end-to-end latency, and evaluate against declared service objectives.
Capacity modeling
Fastest first token prediction using each replica’s live queue and cache state.
Semantic classification
Tag requests by complexity, PII, and other factors, and ensure they’re routed correctly.
Control every step with declarative inference policies
Kinference’s rich, declarative policy language allow you to control exactly how inference requests are handled, based on user, model, prompt, state of the GPUs, cost, and more.
Example policies
Send the simple prompts to less expensive models
Semantically classify prompts by complexity level, and route simpler prompts to cheaper models. But take into account cache usage: if it's faster to send it to an expensive model (because of caching), just do that.
Without Kinference
The model that the user selects is used directly, regardless of suitability for the model.
With Kinference
Inference platform owners can decide exactly under which situations a model is used, including by taking into account semantic classification and GPU cache overlap.
InferencePolicy
apiVersion: kinference.buoyant.io/v1alpha1
kind: InferencePolicy
metadata: { name: smart-semantic-routing,
namespace: inference }
spec:
rules:
- name: low-complexity-consider-qwen
matches:
- model: glm-5.3
tag: complexity-low
route:
consider:
- models: { names: [ qwen-3.6-27b,
glm-5.3 ] }
select: [ { minimize: PredictedTTFT } ]Reading the policy
matches:This rule applies only when the requested model is glm-5.3 and the classifier has tagged the prompt as low complexity.
consider:This rule can route to either glm-5.3 or qwen-3.6-27b.
minimize: PredictedTTFTThis rule picks whichever endpoint is the fastest, based on the predicted TTFT for the prompt. This measure incorporates both KV cache overlpa and GPU queue length.
InferencePolicy flow

Tap the animation to view it full size
Send requests to cloud providers if GPU pool is saturated
When queue depth or cache pressure indicate GPUs are saturated, burst the excess requests to cloud providers.
Without Kinference
Requests to overloaded GPUs go into the pending queue, where they may cause client timeouts
With Kinference
Inference platform owners can decide exactly what happens when GPU pools are saturated, including sending to a third-party inference provider.
InferencePolicy
apiVersion: kinference.buoyant.io/v1alpha1
kind: InferencePolicy
metadata: { name: burst-policy,
namespace: inference }
spec:
rules:
- name: burst-to-cloud
route:
exclude:
- name: drop-overloaded-workers
drop:
backend: { type: Local }
queuedPrompts: { greaterThan: 2 }
select:
- { type: Local,
minimize: PredictedTTFT }
- { type: Provider,
order: [bedrock, databricks] }Reading the policy
exclude:Drop from consideration any workers that have more than an average of 2 pending prompts over the past 30 seconds
select:Select from the remaining endpoints, in order: local endpoints first, then third-party providers.
InferencePolicy flow

Tap the animation to view it full size
Ensure that prompts tagged as containing PII never leave the environment.
A more sophisticated variant of burst-to-cloud. Semantically classify prompts by whether they contain PII, and burst requests to the third-party provider if the GPU pool is saturated -- unless the prompt contains PII, in which case it would be better to fail it entirely.
Without Kinference
No control over whether PII leaves the cluster.
With Kinference
Inference platform owners explicitly decide how inference that contains PII is handled.
InferencePolicy
apiVersion: kinference.buoyant.io/v1alpha1
kind: InferencePolicy
metadata: { name: pii-stays-local,
namespace: inference }
spec:
rules:
- name: serve
route:
exclude:
- name: pii-never-leaves-cluster
if: { tag: pii-detected }
drop: { backend: { type: Provider } }
select:
- { type: Provider,
order: [ databricks ] }
- { type: Local,
minimize: PredictedTTFT }
onEmpty: Reject # fail closedReading the policy
exclude:Remove all provider (third-party) endponits if the request contains PII.
select:Select from the remaining endpoints, in order: local endpoints first, then third-party providers.
onEmpty: RejectFail closed. If nothing acceptable survives the filter, refuse the request rather than fall back to something unsafe.
InferencePolicy flow

Tap the animation to view it full size
Questions that platform teams ask
How does Kinference use vLLM?
vLLM is the data plane; Kinference is the control plane. Kinference builds on top of vLLM to provide a scalable, multi-tenant inference platform. Kinference ties together multiple vLLM instances into a unified platform capable of serving many models, tenants, SLOs, and more.
How is Kinference different from an AI gateway?
AI gateways send traffic across many different external inference providers, handling the peculiarities of each one. Kinference can use an AI gateway when it offloads traffic from the local cluster.
Which models does Kinference support?
Kinference is model agnostic and can support any model that vLLM is capable of running.
Does my AI application need code changes?
No. Simply point your client at the Kinference service instead of the model endpoint. Model selection, fallback, and rate limiting move into Kubernetes resources managed by the platform team.
How does semantic classification work in Kinference?
Kinference can optionally tag inference request with information from semantic classifiers, like prompt complexity and the presence of PII. This information is available to platform policies that you write, alongside deterministic information such as tenant identity and what QoS was requested.
How do I run Kinference?
Kinference is a standard Kubernetes deployment. It contains an enforcement gateway that handles inference traffic. Once you deploy Kinference to your Kubernetes cluster, all you need to do is point inference requests through the Kinference enforcement gateway.
Kinference ships with Helm charts as well as a suite of metrics and dashboards in standard Prometheus and Grafana formats, so plugging it into your existing Kubernetes stack is easy.