The inference serving control plane for vLLM + Kubernetes

Kinference [kay-inference] turns Kubernetes and vLLM into a fully-featured distributed inference platform, with token rate limits, routing policies, cache-aware load balancing, and everything else needed to serve multiple models, tenants, and SLOs.

Solve routing, limiting, and load balancing in one place

Optimize performance, increase GPU utilization, and protect overall system health with Kinference’s intelligent inference routing, token rate limiting, and KV-cache-aware load balancing—all in one unified layer.

Build an inference platform, not a science project

Provide a single endpoint for every inference use case and every model, while tracking capacity, per-tenant costs, and inference-specific metrics.

KV-cache-aware balancing

Cache-aware balancing optimizes inference response times and maximizes GPU utilization using each replica’s live queue and cache state.

Token rate limiting

Prevent tenants from consuming more capacity than they should.

Per-tenant cost attribution

Measure and report actual GPU spend per tenant.

Inference metrics and SLOs

Measure per-tenant TTFT, ITL, TPOT, and end-to-end latency, and evaluate against declared service objectives.

Capacity modeling

Fastest first token prediction using each replica’s live queue and cache state.

Semantic classification

Tag requests by complexity, PII, and other factors, and ensure they’re routed correctly.

Control every step with declarative inference policies

Kinference’s rich, declarative policy language allow you to control exactly how inference requests are handled, based on user, model, prompt, state of the GPUs, cost, and more.

Example policies

Send the simple prompts to less expensive models

Semantically classify prompts by complexity level, and route simpler prompts to cheaper models. But take into account cache usage: if it's faster to send it to an expensive model (because of caching), just do that.

Without Kinference

The model that the user selects is used directly, regardless of suitability for the model.

With Kinference

Inference platform owners can decide exactly under which situations a model is used, including by taking into account semantic classification and GPU cache overlap.

InferencePolicy

apiVersion: kinference.buoyant.io/v1alpha1
kind: InferencePolicy
metadata: { name: smart-semantic-routing,
            namespace: inference }
spec:
  rules:
  - name: low-complexity-consider-qwen
    matches:
    - model: glm-5.3
      tag: complexity-low
    route:
      consider:
      - models: { names: [ qwen-3.6-27b,
                          glm-5.3 ] }
      select: [ { minimize: PredictedTTFT } ]

Reading the policy

matches:

This rule applies only when the requested model is glm-5.3 and the classifier has tagged the prompt as low complexity.

consider:

This rule can route to either glm-5.3 or qwen-3.6-27b.

minimize: PredictedTTFT

This rule picks whichever endpoint is the fastest, based on the predicted TTFT for the prompt. This measure incorporates both KV cache overlpa and GPU queue length.

Questions that platform teams ask

How does Kinference use vLLM?

vLLM is the data plane; Kinference is the control plane. Kinference builds on top of vLLM to provide a scalable, multi-tenant inference platform. Kinference ties together multiple vLLM instances into a unified platform capable of serving many models, tenants, SLOs, and more.

How is Kinference different from an AI gateway?

AI gateways send traffic across many different external inference providers, handling the peculiarities of each one. Kinference can use an AI gateway when it offloads traffic from the local cluster.

Which models does Kinference support?

Kinference is model agnostic and can support any model that vLLM is capable of running.

Does my AI application need code changes?

No. Simply point your client at the Kinference service instead of the model endpoint. Model selection, fallback, and rate limiting move into Kubernetes resources managed by the platform team.

How does semantic classification work in Kinference?

Kinference can optionally tag inference request with information from semantic classifiers, like prompt complexity and the presence of PII. This information is available to platform policies that you write, alongside deterministic information such as tenant identity and what QoS was requested.

How do I run Kinference?

Kinference is a standard Kubernetes deployment. It contains an enforcement gateway that handles inference traffic. Once you deploy Kinference to your Kubernetes cluster, all you need to do is point inference requests through the Kinference enforcement gateway.

Kinference ships with Helm charts as well as a suite of metrics and dashboards in standard Prometheus and Grafana formats, so plugging it into your existing Kubernetes stack is easy.

Get early access to the code.

It’s all open source. Come and kick the tires.

Get early access