Design partner program
Design partners for vLLM and Triton on Kubernetes.
I'm looking for a small number of teams running real inference on Kubernetes to run ifa read-only against real telemetry and real queue behaviour.
Where this stands
What design partners can evaluate today
The vLLM adapter, the Triton adapter, and the full Kubernetes path are validated — the vLLM adapter against a live vLLM 0.28.0 CPU deployment, the Triton adapter against a live Triton 25.12 CPU deployment, and the cluster path end to end in CI against a real Kubernetes 1.31 cluster. The rule engine has been exercised in-cluster at its shipped default thresholds.
Next milestone: GPU-backed validation, for either runtime. Neither has been run yet — a design partner running GPU-backed vLLM or Triton is the most useful conversation available right now.
Who this is for
- Teams running vLLM, Triton, or DCGM Exporter on Kubernetes
- Platform, MLOps, AI infrastructure, and SRE teams
- Teams dealing with queueing, time-to-first-token, KV cache pressure, or unclear GPU utilisation
What the alpha does
What the alpha does not require by default
What the first call looks like
The first conversation is usually 20 to 30 minutes. We discuss your inference stack, runtime, Kubernetes setup, observability tools, and how your team diagnoses latency and queueing issues today. There is no expectation that you share sensitive data in the first conversation.
What a safe pilot looks like
- 1Start in development, staging, or a scoped production-like environment
- 2Review the RBAC role and data boundaries on the security page
- 3Deploy through the Helm chart
- 4Compare findings against your existing dashboards
- 5Decide together whether GPU-backed vLLM or GPU-backed Triton validation makes sense next