Back home

Design partner program

Design partners for vLLM and Triton on Kubernetes.

I'm looking for a small number of teams running real inference on Kubernetes to run ifa read-only against real telemetry and real queue behaviour.

Where this stands

What design partners can evaluate today

The vLLM adapter, the Triton adapter, and the full Kubernetes path are validated — the vLLM adapter against a live vLLM 0.28.0 CPU deployment, the Triton adapter against a live Triton 25.12 CPU deployment, and the cluster path end to end in CI against a real Kubernetes 1.31 cluster. The rule engine has been exercised in-cluster at its shipped default thresholds.

Next milestone: GPU-backed validation, for either runtime. Neither has been run yet — a design partner running GPU-backed vLLM or Triton is the most useful conversation available right now.

Who this is for

  • Teams running vLLM, Triton, or DCGM Exporter on Kubernetes
  • Platform, MLOps, AI infrastructure, and SRE teams
  • Teams dealing with queueing, time-to-first-token, KV cache pressure, or unclear GPU utilisation

What the alpha does

Discovers Kubernetes inference workloads by runtime and model annotation
Reads vLLM, Triton, and DCGM Exporter metrics from Prometheus
Stores telemetry in TimescaleDB or an in-memory fallback
Evaluates 19 rules across 7 families with evidence, not just a symptom
Serves findings through a REST API and the ifa CLI
Holds read-only RBAC — no write verbs on any resource

What the alpha does not require by default

No promptsNo request bodiesNo model outputsNo training dataNo secretsNo request-path proxyingNo autonomous mutation of your cluster

What the first call looks like

The first conversation is usually 20 to 30 minutes. We discuss your inference stack, runtime, Kubernetes setup, observability tools, and how your team diagnoses latency and queueing issues today. There is no expectation that you share sensitive data in the first conversation.

What a safe pilot looks like

  1. 1Start in development, staging, or a scoped production-like environment
  2. 2Review the RBAC role and data boundaries on the security page
  3. 3Deploy through the Helm chart
  4. 4Compare findings against your existing dashboards
  5. 5Decide together whether GPU-backed vLLM or GPU-backed Triton validation makes sense next

Interested in design partnership?

Reach out directly and tell us what you're running.

Email P95 Labs