Design partner program
Design partners for real inference workloads.
We're working with a small number of teams running inference on Kubernetes to validate P95 Labs in read-only mode against real telemetry, real queue behaviour, and real operational workflows.
Phase 2.5 alpha complete
What design partners can evaluate today
Design partners can review and test the current Kubernetes-native alpha in a controlled environment. The current build includes workload discovery, telemetry ingestion, recommendation output, CLI access, Helm deployment, documented RBAC, data collection boundaries, an operations runbook, and air-gapped deployment preparation.
Next milestone: validating recommendations against real inference workloads such as vLLM, Triton, TGI, KServe, or custom serving stacks.
Who this is for
- Teams running vLLM, Triton, TGI, KServe, Ray Serve, or custom runtimes
- Platform, MLOps, AI infrastructure, and SRE teams
- Teams dealing with p95/p99 latency, queue pressure, GPU utilisation issues, autoscaling lag, or manual runtime tuning
What the alpha does
What the alpha does not require by default
What the first call looks like
The first conversation is usually 20 to 30 minutes. We discuss your inference stack, runtime, Kubernetes setup, observability tools, latency issues, GPU utilisation patterns, and how your team diagnoses incidents today. There is no expectation that you share sensitive data in the first conversation.
What a safe pilot looks like
- 1Start with development, staging, synthetic, or scoped production-like environment
- 2Review RBAC and telemetry boundaries
- 3Deploy through Helm
- 4Validate recommendations against existing dashboards
- 5Decide together whether deeper validation makes sense