Your Inference Autoscaler Watches CPU. Your Users Feel the Queue.
You ramp traffic onto an LLM serving deployment and the graph does not add up. Time-to-first-token, flat at 300ms all morning, starts climbing: 900ms, 2s, 5s. Users are waiting. And the HorizontalPodAutoscaler that is supposed to save you is sitting at 2 replicas, perfectly calm, because the metric it watches — CPU — reads 22%. Nothing is broken, exactly. The autoscaler is doing precisely what you told it to do. You just told it to watch the wrong number. The graph that doesn't add up Here is the HPA's view of the world during the incident: $ kubectl describe hpa vllm-llama Name: vllm-llama Reference: Deployment/vllm-llama Metrics: ( current / target ) resource cpu on pods (as a percentage of request): 22% (110m) / 70% Min replicas: 2 Max replicas: 12 Deployment pods: 2 current / 2 desired Conditions: Type Status Reason Message ScalingActive True ValidMetricFound the HPA was able to compute a target ...