Your Inference Autoscaler Watches CPU. Your Users Feel the Queue.
You ramp traffic onto an LLM serving deployment and the graph does not add up. Time-to-first-token, flat at 300ms all morning, starts climbing: 900ms, 2s, 5s. Users are waiting. And the HorizontalPodAutoscaler that is supposed to save you is sitting at 2 replicas, perfectly calm, because the metric it watches — CPU — reads 22%.
Nothing is broken, exactly. The autoscaler is doing precisely what you told it to do. You just told it to watch the wrong number.
The graph that doesn't add up
Here is the HPA's view of the world during the incident:
$ kubectl describe hpa vllm-llama
Name: vllm-llama
Reference: Deployment/vllm-llama
Metrics: ( current / target )
resource cpu on pods (as a percentage of request): 22% (110m) / 70%
Min replicas: 2
Max replicas: 12
Deployment pods: 2 current / 2 desired
Conditions:
Type Status Reason Message
ScalingActive True ValidMetricFound the HPA was able to compute a target
AbleToScale True ReadyForNewScale recommended size matches current size
Events: <none>
CPU at 22% against a 70% target. From where the autoscaler stands, this deployment is nearly idle and there is nothing to do. Now here is the same moment from inside the inference server:
$ curl -s localhost:8000/metrics | grep -E 'vllm:num_requests|vllm:gpu_cache'
vllm:num_requests_running{model_name="llama-3.1-8b"} 48.0
vllm:num_requests_waiting{model_name="llama-3.1-8b"} 137.0
vllm:gpu_cache_usage_perc{model_name="llama-3.1-8b"} 0.94
Forty-eight requests decoding, one hundred and thirty-seven waiting in line, and the KV cache 94% full. The server is saturated and backlogged. The autoscaler has no idea, because none of that shows up as CPU.
Why utilization lies for inference
CPU on a GPU inference pod is close to meaningless — the work happens on the accelerator, and the CPU mostly shuffles tokens in and out. The obvious fix is to scale on GPU utilization instead. That is better, but it still lies, just more subtly.
GPU "utilization" (the DCGM_FI_DEV_GPU_UTIL most people reach for)
measures whether a kernel was executing during a sample window. It says nothing
about whether that work is keeping up. Token generation is memory-bandwidth
bound during decode; a card can sit at 70% "utilization" while a queue builds
behind it, because the bottleneck is KV cache capacity and memory bandwidth, not
raw compute occupancy. You can be 100% utilized and healthy, or 65% utilized and
drowning. Utilization tells you a device is busy. It does not tell you whether
work is piling up.
The thing you actually care about has a name, and it is old: queueing. Little's
Law says the number of requests in the system equals the arrival rate times the
time each one spends there (L = λW). When arrivals outrun
service capacity, the queue — num_requests_waiting — grows, and
it grows before any utilization gauge pins. That queue length is a direct,
early proxy for the latency your users feel. It is the signal your autoscaler
should be chasing.
Scaling on the signal that tracks the SLO
Modern inference servers already export the right metric. vLLM publishes
vllm:num_requests_waiting, vllm:num_requests_running,
and vllm:gpu_cache_usage_perc; Text Generation Inference exposes
tgi_queue_size and tgi_batch_current_size. The job is to
get one of those in front of the autoscaler.
Step 1: scrape it
On GKE with Managed Service for Prometheus, a PodMonitoring
resource is enough. Keep the scrape interval tight — it sets the floor on how
fast your control loop can react.
apiVersion: monitoring.googleapis.com/v1
kind: PodMonitoring
metadata:
name: vllm
spec:
selector:
matchLabels:
app: vllm-llama
endpoints:
- port: metrics
interval: 15s
Step 2: point the HPA at the queue, not the CPU
GKE's HPA can consume Managed Prometheus metrics directly, and recent GKE versions let you express the scaling signal as a native PromQL query rather than a bare metric name. The flap-prevention here doesn't come from smoothing the metric itself — a function like max_over_time holds the peak of its window, so it makes you react to a single noisy scrape immediately and keep reacting to it until the window ages out, which is the opposite of what you want on the way up. The real work is done by the HPA's own behavior block below: stabilizationWindowSeconds: 0 on scale-up means you react to a genuine backlog fast (queued GPU requests cost users, waiting is expensive), while stabilizationWindowSeconds: 300 on scale-down means you don't give that capacity back for five minutes — so a transient spike doesn't cause a scale-up, an immediate scale-down, then another scale-up a minute later.apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
name: vllm-llama
spec:
scaleTargetRef:
apiVersion: apps/v1
kind: Deployment
name: vllm-llama
minReplicas: 2
maxReplicas: 12
metrics:
- type: External
external:
metric:
# PromQL: smooth the gauge so one bad scrape doesn't flap the fleet
name: prometheus.googleapis.com|vllm:num_requests_waiting|gauge
target:
# External + AverageValue divides by replica count:
# "keep ~5 queued requests per replica"
type: AverageValue
averageValue: "5"
behavior:
scaleUp:
stabilizationWindowSeconds: 0 # react to backlog immediately
policies:
- type: Pods
value: 4 # but at most +4 pods per 30s
periodSeconds: 30
scaleDown:
stabilizationWindowSeconds: 300 # cool down slowly; GPUs are expensive to churn
policies:
- type: Pods
value: 1
periodSeconds: 60
The AverageValue target is the key line. With an
External metric, the HPA divides the total by the current replica
count, so "averageValue: 5" means keep roughly five requests queued per
replica. Cross that and it scales up; drop well under it and it scales down.
The target is now expressed in the same units as your problem — backlog —
instead of a device-busy percentage that correlates with your SLO only by
accident.
The asymmetric behavior block matters for GPUs specifically. Scale
up fast, because a backlog turns into SLO burn in seconds. Scale down slowly,
because a GPU pod is expensive to spin up (multi-gigabyte image pulls, weight
loading, CUDA graph capture) and you do not want to pay that cold-start tax again
in a minute because the queue briefly dipped.
What actually changes
The mechanics of the fix are almost boring: same HPA object, same Deployment,
one metric swapped. But the behavior inverts. Under the same traffic ramp, a
CPU-driven HPA holds at its floor while the queue and time-to-first-token climb
unchecked. A queue-driven HPA starts adding replicas the moment
num_requests_waiting crosses roughly five per pod — well before
any utilization gauge would have twitched — and the queue drains as capacity
comes online. You are scaling on cause (work arriving faster than it is served)
instead of on a lagging, ambiguous symptom.
Two caveats worth stating honestly. First, autoscaling cannot outrun cold
starts: if a fresh GPU pod takes 90 seconds to serve traffic, keep a small warm
buffer (a slightly higher minReplicas, or a low-priority placeholder
pod) so the first burst does not wait on a cold node. Second, the queue signal is
only as fresh as your slowest link — scrape interval plus adapter propagation
plus the HPA's own 15-second sync period. Tighten the scrape, and know that your
control loop reacts in tens of seconds, not instantly.
Lesson
Autoscale on the signal that moves with your SLO, not on whichever gauge is easiest to reach. For request-serving systems that signal is almost always queue depth or concurrency — Little's Law, not CPU percent — and for LLM inference the server hands it to you for free. Device utilization answers "is the hardware busy?" Your users are asking a different question: "am I waiting?" Scale on the one they are actually asking.
Hitting something like this in production? I help teams with performance engineering, SRE/observability, and AI-driven root cause analysis — work with me.
Comments
Post a Comment