Your Inference Autoscaler Watches CPU. Your Users Feel the Queue.

You ramp traffic onto an LLM serving deployment and the graph does not add up. Time-to-first-token, flat at 300ms all morning, starts climbing: 900ms, 2s, 5s. Users are waiting. And the HorizontalPodAutoscaler that is supposed to save you is sitting at 2 replicas, perfectly calm, because the metric it watches — CPU — reads 22%.

Nothing is broken, exactly. The autoscaler is doing precisely what you told it to do. You just told it to watch the wrong number.

The graph that doesn't add up

Here is the HPA's view of the world during the incident:

$ kubectl describe hpa vllm-llama
Name:            vllm-llama
Reference:       Deployment/vllm-llama
Metrics:         ( current / target )
  resource cpu on pods (as a percentage of request):  22% (110m) / 70%
Min replicas:    2
Max replicas:    12
Deployment pods: 2 current / 2 desired
Conditions:
  Type            Status  Reason              Message
  ScalingActive   True    ValidMetricFound    the HPA was able to compute a target
  AbleToScale     True    ReadyForNewScale    recommended size matches current size
Events:          <none>

CPU at 22% against a 70% target. From where the autoscaler stands, this deployment is nearly idle and there is nothing to do. Now here is the same moment from inside the inference server:

$ curl -s localhost:8000/metrics | grep -E 'vllm:num_requests|vllm:gpu_cache'
vllm:num_requests_running{model_name="llama-3.1-8b"} 48.0
vllm:num_requests_waiting{model_name="llama-3.1-8b"} 137.0
vllm:gpu_cache_usage_perc{model_name="llama-3.1-8b"} 0.94

Forty-eight requests decoding, one hundred and thirty-seven waiting in line, and the KV cache 94% full. The server is saturated and backlogged. The autoscaler has no idea, because none of that shows up as CPU.

Why utilization lies for inference

CPU on a GPU inference pod is close to meaningless — the work happens on the accelerator, and the CPU mostly shuffles tokens in and out. The obvious fix is to scale on GPU utilization instead. That is better, but it still lies, just more subtly.

GPU "utilization" (the DCGM_FI_DEV_GPU_UTIL most people reach for) measures whether a kernel was executing during a sample window. It says nothing about whether that work is keeping up. Token generation is memory-bandwidth bound during decode; a card can sit at 70% "utilization" while a queue builds behind it, because the bottleneck is KV cache capacity and memory bandwidth, not raw compute occupancy. You can be 100% utilized and healthy, or 65% utilized and drowning. Utilization tells you a device is busy. It does not tell you whether work is piling up.

The thing you actually care about has a name, and it is old: queueing. Little's Law says the number of requests in the system equals the arrival rate times the time each one spends there (L = λW). When arrivals outrun service capacity, the queue — num_requests_waiting — grows, and it grows before any utilization gauge pins. That queue length is a direct, early proxy for the latency your users feel. It is the signal your autoscaler should be chasing.

Scale on GPU utilization Scale on requests-waiting load replicas: 2 (flat) queue ↑ load replicas track load queue ≈ 0 Illustrative. Left: utilization stays "fine" while the queue and latency blow up. Right: queue-driven HPA adds replicas as work backs up, holding latency.

Scaling on the signal that tracks the SLO

Modern inference servers already export the right metric. vLLM publishes vllm:num_requests_waiting, vllm:num_requests_running, and vllm:gpu_cache_usage_perc; Text Generation Inference exposes tgi_queue_size and tgi_batch_current_size. The job is to get one of those in front of the autoscaler.

Step 1: scrape it

On GKE with Managed Service for Prometheus, a PodMonitoring resource is enough. Keep the scrape interval tight — it sets the floor on how fast your control loop can react.

apiVersion: monitoring.googleapis.com/v1
kind: PodMonitoring
metadata:
  name: vllm
spec:
  selector:
    matchLabels:
      app: vllm-llama
  endpoints:
  - port: metrics
    interval: 15s

Step 2: point the HPA at the queue, not the CPU

GKE's HPA can consume Managed Prometheus metrics directly, and recent GKE versions let you express the scaling signal as a native PromQL query rather than a bare metric name. The flap-prevention here doesn't come from smoothing the metric itself — a function like max_over_time holds the peak of its window, so it makes you react to a single noisy scrape immediately and keep reacting to it until the window ages out, which is the opposite of what you want on the way up. The real work is done by the HPA's own behavior block below: stabilizationWindowSeconds: 0 on scale-up means you react to a genuine backlog fast (queued GPU requests cost users, waiting is expensive), while stabilizationWindowSeconds: 300 on scale-down means you don't give that capacity back for five minutes — so a transient spike doesn't cause a scale-up, an immediate scale-down, then another scale-up a minute later.
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
  name: vllm-llama
spec:
  scaleTargetRef:
    apiVersion: apps/v1
    kind: Deployment
    name: vllm-llama
  minReplicas: 2
  maxReplicas: 12
  metrics:
  - type: External
    external:
      metric:
        # PromQL: smooth the gauge so one bad scrape doesn't flap the fleet
        name: prometheus.googleapis.com|vllm:num_requests_waiting|gauge
      target:
        # External + AverageValue divides by replica count:
        # "keep ~5 queued requests per replica"
        type: AverageValue
        averageValue: "5"
  behavior:
    scaleUp:
      stabilizationWindowSeconds: 0      # react to backlog immediately
      policies:
      - type: Pods
        value: 4                          # but at most +4 pods per 30s
        periodSeconds: 30
    scaleDown:
      stabilizationWindowSeconds: 300     # cool down slowly; GPUs are expensive to churn
      policies:
      - type: Pods
        value: 1
        periodSeconds: 60

The AverageValue target is the key line. With an External metric, the HPA divides the total by the current replica count, so "averageValue: 5" means keep roughly five requests queued per replica. Cross that and it scales up; drop well under it and it scales down. The target is now expressed in the same units as your problem — backlog — instead of a device-busy percentage that correlates with your SLO only by accident.

The asymmetric behavior block matters for GPUs specifically. Scale up fast, because a backlog turns into SLO burn in seconds. Scale down slowly, because a GPU pod is expensive to spin up (multi-gigabyte image pulls, weight loading, CUDA graph capture) and you do not want to pay that cold-start tax again in a minute because the queue briefly dipped.

What actually changes

The mechanics of the fix are almost boring: same HPA object, same Deployment, one metric swapped. But the behavior inverts. Under the same traffic ramp, a CPU-driven HPA holds at its floor while the queue and time-to-first-token climb unchecked. A queue-driven HPA starts adding replicas the moment num_requests_waiting crosses roughly five per pod — well before any utilization gauge would have twitched — and the queue drains as capacity comes online. You are scaling on cause (work arriving faster than it is served) instead of on a lagging, ambiguous symptom.

Two caveats worth stating honestly. First, autoscaling cannot outrun cold starts: if a fresh GPU pod takes 90 seconds to serve traffic, keep a small warm buffer (a slightly higher minReplicas, or a low-priority placeholder pod) so the first burst does not wait on a cold node. Second, the queue signal is only as fresh as your slowest link — scrape interval plus adapter propagation plus the HPA's own 15-second sync period. Tighten the scrape, and know that your control loop reacts in tens of seconds, not instantly.

Lesson

Autoscale on the signal that moves with your SLO, not on whichever gauge is easiest to reach. For request-serving systems that signal is almost always queue depth or concurrency — Little's Law, not CPU percent — and for LLM inference the server hands it to you for free. Device utilization answers "is the hardware busy?" Your users are asking a different question: "am I waiting?" Scale on the one they are actually asking.


Hitting something like this in production? I help teams with performance engineering, SRE/observability, and AI-driven root cause analysis — work with me.

Comments

Popular posts from this blog

Performance Testing 102: Little's Law and It's usage in Performance Testing

Performance Testing 104: Workload Modelling Designing & Process

Mastering the Art of Scaling in SaaS Applications