The Cold-Start Tax on GPU Model Serving: Why 'Scale Up' Takes Minutes

Here is a failure mode that looks like an autoscaling bug but isn't. A traffic spike hits your model-serving deployment, the HPA (or KEDA, or your queue-depth scaler) correctly decides it needs three more replicas, and the new pods go Pending → ContainerCreating → eventually Running. The whole time, the queue keeps growing and p99 stays in the toilet. By the time the new replicas actually start serving tokens, the spike is half over.

The scaler did its job in seconds. The pods took minutes to become useful. That gap — from "a replica was requested" to "that replica is serving real traffic at full speed" — is the cold-start tax, and on GPU model serving it is brutal: a multi-gigabyte container image, tens of gigabytes of weights, and a warmup pass all sit between the scale-up decision and the first good response. The usual reaction is to stop trusting the autoscaler and just run a fat idle buffer of hot replicas, which is the most expensive possible workaround on hardware that bills by the GPU-second.

Measure the phases before you fix anything

"Scale-up is slow" is not actionable. Time-to-ready is a sum of distinct phases, and each has a different fix. Break it down first. The pod's own condition transitions and container state give you most of it without any extra tooling:

# When did the pod schedule, start pulling, and go Ready?
kubectl get pod $POD -o jsonpath='{range .status.conditions[*]}{.type}{"\t"}{.lastTransitionTime}{"\n"}{end}'
#   PodScheduled     2026-10-09T18:02:01Z
#   Initialized      2026-10-09T18:02:05Z
#   ContainersReady  2026-10-09T18:08:40Z   <-- the long pole lives in here
#   Ready            2026-10-09T18:08:41Z

# Image pull duration shows up as an event with the pulled size + elapsed time
kubectl get events --field-selector involvedObject.name=$POD \
  | grep -E 'Pulling|Pulled|Started'
#   Pulling   image "myrepo/vllm-serving:cu124"
#   Pulled    ... in 3m12s (3m12s including waiting)   <-- image pull alone

Lay those timestamps on a line and the shape is almost always the same. A cold scale-up onto a brand-new node splits into four stages, and two of them dominate:

time-to-ready, cold scale-up onto a new node (illustrative) 0s ~3 min ~6.5 min node provision image pull multi-GB CUDA image weight load tens of GB from object store warmup CUDA graphs

Node provisioning only appears when the cluster autoscaler has to add a node, but when it does it is unavoidable minutes of VM boot plus GPU driver and device-plugin readiness. The image pull and the weight load are the two you can actually attack, and together they are usually three-quarters of the wall-clock.

Image pull: stop downloading gigabytes you won't touch at startup

A typical inference image — a CUDA base, PyTorch, the serving framework, and assorted system libraries — runs well past 5 GB, often 10 GB+. On a cold node the container runtime pulls and decompresses the whole thing before the main process starts, even though serving touches a small fraction of those files at boot.

The highest-leverage fix is lazy image pulling, where the container starts against a mounted image and pages in file contents on demand. On GKE that is Image Streaming; on EKS it is the SOCI snapshotter; both read from an eStargz/SOCI-indexed image so the runtime fetches only the bytes actually read:

# GKE: Image Streaming is a cluster/node-pool feature, not a pod field.
gcloud container clusters create serving \
  --enable-image-streaming

# The image must live in Artifact Registry; first pull on a node seeds the
# node-local cache, subsequent pods on that node start near-instantly.

Two things that compound with it, and cost nothing: slim the image so there is less to stream in the first place (a multi-stage build that ships the runtime without the full build toolchain routinely halves it), and bake the image onto a secondary boot disk so it is already resident on the node at boot instead of pulled over the network. A preloaded disk image turns the image-pull bar in that waterfall into essentially zero on every new node.

Weight load: the real long pole

Even with the image handled, a 13B model in fp16 is ~26 GB and a quantized 70B is well into the tens of gigabytes. The default path — a single streamed download from object storage, deserialized on the CPU, then copied to the GPU — leaves almost all of your available throughput on the table because it is one connection doing one thing at a time.

Attack it on three fronts: parallelize the transfer, stage onto local NVMe so a replacement pod on the same node never re-downloads, and use a loader that streams tensors straight to the GPU instead of round-tripping through a slow Python deserialization step. First, the parallel pull into a local-SSD scratch volume, run as an init container:

# Local NVMe scratch, shared between the init and main containers.
volumes:
  - name: weights
    emptyDir:
      medium: ""          # node local SSD, NOT tmpfs/RAM
      sizeLimit: 200Gi

initContainers:
  - name: fetch-weights
    image: myrepo/s5cmd:latest
    args:
      # Many parallel connections saturate the NIC instead of trickling
      # one stream; this is often a 5-10x speedup over a single-threaded copy.
      - "--numworkers=64"
      - "cp"
      - "s3://models/llama-70b-awq/*"
      - "/weights/"
    volumeMounts:
      - { name: weights, mountPath: /weights }

Then load from that local copy with a streaming deserializer. Safetensors is already zero-copy via mmap, but under heavy cold-start pressure a purpose-built streamer such as Tensorizer or the run:ai model streamer overlaps the read, the deserialize, and the host-to-device copy so weights land in VRAM as they arrive rather than in three sequential passes:

# safetensors: mmap means the file isn't copied into Python heap first.
from safetensors.torch import load_file
state = load_file("/weights/model.safetensors", device="cuda:0")

# Tensorizer: stream tensors directly to the GPU, overlapping I/O and copy.
from tensorizer import TensorDeserializer
des = TensorDeserializer("/weights/model.tensors", device="cuda:0")
des.load_into_module(model)   # VRAM fills as bytes arrive, not after

Don't let the load balancer route to a pod that isn't warm

The last trap is subtle: a pod can be Running and accepting connections while it is still capturing CUDA graphs or JIT-compiling kernels on its first few requests. If your readiness gate is a plain TCP check, the load balancer starts sending real traffic into a replica whose first responses are pathologically slow — you scaled up and p99 still got worse. Gate readiness on an endpoint that only returns 200 after a warmup inference has completed:

readinessProbe:
  httpGet:
    path: /health/ready     # returns 200 only after one warmup pass
    port: 8000
  initialDelaySeconds: 5
  periodSeconds: 5
  failureThreshold: 60      # allow minutes for cold weight load, then fail
startupProbe:
  httpGet: { path: /health/ready, port: 8000 }
  periodSeconds: 10
  failureThreshold: 60      # keep the slow first boot from being killed

What the fixes add up to

Each lever hits a different bar in the waterfall: image streaming and a preloaded disk collapse the pull, parallel fetch plus a streaming loader collapse the weight load, and the local-SSD cache means the second pod on a node skips both almost entirely. The remaining floor is node provisioning, which you buy down separately — keep a small warm node pool, or run low-priority placeholder pods that a real serving pod can preempt instantly, so a scale-up lands on a node that already exists. The shape of the win, before and after (illustrative, not telemetry from one run):

0 ~3.5 min ~7 min before after node image weights warmup

Lesson

Autoscaling responsiveness is not a property of the autoscaler — it is a property of how fast a new replica becomes useful, and on GPU serving that number is dominated by two things the scaler never sees: the image pull and the weight load. Measure time-to-ready as a waterfall of named phases before you touch anything, because "scale-up is slow" points at the wrong component. Fix the two long poles with lazy image pulling and a streaming, parallel, locally-cached weight load, and the fat idle buffer of hot replicas — the real cost of a slow cold start — is something you can finally shrink instead of pay for around the clock.


Hitting something like this in production? I help teams with performance engineering, SRE/observability, and AI-driven root cause analysis — work with me.

Comments

Popular posts from this blog

Performance Testing 102: Little's Law and It's usage in Performance Testing

Performance Testing 104: Workload Modelling Designing & Process

Mastering the Art of Scaling in SaaS Applications