The Cold-Start Tax on GPU Model Serving: Why 'Scale Up' Takes Minutes
Here is a failure mode that looks like an autoscaling bug but isn't. A traffic spike hits your model-serving deployment, the HPA (or KEDA, or your queue-depth scaler) correctly decides it needs three more replicas, and the new pods go Pending → ContainerCreating → eventually Running . The whole time, the queue keeps growing and p99 stays in the toilet. By the time the new replicas actually start serving tokens, the spike is half over. The scaler did its job in seconds. The pods took minutes to become useful. That gap — from "a replica was requested" to "that replica is serving real traffic at full speed" — is the cold-start tax, and on GPU model serving it is brutal: a multi-gigabyte container image, tens of gigabytes of weights, and a warmup pass all sit between the scale-up decision and the first good response. The usual reaction is to stop trusting the autoscaler and just run a fat idle buffer of hot replicas, which is the most expensive possible workarou...