The Cluster Is 60% Idle and the Important Job Has Waited Three Days: Gang Scheduling for GPU Training
The cluster is 60% idle and the important job has waited three days
Here is a combination of facts that should not be able to coexist, but
does, on GPU clusters everywhere. The accelerator bill is enormous. The
"GPU allocated" line on the dashboard sits around 60%. And the one training
run the research lead actually cares about has been stuck in Pending
for three days while a dozen tiny experiments churn away on the hardware it
needs.
Everyone's first instinct is to buy more GPUs. But this is almost never a capacity problem. A cluster that is 60% allocated has room; the important job can't get in anyway. This is a scheduling problem, and you can usually fix it without adding a single accelerator. AI21, writing about their move to Google Cloud's AI Hypercomputer, reported cutting high-priority job wait times from 72 hours to 12 and manual scheduling interventions from roughly 20 per week to zero — the kind of swing you get from fixing how work is admitted, not from more silicon.
What's actually happening in the queue
Start with the symptom the right way. "GPU allocated" and "GPU utilized" are different numbers, and the gap tells you which problem you have.
# What the scheduler thinks is in use (requests on running pods)
kubectl describe nodes | grep -A4 "Allocated resources" | grep nvidia
# What the hardware is actually doing, sampled across the fleet
nvidia-smi --query-gpu=utilization.gpu,memory.used --format=csv -l 5
If allocation is high but utilization is low, you have the inference-style starvation problem — small batches, bad input pipelines, whole GPUs pinned by tiny models. That's a different post. The situation here is the opposite and more frustrating: utilization is reasonable on the jobs that are running, but a large, high-value job can't start even though the node count on paper is sufficient.
Look at why the big job is pending:
kubectl get pods -l job-name=pretrain-run-7 -o wide
# NAME READY STATUS NODE
# pretrain-run-7-0 0/1 Pending <none>
# pretrain-run-7-1 0/1 Pending <none>
# ... 14 more Pending
kubectl get events --field-selector reason=FailedScheduling | head
# 0/40 nodes available: 37 Insufficient nvidia.com/gpu, 3 node(s) had taint...
The job needs 16 GPUs placed together. Across the cluster there are 14
free, scattered three-here, two-there across nodes that are otherwise full of
single-GPU experiments. The default Kubernetes scheduler is a
pod-by-pod scheduler: it will happily place 14 of the 16 pods, hold
two in Pending, and wait. Those 14 pods now sit on GPUs doing
nothing — a distributed training job does no useful work until all
ranks are up — while they block the very capacity that might let another job
finish and free the last two. That's the idle-but-allocated paradox, and it's
a textbook partial-allocation deadlock.
Root cause: a batch workload on a service scheduler
The default scheduler was designed for long-running services where any pod can land on any node whenever a slot opens. Batch ML training breaks three of its assumptions at once:
- All-or-nothing. 15 of 16 ranks is worth exactly zero. The scheduler has no concept of "admit the whole thing or none of it."
- Fair-share and priority. Twenty small jobs submitted first will hold the hardware indefinitely; there's no quota or preemption model that says "this pretraining run outranks a hyperparameter sweep."
- Topology. 16 GPUs spread across 16 nodes connected by regular networking will train a large model far slower than 16 GPUs on two NVLink-connected hosts. The scheduler treats them as interchangeable.
You don't fix this by tuning the default scheduler. You put a batch admission layer in front of it. On Kubernetes the increasingly standard answer is Kueue, a job queueing controller that holds workloads until their entire resource footprint can be reserved, enforces quotas, and does priority preemption.
The fix: quota, gang admission, and priority
Kueue's model is three objects. A ResourceFlavor describes a
class of hardware. A ClusterQueue owns a quota of that hardware
and sets preemption rules. A LocalQueue is the namespaced handle
teams submit against. Jobs don't get scheduled directly — they get
admitted as a unit, all-or-nothing, against the quota.
apiVersion: kueue.x-k8s.io/v1beta1
kind: ResourceFlavor
metadata:
name: a100-80gb
spec:
nodeLabels:
cloud.google.com/gke-accelerator: nvidia-a100-80gb
---
apiVersion: kueue.x-k8s.io/v1beta1
kind: ClusterQueue
metadata:
name: training
spec:
namespaceSelector: {}
cohort: shared-gpu # queues in a cohort can borrow each other's idle quota
preemption:
withinClusterQueue: LowerPriority # a big job can evict smaller ones here
reclaimWithinCohort: Any # reclaim quota lent to other queues
resourceGroups:
- coveredResources: ["nvidia.com/gpu"]
flavors:
- name: a100-80gb
resources:
- name: "nvidia.com/gpu"
nominalQuota: 32
Now give the pretraining run a priority that outranks sweeps, and submit it against the queue:
apiVersion: kueue.x-k8s.io/v1beta1
kind: WorkloadPriorityClass
metadata:
name: high-priority
value: 10000
---
apiVersion: batch/v1
kind: Job
metadata:
name: pretrain-run-7
labels:
kueue.x-k8s.io/queue-name: training
kueue.x-k8s.io/priority-class: high-priority
spec:
parallelism: 16
completions: 16
suspend: true # Kueue unsuspends only once the whole job is admitted
template:
spec:
containers:
- name: trainer
image: your-registry/trainer:latest
resources:
limits:
nvidia.com/gpu: 1
The suspend: true field is the mechanism that kills the
partial-allocation deadlock. Kueue keeps the Job suspended — zero pods
created — until it can reserve all 16 GPUs of quota at once. No more 14 ranks
sitting idle waiting for two. Watch it work from the queue's point of view,
not the pod's:
kubectl get workloads
# NAME QUEUE RESERVED ADMITTED AGE
# pretrain-run-7-a1b2 training False 20s # waiting for quota
# sweep-trial-41-c3d4 training training True 4m # candidate for preemption
kubectl get clusterqueue training -o jsonpath='{.status.flavorsUsage}'
Because the ClusterQueue allows withinClusterQueue: LowerPriority
preemption, the high-priority workload doesn't wait behind the sweeps — Kueue
evicts enough lower-priority workloads (they get re-suspended and requeued,
not killed-for-good) to free the 16 GPUs, then admits the pretraining run as a
unit. The 72-hour wait becomes minutes because the question changed from "is
there a free GPU right now?" to "which running work is least important?"
Two refinements matter in production. First, turn on waitForPodsReady
in the Kueue config so that if the underlying pods can't all become ready
(image pull failure, a bad node), the whole workload is requeued instead of
limping along half-up. Second, for the topology problem, newer Kueue supports
Topology Aware Scheduling: annotate the pod set so ranks are packed onto
close-together nodes rather than scattered.
template:
metadata:
annotations:
# pack this job's pods within one rack/block where possible
kueue.x-k8s.io/podset-preferred-topology: "cloud.google.com/gce-topology-block"
Before and after, in one picture
Lesson
When a GPU cluster feels full but the important work won't start, measure allocation against utilization before you sign a hardware PO. If the gap is "allocated-but-idle ranks waiting on a partial placement," you have a batch scheduler missing from an infrastructure that needs one. Gang admission, quota, and priority preemption turn "is a GPU free right now?" into "what is the least important thing running?" — and that single reframing is usually worth more than the next tranche of accelerators.
Hitting something like this in production? I help teams with performance engineering, SRE/observability, and AI-driven root cause analysis — work with me.
Comments
Post a Comment