The Cluster Is 60% Idle and the Important Job Has Waited Three Days: Gang Scheduling for GPU Training

The cluster is 60% idle and the important job has waited three days

Here is a combination of facts that should not be able to coexist, but does, on GPU clusters everywhere. The accelerator bill is enormous. The "GPU allocated" line on the dashboard sits around 60%. And the one training run the research lead actually cares about has been stuck in Pending for three days while a dozen tiny experiments churn away on the hardware it needs.

Everyone's first instinct is to buy more GPUs. But this is almost never a capacity problem. A cluster that is 60% allocated has room; the important job can't get in anyway. This is a scheduling problem, and you can usually fix it without adding a single accelerator. AI21, writing about their move to Google Cloud's AI Hypercomputer, reported cutting high-priority job wait times from 72 hours to 12 and manual scheduling interventions from roughly 20 per week to zero — the kind of swing you get from fixing how work is admitted, not from more silicon.

What's actually happening in the queue

Start with the symptom the right way. "GPU allocated" and "GPU utilized" are different numbers, and the gap tells you which problem you have.

# What the scheduler thinks is in use (requests on running pods)
kubectl describe nodes | grep -A4 "Allocated resources" | grep nvidia

# What the hardware is actually doing, sampled across the fleet
nvidia-smi --query-gpu=utilization.gpu,memory.used --format=csv -l 5

If allocation is high but utilization is low, you have the inference-style starvation problem — small batches, bad input pipelines, whole GPUs pinned by tiny models. That's a different post. The situation here is the opposite and more frustrating: utilization is reasonable on the jobs that are running, but a large, high-value job can't start even though the node count on paper is sufficient.

Look at why the big job is pending:

kubectl get pods -l job-name=pretrain-run-7 -o wide
# NAME                 READY   STATUS    NODE
# pretrain-run-7-0     0/1     Pending   <none>
# pretrain-run-7-1     0/1     Pending   <none>
# ... 14 more Pending

kubectl get events --field-selector reason=FailedScheduling | head
# 0/40 nodes available: 37 Insufficient nvidia.com/gpu, 3 node(s) had taint...

The job needs 16 GPUs placed together. Across the cluster there are 14 free, scattered three-here, two-there across nodes that are otherwise full of single-GPU experiments. The default Kubernetes scheduler is a pod-by-pod scheduler: it will happily place 14 of the 16 pods, hold two in Pending, and wait. Those 14 pods now sit on GPUs doing nothing — a distributed training job does no useful work until all ranks are up — while they block the very capacity that might let another job finish and free the last two. That's the idle-but-allocated paradox, and it's a textbook partial-allocation deadlock.

Root cause: a batch workload on a service scheduler

The default scheduler was designed for long-running services where any pod can land on any node whenever a slot opens. Batch ML training breaks three of its assumptions at once:

  • All-or-nothing. 15 of 16 ranks is worth exactly zero. The scheduler has no concept of "admit the whole thing or none of it."
  • Fair-share and priority. Twenty small jobs submitted first will hold the hardware indefinitely; there's no quota or preemption model that says "this pretraining run outranks a hyperparameter sweep."
  • Topology. 16 GPUs spread across 16 nodes connected by regular networking will train a large model far slower than 16 GPUs on two NVLink-connected hosts. The scheduler treats them as interchangeable.

You don't fix this by tuning the default scheduler. You put a batch admission layer in front of it. On Kubernetes the increasingly standard answer is Kueue, a job queueing controller that holds workloads until their entire resource footprint can be reserved, enforces quotas, and does priority preemption.

The fix: quota, gang admission, and priority

Kueue's model is three objects. A ResourceFlavor describes a class of hardware. A ClusterQueue owns a quota of that hardware and sets preemption rules. A LocalQueue is the namespaced handle teams submit against. Jobs don't get scheduled directly — they get admitted as a unit, all-or-nothing, against the quota.

apiVersion: kueue.x-k8s.io/v1beta1
kind: ResourceFlavor
metadata:
  name: a100-80gb
spec:
  nodeLabels:
    cloud.google.com/gke-accelerator: nvidia-a100-80gb
---
apiVersion: kueue.x-k8s.io/v1beta1
kind: ClusterQueue
metadata:
  name: training
spec:
  namespaceSelector: {}
  cohort: shared-gpu          # queues in a cohort can borrow each other's idle quota
  preemption:
    withinClusterQueue: LowerPriority      # a big job can evict smaller ones here
    reclaimWithinCohort: Any               # reclaim quota lent to other queues
  resourceGroups:
  - coveredResources: ["nvidia.com/gpu"]
    flavors:
    - name: a100-80gb
      resources:
      - name: "nvidia.com/gpu"
        nominalQuota: 32

Now give the pretraining run a priority that outranks sweeps, and submit it against the queue:

apiVersion: kueue.x-k8s.io/v1beta1
kind: WorkloadPriorityClass
metadata:
  name: high-priority
value: 10000
---
apiVersion: batch/v1
kind: Job
metadata:
  name: pretrain-run-7
  labels:
    kueue.x-k8s.io/queue-name: training
    kueue.x-k8s.io/priority-class: high-priority
spec:
  parallelism: 16
  completions: 16
  suspend: true           # Kueue unsuspends only once the whole job is admitted
  template:
    spec:
      containers:
      - name: trainer
        image: your-registry/trainer:latest
        resources:
          limits:
            nvidia.com/gpu: 1

The suspend: true field is the mechanism that kills the partial-allocation deadlock. Kueue keeps the Job suspended — zero pods created — until it can reserve all 16 GPUs of quota at once. No more 14 ranks sitting idle waiting for two. Watch it work from the queue's point of view, not the pod's:

kubectl get workloads
# NAME                  QUEUE      RESERVED   ADMITTED   AGE
# pretrain-run-7-a1b2   training              False      20s   # waiting for quota
# sweep-trial-41-c3d4   training   training   True       4m    # candidate for preemption

kubectl get clusterqueue training -o jsonpath='{.status.flavorsUsage}'

Because the ClusterQueue allows withinClusterQueue: LowerPriority preemption, the high-priority workload doesn't wait behind the sweeps — Kueue evicts enough lower-priority workloads (they get re-suspended and requeued, not killed-for-good) to free the 16 GPUs, then admits the pretraining run as a unit. The 72-hour wait becomes minutes because the question changed from "is there a free GPU right now?" to "which running work is least important?"

Two refinements matter in production. First, turn on waitForPodsReady in the Kueue config so that if the underlying pods can't all become ready (image pull failure, a bad node), the whole workload is requeued instead of limping along half-up. Second, for the topology problem, newer Kueue supports Topology Aware Scheduling: annotate the pod set so ranks are packed onto close-together nodes rather than scattered.

  template:
    metadata:
      annotations:
        # pack this job's pods within one rack/block where possible
        kueue.x-k8s.io/podset-preferred-topology: "cloud.google.com/gce-topology-block"

Before and after, in one picture

BEFORE: pod-by-pod — 14 ranks idle, big job stuck Pending yellow = big job's ranks, placed but idle, blocking the cluster. gray free slots can't form a set of 16. AFTER: gang admission + priority preemption blue = the 16-rank job admitted as one unit and now training. lower-priority sweeps requeued, not lost. Same hardware. The change is admission policy, not capacity.

Lesson

When a GPU cluster feels full but the important work won't start, measure allocation against utilization before you sign a hardware PO. If the gap is "allocated-but-idle ranks waiting on a partial placement," you have a batch scheduler missing from an infrastructure that needs one. Gang admission, quota, and priority preemption turn "is a GPU free right now?" into "what is the least important thing running?" — and that single reframing is usually worth more than the next tranche of accelerators.


Hitting something like this in production? I help teams with performance engineering, SRE/observability, and AI-driven root cause analysis — work with me.

Comments

Popular posts from this blog

Performance Testing 102: Little's Law and It's usage in Performance Testing

Performance Testing 104: Workload Modelling Designing & Process

Mastering the Art of Scaling in SaaS Applications