Your GPUs Are 100% Allocated and 35% Utilized: Finding the Data-Loading Wall
Here is a billing situation that should make anyone running a training cluster uncomfortable: eight H100s, 100% allocated, reserved for the whole job, and the accelerators themselves working about a third of the time. You pay for every second the hardware is yours. The hardware spends most of each step waiting for data to show up. The reason this goes unnoticed for so long is that the obvious metric lies to you. nvidia-smi prints a column called GPU-Util , and during a stalled training run it will happily read 95%. The job looks healthy. The loss goes down. Nobody opens a ticket. Meanwhile the effective throughput — samples per second per dollar — is less than half of what the same silicon could do. Why the headline utilization number is misleading The GPU-Util field from nvidia-smi is defined as the fraction of the sampling window during which at least one kernel was executing. That is a very low bar. A training step that spends 600ms blocked on the dataloader and 400ms doin...