Classifying 400 million documents: LLM inference is a throughput problem, not a latency one
Scribd recently reported classifying more than 400 million documents with Gemini batch inference. The headline number is impressive, but the more useful thing it tells you is a design constraint: you cannot do 400 million anythings by sending one request at a time and watching a p99 latency graph.
I've watched teams start a job like this the way they'd build a chat endpoint — one HTTP call per document, modest concurrency, retries with backoff, a dashboard tracking per-request latency. The pipeline works on a test set of a few hundred rows. Then someone points it at the full corpus, the ETA comes back measured in weeks or months, and the expensive GPU behind the model sits at 15% utilization the whole time. Nothing is broken. You're just optimizing for the wrong number.
Two jobs that pull the knobs in opposite directions
Interactive serving and batch inference look like the same task — "run text through the model" — but they optimize for opposite metrics.
- Interactive serving cares about time to first token and p99 latency per request. A user is waiting. You keep batch sizes small and queues short so no single request gets stuck behind others.
- Batch classification cares about total tokens per second per dollar and wall-clock across the whole corpus. Nobody is waiting on any individual document. A document that takes 400ms instead of 90ms is irrelevant if it means the GPU is fuller.
Every tuning decision follows from which of those you're doing. The batch job wants the latency you'd fight to avoid in serving, because that latency is the sign the accelerator is full.
Why one-request-at-a-time wastes the GPU
The reason batching helps so much comes straight from the memory roofline. Generating tokens one request at a time is memory-bandwidth bound: to produce each token you stream the entire model's weights out of HBM, do a tiny amount of arithmetic, and throw the compute units away idle. Arithmetic intensity is terrible.
Batching fixes exactly that. You read the weights once and apply them to hundreds of sequences in the same pass. The weight-load cost is amortized; throughput climbs almost linearly with batch size until you hit either the compute ceiling or the memory the KV cache needs. Classification is the friendliest possible case here: the output is short — often a single label token under constrained decoding — so the work is dominated by prefill, which is compute-bound and batches beautifully. You are paying almost entirely to read the prompt, not to generate.
The other half is how you batch. Naive static batching assembles a fixed group of N sequences and waits for the slowest one to finish before starting the next group — so a batch with one long document stalls the other N-1 slots. Continuous batching (iteration-level scheduling, as in vLLM or TGI) instead injects a new sequence into a free slot the moment one finishes. The accelerator stays saturated instead of draining at the end of every batch.
The two roads: run your own engine, or use a batch API
There are two credible ways to get this throughput, and the right choice is mostly about whether you already run GPUs.
Self-hosted: saturate the engine yourself
If you're serving your own model, an inference engine like vLLM already does continuous batching and paged KV-cache management. Your job is to feed it wide and constrain the output. The offline entry point takes the whole list of prompts and schedules them internally — you do not loop and call it once per document:
from vllm import LLM, SamplingParams
from vllm.sampling_params import GuidedDecodingParams
LABELS = ["invoice", "contract", "resume", "manual", "other"]
llm = LLM(
model="Qwen/Qwen2.5-7B-Instruct",
max_num_seqs=256, # sequences resident concurrently
max_num_batched_tokens=8192, # prefill token budget per step
gpu_memory_utilization=0.92, # headroom so the KV cache doesn't OOM
)
params = SamplingParams(
max_tokens=1, # one label token, nothing to generate
temperature=0.0,
guided_decoding=GuidedDecodingParams(choice=LABELS),
)
# Pass the whole shard at once; the scheduler batches it for you.
outputs = llm.generate(prompts, params)
Two of those knobs do most of the work. guided_decoding restricts the output to the label set, which both guarantees a parseable answer and cuts generation to a single token — so you're paying for prefill, not decode. max_num_seqs and max_num_batched_tokens set how wide the engine runs; push them up until GPU utilization is high but the KV cache isn't spilling. Watch nvidia-smi and the engine's own throughput logs, not a per-request latency panel.
Managed: rent throughput at half price
If you don't want to run GPUs, every major provider now has an asynchronous batch tier — Gemini batch mode, the OpenAI Batch API, Anthropic Message Batches, Bedrock batch inference. You upload a file of requests, get results back within a stated window (typically up to 24 hours), and pay roughly 50% of the synchronous price. You submit a manifest instead of streaming calls:
# One line per document, submitted as a single job, not 400M live calls.
{"custom_id": "doc-000001", "method": "POST", "url": "/v1/chat/completions",
"body": {"model": "...", "max_tokens": 1,
"messages": [{"role": "user", "content": "Classify: ..."}]}}
{"custom_id": "doc-000002", "method": "POST", "url": "/v1/chat/completions", ...}
The 50% discount exists precisely because you've told the provider "I don't care about latency" — which lets them pack your work into idle capacity. That's the same trade you're making internally; here you're just buying it.
The part that isn't about the model at all
At 400 million items, the model is the easy part. The pipeline around it is what decides whether the job finishes. A crash 30 hours in must not mean restarting from zero, and a re-run must not double-charge you for work already done. That means a durable work queue and idempotent writes, not a for-loop over a list in memory.
The rules that keep it honest:
- Idempotent writes keyed by
doc_id + prompt_version. Re-running the job skips anything already classified under the current prompt, and a prompt change re-does only what it should. - Checkpoint by shard. Workers claim a shard, process it, mark it done. A crashed worker's shard is reclaimed; finished shards are never reprocessed.
- A dead-letter queue. Documents that exceed the context window, fail repeatedly, or come back with low-confidence labels go to a side queue for a bigger model or a human — they don't stall or silently corrupt the run.
- Backpressure. Cap in-flight requests so you fill the engine without OOMing the KV cache. "As wide as possible" has a ceiling, and crossing it turns into cascading retries.
Pick the smallest model that clears the bar
The last lever is the biggest one for cost, and it's the least technical: use the smallest model that hits your accuracy target. Classification into a known label set is not a frontier-model task. A 7B–8B instruct model, or a small model few-shot-prompted or lightly fine-tuned on your labels, will often match a frontier model on a bounded classification problem at a fraction of the cost per token — and it batches more of itself onto each GPU because its weights and KV cache are smaller. Reserve the expensive model for the dead-letter queue, where the hard 1–2% actually needs it.
Stack the effects and the same corpus can move from "GPU idle at 15%, ETA in months" to "accelerator saturated, done over a weekend" — not from a cleverer model, but from batching wide, constraining the output, right-sizing the model, and buying the throughput at the batch-tier price. (Those framings are illustrative of the shape of the win, not measured figures from a specific run.)
Lesson
Large-scale LLM classification is a throughput problem wearing a serving problem's clothes. The moment nobody is waiting on any individual answer, per-request latency stops being the metric that matters and total tokens per second per dollar takes over — and almost every default from your interactive stack is now working against you. Measure the accelerator's utilization, not the request's latency, and most of the job's cost and wall-clock disappears before you've touched the model itself.
Hitting something like this in production? I help teams with performance engineering, SRE/observability, and AI-driven root cause analysis — work with me.
Comments
Post a Comment