Classifying 400 million documents: LLM inference is a throughput problem, not a latency one
Scribd recently reported classifying more than 400 million documents with Gemini batch inference. The headline number is impressive, but the more useful thing it tells you is a design constraint: you cannot do 400 million anythings by sending one request at a time and watching a p99 latency graph. I've watched teams start a job like this the way they'd build a chat endpoint — one HTTP call per document, modest concurrency, retries with backoff, a dashboard tracking per-request latency. The pipeline works on a test set of a few hundred rows. Then someone points it at the full corpus, the ETA comes back measured in weeks or months, and the expensive GPU behind the model sits at 15% utilization the whole time. Nothing is broken. You're just optimizing for the wrong number. Two jobs that pull the knobs in opposite directions Interactive serving and batch inference look like the same task — "run text through the model" — but they optimize for opposite metrics. ...