Posts

One Endpoint, Every Model: Body-Based Routing for a Self-Hosted AI Gateway

Count the inference endpoints in your org. There's the self-hosted vLLM deployment on GKE serving the model you fine-tuned. There's a managed endpoint — Vertex, Bedrock, whatever — for the one you didn't want to operate. There's OpenAI or Anthropic for the frontier models you haven't brought in-house. Three backends, three URLs, three auth schemes, and every one of them hardcoded into a dozen client services. That sprawl is not a cosmetic problem. It has concrete, recurring failure modes: Adding or renaming a model means a coordinated client redeploy, because the endpoint URL lives in application config. There's no shared rate limit, so one team's overnight batch job saturates the GPU pool and interactive traffic's time-to-first-token falls off a cliff — with no throttle to protect it. Retry, timeout, and fallback logic is reimplemented (badly, inconsistently) in every client. When finance asks which team spent the inference budge...

Node Swap Is GA in Kubernetes. Idle Agent Pods Are Exactly What It's For — and Exactly How to Get Burned.

Memory is the first hard limit most Kubernetes clusters hit. Nodes run out of RAM long before they run out of CPU, and the scheduler stops placing pods the moment requests exhaust allocatable memory — even if the CPUs are half-idle. You end up paying for cores you can't use because there's no RAM left to hand out. Agentic AI workloads make this worse in a specific, annoying way. A typical agent pod loads a language runtime, maybe a model client, maybe a sandbox for executing untrusted code — a few gigabytes of resident footprint — spikes on startup, and then sits idle waiting for the next prompt. That memory is reserved and resident, but it's cold most of the time. Multiply by a fleet of agents and you've got a node that's RAM-bound on memory that nobody is actually touching. Kubernetes node swap just graduated to GA, and this is precisely the workload it was built for. It's also a very good way to wreck your tail latency if you treat it as free memory. Her...

Page-Chunked Classification Is Cheap and Predictable. It's Blind to the Table That Crosses a Page.

The cost model is a page. The document isn't. The last post on batch-classifying 400 million documents got a sharp question in the comments: Gemini's native PDF handling gives you a fixed, predictable token cost per page — 258 tokens per page is the actual number Google publishes — so a 50-page document costs roughly 50x a 1-page one. That's what makes costing 400 million documents a multiplication instead of a guess. The question was: what happens to a table that starts on page 12 and ends on page 13, once you're chunking and classifying by page? Nothing good, by default. And the failure is quiet — it doesn't throw an error, it just returns a confident, wrong label. A page break is a rendering artifact, not a semantic one Nothing in the PDF format guarantees that a table, a clause, or even a sentence finishes before the page does. A contract's indemnification table can run four rows on page 12 and six more on page 13 with no repeated header row — it's...

The Cluster Is 60% Idle and the Important Job Has Waited Three Days: Gang Scheduling for GPU Training

The cluster is 60% idle and the important job has waited three days Here is a combination of facts that should not be able to coexist, but does, on GPU clusters everywhere. The accelerator bill is enormous. The "GPU allocated" line on the dashboard sits around 60%. And the one training run the research lead actually cares about has been stuck in Pending for three days while a dozen tiny experiments churn away on the hardware it needs. Everyone's first instinct is to buy more GPUs. But this is almost never a capacity problem. A cluster that is 60% allocated has room; the important job can't get in anyway. This is a scheduling problem, and you can usually fix it without adding a single accelerator. AI21, writing about their move to Google Cloud's AI Hypercomputer, reported cutting high-priority job wait times from 72 hours to 12 and manual scheduling interventions from roughly 20 per week to zero — the kind of swing you get from fixing how work is admitted, no...

The Agent That Got Slower the Longer It Ran: Two-Tier Memory for Long-Horizon Agents

The agent that remembered everything, and slowed to a crawl Here is a failure mode that shows up on a schedule once you give an agent persistent memory. In week one it feels sharp: it recalls what you told it yesterday, it picks up a multi-day task where it left off, memory lookups are a few milliseconds and nobody thinks about them. By week six the same agent is noticeably slower on every turn, and — more insidiously — its answers have started drifting toward stale or irrelevant facts it dug up from months ago. The reflex is to blame the vector index, or the model, or to move to a bigger database instance. Usually none of those is the real problem. The problem is architectural, and it is almost always the same one: all of the agent's memory lives in a single flat, append-only table that nothing ever consolidates or evicts. That design degrades on two axes at once, and you have to fix both. Why a flat memory table degrades on two axes The naive design is seductive because ...

Your Agent Issued the Refund Twice: Durable Side Effects with the Transactional Outbox

Your support agent decides a customer is owed a refund. It calls the issue_refund tool, the payment provider returns 200, and the agent writes status = "refunded" to its own task row. Clean. Except on about one run in a few thousand, the customer gets refunded twice — or, on a different unlucky run, the agent swears it refunded them and the money never moved. Nothing in the model is broken. The prompt is fine, the tool schema is fine, the eval suite is green. What's broken is older and more boring than anything about LLMs: the agent is doing two writes to two different systems and pretending they happen together. They don't. The symptom You notice it from a reconciliation job, not from the agent. A finance export shows two refund transactions against one order ID, minutes apart, both tagged with the same agent run. Or a customer emails back asking where their money is, and the trace shows the tool call succeeded. When you pull the logs for the double case, yo...

Most of your prompts don't need the frontier model: building a difficulty-routed LLM cascade

The symptom is a line on the billing dashboard that tracks your request count almost perfectly: inference spend climbing in lockstep with traffic, every request costing roughly the same. That flat per-request cost is the tell. It means every prompt — the one-line sentiment check and the twelve-step tool-calling reasoning chain — is being answered by the same expensive model. Your traffic is almost never that uniform. In most production LLM workloads the difficulty distribution is heavily skewed: a large fraction of requests are easy enough that a 7B open model gets them right, and a small tail genuinely needs a frontier model. Paying frontier prices for the whole distribution is the expensive default. A cascade is how you stop. Why one model for everything is the wrong shape The instinct to standardize on a single strong model is understandable — one integration, one prompt to maintain, one set of behaviors to reason about. But it couples your cost to your worst request rather t...

Classifying 400 million documents: LLM inference is a throughput problem, not a latency one

Scribd recently reported classifying more than 400 million documents with Gemini batch inference. The headline number is impressive, but the more useful thing it tells you is a design constraint: you cannot do 400 million anythings by sending one request at a time and watching a p99 latency graph. I've watched teams start a job like this the way they'd build a chat endpoint — one HTTP call per document, modest concurrency, retries with backoff, a dashboard tracking per-request latency. The pipeline works on a test set of a few hundred rows. Then someone points it at the full corpus, the ETA comes back measured in weeks or months, and the expensive GPU behind the model sits at 15% utilization the whole time. Nothing is broken. You're just optimizing for the wrong number. Two jobs that pull the knobs in opposite directions Interactive serving and batch inference look like the same task — "run text through the model" — but they optimize for opposite metrics. ...

Your Inference Autoscaler Watches CPU. Your Users Feel the Queue.

You ramp traffic onto an LLM serving deployment and the graph does not add up. Time-to-first-token, flat at 300ms all morning, starts climbing: 900ms, 2s, 5s. Users are waiting. And the HorizontalPodAutoscaler that is supposed to save you is sitting at 2 replicas, perfectly calm, because the metric it watches — CPU — reads 22%. Nothing is broken, exactly. The autoscaler is doing precisely what you told it to do. You just told it to watch the wrong number. The graph that doesn't add up Here is the HPA's view of the world during the incident: $ kubectl describe hpa vllm-llama Name: vllm-llama Reference: Deployment/vllm-llama Metrics: ( current / target ) resource cpu on pods (as a percentage of request): 22% (110m) / 70% Min replicas: 2 Max replicas: 12 Deployment pods: 2 current / 2 desired Conditions: Type Status Reason Message ScalingActive True ValidMetricFound the HPA was able to compute a target ...