Posts

Your Inference Autoscaler Watches CPU. Your Users Feel the Queue.

You ramp traffic onto an LLM serving deployment and the graph does not add up. Time-to-first-token, flat at 300ms all morning, starts climbing: 900ms, 2s, 5s. Users are waiting. And the HorizontalPodAutoscaler that is supposed to save you is sitting at 2 replicas, perfectly calm, because the metric it watches — CPU — reads 22%. Nothing is broken, exactly. The autoscaler is doing precisely what you told it to do. You just told it to watch the wrong number. The graph that doesn't add up Here is the HPA's view of the world during the incident: $ kubectl describe hpa vllm-llama Name: vllm-llama Reference: Deployment/vllm-llama Metrics: ( current / target ) resource cpu on pods (as a percentage of request): 22% (110m) / 70% Min replicas: 2 Max replicas: 12 Deployment pods: 2 current / 2 desired Conditions: Type Status Reason Message ScalingActive True ValidMetricFound the HPA was able to compute a target ...

The Agent Migrated Your YAML. It Didn't Migrate Your p99.

The green cutover with a red p99 The migration goes suspiciously well. An agent reads your EKS cluster, translates every manifest, generates the Terraform for the GKE side, and opens a change for review. You apply it. Every Deployment reports Ready , every synthetic check is green, the smoke tests pass. You cut traffic over. Then real traffic arrives and p99 latency climbs. Not by a rounding error — it roughly doubles on a couple of services and stays there. Nothing is crashing. No pod is CrashLoopBackOff . The dashboards the agent knew how to reason about are all green. The problem is that a Kubernetes manifest describes intent, and the agent proved manifest-equivalence, not behavior-equivalence. Those are different claims, and the gap between them is exactly the part of a migration that only shows up under load. What the agent is genuinely good at This is not a "don't use the tool" post. Agentic migration tooling — GKE's own EKS-to-GKE assistant, or a home-g...

Give Your Agent an Error Budget: SLOs for Hallucination and Tool-Call Failure

"It works" is not a number An LLM agent ships. It passed the demo, it passed the handful of prompts in the eval notebook, and for two weeks nobody complains. Then a support ticket arrives: the agent confidently told a customer about a refund policy that doesn't exist. You go looking for a dashboard that tells you how often that happens and there isn't one. You have latency graphs, you have token-cost graphs, and for the thing that actually matters — is the agent right — you have a vibe. That gap is the whole problem. We instrument agents like web services (RPS, p99, error rate) and then act surprised that none of those numbers move when the agent hallucinates. A 200 response carrying a made-up fact is still a 200. The reliability machinery you already run is fine; it's just pointed at the wrong signals. Reframe: agent quality is a reliability problem An error budget is just the inverse of a target: if you promise 99%, you have a 1% budget to spend. That f...

Round-Robin Is Wrecking Your LLM Throughput: Route by KV Cache, Not Connection Count

You went from 2 vLLM replicas to 8 because your queue was backing up. Eight GPUs now, four times the hardware. Aggregate throughput went up maybe 30%, and p99 time-to-first-token (TTFT) got worse . The per-GPU utilization graphs look busy, the autoscaler is happy, and the bill quadrupled for a fraction of the goodput you paid for. Before you reach for bigger GPUs, look at the thing in front of them. A plain round-robin or least-request load balancer is the wrong tool for LLM inference, and it fails in a way that's invisible on a CPU/GPU utilization dashboard. Why a normal load balancer is wrong here Web load balancing works because requests are roughly interchangeable: each one is short-lived, stateless, and costs about the same. Spraying them round-robin across identical replicas is close to optimal. LLM serving violates every one of those assumptions. Requests are not uniform. A request that generates 2,000 tokens occupies KV cache in GPU HBM for its entire decode l...

You Gave an Agent Your Database via MCP, and Now the Connection Pool Is on Fire

The traffic graph is flat. Nobody deployed a schema change, nobody launched a campaign, the request rate to your API looks exactly like it did last week. And yet the database is throwing FATAL: sorry, too many clients already , and pg_stat_activity is full of sessions parked in the idle in transaction state. The one thing that changed: yesterday you wired an AI agent to the database through an MCP server so it could answer questions over your data. This is becoming a common way to melt a perfectly healthy Postgres. The Model Context Protocol makes it a few lines of config to hand an agent a query tool backed by AlloyDB, Cloud SQL, or plain Postgres. What the tutorial doesn't mention is that an LLM is a very strange database client, and the connection pool you sized for a web app makes assumptions the agent quietly violates. I have seen this movie before, without the AI Early in my career I ran a web API backend on GCP that behaved fine in staging and then, under real load,...

Your Vector Search Got Faster and Recall Fell Off a Cliff — Nobody Got Paged

The latency graph looked like a win. Someone had tuned the vector index over the weekend, and p99 for semantic search dropped from around 200ms to 25ms. The dashboard was green, the change shipped, everyone moved on. Two weeks later a different graph started climbing: support tickets. Search "felt dumb." The RAG assistant kept missing answers that were obviously in the knowledge base — type in almost the exact wording of a document and it still wouldn't surface. Nothing had errored. Nothing had paged. The index was returning ten results for every query, fast, every time. They were just increasingly the wrong ten. Why approximate search fails quietly Any production vector search runs on an approximate nearest neighbor (ANN) index, because exact nearest-neighbor over millions of high-dimensional embeddings means scanning every row. ANN indexes buy their speed by not looking at most of your data on each query — and every one of them exposes a knob that controls exact...

Your AI Agent Spends 8 Seconds Booting a Sandbox to Run 300ms of Code

An agent answers a question that requires it to run a three-line Python snippet — parse a CSV, sum a column, return the total. The code runs in about 300 milliseconds. The user waits eight seconds. Trace the request and the actual computation is a rounding error; everything else is the agent waiting for somewhere to run the code. The bill tells the same story from the other side. You provisioned a generously sized pool of execution environments so agents never queue, and the utilization dashboard shows them sitting idle most of the day, warm and billable, waiting for the next tool call. You are paying full compute prices to keep empty rooms lit. Both symptoms — the latency and the cost — come from the same design decision: one fresh, full-sized sandbox per agent, cold on the way in and idle on the way out. Where the eight seconds go An agent that executes code or calls tools needs an isolated environment to do it in — you cannot run model-generated code in ...

Your Inference GPUs Are Starved, Not Slow: Finding the Idle Time in Your AI Bill

It usually shows up as a billing question, not an alert. The GPU line on the cloud bill has climbed quarter over quarter while throughput has stayed flat, and someone from finance wants to know why. You open the accelerator dashboard expecting pegged GPUs and instead find them hovering around 25–30% utilization. You are renting some of the most expensive compute available and using a third of it. This is the most common failure mode in production AI infrastructure right now, and it is rarely a "buy fewer GPUs" problem. The GPUs are not slow. They are starved, fragmented, or idle — and every one of those is fixable without touching the model. Step 1: distrust the utilization number you have The first mistake is trusting nvidia-smi . Its headline GPU-Util field does not mean what most people assume. The NVIDIA docs define it as the percent of time over the sample window during which one or more kernels was executing. A single tiny kernel copying data counts the same a...

Your Dual-Write Migration Is Lying to You: Catching Silent Divergence Before Cutover

You flip the read path to the new database. Within a minute, a fraction of a percent of lookups start returning 404 , and a handful return data that's a few minutes stale. The write path has been dual-writing to both stores for three weeks. Every write returned 200 . The backfill job logged "complete." By every dashboard you had, the migration was done. It wasn't. The old store and the new store had been quietly drifting apart the entire time, and the cutover is exactly the moment you discover it — when the new store becomes the source of truth for reads and its gaps become user-visible. Why "both writes succeeded" doesn't mean "both stores match" The seductive thing about dual-write is that it looks correct locally. Your write path does something like this: def save_order(order): old_db.upsert(order) # legacy Postgres new_db.upsert(order) # Spanner return 200 Both calls return, you return 200 , and the r...