Posts

The Agent That Got Slower the Longer It Ran: Two-Tier Memory for Long-Horizon Agents

The agent that remembered everything, and slowed to a crawl Here is a failure mode that shows up on a schedule once you give an agent persistent memory. In week one it feels sharp: it recalls what you told it yesterday, it picks up a multi-day task where it left off, memory lookups are a few milliseconds and nobody thinks about them. By week six the same agent is noticeably slower on every turn, and — more insidiously — its answers have started drifting toward stale or irrelevant facts it dug up from months ago. The reflex is to blame the vector index, or the model, or to move to a bigger database instance. Usually none of those is the real problem. The problem is architectural, and it is almost always the same one: all of the agent's memory lives in a single flat, append-only table that nothing ever consolidates or evicts. That design degrades on two axes at once, and you have to fix both. Why a flat memory table degrades on two axes The naive design is seductive because ...

Your Agent Issued the Refund Twice: Durable Side Effects with the Transactional Outbox

Your support agent decides a customer is owed a refund. It calls the issue_refund tool, the payment provider returns 200, and the agent writes status = "refunded" to its own task row. Clean. Except on about one run in a few thousand, the customer gets refunded twice — or, on a different unlucky run, the agent swears it refunded them and the money never moved. Nothing in the model is broken. The prompt is fine, the tool schema is fine, the eval suite is green. What's broken is older and more boring than anything about LLMs: the agent is doing two writes to two different systems and pretending they happen together. They don't. The symptom You notice it from a reconciliation job, not from the agent. A finance export shows two refund transactions against one order ID, minutes apart, both tagged with the same agent run. Or a customer emails back asking where their money is, and the trace shows the tool call succeeded. When you pull the logs for the double case, yo...

Most of your prompts don't need the frontier model: building a difficulty-routed LLM cascade

The symptom is a line on the billing dashboard that tracks your request count almost perfectly: inference spend climbing in lockstep with traffic, every request costing roughly the same. That flat per-request cost is the tell. It means every prompt — the one-line sentiment check and the twelve-step tool-calling reasoning chain — is being answered by the same expensive model. Your traffic is almost never that uniform. In most production LLM workloads the difficulty distribution is heavily skewed: a large fraction of requests are easy enough that a 7B open model gets them right, and a small tail genuinely needs a frontier model. Paying frontier prices for the whole distribution is the expensive default. A cascade is how you stop. Why one model for everything is the wrong shape The instinct to standardize on a single strong model is understandable — one integration, one prompt to maintain, one set of behaviors to reason about. But it couples your cost to your worst request rather t...

Classifying 400 million documents: LLM inference is a throughput problem, not a latency one

Scribd recently reported classifying more than 400 million documents with Gemini batch inference. The headline number is impressive, but the more useful thing it tells you is a design constraint: you cannot do 400 million anythings by sending one request at a time and watching a p99 latency graph. I've watched teams start a job like this the way they'd build a chat endpoint — one HTTP call per document, modest concurrency, retries with backoff, a dashboard tracking per-request latency. The pipeline works on a test set of a few hundred rows. Then someone points it at the full corpus, the ETA comes back measured in weeks or months, and the expensive GPU behind the model sits at 15% utilization the whole time. Nothing is broken. You're just optimizing for the wrong number. Two jobs that pull the knobs in opposite directions Interactive serving and batch inference look like the same task — "run text through the model" — but they optimize for opposite metrics. ...

Your Inference Autoscaler Watches CPU. Your Users Feel the Queue.

You ramp traffic onto an LLM serving deployment and the graph does not add up. Time-to-first-token, flat at 300ms all morning, starts climbing: 900ms, 2s, 5s. Users are waiting. And the HorizontalPodAutoscaler that is supposed to save you is sitting at 2 replicas, perfectly calm, because the metric it watches — CPU — reads 22%. Nothing is broken, exactly. The autoscaler is doing precisely what you told it to do. You just told it to watch the wrong number. The graph that doesn't add up Here is the HPA's view of the world during the incident: $ kubectl describe hpa vllm-llama Name: vllm-llama Reference: Deployment/vllm-llama Metrics: ( current / target ) resource cpu on pods (as a percentage of request): 22% (110m) / 70% Min replicas: 2 Max replicas: 12 Deployment pods: 2 current / 2 desired Conditions: Type Status Reason Message ScalingActive True ValidMetricFound the HPA was able to compute a target ...

The Agent Migrated Your YAML. It Didn't Migrate Your p99.

The green cutover with a red p99 The migration goes suspiciously well. An agent reads your EKS cluster, translates every manifest, generates the Terraform for the GKE side, and opens a change for review. You apply it. Every Deployment reports Ready , every synthetic check is green, the smoke tests pass. You cut traffic over. Then real traffic arrives and p99 latency climbs. Not by a rounding error — it roughly doubles on a couple of services and stays there. Nothing is crashing. No pod is CrashLoopBackOff . The dashboards the agent knew how to reason about are all green. The problem is that a Kubernetes manifest describes intent, and the agent proved manifest-equivalence, not behavior-equivalence. Those are different claims, and the gap between them is exactly the part of a migration that only shows up under load. What the agent is genuinely good at This is not a "don't use the tool" post. Agentic migration tooling — GKE's own EKS-to-GKE assistant, or a home-g...

Give Your Agent an Error Budget: SLOs for Hallucination and Tool-Call Failure

"It works" is not a number An LLM agent ships. It passed the demo, it passed the handful of prompts in the eval notebook, and for two weeks nobody complains. Then a support ticket arrives: the agent confidently told a customer about a refund policy that doesn't exist. You go looking for a dashboard that tells you how often that happens and there isn't one. You have latency graphs, you have token-cost graphs, and for the thing that actually matters — is the agent right — you have a vibe. That gap is the whole problem. We instrument agents like web services (RPS, p99, error rate) and then act surprised that none of those numbers move when the agent hallucinates. A 200 response carrying a made-up fact is still a 200. The reliability machinery you already run is fine; it's just pointed at the wrong signals. Reframe: agent quality is a reliability problem An error budget is just the inverse of a target: if you promise 99%, you have a 1% budget to spend. That f...

Round-Robin Is Wrecking Your LLM Throughput: Route by KV Cache, Not Connection Count

You went from 2 vLLM replicas to 8 because your queue was backing up. Eight GPUs now, four times the hardware. Aggregate throughput went up maybe 30%, and p99 time-to-first-token (TTFT) got worse . The per-GPU utilization graphs look busy, the autoscaler is happy, and the bill quadrupled for a fraction of the goodput you paid for. Before you reach for bigger GPUs, look at the thing in front of them. A plain round-robin or least-request load balancer is the wrong tool for LLM inference, and it fails in a way that's invisible on a CPU/GPU utilization dashboard. Why a normal load balancer is wrong here Web load balancing works because requests are roughly interchangeable: each one is short-lived, stateless, and costs about the same. Spraying them round-robin across identical replicas is close to optimal. LLM serving violates every one of those assumptions. Requests are not uniform. A request that generates 2,000 tokens occupies KV cache in GPU HBM for its entire decode l...

You Gave an Agent Your Database via MCP, and Now the Connection Pool Is on Fire

The traffic graph is flat. Nobody deployed a schema change, nobody launched a campaign, the request rate to your API looks exactly like it did last week. And yet the database is throwing FATAL: sorry, too many clients already , and pg_stat_activity is full of sessions parked in the idle in transaction state. The one thing that changed: yesterday you wired an AI agent to the database through an MCP server so it could answer questions over your data. This is becoming a common way to melt a perfectly healthy Postgres. The Model Context Protocol makes it a few lines of config to hand an agent a query tool backed by AlloyDB, Cloud SQL, or plain Postgres. What the tutorial doesn't mention is that an LLM is a very strange database client, and the connection pool you sized for a web app makes assumptions the agent quietly violates. I have seen this movie before, without the AI Early in my career I ran a web API backend on GCP that behaved fine in staging and then, under real load,...