Posts

Showing posts from October, 2026

The Agent That Got Slower the Longer It Ran: Two-Tier Memory for Long-Horizon Agents

The agent that remembered everything, and slowed to a crawl Here is a failure mode that shows up on a schedule once you give an agent persistent memory. In week one it feels sharp: it recalls what you told it yesterday, it picks up a multi-day task where it left off, memory lookups are a few milliseconds and nobody thinks about them. By week six the same agent is noticeably slower on every turn, and — more insidiously — its answers have started drifting toward stale or irrelevant facts it dug up from months ago. The reflex is to blame the vector index, or the model, or to move to a bigger database instance. Usually none of those is the real problem. The problem is architectural, and it is almost always the same one: all of the agent's memory lives in a single flat, append-only table that nothing ever consolidates or evicts. That design degrades on two axes at once, and you have to fix both. Why a flat memory table degrades on two axes The naive design is seductive because ...

Your Agent Issued the Refund Twice: Durable Side Effects with the Transactional Outbox

Your support agent decides a customer is owed a refund. It calls the issue_refund tool, the payment provider returns 200, and the agent writes status = "refunded" to its own task row. Clean. Except on about one run in a few thousand, the customer gets refunded twice — or, on a different unlucky run, the agent swears it refunded them and the money never moved. Nothing in the model is broken. The prompt is fine, the tool schema is fine, the eval suite is green. What's broken is older and more boring than anything about LLMs: the agent is doing two writes to two different systems and pretending they happen together. They don't. The symptom You notice it from a reconciliation job, not from the agent. A finance export shows two refund transactions against one order ID, minutes apart, both tagged with the same agent run. Or a customer emails back asking where their money is, and the trace shows the tool call succeeded. When you pull the logs for the double case, yo...

Most of your prompts don't need the frontier model: building a difficulty-routed LLM cascade

The symptom is a line on the billing dashboard that tracks your request count almost perfectly: inference spend climbing in lockstep with traffic, every request costing roughly the same. That flat per-request cost is the tell. It means every prompt — the one-line sentiment check and the twelve-step tool-calling reasoning chain — is being answered by the same expensive model. Your traffic is almost never that uniform. In most production LLM workloads the difficulty distribution is heavily skewed: a large fraction of requests are easy enough that a 7B open model gets them right, and a small tail genuinely needs a frontier model. Paying frontier prices for the whole distribution is the expensive default. A cascade is how you stop. Why one model for everything is the wrong shape The instinct to standardize on a single strong model is understandable — one integration, one prompt to maintain, one set of behaviors to reason about. But it couples your cost to your worst request rather t...