Give Your Agent an Error Budget: SLOs for Hallucination and Tool-Call Failure
"It works" is not a number An LLM agent ships. It passed the demo, it passed the handful of prompts in the eval notebook, and for two weeks nobody complains. Then a support ticket arrives: the agent confidently told a customer about a refund policy that doesn't exist. You go looking for a dashboard that tells you how often that happens and there isn't one. You have latency graphs, you have token-cost graphs, and for the thing that actually matters — is the agent right — you have a vibe. That gap is the whole problem. We instrument agents like web services (RPS, p99, error rate) and then act surprised that none of those numbers move when the agent hallucinates. A 200 response carrying a made-up fact is still a 200. The reliability machinery you already run is fine; it's just pointed at the wrong signals. Reframe: agent quality is a reliability problem An error budget is just the inverse of a target: if you promise 99%, you have a 1% budget to spend. That f...