Give Your Agent an Error Budget: SLOs for Hallucination and Tool-Call Failure
"It works" is not a number
An LLM agent ships. It passed the demo, it passed the handful of prompts in the eval notebook, and for two weeks nobody complains. Then a support ticket arrives: the agent confidently told a customer about a refund policy that doesn't exist. You go looking for a dashboard that tells you how often that happens and there isn't one. You have latency graphs, you have token-cost graphs, and for the thing that actually matters — is the agent right — you have a vibe.
That gap is the whole problem. We instrument agents like web services (RPS, p99, error rate) and then act surprised that none of those numbers move when the agent hallucinates. A 200 response carrying a made-up fact is still a 200. The reliability machinery you already run is fine; it's just pointed at the wrong signals.
Reframe: agent quality is a reliability problem
An error budget is just the inverse of a target: if you promise 99%, you have a 1% budget to spend. That framing works for "correct" as readily as it works for "available" — the only hard part is defining a Service Level Indicator you can measure automatically and cheaply enough to run in production. For an agent, a few SLIs carry most of the weight:
- Tool-call validity — of the tool/function calls the agent emitted, what fraction had schema-valid arguments and returned successfully? A malformed argument or a 500 from the tool is a deterministic, un-arguable failure.
- Groundedness — of the answers that cite retrieved context, what fraction make no claim that the context doesn't support? This is the operational definition of "didn't hallucinate."
- Task success — did the run accomplish what it was asked to? Often only observable from a downstream signal (the user didn't re-ask, the ticket closed, the code compiled).
The first is a cheap deterministic check. The second needs judgment. The third needs a downstream event. Treat them differently rather than pretending one number covers all three.
Measuring the un-assertable
Deterministic signals first
Before reaching for a model to grade a model, extract everything you can check with code. Tool arguments validate against a JSON schema or they don't. A tool call returns 2xx or it doesn't. These are free, they run on 100% of traffic, and they catch a surprising amount — a large share of "the agent did something weird" traces back to a malformed tool call, not a subtle factual error.
import random, time, json
from prometheus_client import Counter, Histogram
RUNS = Counter("agent_runs_total", "agent runs", ["outcome"])
TOOLS = Counter("agent_tool_calls_total", "tool calls", ["tool", "ok"])
JUDGED = Counter("agent_groundedness_total", "judged answers", ["verdict"])
LAT = Histogram("agent_run_seconds", "end-to-end latency")
SAMPLE_RATE = 0.05 # judge 5% of eligible answers
def run_agent(request):
t0 = time.perf_counter()
result = agent.invoke(request)
LAT.observe(time.perf_counter() - t0)
# deterministic SLIs — cheap, 100% of traffic
for call in result.tool_calls:
ok = call.args_valid and call.status < 400
TOOLS.labels(call.name, str(ok)).inc()
RUNS.labels("error" if result.failed else "ok").inc()
# expensive SLI — sample, judge off the hot path
if result.retrieved_context and random.random() < SAMPLE_RATE:
enqueue_groundedness_judge(request, result)
return result # user is never blocked on judging
LLM-as-judge for the fuzzy ones
Groundedness can't be asserted with an equality check, so grade it with a second, cheaper model. The discipline that makes this trustworthy: give the judge exactly one narrow job, feed it only the context the agent was given, and force a structured verdict. Do not ask it whether the answer is "good."
JUDGE_PROMPT = """You are grading an assistant answer for GROUNDEDNESS only.
Not style, not helpfulness, not whether you agree.
CONTEXT (the only facts the assistant was given):
{context}
ANSWER:
{answer}
Return JSON: {{"verdict": "grounded" | "unsupported",
"unsupported_claims": ["..."]}}
Mark "unsupported" if ANY factual claim in the answer is not
entailed by CONTEXT. Ignore hedged or clearly general-knowledge
statements."""
def judge_groundedness(request, result):
raw = judge_model.complete(
JUDGE_PROMPT.format(context=result.retrieved_context,
answer=result.answer),
temperature=0,
response_format={"type": "json_object"},
)
verdict = json.loads(raw)["verdict"]
JUDGED.labels(verdict).inc()
Two things keep this honest. First, you don't judge every request — sampling 5% keeps the judge's own token cost bounded while still giving you thousands of graded answers a day at any real traffic. Second, and this is the part people skip: the judge is itself fallible, so calibrate it. Take a few hundred answers, have a human label them grounded/unsupported, and measure how often the judge agrees. If it agrees 96% of the time, your measured groundedness SLI carries roughly ±4 points of slop — which means a 99.9% target is noise and a 98% target is meaningful. Re-run that calibration whenever you change the judge model or prompt.
From SLI to error budget
Once the counters exist, the SLI is a ratio and the error budget is one minus your target. Recording rules keep the math in one place and make alert expressions cheap:
groups:
- name: agent-slo
rules:
- record: agent:groundedness:ratio_rate5m
expr: |
sum(rate(agent_groundedness_total{verdict="grounded"}[5m]))
/
sum(rate(agent_groundedness_total[5m]))
- record: agent:groundedness:ratio_rate1h
expr: |
sum(rate(agent_groundedness_total{verdict="grounded"}[1h]))
/
sum(rate(agent_groundedness_total[1h]))
- record: agent:toolcalls:ratio_rate5m
expr: |
sum(rate(agent_tool_calls_total{ok="True"}[5m]))
/
sum(rate(agent_tool_calls_total[5m]))
Say you set a 99% groundedness target over a 30-day window. The budget is 1% of judged answers. Because you're sampling, watch the volume too: at 5% you need enough traffic that the denominator over your alert window isn't a handful of requests, or the ratio will swing wildly on sampling noise alone. A useful guardrail is to suppress budget alerts when the judged-answer rate over the window is below some floor — better a gap in the graph than a page fired by three unlucky samples.
Alert on burn rate, not on one bad answer
A single hallucination is not an incident; it's the tax you agreed to pay when you set the target below 100%. What deserves a page is burning the budget fast — the same multi-window, multi-burn-rate pattern from the Google SRE workbook applies unchanged. Require both a long and a short window to fire, so a brief blip doesn't wake anyone and a real regression is caught in minutes:
- alert: AgentGroundednessFastBurn
expr: |
(1 - agent:groundedness:ratio_rate1h) > (14.4 * 0.01)
and
(1 - agent:groundedness:ratio_rate5m) > (14.4 * 0.01)
for: 2m
labels:
severity: page
annotations:
summary: "Agent burning groundedness budget 14.4x too fast"
- alert: AgentGroundednessSlowBurn
expr: |
(1 - agent:groundedness:ratio_rate1h) > (3 * 0.01)
and
(1 - agent:groundedness:ratio_rate6h) > (3 * 0.01)
for: 15m
labels:
severity: ticket
The 14.4 factor is the standard fast-burn multiplier: at that
rate you'd exhaust a 30-day budget in about two days, so it's worth a page.
The slow-burn rule catches the quiet degradation — a retriever that started
returning slightly-off chunks after an index rebuild — that never trips a fast
alert but steadily eats the month's budget. That slow drift is exactly the
failure mode a demo and a static eval set will never show you.
Lesson
Putting an agent on SLOs doesn't require new infrastructure — it requires deciding that "correct" is a measurable service level and then spending the effort to measure it: deterministic checks on everything cheap, a calibrated judge on a sample of everything expensive, and burn-rate alerts instead of per-answer panic. The honest caveat is that your quality number is only as trustworthy as the judge behind it, so treat the judge as a monitored component in its own right. Once you have the number, the oldest idea in reliability engineering does the rest: you stop hoping the agent is right and start knowing how often it isn't.
Hitting something like this in production? I help teams with performance engineering, SRE/observability, and AI-driven root cause analysis — work with me.
Comments
Post a Comment