Give Your Agent an Error Budget: SLOs for Hallucination and Tool-Call Failure

"It works" is not a number

An LLM agent ships. It passed the demo, it passed the handful of prompts in the eval notebook, and for two weeks nobody complains. Then a support ticket arrives: the agent confidently told a customer about a refund policy that doesn't exist. You go looking for a dashboard that tells you how often that happens and there isn't one. You have latency graphs, you have token-cost graphs, and for the thing that actually matters — is the agent right — you have a vibe.

That gap is the whole problem. We instrument agents like web services (RPS, p99, error rate) and then act surprised that none of those numbers move when the agent hallucinates. A 200 response carrying a made-up fact is still a 200. The reliability machinery you already run is fine; it's just pointed at the wrong signals.

Reframe: agent quality is a reliability problem

An error budget is just the inverse of a target: if you promise 99%, you have a 1% budget to spend. That framing works for "correct" as readily as it works for "available" — the only hard part is defining a Service Level Indicator you can measure automatically and cheaply enough to run in production. For an agent, a few SLIs carry most of the weight:

  • Tool-call validity — of the tool/function calls the agent emitted, what fraction had schema-valid arguments and returned successfully? A malformed argument or a 500 from the tool is a deterministic, un-arguable failure.
  • Groundedness — of the answers that cite retrieved context, what fraction make no claim that the context doesn't support? This is the operational definition of "didn't hallucinate."
  • Task success — did the run accomplish what it was asked to? Often only observable from a downstream signal (the user didn't re-ask, the ticket closed, the code compiled).

The first is a cheap deterministic check. The second needs judgment. The third needs a downstream event. Treat them differently rather than pretending one number covers all three.

user request agent LLM + tools answer to user serving path — never blocked deterministic checks arg schema, HTTP status 100% of traffic LLM judge groundedness score ~5% sample, async Prometheus SLIs + error budget burn-rate alert

Measuring the un-assertable

Deterministic signals first

Before reaching for a model to grade a model, extract everything you can check with code. Tool arguments validate against a JSON schema or they don't. A tool call returns 2xx or it doesn't. These are free, they run on 100% of traffic, and they catch a surprising amount — a large share of "the agent did something weird" traces back to a malformed tool call, not a subtle factual error.

import random, time, json
from prometheus_client import Counter, Histogram

RUNS   = Counter("agent_runs_total", "agent runs", ["outcome"])
TOOLS  = Counter("agent_tool_calls_total", "tool calls", ["tool", "ok"])
JUDGED = Counter("agent_groundedness_total", "judged answers", ["verdict"])
LAT    = Histogram("agent_run_seconds", "end-to-end latency")

SAMPLE_RATE = 0.05  # judge 5% of eligible answers

def run_agent(request):
    t0 = time.perf_counter()
    result = agent.invoke(request)
    LAT.observe(time.perf_counter() - t0)

    # deterministic SLIs — cheap, 100% of traffic
    for call in result.tool_calls:
        ok = call.args_valid and call.status < 400
        TOOLS.labels(call.name, str(ok)).inc()
    RUNS.labels("error" if result.failed else "ok").inc()

    # expensive SLI — sample, judge off the hot path
    if result.retrieved_context and random.random() < SAMPLE_RATE:
        enqueue_groundedness_judge(request, result)

    return result  # user is never blocked on judging

LLM-as-judge for the fuzzy ones

Groundedness can't be asserted with an equality check, so grade it with a second, cheaper model. The discipline that makes this trustworthy: give the judge exactly one narrow job, feed it only the context the agent was given, and force a structured verdict. Do not ask it whether the answer is "good."

JUDGE_PROMPT = """You are grading an assistant answer for GROUNDEDNESS only.
Not style, not helpfulness, not whether you agree.

CONTEXT (the only facts the assistant was given):
{context}

ANSWER:
{answer}

Return JSON: {{"verdict": "grounded" | "unsupported",
              "unsupported_claims": ["..."]}}
Mark "unsupported" if ANY factual claim in the answer is not
entailed by CONTEXT. Ignore hedged or clearly general-knowledge
statements."""

def judge_groundedness(request, result):
    raw = judge_model.complete(
        JUDGE_PROMPT.format(context=result.retrieved_context,
                            answer=result.answer),
        temperature=0,
        response_format={"type": "json_object"},
    )
    verdict = json.loads(raw)["verdict"]
    JUDGED.labels(verdict).inc()

Two things keep this honest. First, you don't judge every request — sampling 5% keeps the judge's own token cost bounded while still giving you thousands of graded answers a day at any real traffic. Second, and this is the part people skip: the judge is itself fallible, so calibrate it. Take a few hundred answers, have a human label them grounded/unsupported, and measure how often the judge agrees. If it agrees 96% of the time, your measured groundedness SLI carries roughly ±4 points of slop — which means a 99.9% target is noise and a 98% target is meaningful. Re-run that calibration whenever you change the judge model or prompt.

From SLI to error budget

Once the counters exist, the SLI is a ratio and the error budget is one minus your target. Recording rules keep the math in one place and make alert expressions cheap:

groups:
  - name: agent-slo
    rules:
      - record: agent:groundedness:ratio_rate5m
        expr: |
          sum(rate(agent_groundedness_total{verdict="grounded"}[5m]))
          /
          sum(rate(agent_groundedness_total[5m]))
      - record: agent:groundedness:ratio_rate1h
        expr: |
          sum(rate(agent_groundedness_total{verdict="grounded"}[1h]))
          /
          sum(rate(agent_groundedness_total[1h]))
      - record: agent:toolcalls:ratio_rate5m
        expr: |
          sum(rate(agent_tool_calls_total{ok="True"}[5m]))
          /
          sum(rate(agent_tool_calls_total[5m]))

Say you set a 99% groundedness target over a 30-day window. The budget is 1% of judged answers. Because you're sampling, watch the volume too: at 5% you need enough traffic that the denominator over your alert window isn't a handful of requests, or the ratio will swing wildly on sampling noise alone. A useful guardrail is to suppress budget alerts when the judged-answer rate over the window is below some floor — better a gap in the graph than a page fired by three unlucky samples.

Alert on burn rate, not on one bad answer

A single hallucination is not an incident; it's the tax you agreed to pay when you set the target below 100%. What deserves a page is burning the budget fast — the same multi-window, multi-burn-rate pattern from the Google SRE workbook applies unchanged. Require both a long and a short window to fire, so a brief blip doesn't wake anyone and a real regression is caught in minutes:

      - alert: AgentGroundednessFastBurn
        expr: |
          (1 - agent:groundedness:ratio_rate1h) > (14.4 * 0.01)
          and
          (1 - agent:groundedness:ratio_rate5m) > (14.4 * 0.01)
        for: 2m
        labels:
          severity: page
        annotations:
          summary: "Agent burning groundedness budget 14.4x too fast"

      - alert: AgentGroundednessSlowBurn
        expr: |
          (1 - agent:groundedness:ratio_rate1h) > (3 * 0.01)
          and
          (1 - agent:groundedness:ratio_rate6h) > (3 * 0.01)
        for: 15m
        labels:
          severity: ticket

The 14.4 factor is the standard fast-burn multiplier: at that rate you'd exhaust a 30-day budget in about two days, so it's worth a page. The slow-burn rule catches the quiet degradation — a retriever that started returning slightly-off chunks after an index rebuild — that never trips a fast alert but steadily eats the month's budget. That slow drift is exactly the failure mode a demo and a static eval set will never show you.

Lesson

Putting an agent on SLOs doesn't require new infrastructure — it requires deciding that "correct" is a measurable service level and then spending the effort to measure it: deterministic checks on everything cheap, a calibrated judge on a sample of everything expensive, and burn-rate alerts instead of per-answer panic. The honest caveat is that your quality number is only as trustworthy as the judge behind it, so treat the judge as a monitored component in its own right. Once you have the number, the oldest idea in reliability engineering does the rest: you stop hoping the agent is right and start knowing how often it isn't.


Hitting something like this in production? I help teams with performance engineering, SRE/observability, and AI-driven root cause analysis — work with me.

Comments

Popular posts from this blog

Performance Testing 102: Little's Law and It's usage in Performance Testing

Performance Testing 104: Workload Modelling Designing & Process

Mastering the Art of Scaling in SaaS Applications