Most of your prompts don't need the frontier model: building a difficulty-routed LLM cascade

The symptom is a line on the billing dashboard that tracks your request count almost perfectly: inference spend climbing in lockstep with traffic, every request costing roughly the same. That flat per-request cost is the tell. It means every prompt — the one-line sentiment check and the twelve-step tool-calling reasoning chain — is being answered by the same expensive model.

Your traffic is almost never that uniform. In most production LLM workloads the difficulty distribution is heavily skewed: a large fraction of requests are easy enough that a 7B open model gets them right, and a small tail genuinely needs a frontier model. Paying frontier prices for the whole distribution is the expensive default. A cascade is how you stop.

Why one model for everything is the wrong shape

The instinct to standardize on a single strong model is understandable — one integration, one prompt to maintain, one set of behaviors to reason about. But it couples your cost to your worst request rather than your typical one. If 70% of your traffic is easy (an illustrative split — measure your own), you are paying the hard-request price 70% of the time for no quality benefit, because an easy request gets the right answer from a much cheaper model.

The idea isn't new. Google's FrugalGPT work (2023) laid out the LLM-cascade pattern: try a cheap model first, and only call a more expensive one when the cheap answer isn't good enough. What's changed since is the cheap tier itself. Self-hosting a capable open model — Qwen2.5, Llama 3.x, Mistral — on vLLM or TGI is now routine, and those models are strong enough that the cheap tier resolves a real share of production traffic instead of being a toy. That's the practical core of the "open models alongside frontier APIs" argument: the open model isn't a replacement for the frontier API, it's the first stage of a cascade in front of it.

Two ways to route: classify up front, or cascade

There are two families of routing, and they have genuinely different failure modes.

  • Up-front classification. A small, fast classifier looks at the incoming request and predicts which tier it needs, then routes directly. One model call on the happy path. The problem: you're predicting difficulty before seeing any answer, so the classifier has to be good, and when it's wrong it routes a hard request to the weak model and ships a bad answer with no second chance.
  • Cascade. Run the cheap tier first, measure how good its answer actually is, and escalate only if it falls short. You get a real signal — an actual attempt — instead of a prediction. The cost is latency and compute on escalated requests: you ran the small model and the big one.

Cascades are usually the better starting point because the escalation decision is grounded in an observed attempt rather than a guess. The entire game is then the quality of your confidence gate — the function that decides "is this cheap answer good enough, or do I escalate?"

request small tier open model, vLLM confidence >= threshold? frontier API escalate response cheap answer no yes frontier answer

Building the cascade

The confidence signal matters more than the plumbing. A tempting but weak option is to ask the small model to rate its own confidence — self-reported confidence is poorly calibrated and easy to game. Better signals, in rough order of effort:

  • Sequence log-probability. Open models served on vLLM expose token logprobs. The mean token logprob of the generated answer is a cheap, surprisingly usable proxy for "the model found this easy." This is essentially FrugalGPT's scoring approach.
  • A structural/semantic validator. If the task has a checkable shape — valid JSON against a schema, a SQL query that parses and runs against an explain plan, a tool call whose arguments type-check — a failed check is an unambiguous escalation trigger.
  • A small judge model. A second cheap call that scores the answer. More signal, but it adds cost and latency to the cheap path, so use it only when logprobs and validators aren't enough.

Here's the core of a cascade using logprob-based gating. The small tier is a self-hosted open model behind an OpenAI-compatible vLLM endpoint; the frontier tier is a managed API.

import math
from openai import OpenAI

small = OpenAI(base_url="http://vllm:8000/v1", api_key="local")
# frontier tier: a managed API client (Anthropic, OpenAI, etc.)

ESCALATION_LOGPROB = -0.55   # calibrate this on a labeled eval set, not by vibes

def mean_logprob(resp) -> float:
    toks = resp.choices[0].logprobs.content
    return sum(t.logprob for t in toks) / max(len(toks), 1)

def answer(prompt: str) -> dict:
    # Stage 1: cheap tier, with logprobs
    r = small.chat.completions.create(
        model="qwen2.5-7b-instruct",
        messages=[{"role": "user", "content": prompt}],
        logprobs=True,
        temperature=0,
    )
    conf = mean_logprob(r)           # ~0 is confident, more negative is shakier
    text = r.choices[0].message.content

    if conf >= ESCALATION_LOGPROB and valid(text):
        return {"text": text, "tier": "small", "confidence": conf}

    # Stage 2: escalate to frontier
    big = frontier_complete(prompt)  # managed API call
    return {"text": big, "tier": "frontier", "confidence": conf}

def valid(text: str) -> bool:
    # task-specific structural check; return False to force escalation
    return True

Note what the gate combines: a calibrated logprob threshold and a structural validator. The validator catches the confident-but-wrong case — a model that's fluent and certain while producing malformed output — which a probability threshold alone will happily pass through.

The one metric that decides whether this works

A cascade's economics live and die on the escalation rate — the fraction of requests that fall through to the frontier tier. It is your single most important SLI here, and you should emit it as a first-class metric alongside latency and error rate.

The math is simple enough to reason about on a napkin. Say the frontier tier costs 10× the small tier per request (an illustrative ratio — plug in your own). You always pay the small-tier cost, plus the frontier cost on the escalated fraction e:

blended_cost_per_req = 1  +  e * 10          # small always runs; frontier on escalations

e = 0.10  ->  1 + 1.0  =  2.0   units   (vs 10 all-frontier  -> ~80% cheaper)
e = 0.30  ->  1 + 3.0  =  4.0   units   (vs 10 all-frontier  -> ~60% cheaper)
e = 0.60  ->  1 + 6.0  =  7.0   units   (vs 10 all-frontier  -> ~30% cheaper)
e = 0.90  ->  1 + 9.0  = 10.0   units   (break-even — you're paying for both)

(Figures illustrative.) The shape is the point: savings are real at low escalation rates and evaporate as e climbs, because on every escalated request you paid for both tiers. A cascade with a 90% escalation rate is strictly worse than just calling the frontier model directly — you've added latency and the small-tier bill for nothing.

Blended cost per request (illustrative, frontier = 10x small) e=10% 2.0 e=30% 4.0 e=60% 7.0 e=90% 10.0 all-frontier 10.0

What goes wrong in production

Two failure modes, mirror images of each other, both caused by a mis-set gate:

  • Gate too strict. You escalate almost everything to be safe. Quality is fine, but the escalation rate creeps toward 1 and your bill is higher than going frontier-only. This one hides because the dashboards look healthy — you have to be watching escalation rate specifically to catch it.
  • Gate too loose. You pass borderline cheap answers through, the escalation rate looks great, and the savings look enormous — but answer quality quietly degrades. This is the dangerous one, because cost metrics reward you for exactly the behavior that's hurting quality.

Because the two failure modes pull in opposite directions, you cannot tune the threshold by watching cost alone. Calibrate it on a labeled eval set: for a range of thresholds, measure both the escalation rate and the end-to-end answer quality (against the frontier model as reference, or against ground truth), and pick the knee where quality is still acceptable and escalation is as low as it goes. Then keep a canary eval running in production, because the request mix drifts — a new feature can shift your difficulty distribution overnight and quietly push the escalation rate somewhere you didn't calibrate for.

Lesson

A cascade doesn't save money because the small model is cheap; it saves money because most of your traffic is easy and you stopped paying the hard-request price for it. That makes the win entirely dependent on two numbers you have to actually measure — the shape of your difficulty distribution and the escalation rate your confidence gate produces on it. Treat the escalation rate as an SLI, calibrate the gate against quality rather than vibes, and keep watching both after launch. Skip that and a cascade is just a more complicated way to pay for two models at once.


Hitting something like this in production? I help teams with performance engineering, SRE/observability, and AI-driven root cause analysis — work with me.

Comments

Popular posts from this blog

Performance Testing 102: Little's Law and It's usage in Performance Testing

Performance Testing 104: Workload Modelling Designing & Process

Mastering the Art of Scaling in SaaS Applications