One Endpoint, Every Model: Body-Based Routing for a Self-Hosted AI Gateway

Count the inference endpoints in your org. There's the self-hosted vLLM deployment on GKE serving the model you fine-tuned. There's a managed endpoint — Vertex, Bedrock, whatever — for the one you didn't want to operate. There's OpenAI or Anthropic for the frontier models you haven't brought in-house. Three backends, three URLs, three auth schemes, and every one of them hardcoded into a dozen client services.

That sprawl is not a cosmetic problem. It has concrete, recurring failure modes:

  • Adding or renaming a model means a coordinated client redeploy, because the endpoint URL lives in application config.
  • There's no shared rate limit, so one team's overnight batch job saturates the GPU pool and interactive traffic's time-to-first-token falls off a cliff — with no throttle to protect it.
  • Retry, timeout, and fallback logic is reimplemented (badly, inconsistently) in every client.
  • When finance asks which team spent the inference budget, nobody can answer, because token usage is scattered across three vendors' billing consoles and zero of your own logs.

The fix is an inference gateway: one endpoint in front of all the backends. The catch is that the usual way of building one — an L7 load balancer routing by host or path — cannot do the one job that matters here.

The model name is in the body, not the URL

Every OpenAI-compatible client, regardless of which backend ultimately serves it, sends the same request shape to the same path:

POST /v1/chat/completions HTTP/1.1
Host: ai-gateway.internal
Authorization: Bearer sk-team-checkout-...
Content-Type: application/json

{
  "model": "llama-3.1-70b",
  "messages": [{"role": "user", "content": "..."}],
  "stream": true
}

The only thing distinguishing a request for your self-hosted Llama from a request for a managed Gemini endpoint is the model field — and that field is in the JSON body. The method is POST, the path is /v1/chat/completions, and the host is your gateway, for every request. A standard HTTPRoute or Ingress that matches on host, path, or method has nothing to grab onto. You could force clients to use per-model paths like /llama/v1/..., but then you've thrown away OpenAI compatibility — the whole reason the SDKs work out of the box — and you're back to clients hardcoding routes.

So the gateway has to read the request body, extract the model, and make the routing decision from that. This is "body-based routing," and it's the one piece a generic load balancer doesn't give you for free.

The architecture

The shape is a single OpenAI-compatible front door, a small processor that inspects the body and tags the request with a routing header, and a set of backends the gateway's normal header-matching can then route to.

checkout svc search svc batch job POST /v1/chat/completions {"model": "..."} AI Gateway 1. authn / api-key 2. read body -> model 3. set X-Model-Route 4. rate limit + log 5. route on header vLLM InferencePool llama-3.1-70b (GKE, GPU) managed endpoint gemini-* (Vertex) third-party API gpt-* (egress)

Steps 1, 4, and 5 are ordinary gateway features. Step 2 — reading the body and extracting the model — is the part you add. On GKE, the Gateway API Inference Extension ships this as a "body-based routing" component that injects a header for you; generically, it's an Envoy external processor (ext_proc). I'll show the explicit version, because once you see the mechanism the managed one is obvious.

Body-based routing, concretely

An ext_proc service is a gRPC server Envoy streams each request through. For body inspection you ask Envoy to buffer the request body and hand it to you; you parse the model and return a header mutation. The handler is small:

class ModelRouter(ext_proc_pb2_grpc.ExternalProcessorServicer):
    # map a requested model name to a backend routing label
    ROUTES = {
        "llama-3.1-70b": "vllm-pool",
        "gemini-1.5-pro": "vertex",
        "gpt-4o":         "openai",
    }

    def Process(self, request_iterator, context):
        for req in request_iterator:
            if req.HasField("request_body"):
                body = json.loads(req.request_body.body or b"{}")
                model = body.get("model", "")
                target = self.ROUTES.get(model, "unknown")
                # add a header Envoy's HTTPRoute can match on
                yield ext_proc_pb2.ProcessingResponse(
                    request_body=ext_proc_pb2.BodyResponse(
                        response=ext_proc_pb2.CommonResponse(
                            header_mutation=ext_proc_pb2.HeaderMutation(
                                set_headers=[header("x-model-route", target),
                                             header("x-model-name", model)]))))
            else:
                yield ext_proc_pb2.ProcessingResponse()

Now the model identity is a header, and routing is back in familiar territory. The HTTPRoute matches on x-model-route and sends each class of request to its own backend service:

apiVersion: gateway.networking.k8s.io/v1
kind: HTTPRoute
metadata:
  name: inference-router
spec:
  parentRefs:
    - name: ai-gateway
  rules:
    - matches:
        - headers:
            - name: x-model-route
              value: vllm-pool
      backendRefs:
        - group: inference.networking.k8s.io
          kind: InferencePool      # KV-cache/queue-aware picker lives here
          name: llama-70b-pool
    - matches:
        - headers:
            - name: x-model-route
              value: vertex
      backendRefs:
        - name: vertex-proxy
          port: 443
    - matches:
        - headers: [{ name: x-model-route, value: unknown }]
      filters:
        - type: ExtensionRef       # return 404 for unrouteable models
          extensionRef: { group: "", kind: Service, name: reject-unknown }

Note the division of labor. Body-based routing picks the backend for a given model. Inside the self-hosted pool, a smarter picker still decides which replica — by KV-cache pressure and queue depth, not round-robin. Those are two different layers, and you want both: the gateway routes "llama vs gemini," the InferencePool routes "replica 3 vs replica 7."

Rate limiting and cost attribution, where they belong

Once every request flows through one place, carrying both an identity (the API key you authenticated in step 1) and a model (the header you set in step 2), two things that were impossible become easy.

Rate limiting keys on both dimensions, so the overnight batch job can't starve interactive callers of the same GPU pool:

# Envoy global rate-limit descriptors
descriptors:
  - entries:
      - key: api_key_team         # from the Authorization header
      - key: x-model-route
        value: vllm-pool          # the scarce, self-hosted GPU backend
    rate_limit: { requests_per_unit: 60, unit: minute }

Cost attribution falls out of logging the response. OpenAI-compatible responses carry a usage block with prompt_tokens and completion_tokens; log that alongside the team and model you already have in headers, and a daily rollup answers "who spent the budget" without touching any vendor console.

The streaming gotcha

There's a real cost to reading the body: ext_proc in buffered mode makes Envoy hold the full request body before forwarding. For chat requests that's fine — a few KB — but set a sane max_request_bytes so a pathological 10 MB prompt can't pin memory, and return 413 past the limit rather than buffering blindly.

The bigger trap is on the way back. Almost all of this traffic is stream: true — server-sent events, tokens dribbling out over a long-lived connection. If you leave response buffering on, the gateway holds the entire generation and hands the client its "streaming" answer in one lump after 20 seconds, destroying the time-to-first-token you deployed GPUs to protect. Disable response buffering on these routes and set timeouts long enough for a full generation (a route-level timeout of a few seconds, copied from an HTTP default, will sever long completions mid-stream):

route:
  timeout: 0s                 # don't cap on total response time
  idle_timeout: 300s          # cap on silence between chunks instead
  # and: do not enable response-body buffering on streaming routes

An idle timeout — time since the last chunk — is the right control for streaming, not a total-duration timeout. A long answer is healthy; a backend that went silent for five minutes is not.

What this doesn't solve

A gateway centralizes control; it doesn't conjure capacity. If the vLLM pool is undersized, body-based routing just gives you a tidy place to watch it saturate. Cross-backend fallback (spill from the self-hosted pool to a managed endpoint when queue depth is high) is genuinely useful but adds correctness questions — different models produce different outputs, and a silent failover can change behavior under the client's feet — so gate it behind an explicit policy, not a blanket retry. And the gateway is now a critical path: it needs its own autoscaling and health checks, because every inference call in the org depends on it.

Lesson

The reason inference endpoints sprawl is that the usual routing tools can't see what distinguishes one model from another — the identity is in the request body, where host and path matching never look. Once you accept that and route on the body, the hard problems (rate limiting the scarce GPU pool, attributing token cost, adding a model without a client redeploy) turn into ordinary header-matching and logging. The gateway isn't infrastructure for its own sake; it's the one place where "which model, for whom, at what cost" is finally a question you can answer.


Hitting something like this in production? I help teams with performance engineering, SRE/observability, and AI-driven root cause analysis — work with me.

Comments

Popular posts from this blog

Performance Testing 102: Little's Law and It's usage in Performance Testing

Performance Testing 104: Workload Modelling Designing & Process

Mastering the Art of Scaling in SaaS Applications