One Endpoint, Every Model: Body-Based Routing for a Self-Hosted AI Gateway
Count the inference endpoints in your org. There's the self-hosted vLLM deployment on GKE serving the model you fine-tuned. There's a managed endpoint — Vertex, Bedrock, whatever — for the one you didn't want to operate. There's OpenAI or Anthropic for the frontier models you haven't brought in-house. Three backends, three URLs, three auth schemes, and every one of them hardcoded into a dozen client services. That sprawl is not a cosmetic problem. It has concrete, recurring failure modes: Adding or renaming a model means a coordinated client redeploy, because the endpoint URL lives in application config. There's no shared rate limit, so one team's overnight batch job saturates the GPU pool and interactive traffic's time-to-first-token falls off a cliff — with no throttle to protect it. Retry, timeout, and fallback logic is reimplemented (badly, inconsistently) in every client. When finance asks which team spent the inference budge...