LLM gateway
An LLM gateway is one endpoint that sits in front of all your models. Your app calls the gateway, and the gateway decides which model actually answers, keeps a log, and enforces limits. The part people miss: it turns "which model, which key, which provider" into a config change instead of a code change.
Why it matters
Without it, every model call is wired straight into your code, so swapping providers or adding a fallback means editing and redeploying. When one provider goes down, you have nowhere to reroute. When the bill spikes or someone leaks a key, you have no single place to see spend, set a limit, or shut it off. The gateway gives you one control point for all of that.
How it works
It exposes a standard API and translates each request into whatever the target provider expects, so your code talks to one shape no matter what runs underneath. Routing rules pick the model by cost, latency, or task, and fall back to a backup when the primary fails or times out. It logs every request and response, counts tokens for spend tracking, and enforces per-key rate limits and budgets. Popular ones are LiteLLM, OpenRouter, and Portkey.
A support bot sends every question to your gateway. Cheap FAQ-style questions go to a small fast model, while billing disputes go to a stronger one. When the main provider starts throwing errors at 2am, the gateway quietly fails over to a backup. The next morning you open one dashboard and see exactly what each model cost overnight.