LLM Gateway· LLM infrastructure
LLM Gateway — One Door Between Your Chatbot and the Models
Quick answer: A chatbot that calls one model from one provider can call it directly. An LLM gateway becomes useful when the chatbot depends on more than one model or provider, or when you need one place to cap spend, switch to a backup when a provider is down, and read a single log. The application talks to the gateway in one format, usually the OpenAI request format, and the gateway talks to each provider. The trade-off is one more service in the path of every answer.
What a gateway does
LiteLLM's repository describes its open-source project as one interface to 100+ model providers, including OpenAI, Anthropic, Gemini, Bedrock and Azure, using the OpenAI request format. Portkey describes its gateway as routing requests to many models through one API, and Helicone describes a unified API that translates requests to each provider's format while logging cost and errors. The shared idea has five parts:
- One endpoint and one format. The application is written once. Changing the model or the provider is a configuration change, not a rewrite.
- Reliability. Retries on a failed call, and a fallback to another model or provider when a call keeps failing.
- Control. Keys, budgets and rate limits set per team, project or user instead of one shared provider key.
- Cost levers. Caching, including semantic caching in some products, and routing a request to a cheaper model.
- One log. Every request, its token use and its cost in one place, which feeds LLM observability.
Gateway, API gateway and router are different things
| What it manages | Understands model calls? | Typical job | |
|---|---|---|---|
| API gateway | Any HTTP service | No | Authentication, rate limits and routing for your own APIs |
| LLM gateway | Calls to language-model providers | Yes: tokens, models, streaming, provider formats | Failover, spend limits, one format, one log |
| LLM router | The choice of which model answers | Yes | Picks a model per request, for example small or large |
The terms overlap in practice. Many gateways include a router, and OpenRouter, for example, offers a list of models tried in order. Treat "router" as a function a gateway may perform, not as a separate product category.
What the documentation says about four options
These are vendor statements, not our test results, and we have not benchmarked any of them.
- LiteLLM. An open-source project with a Python library and a proxy server. Its README lists, for the proxy, authentication, spend tracking per project or user, guardrails, caching and virtual keys. Its documentation sets retries and fallbacks in the proxy configuration, with fallbacks tried in order and a deployment put into cooldown after repeated failures. Its virtual keys accept a budget, a budget period and request-rate limits, and the key management page lists a Postgres database as a requirement. The repository's license file is MIT for most code, with a separate license for an
enterprisedirectory. - Portkey. An open-source gateway under the MIT license, per its repository, with fallbacks, load balancing, automatic retries (up to five, with exponential backoff), guardrails and caching. The page marks semantic caching as a hosted and enterprise feature. Its page gives inconsistent counts of supported models, so check the current list for the one you need.
- OpenRouter. A hosted service. Its documentation says you can pass an ordered list of models and it tries the next one on errors such as rate limiting, downtime or a context-length failure, and that you are priced for the model that actually served the request.
- Helicone. Describes a unified API with automatic failover across providers and one dashboard for cost and errors. Its overview does not state a license or self-hosting terms, so confirm both on its repository.
When a chatbot needs one, and when it does not
A gateway earns its place when at least one of these is true: you want a backup provider so an outage does not stop the chatbot; you run several models and want to move traffic between them, as in testing a smaller model; several teams or clients share a provider account and each needs its own cap; or you need one log of cost across everything. It is a poor fit for a single low-volume chatbot on one model with a provider budget alert already set. The simpler path is to call the provider directly and revisit when the second model arrives. If you let customers bring their own model key, a gateway is also a convenient place to hold and limit those keys.
Limits and risks to plan for
- One more thing that can fail. A gateway in front of the model is a single point of failure unless you run it redundantly or keep a direct path as a last resort.
- Latency. Every request passes through an extra hop. LiteLLM states 8 ms at the 95th percentile at 1,000 requests per second, and Portkey states under 1 ms; both are vendor figures measured under their own conditions, so measure your own.
- Provider features arrive later. A new provider feature may not be available through the common format until the gateway adds it. Check streaming and function calling through the gateway before relying on either.
- Behavior differs between models. A fallback sends the same prompt to a different model. Its answer style, refusals and tool-call habits may differ, and LiteLLM's documentation notes that encrypted reasoning from a failed deployment cannot be carried to another provider. Test the backup model with your own questions.
- Customer data passes through it. A hosted gateway is one more party handling conversations. Review its retention and security terms alongside your data retention policy and LLM security practice.
- Guardrails are not a substitute for design. A gateway's content checks supplement, and do not replace, the controls described under AI guardrails.
FAQ
What is an LLM gateway?
A service between your application and model providers that offers one endpoint and adds fallbacks, retries, limits, caching and logging, so the application does not call each provider separately.
Is an LLM gateway the same as an API gateway?
No. An API gateway manages any HTTP service and does not understand tokens, models or streaming. An LLM gateway is built for model calls and knows about those things.
What is the difference between an LLM gateway and an LLM router?
A router chooses which model answers a request. A gateway is the wider layer that also handles keys, limits, fallbacks and logs, and it often includes a router.
Do I need one for a customer support chatbot?
Not for one model from one provider at modest volume. You likely need one when you want automatic failover to a second provider, when several models are in use, or when you need per-team spend caps.
Does a gateway make answers cheaper?
It can, through caching and routing to a cheaper model, but only if you configure those. A gateway that only forwards requests adds a hop and no saving.
Is LiteLLM free?
Its repository license file says MIT for most of the code, with a separate license for an enterprise directory. Running it still costs hosting and your time, so check the current licensing and pricing pages.
Related terms
- Large language model — what the gateway forwards requests to.
- BYOLLM — letting customers attach their own model key.
- Prompt caching and semantic caching — two cost levers a gateway may offer.
- LLM observability — what the gateway's log feeds.
- AI guardrails — the checks some gateways add.
- LLM token — the unit the gateway meters.
Sources
- LiteLLM, GitHub repository — github.com/BerriAI/litellm, read 10 October 2026. Source of the project description, the library-versus-proxy split, the proxy feature list and the 8 ms P95 figure (vendor claim).
- LiteLLM, Fallbacks, retries and cooldowns — docs.litellm.ai/docs/proxy/reliability, read 10 October 2026. Source of the ordered fallbacks, cooldown behavior and the note on reasoning across providers.
- LiteLLM, Virtual keys — docs.litellm.ai/docs/proxy/virtual_keys, read 10 October 2026. Source of budgets, rate limits and the database requirement.
- LiteLLM, LICENSE — github.com/BerriAI/litellm/blob/main/LICENSE, read 10 October 2026.
- Portkey, AI Gateway repository — github.com/Portkey-AI/gateway, read 10 October 2026. Source of the license, feature list and latency claim (vendor claim).
- OpenRouter, Model fallbacks — openrouter.ai/docs/guides/routing/model-fallbacks, read 10 October 2026.
- Helicone, AI Gateway overview — docs.helicone.ai/gateway/overview, read 10 October 2026.
- Chatbotscape evaluation methodology. /methodology (continuously updated).