Skip to content
Chatbotscape
The LiteLLM, Portkey, OpenRouter and Helicone documentation pages cited below were read on 10 October 2026. Gateways change quickly, so check each project's current documentation, license and maintenance status before you depend on one.
LLM Gateway· LLM infrastructure
An LLM gateway is a service that sits between an application and one or more language-model providers. The application sends every model request to one endpoint, and the gateway forwards it to a provider while adding controls such as fallbacks, retries, usage limits, caching and logging.
By Chatbotscape Editorial· Methodology· Published 11 October 2026· Updated 11 October 2026

LLM Gateway — One Door Between Your Chatbot and the Models

Quick answer: A chatbot that calls one model from one provider can call it directly. An LLM gateway becomes useful when the chatbot depends on more than one model or provider, or when you need one place to cap spend, switch to a backup when a provider is down, and read a single log. The application talks to the gateway in one format, usually the OpenAI request format, and the gateway talks to each provider. The trade-off is one more service in the path of every answer.

What a gateway does

LiteLLM's repository describes its open-source project as one interface to 100+ model providers, including OpenAI, Anthropic, Gemini, Bedrock and Azure, using the OpenAI request format. Portkey describes its gateway as routing requests to many models through one API, and Helicone describes a unified API that translates requests to each provider's format while logging cost and errors. The shared idea has five parts:

  • One endpoint and one format. The application is written once. Changing the model or the provider is a configuration change, not a rewrite.
  • Reliability. Retries on a failed call, and a fallback to another model or provider when a call keeps failing.
  • Control. Keys, budgets and rate limits set per team, project or user instead of one shared provider key.
  • Cost levers. Caching, including semantic caching in some products, and routing a request to a cheaper model.
  • One log. Every request, its token use and its cost in one place, which feeds LLM observability.

Gateway, API gateway and router are different things

What it managesUnderstands model calls?Typical job
API gatewayAny HTTP serviceNoAuthentication, rate limits and routing for your own APIs
LLM gatewayCalls to language-model providersYes: tokens, models, streaming, provider formatsFailover, spend limits, one format, one log
LLM routerThe choice of which model answersYesPicks a model per request, for example small or large

The terms overlap in practice. Many gateways include a router, and OpenRouter, for example, offers a list of models tried in order. Treat "router" as a function a gateway may perform, not as a separate product category.

What the documentation says about four options

These are vendor statements, not our test results, and we have not benchmarked any of them.

  • LiteLLM. An open-source project with a Python library and a proxy server. Its README lists, for the proxy, authentication, spend tracking per project or user, guardrails, caching and virtual keys. Its documentation sets retries and fallbacks in the proxy configuration, with fallbacks tried in order and a deployment put into cooldown after repeated failures. Its virtual keys accept a budget, a budget period and request-rate limits, and the key management page lists a Postgres database as a requirement. The repository's license file is MIT for most code, with a separate license for an enterprise directory.
  • Portkey. An open-source gateway under the MIT license, per its repository, with fallbacks, load balancing, automatic retries (up to five, with exponential backoff), guardrails and caching. The page marks semantic caching as a hosted and enterprise feature. Its page gives inconsistent counts of supported models, so check the current list for the one you need.
  • OpenRouter. A hosted service. Its documentation says you can pass an ordered list of models and it tries the next one on errors such as rate limiting, downtime or a context-length failure, and that you are priced for the model that actually served the request.
  • Helicone. Describes a unified API with automatic failover across providers and one dashboard for cost and errors. Its overview does not state a license or self-hosting terms, so confirm both on its repository.

When a chatbot needs one, and when it does not

A gateway earns its place when at least one of these is true: you want a backup provider so an outage does not stop the chatbot; you run several models and want to move traffic between them, as in testing a smaller model; several teams or clients share a provider account and each needs its own cap; or you need one log of cost across everything. It is a poor fit for a single low-volume chatbot on one model with a provider budget alert already set. The simpler path is to call the provider directly and revisit when the second model arrives. If you let customers bring their own model key, a gateway is also a convenient place to hold and limit those keys.

Limits and risks to plan for

  • One more thing that can fail. A gateway in front of the model is a single point of failure unless you run it redundantly or keep a direct path as a last resort.
  • Latency. Every request passes through an extra hop. LiteLLM states 8 ms at the 95th percentile at 1,000 requests per second, and Portkey states under 1 ms; both are vendor figures measured under their own conditions, so measure your own.
  • Provider features arrive later. A new provider feature may not be available through the common format until the gateway adds it. Check streaming and function calling through the gateway before relying on either.
  • Behavior differs between models. A fallback sends the same prompt to a different model. Its answer style, refusals and tool-call habits may differ, and LiteLLM's documentation notes that encrypted reasoning from a failed deployment cannot be carried to another provider. Test the backup model with your own questions.
  • Customer data passes through it. A hosted gateway is one more party handling conversations. Review its retention and security terms alongside your data retention policy and LLM security practice.
  • Guardrails are not a substitute for design. A gateway's content checks supplement, and do not replace, the controls described under AI guardrails.

FAQ

What is an LLM gateway?

A service between your application and model providers that offers one endpoint and adds fallbacks, retries, limits, caching and logging, so the application does not call each provider separately.

Is an LLM gateway the same as an API gateway?

No. An API gateway manages any HTTP service and does not understand tokens, models or streaming. An LLM gateway is built for model calls and knows about those things.

What is the difference between an LLM gateway and an LLM router?

A router chooses which model answers a request. A gateway is the wider layer that also handles keys, limits, fallbacks and logs, and it often includes a router.

Do I need one for a customer support chatbot?

Not for one model from one provider at modest volume. You likely need one when you want automatic failover to a second provider, when several models are in use, or when you need per-team spend caps.

Does a gateway make answers cheaper?

It can, through caching and routing to a cheaper model, but only if you configure those. A gateway that only forwards requests adds a hop and no saving.

Is LiteLLM free?

Its repository license file says MIT for most of the code, with a separate license for an enterprise directory. Running it still costs hosting and your time, so check the current licensing and pricing pages.

Sources