
How to Set Up an LLM Gateway for a Support Chatbot (With a Backup Provider)
Quick answer: Put the gateway between your chatbot and the model, define a backup model on a different provider, set a short retry count and a timeout, give the chatbot its own key with a monthly budget, and then deliberately break the main provider in a test to prove the backup answers. In our invented example the expected saving is about $81 a month, almost all of it from avoiding a rare outage, which is why the first step is deciding whether you need a gateway at all.
Step 1: Decide whether you need one
Answer four questions. Two or more yeses make a gateway worth trying.
- Would an hour of the model provider being down stop your chatbot, and does that cost you real money or customers?
- Do you already use, or plan to test, more than one model, as in testing a smaller model?
- Do several teams, clients or customer-supplied keys need separate spend caps?
- Do you want one record of cost across every model call?
If all four are no, call the provider directly, set a budget alert in the provider's console and revisit later. A gateway you do not need is one more service to secure, update and monitor.
Step 2: Pick a gateway and check the basics
Options differ in whether you host them or rent them. From their own documentation, read on 10 October 2026: LiteLLM and Portkey publish open-source gateways you can run yourself; OpenRouter is a hosted service where you pass an ordered list of models; Helicone describes a hosted unified API with failover. Before choosing, check four things on the project's own pages: the license (LiteLLM's license file is MIT with a separate license for an enterprise directory), whether the project is actively maintained, whether it supports the features you use through the gateway (streaming, function calling), and how it handles customer data. A hosted gateway is one more party handling your conversations, so review its terms against your data privacy guide before you send real traffic.
Step 3: Define the main model and a backup on another provider
A backup on the same provider does not help when the provider itself is down, so choose one on a different provider. Choose a model that is good enough, not necessarily equal; you are buying continuity. Then test the backup with your own questions, because it will read the same prompt differently. Reuse the 100-question test and pass bar from the smaller-model guide, and keep the backup's system prompt behavior in mind: refusals, tone and tool-call habits can change.
Step 4: Write the configuration
The shape below follows the keys on LiteLLM's reliability page. The model names are placeholders, and this snippet is our illustration, not a tested configuration, so check it against the current documentation.
model_list:
- model_name: support-main
litellm_params:
model: <provider-a>/<main-model>
api_key: os.environ/PROVIDER_A_KEY
- model_name: support-backup
litellm_params:
model: <provider-b>/<backup-model>
api_key: os.environ/PROVIDER_B_KEY
litellm_settings:
num_retries: 2
request_timeout: 10
fallbacks: [{"support-main": ["support-backup"]}]
allowed_fails: 3
cooldown_time: 30
What the keys do, per LiteLLM's page: num_retries retries a group before failing over, request_timeout raises a timeout after that many seconds, fallbacks lists backup groups tried in order, and allowed_fails with cooldown_time set how many failures put a deployment into a cooldown and for how many seconds. The chatbot then calls support-main and never names a provider. Choose the timeout from your own latency: it must be longer than a normal answer, or healthy calls will fail over, and short enough that a customer does not wait through a retry and a fallback. Gateways that work differently, such as OpenRouter, take an ordered list of models in the request, and its documentation says you are priced for the model that actually served the request.
Step 5: Give the chatbot its own key, with a budget
Do not give the chatbot your master provider key. In LiteLLM, virtual keys are created with parameters such as max_budget, budget_duration, rpm_limit and tpm_limit, and its documentation lists a Postgres database as a requirement for key management. Create one key per chatbot or per client, restrict it to the models it needs, and set a monthly budget a little above your expected bill (see the token usage and cost estimate for how to size it). A budget that stops service is a risk of its own, so decide in advance who is alerted before the cap and what the chatbot says to customers if the cap is hit. Keep the master key out of the application and out of chat transcripts, per the security checklist.
Step 6: Force a failure, then measure
This is our test, not a vendor's, and it is the step that proves the setup.
- Break the main model on purpose in a test environment: a wrong key, or the provider's address blocked. Do not do this in production.
- Send 20 test questions. All 20 should be answered by the backup, within your timeout.
- Check the log: it should show the failed attempts, the fallback and the model that answered. If you cannot tell which model answered, fix that first.
- Restore the main model and confirm traffic returns to it after the cooldown.
- Send the same questions with streaming on, and with a tool call if your bot uses them. Gateways can handle these differently.
- Time a normal answer directly and through the gateway. The difference is your added delay; vendor latency claims are measured under their own conditions.
Then add the gateway's log to your production monitoring routine: fallback rate, error rate and spend per key. A fallback rate that stays above zero means the main provider or your timeout needs attention.
Check the arithmetic
Every number below is invented for illustration, not a measurement or a vendor price.
| Item | Value |
|---|---|
| Conversations per month | 28,800 (40 per hour, around the clock) |
| Main-provider outage assumed | one 4-hour outage every 3 months |
| Conversations hit by one outage | 4 × 40 = 160 |
| Share of those customers who then contact a human | 30% = 48 |
| Cost of each human contact (assumed) | $6.00 |
| Cost of one outage without a gateway | 48 × $6.00 = $288.00 |
| Extra model cost on backup (assumed $0.002 more per conversation) | 160 × $0.002 = $0.32 |
| Expected monthly benefit | $288.00 ÷ 3 = $96.00 |
| Expected extra backup cost per month | $0.32 ÷ 3 ≈ $0.11 |
| Gateway hosting (assumed) | $15.00 |
| Expected net monthly benefit | $96.00 − $0.11 − $15.00 ≈ $80.89 |
Two lessons. The benefit is a probability, not a saving you see on an invoice: in a month with no outage you pay $15 and gain nothing visible. And the result flips with your inputs: at a tenth of the traffic, the same outage costs about $28.80 and the expected benefit is $9.60 a month, below the hosting line. The rule we use: if an outage would not noticeably cost you, or if you will not run the forced-failure test, skip the gateway.
Common mistakes
- A backup on the same provider. It fails with the main one.
- Never testing the backup. It may refuse or phrase things differently from the main model.
- A timeout shorter than a normal answer. Healthy calls fail over and you pay twice.
- Using one master key everywhere. One leak or runaway loop spends the whole account.
- Treating the gateway as the whole safety layer. Its checks supplement your guardrails.
- Forgetting the gateway can fail. Monitor it, and keep a documented way to point the chatbot straight at a provider.
FAQ
How do I set up an LLM gateway for a chatbot?
Run or rent a gateway, define a main model and a backup on another provider, set retries and a timeout, issue the chatbot its own budgeted key, point the chatbot at the gateway, and test with a forced failure.
How do I fail over between OpenAI, Anthropic and Google?
Define each as a separate model group in the gateway and list the backups in order in its fallback setting. The chatbot calls one group name and the gateway handles the switch.
Which open-source gateway should I choose?
Compare LiteLLM and Portkey on license, maintenance, the features you use and how you will host them. We have not benchmarked either, so run your own test.
Does a gateway add latency?
Yes, one extra hop. Vendors publish low figures under their own conditions, so time a normal answer directly and through the gateway.
Does the backup model give the same answers?
No. Test it with your own questions, as the same prompt can produce different style, refusals and tool use.
Can a gateway also cut my bill?
Only if you configure caching or routing to a cheaper model. See the semantic cache guide for one such lever.
Related guides
- Reduce chatbot costs guide — all the levers, ranked.
- Test a smaller model for your chatbot — the test that gives a gateway something to route.
- Add a semantic cache — a cost lever some gateways offer.
- Monitor an LLM chatbot in production — where fallback rate belongs.
- LLM gateway, prompt caching and LLM observability — the glossary entries behind the terms used here.
Sources
- LiteLLM, Fallbacks, retries and cooldowns — docs.litellm.ai/docs/proxy/reliability, read 10 October 2026: num_retries, request_timeout, fallbacks, allowed_fails, cooldown_time.
- LiteLLM, Virtual keys — docs.litellm.ai/docs/proxy/virtual_keys, read 10 October 2026: budgets, rate limits, database requirement.
- LiteLLM, GitHub repository and LICENSE — github.com/BerriAI/litellm, read 10 October 2026.
- Portkey, AI Gateway repository — github.com/Portkey-AI/gateway, read 10 October 2026.
- OpenRouter, Model fallbacks — openrouter.ai/docs/guides/routing/model-fallbacks, read 10 October 2026.
- Helicone, AI Gateway overview — docs.helicone.ai/gateway/overview, read 10 October 2026.
- Chatbotscape evaluation methodology. /methodology (continuously updated).
About this guide
Chatbotscape launched in 2026 as an independent review site for chatbot platforms. This guide is editorial guidance built from vendor documentation; the forced-failure test, the questions in Step 1 and the worked example are our own practice with invented numbers, and we have not benchmarked any gateway.
Last updated
11 October 2026 — first published.