Skip to content
Chatbotscape
Editorial flat-vector illustration for How to Add a Semantic Cache to a Support Chatbot Without Serving Wrong Answers
Published 8 min read

How to Add a Semantic Cache to a Support Chatbot Without Serving Wrong Answers

Quick answer: Add the cache only for the first message of a conversation, only for general questions that have the same answer for every customer, and start with a strict similarity threshold. Test it on pairs of real questions that should and should not match, loosen it only while wrong matches stay at zero, give every entry an expiry, and then measure your own hit rate. In our invented example below the saving is about $14 a month, which is why the last step is deciding whether to do it at all.

Step 1: Check that your traffic is repetitive enough

A cache only pays when the same questions come back. Pull a month of first customer messages (anonymized; see the data privacy guide) and sort them by topic: opening hours, shipping times, password reset, return policy. If a small number of topics covers a large share of the volume, caching can help. If most conversations are about individual orders and accounts, it cannot, because those answers differ per customer.

Write down two numbers: first messages per month (N) and the share that are general, repeatable questions. That share is the ceiling on your hit rate.

Step 2: Decide what may be cached

Use these rules. The first three follow Redis's published guidance for its semantic cache; the last two are our own.

  1. First message only. Redis's documentation advises caching only the first message of a conversation, with its first_message_only=True setting, because later messages depend on the earlier turns.
  2. Never cache errors or tool-call output. The Redis page says function-call responses and errors are excluded automatically.
  3. Only cache tools that are safe to repeat. Its example is a weather lookup, not sending an email.
  4. Never cache anything personal such as order status or balance, because the stored answer would be returned to another customer.
  5. Only cache answers a person has approved. The cache repeats whatever it stored, including an AI hallucination. Seed it from reviewed answers, or cache only responses that were grounded in your knowledge base and spot-checked.

Step 3: Build the cache and start strict

The flow is the same in every tool. The Redis page describes it as: check for a similar prompt before the model call; on a match return the stored response; otherwise call the model and store the result. The GPTCache README breaks it into an embedding step, a vector store lookup and a similarity evaluation. You can build it with a library such as GPTCache or Redis's tools, with a vector database you already run, or with an AI gateway that offers it as a setting. Check the maintenance status first: the GPTCache README says its maintainers no longer add support for new model APIs.

Set the threshold strict. In Redis's example the distance scale runs from 0.0 (exact match only) to 0.1 (small phrasing variations match), and higher values bring more hits and more wrong answers. Its page gives no default. Begin at the strict end for your tool and treat the number as something you will tune, not copy.

Step 4: Test the threshold with real question pairs

This is the step most teams skip, and it is the one that prevents wrong answers. It is our method, not a vendor's.

  1. Write 40 pairs of questions that should match, such as "How do I reset my password?" and "I forgot my password, what do I do?"
  2. Write 40 pairs that look alike but must not match: "Can I cancel my order?" and "Can I cancel my subscription?", or hours on a Sunday and hours on a Monday, or two different return windows.
  3. Run all 80 through your embedding and record the distance for each pair.
  4. Pick the loosest threshold at which none of the must-not-match pairs match. If the two groups overlap, no threshold is safe and you need a verification step (below) or a stricter cache.

Keep these pairs. They are a small golden dataset you rerun whenever you change the embedding model or the threshold, as in the regression testing guide. If the pairs overlap, a common fix is to add a second check after the lookup, such as a small model or a rule that confirms the candidate really answers the new question. That adds cost and delay, so recheck the saving in Step 6.

Step 5: Set expiry and invalidation

A cached answer about prices, hours or policy goes stale. Give every entry a time to live; Redis's example sets ttl=3600, though its page does not document the units, so confirm in the reference for your tool. Also clear entries whenever the source document changes: if you update the returns policy in the knowledge base, delete every cached answer that came from it. If your tool returns no entry ID for a cached item, Redis warns you cannot delete that entry individually, so confirm you can purge by topic or clear the whole cache.

Step 6: Measure, then check the arithmetic

Log three numbers every week: hit rate, wrong hits found in a sample review, and response time for hits against misses. GPTCache's README recommends tracking hit ratio, latency and recall. Feed them into the monitoring routine in the production monitoring guide.

Then do the sum. Every number below is invented for illustration, not a measurement or a vendor price.

ItemValue
First messages per month (N)20,000
Model cost per answered first message$0.004
Cost without cache20,000 × $0.004 = $80.00
Hit rate (assumed)30% = 6,000 hits
Model cost avoided6,000 × $0.004 = $24.00
Embedding cost for 20,000 lookups at $0.00001 each$0.20
Cache hosting (assumed)$10.00
Net monthly saving$24.00 − $0.20 − $10.00 = $13.80
If 2% of hits are wrong6,000 × 0.02 = 120 wrong answers a month

Two lessons from the invented table. First, the saving is small at this volume, and the hosting line eats nearly half of it; the case improves with volume, longer answers or costlier models. Second, the wrong-answer line is the real cost: 120 wrong replies a month is a support burden and a trust problem that the $13.80 does not cover. The rule we use: if the threshold test in Step 4 cannot reach zero wrong matches, or the net saving is small, skip the cache. Cheaper first steps, described in the reduce chatbot costs guide, include shortening the prompt and using provider-side prompt caching. To size your own inputs, use the method in the token usage and cost estimate guide.

Common mistakes

  • Caching every turn. Mid-conversation messages depend on context. Cache first messages only.
  • Copying a threshold from a tutorial. It depends on your embedding model and your questions. Test it.
  • Never expiring entries. A stale answer is served confidently until someone notices.
  • Caching personal answers. One customer's order status goes to another.
  • Counting the hit rate as the win. A high hit rate with wrong hits is worse than a low one with none.
  • Reporting vendor or paper numbers as your own. One arXiv paper reports up to 68.8 percent fewer calls on its query sets; your repeat rate is whatever your logs show.

FAQ

How do I add semantic caching to a chatbot?

Put a cache in front of the model call: embed the first customer message, search for a close stored question, return the stored answer if it is within a strict threshold, otherwise call the model and store the result. Then test the threshold and set expiry.

What similarity threshold should I use?

Start at the strict end for your tool and test with pairs of questions that should and should not match. Loosen only while wrong matches stay at zero in your test set.

How do I prevent false positives?

Cache only first messages and general answers, test with look-alike question pairs, keep the threshold strict, and add a verification check if the pairs overlap.

Is semantic caching the same as prompt caching?

No. Prompt caching is a provider discount on a repeated prompt prefix and still generates a fresh answer. A semantic cache skips the model call by returning a stored answer.

Does it work in a RAG chatbot?

Yes, but clear cached answers when the underlying documents change, otherwise the cache keeps serving the old version. See the RAG chatbot build guide.

When should I not use one?

When your questions are mostly about individual orders or accounts, when your volume is low, or when you cannot get to zero wrong matches in your test.

Sources

  • Redis, Semantic caching for Google ADK — redis.io/docs/latest/integrate/google-adk/semantic-caching/, read 8 October 2026: the check-then-store flow, the distance threshold scale, TTL, first-message-only, excluded errors, idempotent tools, entry ID caveat.
  • GPTCache (Zilliz), README — github.com/zilliztech/GPTCache, read 8 October 2026: embedding, vector store and similarity evaluation parts; false positives and negatives; metrics to track; maintenance note.
  • Regmi and Pun, GPT Semantic Cache — arxiv.org/abs/2411.05276, read 8 October 2026: reported call reduction of up to 68.8 percent, the authors' results.
  • Chatbotscape evaluation methodology. /methodology (continuously updated).

About this guide

Chatbotscape launched in 2026 as an independent review site for chatbot platforms. This guide is editorial guidance built from vendor and project documentation; the question-pair test, the rules on what to cache and the worked example are our own practice with invented numbers, and we have not benchmarked any caching tool.

Last updated

9 October 2026 — first published.