Semantic Caching· LLM cost and latency
Semantic Caching — Reusing an Answer When the Question Means the Same Thing
Quick answer: A normal cache needs the same key to get a hit, so "How do I reset my password?" and "I forgot my password, what now?" are two misses. A semantic cache turns each question into an embedding, looks for a stored question whose meaning is close, and if one is close enough returns the stored answer without calling the large language model. It saves the model's cost and waiting time on repeated questions. The price is a new kind of error: a question that only looked similar can get the wrong cached answer.
How a semantic cache works
Redis's documentation for Google's Agent Development Kit describes the loop in three steps: before each model call the cache checks for a semantically similar prompt; if it finds a match, the cached response is returned immediately; if not, the call goes ahead and the response is stored for future lookups. The GPTCache project, an open-source library built for this, splits the same idea into three parts:
- Embeddings. Each query is converted into a vector, a list of numbers that places similar meanings close together.
- A vector store. A similarity search over those vectors finds the nearest stored queries. This is the same technique that powers semantic search over a knowledge base.
- Similarity evaluation. A final check decides whether the nearest stored query is truly a match for the new one.
The agent or chatbot logic does not change. The cache sits in front of the model and answers when it can.
Three kinds of caching that get confused
| What must match | Who runs it | What it saves | |
|---|---|---|---|
| Exact-match cache | The whole request, character for character | Your application | The whole model call, but rarely hits on free text |
| Prompt caching | The beginning (prefix) of the prompt, unchanged | The model provider | A discount on the repeated input; the model still generates a fresh answer |
| Semantic cache | The meaning of the question, within a threshold | Your application or gateway | The whole model call, including the output |
The practical difference: prompt caching makes a repeated prefix cheaper but still produces a new answer every time, so it never returns something stale or mismatched. A semantic cache skips generation entirely, which is why it can save more and why it can go wrong in a way prompt caching cannot. The two can be combined, because a miss in the semantic cache still goes to a model call that prompt caching can discount.
The similarity threshold: the setting that decides everything
A semantic cache has one setting that matters more than the rest: how close is close enough. Redis's documentation calls it distance_threshold and gives the scale for its example: 0.0 means exact match only, 0.1 lets small phrasing variations match, and higher values give more hits but a greater risk of returning a wrong cached response. The page does not state a default, and the right value depends on your embedding model and your questions, so treat any number from a tutorial as a starting point to test, not an answer.
The reason the risk is real: two questions can be near neighbours in embedding space and still need different answers. "Can I cancel my order?" and "Can I cancel my subscription?" read alike; the answers differ. "What are your hours on Sunday?" and "What are your hours on Monday?" may be almost identical to an embedding. GPTCache's own README warns that semantic caches can produce false positives (wrong hits) and false negatives (missed hits), and advises tracking hit ratio, latency and recall.
What is safe to cache, and what is not
Redis's page gives guardrails that apply well beyond its own product:
- Cache the first message of a conversation only. Later messages depend on the earlier turns, so a match on the words alone is unreliable. Its setting for this is
first_message_only=True. - Do not cache errors or function-call responses. The page says these are excluded automatically.
- Cache only tools that are safe to repeat. Its example is a weather lookup, yes; sending an email, no.
Add two rules of our own. Do not cache anything personal to one customer, such as an order status or an account balance, because the cached answer would be returned to someone else. And give every entry an expiry (a time to live, or TTL), because an answer about prices, hours or policy goes out of date. Redis's example sets a ttl of 3600; its page does not document the units, so check the reference for your tool.
What the research and vendors report
An arXiv paper, "GPT Semantic Cache," stores query embeddings in Redis and reports, in its abstract, up to 68.8 percent fewer API calls, cache hit rates between 61.6 and 68.8 percent across query categories, and positive hit rates above 97 percent. Those are one team's results on their own query sets. A support chatbot's repeat rate depends on how repetitive its traffic is, so measure your own before counting on a figure like that. GPTCache's README lists lower cost, faster responses and protection against rate-limit outages as the benefits, and notes that the project is under heavy development and that its maintainers no longer add support for new model APIs, so check the maintenance status of any library before depending on it.
FAQ
What is semantic caching?
A cache for language-model applications that matches new questions to stored ones by meaning rather than exact wording, and returns the stored answer when the match is close enough, skipping the model call.
How is semantic caching different from prompt caching?
Prompt caching is run by the model provider and discounts a repeated, unchanged prompt prefix, but the model still writes a new answer. Semantic caching is run by you and skips the model call by returning a stored answer.
Does semantic caching reduce hallucinations?
No. It returns an answer that was generated earlier, right or wrong. If the original was an AI hallucination, the cache repeats it. Check answers before they are eligible for caching.
What similarity threshold should I use?
There is no universal value. Start strict, test with real pairs of questions that should and should not match, and loosen only while wrong hits stay at zero in your test set.
Can a semantic cache return another customer's information?
Yes, if you cache personalized answers. Cache only general answers that are the same for every customer.
Does it work with RAG?
It can, but cached answers go stale when the knowledge base changes. Expire entries or clear them when the underlying documents are updated.
Related terms
- Prompt caching — the provider-side discount on a repeated prompt prefix.
- Vector embeddings — the numeric form of a question that makes meaning comparable.
- Semantic search — the same similarity lookup, used to find knowledge-base passages.
- LLM token — the unit that a skipped model call stops billing.
- Cost per conversation — the figure a cache hit lowers.
- AI hallucination — the error a cache can repeat.
Sources
- Redis, Semantic caching for Google ADK — redis.io/docs/latest/integrate/google-adk/semantic-caching/, read 8 October 2026. Source of the three-step loop, the distance threshold scale, TTL, first-message-only and idempotent-tool guidance.
- GPTCache (Zilliz), README — github.com/zilliztech/GPTCache, read 8 October 2026. Source of the embedding, vector store and similarity evaluation description, the false-positive caveat and the maintenance note.
- Regmi and Pun, GPT Semantic Cache: Reducing LLM Costs and Latency via Semantic Embedding Caching — arxiv.org/abs/2411.05276, read 8 October 2026. Source of the reported call reduction and hit rates, which are the authors' results.
- Chatbotscape evaluation methodology. /methodology (continuously updated).