Skip to content
Chatbotscape
Editorial flat-vector illustration for Hybrid Search and Reranking for a RAG Chatbot: Fix Wrong-Document Answers in the Right Order
9 min read

Hybrid Search and Reranking for a RAG Chatbot

Fix Wrong-Document Answers in the Right Order

Quick answer: When a RAG chatbot answers from the wrong document, fix retrieval in three steps, cheapest first. First, measure: for 20 to 30 real questions, check whether the correct passage appears in the top results at all. Second, if it is missing on questions containing exact codes, names or jargon, add keyword search alongside embeddings — hybrid search — and merge the two result lists. Third, only if the correct passage is being found but ranked too low, add a reranker. In Anthropic's published tests, contextual embeddings plus keyword search cut failed retrievals from 5.7% to 2.9%, and adding reranking took that to 1.9%. On a platform where you cannot change these settings, your lever is the question you ask the vendor and the content you write.

Step 1: Find out whether it is a retrieval problem

A wrong answer has two very different causes: the model was shown the wrong passages, or it was shown the right one and answered badly. Only the first is a search problem, and mixing them up is how teams spend a month tuning the wrong layer.

Build a small test set: 20 to 30 real customer questions, each paired with the article or passage that should answer it. Most platforms show the sources behind an answer; if yours does, log for each question whether the correct source was present, present but not first, or absent. Three patterns follow:

Absent. The right passage never reached the model. That is a retrieval or content problem: the passage was cut badly (see chunking), the wording does not match how customers ask, or the query contained an exact string the search could not match. Steps 2 and the content fixes below apply.

Present but not first. Search found it and something else outranked it. That is what a reranker is for (Step 3).

Present and first, answer still wrong. Retrieval worked. Look at the prompt, the grounding rules and the hallucination playbook; more search tuning will not help.

Write the tally down. It is the baseline you will compare against after every change, and it stops "it feels better" from standing in for evidence.

Step 2: Add keyword search when exact strings fail

Embedding search matches meaning, which makes it good at paraphrase and poor at exact identifiers. Microsoft's Azure AI Search documentation says so directly: scenarios such as "querying over product codes, highly specialized jargon, dates, and people's names" perform better with keyword search "because it can identify exact matches." If your absent-passage failures cluster on order numbers, error codes, SKUs, plan names or people's names, this is your fix.

Hybrid search runs both searches for each question and merges the results. Microsoft's implementation runs full-text and vector search in parallel and returns one result set. The interesting part is the merge. A keyword engine and an embedding engine score on completely different scales, so you cannot just add the scores. Reciprocal Rank Fusion, RRF, sidesteps that by using only each result's position. Elastic's documentation gives the formula: each result earns 1 divided by a constant plus its rank in each list, and the scores are summed. The constant defaults to 60, and Elastic notes the approach "requires no tuning," which is why it is the usual starting point rather than hand-weighting the two engines.

A worked example with that formula: a passage ranked 1st by keywords and 3rd by embeddings scores 1/61 + 1/63, about 0.0323. A passage ranked 2nd by both scores 1/62 + 1/62, about 0.0323 as well. A passage ranked first by only one engine and absent from the other scores about 0.0164 — below both. Agreement between the two engines beats a single strong vote, and that is the mechanism that rescues the exact-code question while leaving paraphrase questions alone.

When hybrid is not worth it. If your content has no identifiers customers type, and your failures are all paraphrase misses, keyword search adds little; fix the wording gap first (add the phrases customers use, which the semantic search guide walks through). Hybrid also adds a second index to maintain if you run your own stack.

Anthropic's numbers show the size of the gain on their test data: contextual embeddings alone reduced the top-20 failure rate from 5.7% to 3.7%; adding a BM25 keyword layer made it 2.9%, a 49% reduction against the baseline once contextual retrieval was included. These are Anthropic's datasets, not yours, so treat them as a reason to test, not a forecast.

Step 3: Add a reranker only for "present but not first"

A reranker re-scores the shortlist with a model that reads the question and each candidate together; the reranking entry covers the mechanism. Anthropic's configuration retrieved 150 chunks, reranked them, and passed the top 20 on, taking failures from 2.9% to 1.9%. The Sentence Transformers documentation is explicit that this second stage exists because the more accurate cross-encoder is too expensive to run over everything.

Add one when your Step 1 tally shows many "present but not first" results, or when you pass only a few passages to the model and rank matters a lot. Skip it when failures are mostly "absent": a reranker cannot promote a passage that was never in the shortlist. It also costs an extra model call per question, and neither Anthropic nor Microsoft publishes a universal latency or price figure to plan around, so measure it on your own traffic before committing.

Microsoft's guidance for its managed reranking layer is a useful practical detail: set the number of nearest-neighbor results to 50 "to maximize" the semantic ranker's inputs. Whatever service you use, the shortlist you hand a reranker sets its ceiling.

The order matters, and so does re-testing

The sequence is deliberate: measure, fix chunking and content, add keyword search for identifier failures, then rerank for ordering failures. Each step addresses a different failure type, so adding all three at once makes it impossible to tell which one helped. Change one, re-run the same 20 to 30 questions, log the tally, and keep the change only if it moved.

Two cheaper fixes belong before any of this: restructuring how your documents are cut, and rewriting passages so each one carries the context to stand alone. Anthropic's Contextual Retrieval, which prepends a short explanation of where a chunk sits in its document, is the best-documented version of that idea.

If you use a platform and cannot change any of this

Most no-code platforms embed the whole retrieval stack and expose no dials. Your levers are then the test set, the content, and the vendor conversation. Ask, in writing, before you buy: does retrieval combine keyword and semantic search? Is there a reranking step? Can I see which sources were retrieved for each answer? Our vendor evaluation checklist explains why to verify these answers in a pilot rather than accept them from a demo.

Disclosure varies. Our hands-on Intercom review records a named retrieval model and a named reranker in Fin's stack and scored its disclosure highly, and measured 80–87% answer accuracy on a 15-question knowledge-base test; we did not isolate what the reranker contributed. For Botpress and most others in our corpus, the review could not confirm from documentation whether either layer is present, which is itself a finding worth acting on: run your own test set instead of assuming.

Common mistakes

Tuning search when the content is the problem. Two overlapping articles or a heading that promises more than the body delivers will defeat any search stack.

Judging by one impressive demo question. Twenty to thirty real questions, including exact-code ones, is the minimum that means anything.

Adding a reranker to fix hallucination. A better-ordered shortlist does not stop a model from inventing an answer when nothing relevant was found.

Skipping the re-test. Without the tally, you cannot tell an improvement from noise.

FAQ

What is hybrid search in a RAG chatbot?

Running keyword search and embedding search together for each question and merging the two ranked lists into one, commonly with Reciprocal Rank Fusion. Keyword search catches exact codes and names; embeddings catch paraphrases.

Is hybrid search better than vector search?

For content and questions with exact identifiers, jargon or names, usually yes. For pure paraphrase-style questions over plain prose, the gain may be small. Test both on your own questions.

What is Reciprocal Rank Fusion?

A way of merging ranked lists that uses only each result's position, not its score. Elastic's implementation adds 1 divided by (60 plus rank) for each list a result appears in, so results both engines like rise to the top without any score normalization.

Do I need a reranker if I already use hybrid search?

Only if your test set shows the correct passage is retrieved but ranked below others. In Anthropic's tests the reranker added a smaller improvement (2.9% to 1.9% failures) on top of an already strong hybrid setup.

Can I do any of this on a no-code chatbot platform?

Rarely. You can test the results, improve the content, and ask the vendor whether keyword search and reranking are part of its retrieval. If they are not configurable and your failures are severe, that is a reason to consider a custom pipeline; see the RAG build guide.

Sources

About this guide

Chatbotscape launched in 2026 as an independent review site for chatbot platforms. This guide is editorial guidance built from vendor documentation and published research plus patterns in our own review corpus; the worked RRF example is arithmetic on Elastic's published formula, not a benchmark. Intercom has a Chatbotscape review that carries affiliate links; the ordering advice here does not depend on any platform. To flag an error, write to editorial@chatbotscape.com.

Last updated

29 September 2026 — first published.