Skip to content
Chatbotscape
Anthropic's Contextual Retrieval research and the Sentence Transformers retrieve-and-rerank documentation were read on 28 September 2026, and every figure below carries that date. Reranking models and their pricing change quickly; the pattern described here is stable, the model names are not.
Reranking· AI infrastructure
Reranking is a second scoring pass inside a retrieval-augmented generation (RAG) system. The first search — usually embeddings, keywords, or both — quickly pulls back a generous shortlist of candidate passages; a reranker then reads the question and each candidate together, re-orders the shortlist by how well each passage actually answers it, and keeps only the top few for the language model to read. It trades a little speed for a lot of precision, and it only ever re-orders what the first search found.
By Chatbotscape Editorial· Methodology· Published 29 September 2026· Updated 29 September 2026

Reranking — Re-Scoring the Shortlist Before a Chatbot Reads It

Quick answer: Reranking is the step that sits between retrieval and the language model in a RAG chatbot. A fast first search returns, say, 150 loosely relevant passages; a slower, more accurate reranking model scores each one against the question and hands the model only the best 20 or fewer. In Anthropic's published tests, adding a reranker to a pipeline that already combined contextual embeddings and keyword search cut the rate of failed retrievals from 2.9% to 1.9%, and the full technique stack reduced failures by 67% against a plain embedding baseline of 5.7%.

Why a second pass exists at all

The first search has to be fast, because it runs over everything you have ever uploaded. Embedding search achieves that by encoding the question and every passage separately, ahead of time, and comparing the resulting vectors. The Sentence Transformers documentation calls this a bi-encoder: it "independently" encodes queries and documents, which is what makes searching a large collection cheap — and also what limits it, because the question and the passage never get read side by side.

A reranker gets that side-by-side read. The same documentation describes the cross-encoder, which processes the query and a candidate passage jointly and so produces a more accurate relevance score. The cost is that it must run once per candidate, which is why it is "computationally expensive" if pointed at a whole knowledge base and why it is used only as the second stage, on a shortlist. Two stages, then: cheap-and-broad to find candidates, expensive-and-precise to order them.

What it looks like in a pipeline

Anthropic's Contextual Retrieval write-up gives one concrete, published configuration. It performs initial retrieval to get the top 150 potentially relevant chunks, passes them through a reranking model along with the user's query, and then selects the top 20 chunks to pass to the language model. Those numbers are one team's choice, not a standard: the Sentence Transformers documentation, for one, does not prescribe how many candidates to rerank, and the right shortlist size depends on your content, your latency budget, and what your reranker charges per document scored.

Microsoft's Azure AI Search documentation shows the same shape from a managed-service angle: its hybrid queries return a fused result set, and it advises setting the number of nearest-neighbor results to 50 "to maximize" the inputs of its semantic ranker, the reranking layer that scores the merged results and returns a separate reranker score alongside the original search score.

What the measured gain looks like

Anthropic's published failure rates measure how often the correct passage is missing from the top 20 retrieved. Its plain-embedding baseline failed 5.7% of the time; contextual embeddings alone brought that to 3.7%; combining contextual embeddings with a keyword index (BM25) brought it to 2.9%; and adding reranking brought it to 1.9%. In Anthropic's words, reranked contextual embeddings and contextual BM25 "reduced the top-20-chunk retrieval failure rate by 67%." Two cautions belong next to that figure. It comes from Anthropic's own evaluation datasets, not from your documents, and the improvement from the reranking step alone (2.9% to 1.9%) is smaller than the headline suggests, because most of the gain came from the layers beneath it.

What our reviews show about who discloses a reranker

Most platforms hide this layer completely. Our hands-on Intercom review is the exception in our corpus: it records that Intercom's Fin AI Agent pairs a proprietary retrieval model, "fin-cx-retrieval," with a proprietary reranker, "fin-cx-reranker," and scores that disclosure as the most transparent RAG stack in the helpdesk-with-bot category we tested, while noting the underlying foundation model is not publicly named. In the review's 15-question knowledge-base test, Intercom's answer accuracy landed at 80–87%. We did not isolate the reranker's contribution, so that result should not be read as proof that a reranker caused it. For the other platforms we cover, including Botpress, retrieval internals are not documented in enough detail to say whether a reranker is present.

What reranking is not

It is not retrieval. A reranker can only re-order what the first search returned. If the right passage is not in the shortlist, no reranker brings it back — that is a chunking, hybrid-search, or content problem.

It is not a fix for hallucination. A better-ordered shortlist raises the odds that the model sees the right passage; it does not stop the model from inventing an answer when the passage is absent. That is the AI hallucination problem, handled by grounding rules and refusal paths.

It is not the same as sorting by the first score. Reordering by the original similarity score is just retrieval. Reranking means a separate model re-scored each candidate against the question.

FAQ

What is reranking in RAG?

A second scoring stage that re-orders the passages a first search returned, using a model that reads the question and each passage together. The top few after reordering are what the language model sees. It improves precision without changing what was found.

Do I need a reranker for my chatbot?

Usually only if you control your own retrieval pipeline and have measured that the right passage is in your candidates but ranked too low. If you use a platform's built-in knowledge-base feature, you cannot add one yourself; you can only ask the vendor whether one exists.

What is the difference between a retriever and a reranker?

The retriever searches everything fast and returns a broad shortlist. The reranker scores only that shortlist, slowly and more accurately. Retrievers optimize recall (do not miss it); rerankers optimize precision (put the best first).

What is a cross-encoder?

A model that takes the question and one candidate passage as a single input and outputs a relevance score. It is more accurate than comparing separately computed embeddings but too slow to run over a whole knowledge base, which is why it is used as the second stage.

Does reranking add latency and cost?

Yes: it is an extra model call per query, scoring every candidate in the shortlist. Prices and speeds vary by provider and change often, so check your provider's current documentation rather than a figure printed here.

Sources

  • Anthropic, Introducing Contextual Retrieval — anthropic.com/news/contextual-retrieval, read 28 September 2026. Source of the 5.7% / 3.7% / 2.9% / 1.9% failure rates, the 67% reduction, and the top-150-to-top-20 reranking setup.
  • Sentence Transformers, Retrieve & Re-Rank — sbert.net/examples/applications/retrieve_rerank, read 28 September 2026. Source of the bi-encoder versus cross-encoder description and the cost rationale for a two-stage pipeline.
  • Microsoft Learn, Hybrid search overview — learn.microsoft.com/en-us/azure/search/hybrid-search-overview, read 28 September 2026. Source of the guidance to set k to 50 for the semantic ranker and the separate reranker score in results.
  • Chatbotscape, Intercom review — /reviews/intercom-review, re-read 28 September 2026: the fin-cx-retrieval and fin-cx-reranker disclosure, the undisclosed foundation model, and the 80–87% knowledge-base accuracy result.
  • Chatbotscape evaluation methodology. /methodology (continuously updated).