
How to Evaluate a RAG Chatbot
Score Retrieval and Answers Separately, Then Calibrate the Judge
Quick answer: Evaluate a RAG chatbot in two layers, because it fails in two different places. Layer one is retrieval: did the passages the bot was shown contain what the question needed? Layer two is generation: did the answer stick to those passages (faithfulness) and actually address the question (answer relevance)? Build a set of 30 to 50 real questions, score each layer separately, let a language model do the reading only after you have compared its verdicts with your own on ten answers, and read the result as a map of which layer to fix.
Step 1: Split the system into the two things that can go wrong
The Ragas paper (Es et al.), which introduced a widely used open-source evaluation framework, divides a RAG pipeline into three dimensions: "the ability of the retrieval system to identify relevant and focused context passages," "the ability of the LLM to exploit such passages in a faithful way," and "the quality of the generation itself." TruLens frames the same split as the RAG triad: context relevance, groundedness and answer relevance.
For a business owner the practical version is two questions. Was the right material in front of the model? If it was, did the model use it honestly and answer what was asked? A single "accuracy" number blends both, so a drop in it cannot tell you whether to touch your documents, your search settings or your prompt.
Step 2: Build a question set that looks like your customers
Take 30 to 50 questions from real chat logs or your inbox, not ones you invented. For each, write down the passage or article that should answer it, and a one-line correct answer. Include the awkward ones: questions with exact order numbers or plan names, questions your documents do not answer at all, and questions phrased the way customers actually type them.
This is the same artifact our golden dataset entry describes, and if you already keep one, reuse it. The evaluation here adds one column to it: for each question, which passage was the right source. Without that column you can grade answers but not retrieval.
Step 3: Score retrieval (layer one)
For each question, capture the passages the bot retrieved. Most platforms show the sources behind an answer; if yours does not, that is itself a finding, and you can only grade the answer layer.
Two published metrics cover this layer. TruLens's context relevance asks whether "each chunk of context is relevant to the input query." Ragas's context precision measures "the retriever's ability to rank relevant chunks higher than irrelevant ones," on a 0 to 1 scale where 1.0 means every retrieved context is relevant.
You do not need the formula to use the idea. For each question, mark whether the right passage was retrieved, and whether it was first. Three outcomes result: absent, present but not first, present and first. The same tally opens our hybrid search and reranking guide, which is the next stop if many questions land in the first two.
Step 4: Score the answer (layer two)
Two metrics, and they check different things.
Faithfulness asks whether the answer's claims are supported by the retrieved passages. Ragas defines it as how "factually consistent a response is with the retrieved context," computed as the number of claims supported by the context divided by the total claims in the response. A worked example on that formula: an answer that makes five separate claims, four of which appear in the retrieved passages, scores 4 / 5 = 0.8. The unsupported fifth claim is the hallucinated one, and scoring per claim points you straight at it. TruLens's groundedness works the same way, separating the response "into individual claims" and searching for evidence for each.
Answer relevance asks whether the answer addresses the question. Ragas computes it by generating artificial questions from the answer and measuring how close they are to the real question; the documentation is explicit that it does not assess factual correctness, and that it penalizes incomplete or overly detailed answers. An answer can be perfectly relevant and completely wrong, which is why it never stands alone.
A third check is worth adding that neither framework replaces: correctness against your reference answer, from the one-line answer you wrote in Step 2. Faithfulness only proves the answer matches what was retrieved; if the wrong passage was retrieved, a faithful answer is faithfully wrong.
Step 5: Let a judge do the reading, after you check it
Grading 40 answers on three measures by hand is around 120 judgments, so most teams delegate to a language model. Three rules from the published guidance keep that honest. Use a rubric specific enough to apply mechanically. Ask the judge to reason before it gives a verdict; Anthropic's evaluation documentation says this "increases evaluation performance, particularly for tasks requiring complex judgment." And grade with a different model from the one that wrote the answers; the same page calls that generally best practice, and the MT-Bench paper documents why: judges favor their own model's writing, favor longer answers, and favor whichever answer they see first.
Then calibrate. Grade ten of the answers yourself and compare with the judge's verdicts. This is the rule from our golden dataset entry, and it is our rule, not a published standard: if the judge agrees with you on fewer than eight of ten, fix the criteria before trusting it on the rest. If you compare two versions of the bot with the judge, run each comparison twice with the answers in swapped order and count a win only when the same version wins both times, the conservative rule the MT-Bench authors suggest for position bias.
Step 6: Read the result as a map
Put each question into one cell of a two-by-two grid.
| Answer faithful to what was retrieved | Answer not faithful | |
|---|---|---|
| Right passage retrieved | Healthy. Check answer relevance and correctness for polish | The model had the answer and ignored it. Fix the prompt and grounding rules: reduce chatbot hallucinations |
| Right passage missing | Faithfully wrong: the model answered honestly from bad material. Fix retrieval, chunking or content: hybrid search guide, knowledge base guide | Both layers failed. Fix retrieval first; the answer cannot be trusted until the passage is right |
Count questions per cell and write the counts down. That table is the baseline. After any change, re-run the same questions and compare cells, not one overall number.
If you use a platform and cannot see the retrieved passages
Then you can grade only the answer layer, and only against your reference answers. Ask the vendor in writing whether the retrieved sources are exposed per answer, and check in a pilot rather than from a demo, as our vendor evaluation checklist explains. Two platforms in our corpus ship built-in judges: per our golden dataset entry, Voiceflow and Intercom return verdicts from a model reading against criteria you write, so the calibration in Step 5 applies to them as much as to a script you write yourself.
Common mistakes
One blended score. It cannot say which layer to fix. Keep retrieval and answer results in separate columns.
Trusting a judge you never checked. A model that agrees with itself is not a model that agrees with your team.
Grading answers without the sources. A judge shown only the answer grades fluency. Give it the retrieved passages for faithfulness.
Only testing questions you can answer. Include questions your documents do not cover; the correct behavior there is a refusal or a handoff, and that is where hallucination shows up.
Changing three things and re-testing once. Change one thing, re-run the same set, log the cells.
FAQ
What is RAG evaluation?
Measuring how well a retrieval-augmented chatbot performs by scoring its two stages separately: whether retrieval found the right passages, and whether the generated answer is faithful to them and answers the question.
What is faithfulness in RAG?
The share of an answer's claims that the retrieved passages support. Ragas defines it as claims supported by the context divided by total claims, on a 0 to 1 scale.
Do I need a golden dataset to evaluate RAG?
You need a set of real questions with a known right answer; that is a golden dataset by another name. Ragas can score some metrics without reference answers, but a reference answer is what catches a faithfully wrong result.
How many questions is enough?
30 to 50 real questions is a workable start for an SMB chatbot. That is a practical starting point, not a statistical guarantee; small differences between runs on a set that size should not be over-read.
Can an LLM grade its own chatbot?
It can, but it is biased toward its own writing. Use a different model as the judge and calibrate it on ten answers you graded yourself.
Which tools compute these metrics?
Ragas and TruLens are open-source frameworks that define and compute them, and some chatbot platforms ship their own judge features. Whatever tool you use, check its definition of each metric before comparing numbers across tools.
Related guides
- Chatbot regression testing guide — turning this question set into a check you run on every change.
- Hybrid search and reranking for a RAG chatbot — the fix when retrieval is the failing layer.
- Reduce chatbot hallucinations — the fix when the answer ignores good passages.
- Build a RAG chatbot — the build decisions behind the pipeline you are grading.
- Chatbot QA testing protocol — the wider pre-launch test plan.
- LLM-as-a-judge and golden dataset — the glossary entries behind the terms used here.
Sources
- Es et al., Ragas: Automated Evaluation of Retrieval Augmented Generation — arxiv.org/abs/2309.15217, read 29 September 2026: the three evaluation dimensions and the reference-free design.
- Ragas documentation, Faithfulness, Context Precision, Answer Relevancy — docs.ragas.io, read 29 September 2026: the faithfulness formula, the context precision definition and 0 to 1 range, and the answer relevancy computation and limitation.
- TruLens, The RAG Triad — trulens.org/getting_started/core_concepts/rag_triad, read 29 September 2026: context relevance, groundedness and answer relevance definitions.
- Zheng et al., Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena — arxiv.org/abs/2306.05685, read 29 September 2026: judge biases and the swap-and-agree rule.
- Anthropic, Define success criteria and build evaluations — platform.claude.com/docs/en/test-and-evaluate/develop-tests, read 29 September 2026: rubric, different-model and reason-first guidance.
- Chatbotscape, Golden dataset entry — /glossary/golden-dataset, re-read 29 September 2026: the eight-of-ten calibration rule and the Voiceflow and Intercom judge features.
- Chatbotscape evaluation methodology. /methodology (continuously updated).
About this guide
Chatbotscape launched in 2026 as an independent review site for chatbot platforms. This guide is editorial guidance built from published research and framework documentation plus patterns in our own corpus; the faithfulness example is arithmetic on Ragas' published formula, not a benchmark. Voiceflow and Intercom have Chatbotscape reviews that carry affiliate links; the method here does not depend on any platform. To flag an error, write to editorial@chatbotscape.com.
Last updated
30 September 2026 — first published.