LLM-as-a-Judge· AI evaluation
LLM-as-a-Judge — One Model Grading Another
Quick answer: LLM-as-a-judge is a model used as a grader. You give it the customer's question, the chatbot's answer, and a rubric, and it returns a verdict. The 2023 MT-Bench paper found that a strong judge such as GPT-4 can match human preferences with "over 80% agreement", which is why the method spread. The same paper documents four weaknesses: position bias, verbosity bias, self-enhancement bias and limited reasoning. A judge is therefore a measuring instrument you calibrate against your own people's grades before you trust its numbers.
Why anyone grades chatbot answers with a model
A support chatbot produces free text, and free text has no single correct string to compare against. "You can return it within 30 days" and "Returns are accepted for a month" are the same answer, and a string match fails one of them. Human graders handle this well but are slow and expensive; Anthropic's evaluation documentation calls human grading "most flexible and high quality, but slow and expensive," code-based grading "fastest and most reliable" but lacking nuance for complex judgments, and LLM-based grading "fast and flexible, scalable and suitable for complex judgment," with the instruction to "test to ensure reliability first then scale."
That last clause is the whole discipline. A judge is the middle option, and its speed is only useful once you know how far to trust it.
The three ways to run a judge
The MT-Bench paper (Zheng et al.) names three variants, and most tools you will meet are one of them.
Pairwise comparison. The judge sees a question and two answers and must pick the better one or call a tie. Useful when you change something, say a new prompt or a new knowledge-base article, and want to know whether the new answers beat the old ones.
Single-answer grading. The judge scores one answer directly, on a scale or as pass or fail. This is what a nightly regression run uses, because it does not need an old answer to compare against.
Reference-guided grading. The judge is given a reference solution to check the answer against. For a chatbot, the reference is the correct answer from your golden dataset or the policy paragraph it should have used.
The four documented biases, and the fix for each
The same paper defines four weaknesses, and they map onto cheap checks.
| Bias | What the paper says it is | Cheap check |
|---|---|---|
| Position bias | The model "exhibits a propensity to favor certain positions over others" | Run pairwise judgments twice with the answers swapped; the paper's conservative rule is to declare a win only when an answer is preferred in both orders |
| Verbosity bias | The judge "favors longer, verbose responses, even if they are not as clear, high-quality, or accurate" | Put length and directness in the rubric; spot-check that long answers are not winning by length alone |
| Self-enhancement bias | Judges "may favor the answers generated by themselves" | Grade with a different model from the one that wrote the answers; Anthropic's documentation calls that generally best practice |
| Limited reasoning | Judges can be "misled by the provided answers", for example on math they could solve themselves | Give a reference answer, and for numbers or dates have code check the arithmetic |
G-Eval (Liu et al., 2023) reports a Spearman correlation of 0.514 with human judgments on summarization and, in the same abstract, flags "the potential issue of LLM-based evaluators having a bias towards the LLM-generated texts", the same self-enhancement problem seen from another paper.
What a usable judge prompt contains
Three things, all from Anthropic's grading guidance. A rubric that is specific enough to be applied mechanically; its example is that the answer "should always mention 'Acme Inc.' in the first sentence", and if it does not, it is graded incorrect. One rubric per success criterion, because Anthropic notes a use case "might require several rubrics for holistic evaluation." And an instruction to reason before the verdict: "Ask the LLM to reason first before producing an evaluation score, and then discard the reasoning," which the same page says increases evaluation performance on complex judgments.
A support-bot rubric in that spirit: "The answer is correct only if it states the 30-day return window and does not promise a refund timeline that is absent from the policy text. Otherwise incorrect. Think first, then output correct or incorrect."
Judging retrieval-based chatbots: the RAG triad
For a chatbot that answers from your documents, the judge is usually asked three separate questions, which TruLens calls the RAG triad. Context relevance checks that "each chunk of context is relevant to the input query"; groundedness separates the response "into individual claims" and searches for evidence for each in the retrieved context; answer relevance evaluates "the relevance of the final response to the user input." Splitting the verdict this way tells you which layer failed. The full method, and what to do with each failure, is in the RAG evaluation guide.
What our reviews and entries already show
Two platforms in our corpus ship a judge as a product feature. Our golden dataset entry records that Voiceflow's Tests and Evaluations and Intercom's Simulations return a verdict without a person reading the answer, and that apart from Voiceflow's exact-response, routing and tool-call checks these verdicts come from a model reading against criteria you wrote. The consequence stated there holds for every judge: the quality of the verdict is the quality of the criteria. See the Voiceflow and Intercom reviews for the platforms themselves.
What LLM-as-a-judge is not
It is not ground truth. A judge is another model with its own errors. A run that scores 92% tells you the judge agrees with your rubric 92% of the time, not that 92% of answers are right, until you have measured the judge against your own graders.
It is not a replacement for a golden dataset. It grades answers; it does not decide which questions matter. A judge with no fixed question set gives numbers you cannot compare from one week to the next.
It is not a hallucination detector on its own. A judge shown only the answer cannot know whether the answer is supported. Give it the retrieved sources, or it grades fluency instead of truth.
FAQ
What is LLM-as-a-judge?
A method where one language model grades another model's output against written criteria, returning a score, a pass or fail, or a preference between two answers. It was popularized by the 2023 MT-Bench paper as a scalable alternative to human raters.
How accurate is an LLM judge?
The MT-Bench paper reports that strong judges can match human preferences with over 80% agreement on its benchmarks. That figure is from a general chat benchmark, not your support content, so measure agreement on your own questions before relying on it.
Which biases affect LLM judges?
Four are documented: position bias (favoring an answer because of where it appears), verbosity bias (favoring longer answers), self-enhancement bias (favoring its own model's writing) and limited reasoning (being misled by the answer it is grading).
Should the judge be the same model as the chatbot?
Usually not. Anthropic's documentation states it is generally best practice to use a different model to evaluate than the one that produced the output, which reduces self-preference.
Can I use it without writing code?
Yes, on some platforms. Voiceflow and Intercom ship judge-style checks, described in our golden dataset entry. Whatever tool you use, grade ten answers yourself first and compare.
Related terms
- Golden dataset — the fixed question set a judge is run against.
- AI hallucination — the failure a groundedness judge exists to catch.
- Retrieval-augmented generation — the architecture the RAG triad evaluates.
- Chatbot knowledge base — the source material a judge should be shown.
- Reranking — a retrieval stage whose effect the context-relevance check can measure.
Sources
- Zheng et al., Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena — arxiv.org/abs/2306.05685, read 29 September 2026. Source of the three judge variants, the four bias definitions, the swap-and-agree mitigation and the "over 80% agreement" claim.
- Liu et al., G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment — arxiv.org/abs/2303.16634, read 29 September 2026. Source of the 0.514 Spearman figure and the warning about bias toward LLM-generated text.
- Anthropic, Define success criteria and build evaluations — platform.claude.com/docs/en/test-and-evaluate/develop-tests, read 29 September 2026. Source of the three grading methods, the rubric example, the different-model guidance and the reason-first instruction.
- TruLens, The RAG Triad — trulens.org/getting_started/core_concepts/rag_triad, read 29 September 2026. Source of the context relevance, groundedness and answer relevance definitions.
- Chatbotscape, Golden dataset entry — /glossary/golden-dataset, re-read 29 September 2026: the Voiceflow and Intercom judge features and the "quality of the verdict is the quality of the criteria" point.
- Chatbotscape evaluation methodology. /methodology (continuously updated).