Skip to content
Chatbotscape
Editorial flat-vector illustration for Agentic RAG for a Chatbot: Decide If You Need It, Then Add It in Four Steps
7 min read

Agentic RAG for a Chatbot

Decide If You Need It, Then Add It in Four Steps

Quick answer: Add agentic RAG only when your test questions fail because one search cannot answer them: several asks in one message, answers spread over multiple sources, or follow-ups that depend on earlier turns. If your failures are wrong-passage or missing-content problems, fix those first. When you do add it, go in steps: rewrite the question, split multi-part questions, check the retrieved passages and retry once, and only then let the model route between sources. Anthropic's own advice is to find "the simplest solution possible" and add complexity only when needed.

Step 0: Prove you need it

Take 30 real customer questions and run them through your current bot. Log every failure with a cause:

Wrong or missing passage. The right article exists but was not retrieved, or the content is missing. This is not an agentic problem. Work through hybrid search and reranking and content fixes first.

Right passages, wrong answer. Retrieval worked. Use the hallucination playbook.

One search is not enough. The question has several parts ("Can I change plan mid-cycle, and does the refund apply to add-ons?"), relies on earlier turns ("what about the annual one?"), or needs facts from two systems, such as a help center and an order lookup. These are the failures an agentic loop targets. Microsoft's documentation for its agentic retrieval feature lists the same shapes: questions with multiple asks, questions dependent on conversation context, and queries that benefit from rewriting.

If fewer than roughly one in five failures fall in the third group, stop here. The extra latency will cost you more than it returns. That threshold is a rule of thumb, not a published figure; set your own after seeing your tally.

The price of the loop

Every added step is another model call or search. Anthropic says agentic systems "often trade latency and cost for better task performance," and Microsoft says its agentic retrieval "adds latency compared to a single-query pipeline," billed by tokens from both the search service and the language model. Before building, measure your current median and slowest-tenth response time and cost per conversation, so you can see what each step costs. Decide in advance the slowest reply a customer will tolerate on your channel; chat widgets tolerate less than email.

Step 1: Rewrite the question

Cheapest step: one short model call turns the customer's message plus recent turns into a standalone search query ("what about the annual one?" becomes "annual plan refund policy"). It fixes the follow-up failures and most typo problems without any loop. Re-run your 30 questions. If the context-dependent failures disappear, you may not need the next steps.

Step 2: Split multi-part questions

Add a planner call that decides whether the message contains one question or several and, if several, produces one query each. Search them separately and pass all the passages to the answering model. Microsoft's managed version does exactly this, running the subqueries in parallel and reranking each. Cap the number of sub-questions (three is a sensible start) so one rambling message cannot trigger ten searches.

Step 3: Check the results and retry once

Add a judge call: given the question and retrieved passages, is this enough to answer? If yes, answer. If not, rewrite the query once and search again; if still not, hand off to a human or say you do not know. This is the idea behind Corrective RAG, which scores retrieved documents with an evaluator and changes action when confidence is low, and Self-RAG, which has the model reflect on retrieved passages. The retry limit matters: an uncapped loop is how a chatbot runs up a bill and stalls on a customer. One retry, then escalate.

A judge can be wrong in both directions. Sample its decisions weekly: read 20 cases where it approved the passages and 20 where it rejected them.

Step 4: Route between sources

Only now let the model choose which source to query: help center, product catalog, order system, or web. Expose each as a separate tool with a clear description and narrow permissions, and log every tool call. This is the step that adds real operational risk, since the model can now act on live systems; read-only tools first. Our agent orchestration guide covers the wider design questions.

Guardrails that apply to every step

Hard cap on searches and model calls per question. A timeout that falls back to single-pass RAG instead of making the customer wait. Logging of every sub-query, the passages returned and the judge's verdict, so a bad answer can be traced. And a confidence policy that decides what the bot does when the loop ends without a good answer.

Test that it helped

Keep the 30-question set fixed and re-run it after each step, scoring retrieval and answer separately as the RAG evaluation guide explains. Keep a step only if it fixes failures without raising the slow-tenth response time past your limit or the cost per conversation past your budget. Add a regression check on the easy questions: agentic loops sometimes overthink simple lookups. The QA testing protocol turns this into a routine.

If you use a platform

Most no-code platforms do not expose the loop. Ask the vendor in writing whether retrieval is single-pass or multi-step, what the per-question cost and latency ceiling is, and whether you can see each sub-query. We have not hands-on tested a platform's agentic retrieval; our Intercom review documents a named retrieval model and reranker but no multi-step loop, and our Botpress review does not document retrieval internals in enough detail to say. Verify in a pilot, using the vendor evaluation checklist. For the money side, see reducing chatbot costs.

Common mistakes

Starting with the framework. Choosing LangGraph or LlamaIndex before you have a tally of failures. Frameworks package the loop; they do not tell you whether you need one.

No cap on loops. One confusing message should never cost ten times the average.

Letting the judge grade itself unchecked. A model that marks weak passages as sufficient is worse than no judge.

Expecting it to fix content gaps. The loop can only find what is in your knowledge base.

FAQ

When should I use agentic RAG instead of standard RAG?

When your failures come from complex questions: multiple asks, context-dependent follow-ups or several sources. For simple lookups, standard RAG with hybrid search and reranking is faster and cheaper.

How do I build an agentic RAG chatbot?

Add steps in order: question rewriting, splitting, a retrieval check with one retry, then source routing. Re-test after each, and cap calls per question.

How much slower and costlier is it?

It depends on how many extra calls each question triggers. Neither Anthropic nor Microsoft publishes a universal figure, only that latency and token cost rise. Measure your own before and after.

Do I need LangGraph or LlamaIndex?

No. They help package common patterns but the loop can be written directly against a model API with tool calling.

Is agentic RAG safe for customer-facing bots?

With caps, read-only tools first, logging and a human fallback, yes. Without them, loops and wrong tool calls become customer-visible.

Sources

About this guide

Chatbotscape launched in 2026 as an independent review site for chatbot platforms. This guide is editorial guidance built from vendor documentation and published research; the one-in-five threshold is a rule of thumb, not a benchmark. Intercom has a Chatbotscape review that carries affiliate links; the advice here does not depend on it.

Last updated

1 October 2026 — first published.