Agentic RAG· AI infrastructure
Agentic RAG — Retrieval With a Decision-Maker in the Loop
Quick answer: In ordinary retrieval-augmented generation, a chatbot turns the question into one search, takes the top passages and writes an answer from them. In agentic RAG, an AI agent sits in front of retrieval: it can rewrite the question, break a multi-part question into sub-questions, choose between sources, check the results and search again. You get better coverage on complex questions and pay for it in latency and tokens. Anthropic's general guidance on agents says such systems "often trade latency and cost for better task performance."
How it differs from standard RAG
Standard RAG is a fixed pipeline: question in, one retrieval, answer out. The code decides the path; the model only writes the final text. Anthropic's agent guide draws the general line: workflows are "systems where LLMs and tools are orchestrated through predefined code paths," while agents are "systems where LLMs dynamically direct their own processes and tool usage." Agentic RAG moves retrieval from the first category toward the second.
| Standard RAG | Agentic RAG | |
|---|---|---|
| Searches per question | One | One or several, decided at run time |
| Question handling | Used as typed | Rewritten, split or expanded |
| Sources | Usually one index | Can choose among indexes, tools or the web |
| Quality check | None before answering | Judges retrieved passages, may retry |
| Latency and cost | Low, predictable | Higher, variable |
| Failure mode | Wrong passage, confident answer | Loops, extra spend, harder debugging |
The loop
A typical agentic RAG turn has four moves. The model first plans: is a search needed at all, and is the question one question or three? It then retrieves, possibly with several rewritten queries. It then judges: are these passages relevant and sufficient? Finally it either answers or goes round again with a better query or another source. Microsoft's Azure AI Search describes its managed version as "a multi-query pipeline designed for complex questions": an LLM breaks a complex query into smaller subqueries, runs them in parallel, semantically reranks each, and merges the best results for the answering model.
Two named patterns
Self-RAG. Asai and colleagues train a model that "adaptively retrieves passages on-demand" and reflects on both the passages and its own output using special "reflection tokens." The paper's argument is that indiscriminate retrieval, pulling passages whether or not they are needed, can reduce usefulness.
Corrective RAG (CRAG). A lightweight evaluator scores the retrieved documents and produces "a confidence degree" that triggers different actions. When confidence is low, the method falls back to large-scale web search to extend the results, and a decompose-then-recompose step filters the retrieved text down to the key information.
Both papers report gains on their own benchmarks; those numbers come from research datasets, not from a support knowledge base, so this entry quotes none.
What it costs
Each extra step is another model call or search. Microsoft states plainly that agentic retrieval "adds latency compared to a single-query pipeline, but it handles query complexity that a single query can't," and that it is billed by tokens from both the search service and the language model, unlike classic query-based billing. Expect slower first replies, a higher and less predictable cost per conversation, and more to test. Anthropic's advice is to start with the simplest solution and add complexity only when needed, noting that for many applications "optimizing single LLM calls with retrieval" is usually enough.
When it fits and when it does not
It fits questions with several asks in one message, questions that depend on earlier turns, and corpora spread across sources. It is a poor fit for short FAQ-style lookups, where one good search plus a reranker answers faster and cheaper. It does not fix missing content, and it does not remove hallucination: a loop that judges its own passages can still judge wrongly.
What our reviews show
We have not hands-on tested any platform's agentic retrieval in our review corpus. Our Intercom review documents a named retrieval model and reranker for Fin but does not describe a multi-step retrieval loop, and our Botpress review does not document retrieval internals in enough detail to say. Treat vendor claims of "agentic" retrieval as something to verify in a pilot.
FAQ
What is agentic RAG?
RAG where the language model controls retrieval: it plans searches, rewrites or splits the question, judges the results and can retry before answering.
How is agentic RAG different from regular RAG?
Regular RAG searches once and answers. Agentic RAG runs a decision loop with possibly several searches and a quality check, at higher latency and cost.
Is agentic RAG better than RAG?
Only for complex, multi-part or multi-source questions. For simple lookups standard RAG is usually faster, cheaper and easier to test.
Does agentic RAG stop hallucinations?
No. A better retrieval loop raises the odds that the right passage is found, but the model can still answer wrongly or judge poor passages as good.
Do I need a framework to build it?
Not necessarily; the loop can be written directly against a model API with tool calling. Frameworks such as LangGraph and LlamaIndex package common patterns.
Related terms
- Retrieval-augmented generation — the single-pass pattern this extends.
- AI agent — the decision-making component in the loop.
- Agentic AI — the wider idea of goal-directed systems.
- Reranking — the precision step both designs can use.
- Chunking — how passages are cut before any search.
Sources
- Asai et al., Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection — arxiv.org/abs/2310.11511, read 30 September 2026.
- Yan et al., Corrective Retrieval Augmented Generation — arxiv.org/abs/2401.15884, read 30 September 2026: retrieval evaluator, confidence-triggered actions, web-search fallback.
- Microsoft Learn, Agentic retrieval overview — learn.microsoft.com/en-us/azure/search/agentic-retrieval-overview, read 30 September 2026: subquery decomposition, parallel execution, latency and token billing.
- Anthropic, Building Effective Agents — anthropic.com/engineering/building-effective-agents, read 30 September 2026: workflow versus agent definitions and the latency-cost trade-off.
- Chatbotscape evaluation methodology. /methodology (continuously updated).