Agentic RAG: Retrieval That Checks Its Own Work

A fixed pipeline retrieves once and answers whatever it finds. An agentic system lets an LLM make decisions inside the pipeline: grade what was retrieved, rewrite the question and try again, or choose a different tool. This chapter covers the corrective loop and the state-graph model used to build it.

Advanced RAG Masterclass

Every pipeline so far has been a straight line: question in, retrieve, generate, answer out. The line has no way to notice that it went wrong. If retrieval returned five irrelevant chunks, the model receives five irrelevant chunks and does its best with them, which may mean a hallucination or a vague non-answer. Nothing in the system ever asks, "Did that retrieval actually work?"

Agentic RAG adds that question, and the ability to act on the answer.

From pipeline to agent

An agent is a system that uses a model to decide what to do next based on what it observes, rather than following a fixed sequence. Applied to RAG, the LLM is no longer only the final writer. It also makes decisions at points inside the flow:

  • Is the retrieved context relevant enough to answer from?
  • If not, how should the question be rephrased?
  • Is this even a question for the document index, or for a database, a web search or an API?

A plain retrieve-then-generate pipeline is often called vanilla RAG. Once the system can make these choices (call tools, loop, route), it is agentic RAG.

The corrective loop

The most useful agentic pattern is a self-correcting retrieval loop:

  1. Retrieve chunks for the current question.
  2. Grade them: an LLM reads the question and the retrieved chunks and returns a structured verdict, relevant: yes or no.
  3. If yes, generate the answer from those chunks and finish.
  4. If no, rewrite the question (keeping its meaning but changing its wording, or adding the specifics the documents are likely to use) and go back to step 1.

The rewrite is the same move a person makes after a failed search. "Code of wages complaints" returns nothing useful, so you try "grievance redressal procedure under the Code on Wages". The rewritten query may match the document's vocabulary far better, and therefore have a much higher similarity to the right chunks.

This pattern appears in the literature as Corrective RAG (CRAG) and Self-RAG, with variations: grading each chunk individually and keeping only the relevant ones, falling back to web search when the corpus has nothing relevant, or also checking the answer for hallucinations before returning it.

State graph: START → retrieve → grade documents (LLM, structured yes/no) → if yes: generate → END; if no: rewrite question → back to retrieve; a counter limits retries
The corrective loop as a state graph. Nodes do work; the conditional edge after grading decides whether to answer or to rewrite and retry. A retry limit guarantees the loop ends.

Building it as a state graph

Frameworks such as LangGraph model agentic flows as a graph, which maps directly onto the diagram:

  • Nodes are units of work: retrieve, grade, rewrite, generate. Each is an ordinary function, which may call an LLM, a database or a tool.
  • Edges say what runs next. A normal edge always goes to the same node (retrieve → grade). A conditional edge is a function that inspects the current state and returns the name of the next node. Here it reads the grade and returns "generate" or "rewrite".
  • State is the shared memory of one run: the original question, the current rewritten question, the retrieved documents, the grade, the answer and the message history. Each node reads what it needs from the state and writes its result back, which is how the grader sees the retriever's output and the generator sees both the question and the documents.

Thinking in nodes, edges and state is useful even without a framework. It forces you to make every decision point explicit, which also makes the flow easy to trace (Module 5) and to test node by node.

Structured output: making decisions reliable

A decision step must return something the program can branch on, not a paragraph. "Well, the documents are partially relevant, although…" cannot drive a conditional edge. Use the model's structured output support: define a schema ({relevant: "yes" | "no"}, {route: "sql" | "rag"}) and have the model fill it in. The edge then branches on a guaranteed value. Every routing, grading and classification step in agentic RAG relies on this.

Guardrails for loops

A loop that rewrites forever is a bug that costs money. Production agentic RAG needs:

  • A retry cap. Keep a counter in the state, and after two or three rewrites, stop. Then either answer with what you have (and say the evidence is weak) or say plainly that the documents don't cover the question.
  • An honest exit. "I couldn't find this in the HR policies" is a correct answer when the corpus doesn't contain it, and far better than a confident invention.
  • A latency budget. Each loop iteration adds a grading call, a rewrite call and another retrieval, so set a time limit per request.
  • Cheap graders. Grading is a simple classification task, and a small, fast model often does it well.

When it pays off

Agentic loops add latency and cost to every request that triggers them. They help most when queries are diverse and often poorly phrased, when the cost of a wrong answer is high, or when several sources are available. For a narrow FAQ bot with well-phrased queries, a well-tuned linear pipeline (hybrid search, reranking, good chunking) may be all you need. As always, let evaluation decide.

EasyAgentic RAGInterview

What turns vanilla RAG into agentic RAG?

MediumAgentic RAGLangGraph

In a LangGraph-style agent, what are nodes, conditional edges and state, and how do they map onto the corrective RAG loop?

MediumAgentic RAGReliability

Why must a grading step use structured output, and what should happen when the grader says 'no' three times in a row?