The last three chapters gave the system autonomy: it decides what to retrieve, picks among tools, and keeps going until it is satisfied. That loop is what makes agentic retrieval powerful, and it is also what makes it fail in ways a single call never could.
The characteristic failure is this. An agent is asked a research question. It searches, reads a result, searches again with a slightly different phrasing, reads, searches again — and never produces an answer. It is not crashing. Every individual step looks reasonable. It simply never finishes.
This is the most common problem in production agent systems, and it is worth treating the way you would treat any other bug: find the cause, then fix the logic.
Four causes
Looping and stalling almost always trace back to one of four things. They overlap in practice, but each has a different signature and a different fix.
The plan is wrong
The agent decomposes a goal into steps. If that decomposition is flawed, it can be followed faithfully and still go nowhere. A plan of "search, read, and if the answer is not found, search again" has no path to termination when the answer does not exist in the corpus — the agent will keep satisfying the instruction forever.
The familiar version in ordinary code is recursion with no base case, or a loop whose condition can never go false. Nothing errors; the logic is just wrong. The agent will not notice, because its own reasoning keeps telling it there is more to do.
Signature in the trace: steps that are logically equivalent, dressed differently — the same query rephrased, the same subtask re-entered. No step reduces the remaining work.
The tool is being misused
Agents act through tools, and tool failures turn into loops when the agent cannot recognise them.
Two shapes recur. The agent calls a tool incorrectly — a malformed query, a wrong parameter — gets an error or an empty result, treats the absence of data as "nothing found here, try again", and retries the same broken call indefinitely. Or the agent picks the wrong tool: routing a question needing document retrieval to a structured query tool, getting nothing useful, and varying the inputs rather than reconsidering the choice. It is trying the same key in the same lock harder.
In code this is a retry loop around an exception nobody handles. The failure is visible in every iteration and the program never reads it.
Signature in the trace: repeated errors or empty result sets, followed by near-identical calls. Zero results treated as a reason to retry rather than to change approach.
The agent has forgotten what it did
An agent's memory is its context: the history of what it has tried and found. When that record is incomplete, it loses its own progress.
The usual mechanism is mundane. The context fills up, old turns are dropped or summarised away, and a detail the agent retrieved at step three is gone by step ten. So it retrieves it again. Worse, the original task framing can fall out of the window, and the agent drifts from the question it was asked.
This is the bug where a program fails to update the state that tells it what is done.
Signature in the trace: work redone that was already completed, results rediscovered, or a gradual drift away from the original question. Often it begins at a predictable depth, which is a strong clue that context length is the cause.
There is no exit condition
An autonomous loop needs a definition of done. If the condition is missing, never checked, or impossible to satisfy, the loop does not end.
The subtle version is the interesting one. The agent has a stop condition — "finish when you have answered the question" — but it is a judgement the model makes about itself, and a cautious model can always find one more thing to verify. The task is complete and the agent will not say so. Termination arrives only when some external cap trips.
Signature in the trace: the agent reaches an answer and keeps going; or it runs to the iteration limit having had the information for several steps.
Working out which one it is
Read the trace first. Turn on full logging of every thought, action and observation. Nearly every one of these failures is visible on inspection: the repeated query, the ignored error, the rediscovered fact, the answer the agent reached and walked past. What you are looking for is the first step where the pattern begins — not the twentieth repetition, but the decision that started it.
Step through it. Run the loop one iteration at a time, inspecting the state between steps. This is a debugger for an agent, and the goal is the same: find the exact decision point that went wrong. Often it is a single step where the agent had enough to answer and chose to continue instead.
Reduce it to the smallest case that still fails. Agents fail selectively — on questions with no answer in the corpus, on ambiguous requests, on a particular tool. Strip away everything that is not needed to reproduce the loop: fewer tools, a simpler question, a shorter prompt. A minimal reproduction turns a vague "it sometimes hangs" into a test you can iterate against in seconds.
Fixing it
The fix follows from the cause, and almost all of them are changes to the prompt or the loop rather than to the model.
For planning errors, make the structure explicit. State that a step should not be repeated with trivial variations; require that each step name what it adds; ask the agent to track which subgoals are complete. Give it a defined way to conclude that an answer does not exist, so "not found" is an outcome rather than a reason to continue.
For tool errors, handle failure in the loop rather than hoping the model notices. Surface errors and empty results explicitly in the observation, so the agent sees no results rather than nothing. Cap retries per tool. Require a different approach after a repeated failure, and provide a fallback path.
For memory problems, stop relying on raw history. Maintain a compact running record of findings and completed steps, and re-inject it each iteration so the essentials survive truncation. Summarise older turns rather than dropping them. Keep the original question in a position that never gets trimmed.
For missing exits, do both of two things. Add a hard cap — a maximum iteration count or time budget — that forces a final answer or a clean failure, so no run is unbounded. Then fix the real condition, since the cap is a safety net and not a design. Make "done" a question the agent must answer explicitly at each step rather than something it drifts into, and give it permission to answer with what it has: an agent told that a good-enough answer now beats a perfect answer never will stop.
The instinct from traditional engineering is to fix behaviour in code. In an agent, the prompt is the policy — it determines what the agent plans, when it stops and how it handles failure — so most of these bugs are fixed by rewriting instructions, and that is a legitimate fix rather than a workaround. Changing "keep going until you find the answer" to "keep going until you find the answer or you have tried three distinct approaches" closes a whole class of loops. Retraining is almost never the answer. Keep the hard caps in code regardless, because a prompt is a strong suggestion and a loop limit is a guarantee.
A last point, since it is easy to miss while debugging: an agent that loops is at least visibly broken. The more dangerous failure is one that stops confidently with a wrong answer, which looks like success in every log. The caps and stop conditions here keep a run bounded; knowing whether the answer is any good is what Evaluating the Whole Pipeline and Tracing and Debugging are for.