Agents and Tool Use

Giving a model tools and a loop, what that buys, and the four ways it goes wrong — the area where interviewers most often ask you to debug rather than describe.

How to Crack the AI Engineer Interview

An agent is a model that decides what to do next: choose a tool, act, observe the result, and repeat until finished. It is the newest material on this syllabus and the least settled, which is exactly why interviews often pose it as a debugging problem rather than a definition.

What gets asked

What an agent adds. A single call does one retrieval and one generation. An agent can decide whether to retrieve, search several times with different queries, route a question to a database rather than a document store, call a calculator or an API, and judge whether it has enough to answer. It turns a fixed pipeline into a decision process.

The loop. Think, act, observe, repeat. Every agent framework is a variation on it, and most failures are failures of the loop rather than the model.

Tool use. Giving the model a set of callable operations with described inputs, letting it choose. Know that tool descriptions are part of the prompt, and that tool selection is a model decision that can be wrong.

Agentic retrieval. The RAG-specific version: the agent decides when retrieval is needed, reformulates queries after poor results, and chooses between retrieval sources. More capable, less predictable, more expensive.

Why this trade exists. Agents buy capability with reliability. A fixed pipeline does the same thing every time; an agent may take a different path per request, which is the point and also the problem.

The failure modes — expect to debug one

This is the common interview form: your agent sometimes loops forever or never produces an answer. How do you debug it? There are four causes, each with a signature in the trace.

The plan is wrong. No step reduces the remaining work, so the agent faithfully executes a decomposition that cannot terminate. Signature: the same step, reworded — a query rephrased, a subtask re-entered.

The tool is misused or failing. Errors and empty results reach the agent as an absence of information, which reads as "nothing here, try again" rather than "this call is broken". Signature: repeated errors or zero-result calls followed by near-identical retries.

Findings fall out of context. As history grows, older turns are dropped and a result from step three is gone by step ten, so the agent redoes work or drifts from the original question. Signature: rediscovered findings, often starting at a predictable depth.

No workable stop condition. Either absent, or present but impossible to satisfy — a cautious model asked to stop "when satisfied" can always find one more thing to check. Signature: the agent reaches an answer and keeps going, or runs to the iteration cap having had the information for several steps.

The debugging method is the same as for any program: read the trace and find the first step where the pattern starts, step through the loop one iteration at a time, and reduce it to the smallest case that still fails.

The follow-ups that catch people

Is a maximum iteration count a sufficient fix? Wanted: no. It guarantees the run ends, which is necessary, but hitting it means the agent never decided it was done — you have bounded the symptom. Keep the cap and fix the stop condition.

Why are these bugs usually fixed in the prompt? Wanted: the prompt is the agent's policy — it determines planning, stopping and failure handling. Retraining addresses a capability problem, which this generally is not. Hard limits still belong in code, because a prompt is a suggestion and a loop cap is a guarantee.

When would you not use an agent? Wanted: when a fixed pipeline does the job. Agents add latency, cost and variance. If every request needs the same three steps, orchestrate them directly.

What is the worse failure than looping? Wanted: stopping confidently with a wrong answer, which looks like success in every log. Loops are at least visible.

How it gets worded

  • "What does an agent add over a prompt in a loop?"
  • "How do you describe a tool to a model so that it gets called correctly?"
  • "The model picks the right tool and passes the wrong arguments. How do you debug that?"
  • "The agent is looping. Walk me through how you find the cause."
  • "Is an iteration cap a fix?"
  • "How much autonomy would you give this system, and what decides that?"
  • "A long task is overflowing the context. What happens to the findings from step three?"
  • "A tool call fails, or returns something the model did not expect. What should happen next?"
  • "What are the security consequences of letting a model call tools on your behalf?"
  • "How do you evaluate an agent when the path to the answer is different every run?"
  • "When is a fixed pipeline the better engineering choice?"
  • "What is the failure mode worse than looping?"

Reading path

  1. Agentic RAG — what the loop adds over a fixed pipeline.
  2. Routing and Text-to-SQL — choosing between sources, including structured ones.
  3. Retrieval as Tools — exposing retrieval as something the model calls.
  4. When Agents Get Stuck — the four failure modes above in full, with the fixes.
  5. Tracing and Debugging — the observability you need before any of this is diagnosable.
  6. Memory for Assistants and Memory in Practice — the context problem, properly.
  7. Chain-of-Thought Prompting — the reasoning step inside each loop iteration.

What signals experience

Three things read as having actually built one.

Talking about bounds first. Iteration caps, timeouts, cost ceilings per request. Anyone who has run an agent in production has been surprised by a bill or a hang.

Treating the trace as the primary artefact. The answer "I would read the trace and find the first step where it goes wrong" is better than any list of causes, because it is what you actually do.

Being sceptical. Agents are the most over-promised part of this field. A candidate who says that a deterministic pipeline is usually better and that agents earn their variance only when the path genuinely cannot be fixed in advance sounds like someone who has shipped one.