Retrieval as Tools: MCP Servers and Deep Agents

The newest production pattern turns every retrieval strategy (semantic, keyword, hybrid, filtered) into a tool on an MCP server, and lets an agent choose among them. A 'deep agent' adds planning, a scratch workspace and reusable skills, so it can break down a five-part question, retrieve in parallel and answer in a consistent format.

Advanced RAG Masterclass

So far the developer has chosen the retrieval strategy: hybrid with alpha 0.5, rerank the top 50, filter by department. Those choices are fixed in code and applied identically to every question. But questions differ. A query containing an error code wants keyword search. A vague conceptual question wants semantic search. "What changed in the 2026 travel policy?" wants a metadata filter. A question that bundles five sub-questions wants to be split up first.

The pattern that many teams are now adopting hands that choice to an agent. Retrieval strategies become tools, and an agent decides which to call, how many times and in what order.

Retrieval behind an MCP server

The Model Context Protocol (MCP), introduced by Anthropic and now widely supported, is a standard way for an AI application to discover and call external tools. An MCP server exposes a set of named tools, each with a description and an input schema. Any MCP-capable agent can connect to it, list the tools and call them, over HTTP for remote servers or standard input and output for local ones.

For RAG, the server wraps the vector store and exposes each retrieval strategy as a separate tool:

ToolWhat it doesWhen an agent should pick it
semantic_search(query, k)Dense vector similarityConceptual or paraphrased questions
keyword_search(query, k)BM25 / sparseCodes, IDs, names, exact phrases
hybrid_search(query, k, alpha)Both, fusedMixed or unclear queries
metadata_search(query, filters)Filtered searchQuestions naming a year, department or document type

Putting retrieval behind a protocol has architectural benefits that go beyond the agent:

  • Separation of concerns. The data team owns the MCP server (indexes, credentials, filters, access control), and application teams just connect to it.
  • Reuse. The same retrieval server serves the support bot, the internal assistant and a developer's coding agent.
  • Swappability. Change the vector database or add a reranker inside the server, and no client changes.

Why a plain agent is not enough

A simple tool-calling agent receives the question, picks a tool, gets results and answers. Give it a compound question:

"What is the notice period, how does it differ for managers, can it be bought out, what happens to unused leave on exit, and who approves early release?"

It typically makes one search with the whole question, whose embedding is a blur of five topics, gets chunks for one or two of them, and writes an answer that silently skips the rest. In a side-by-side comparison, a reviewer finds whole sub-questions unanswered.

Deep agents

Deep agents (the term comes from LangChain's deepagents library, which was inspired by how Claude Code works) add the structure needed for multi-step work:

  • Planning. The agent first writes a to-do list: split the question into sub-questions, decide which tool suits each, and track progress.
  • A backend (scratch workspace). A file system, local or in a database, where the agent writes intermediate results: findings for sub-question 1, 2, 3 and so on. Long tasks no longer have to keep everything in the context window, and partial work isn't lost.
  • Sub-agents. Delegate a sub-task to a fresh agent with its own clean context, and receive back only its conclusion.
  • Skills. Reusable, written procedures the agent can load when relevant (below).
  • Human-in-the-loop and long-term memory, for approvals and for cross-session knowledge (Module 9).

Skills: procedures the agent loads on demand

A skill is a folder containing a SKILL.md file: a short name and description, followed by step-by-step instructions for a specific job. Like MCP, the format comes from Anthropic. For a RAG deep agent, typical skills might be:

  • query-decomposition: how to split a compound question into independent sub-questions;
  • parallel-retrieval: run the sub-question searches concurrently and collect the results;
  • query-classification: when to prefer keyword, semantic, hybrid or filtered search;
  • answer-synthesis: the required answer format (a summary, one section per sub-question, then bullet-point details with sources).

The important design detail is progressive disclosure. The system prompt lists only each skill's name and one-line description. The agent reads a skill's full instructions only when it decides the skill applies to the current request. Twenty skills cost twenty lines of context, not twenty documents. Skills steer behaviour at specific steps without bloating every prompt.

A deep agent with planner, skills (decomposition, parallel retrieval, classification, synthesis) and a file-system backend connects over MCP to a retrieval server exposing semantic, keyword, hybrid and metadata search tools on top of the vector store; a compound question is split into sub-questions retrieved in parallel and synthesised into a structured answer
Retrieval strategies become MCP tools; the deep agent plans, loads the relevant skills, calls tools in parallel for each sub-question, writes intermediate results to its backend and synthesises one structured answer.

What changes in practice

Run the five-part question through a deep agent and the trace looks different. The agent loads the decomposition skill and writes five sub-questions to its plan. It loads the classification skill and chooses semantic search for four of them, and perhaps keyword search for one that names a specific form. It retrieves all five in parallel, saves the findings to its backend, and loads the synthesis skill to produce a summary followed by one section per sub-question. Every part is answered. The same question through a single-shot pipeline answers two.

The costs are real, though: several LLM calls for planning and synthesis, many tool calls, and more variable latency and spend per question. The trade-off is typically worth it for complex, multi-part or research-style questions and wasteful for simple lookups. A common design routes simple questions to a fast linear pipeline and sends only complex ones to the deep agent.

MediumMCPArchitecture

What are the benefits of exposing retrieval strategies as MCP tools rather than hard-coding one strategy in the application?

MediumDeep agentsQuery decomposition

Why does a plain RAG agent miss parts of a compound question, and which deep-agent capabilities fix it?

HardDeep agentsSkills

What is 'progressive disclosure' in agent skills, and why does it matter?