Query Transformation: Fixing the Question Before the Search

Users ask vague, compound and oddly worded questions, and a single embedding of their exact words often retrieves poorly. Using an LLM to rewrite the query first — multi-query, RAG-Fusion, decomposition, step-back, HyDE and routing — makes retrieval far more robust.

Advanced RAG Masterclass

Everything in this module so far improves the search. But the search can only be as good as the query it receives, and real queries are messy. People type fragments ("leave carry forward?"), use their own vocabulary rather than the document's, pack three questions into one sentence, or ask something so specific that no passage matches it directly. Embedding that raw text and hoping for the best is the weakest link of a basic pipeline.

Query transformation puts a small LLM step before retrieval that turns the user's question into one or more better search queries. It is cheap, it is independent of your index, and it often gives the largest single quality gain in a mature system.

About this chapter

The source lecture series does not have a dedicated video on these techniques. They appear only implicitly, in the agentic rewriting loop of Module 7 and the query decomposition of the deep-agent chapter. They are central to advanced RAG, so we cover them here in full.

Multi-query: cast a wider net

A single phrasing reaches one region of the embedding space. If the document uses different words, that region may miss it. Multi-query asks an LLM to write several alternative phrasings of the same question:

"Can I carry unused leave into next year?" might become:

  • "leave carry-forward policy"
  • "rollover of unused annual leave"
  • "maximum leave days transferable to next calendar year"

Each variant is searched separately and the results are combined, with duplicates removed. Synonyms the user never typed now have a chance to match.

RAG-Fusion: combine the variants by rank

RAG-Fusion is multi-query with a principled merge. Each variant produces its own ranked list, and the lists are combined with reciprocal rank fusion, exactly as in hybrid search. A chunk that ranks well for many phrasings rises to the top, while a chunk that matched one odd phrasing by chance sinks. The final ranking is more robust than that of any single query, at the cost of one extra LLM call and N searches (which can run in parallel).

Decomposition: split compound questions

"What is the notice period, does it differ for managers, and can it be bought out?" is three questions. Its single embedding is an average of all three and matches none of them well. Typically the retriever returns chunks for one part, and the answer silently ignores the others.

Decomposition has an LLM split the question into self-contained sub-questions, retrieves for each independently (in parallel), and then answers using all the evidence together. Sub-questions can also be sequential, where the answer to one is needed to form the next ("Who manages the Fennec project?" → "What is that person's team's on-call policy?"). That case calls for an iterative loop, which leads into the agentic patterns of Module 7. In the deep-agent chapter you will see a system that decomposes a five-part question automatically and answers parts that a plain pipeline missed.

Step-back prompting: ask the general question first

Some questions are too specific to match anything directly: "If I joined on 14 March and resign on 2 October, how many leave days are encashed?" No passage contains that calculation. But one does contain the principle: how leave accrues per month and the encashment rule.

Step-back prompting asks an LLM to produce the more general question behind the specific one ("How does leave accrual and encashment work on resignation?"), retrieves for both the original and the step-back question, and gives the model both the specific details and the general rule. The same idea helps in technical domains. A specific physics problem is best answered after retrieving the governing law it depends on.

HyDE: search with a hypothetical answer

Questions and answers look different: a question is short and interrogative, while a passage is long and declarative. HyDE (Hypothetical Document Embeddings) narrows that gap. First an LLM writes a plausible answer to the question, without retrieval, and possibly wrong in its details. Then you embed that hypothetical answer and search with it. A fake answer is written in the same style as real answer passages, so it often lands closer to them in embedding space than the bare question does. The facts in the hypothetical answer are never shown to the user. It is only a search key.

Routing: send the question to the right place

Not every question should go to the same index. Routing classifies the question first and directs it:

  • Semantic routing: embed the query and compare it with descriptions or example questions for each route ("tax, invoices, reimbursement" → finance index). It is fast and needs no LLM call.
  • LLM routing: a small, fast model reads the question and outputs a route label ("hr_policies", "finance", "sql_database", "web_search") as structured output.

Routing narrows the search (as metadata filters do), and it is the gateway to non-vector sources such as a SQL database, which Module 7 builds out.

A user question enters a query-transformation layer with five options — multi-query, decomposition, step-back, HyDE and routing — producing one or more improved search queries that feed the retriever and fusion before generation
Query transformation sits between the user and the retriever. Each technique fixes a different kind of bad query; their outputs feed the usual retrieve → fuse → rerank → generate pipeline.

Choosing and combining

Symptom in your failing queriesTechnique
Vocabulary mismatch, short or vague queriesMulti-query / RAG-Fusion
Several questions in oneDecomposition
Over-specific, calculation or scenario questionsStep-back
Questions phrased very differently from documentsHyDE
Several distinct knowledge sourcesRouting

Every technique adds at least one LLM call before retrieval, so latency increases by a few hundred milliseconds unless you use a small, fast model, and multi-query multiplies search calls. Apply them where your evaluation shows they help, often only to queries a classifier flags as complex, rather than everywhere by default.

MediumQuery transformationInterview

Why does a compound question often get a partial answer from a basic RAG pipeline, and how does decomposition fix it?

HardHyDE

HyDE embeds an LLM-written answer that may contain wrong facts. Why is that acceptable?

EasyQuery transformation

What is the difference between multi-query retrieval and RAG-Fusion?