Hybrid Search: Meaning Plus Keywords

Dense embeddings understand paraphrase but stumble on product codes, error IDs and jargon. Sparse keyword search is the opposite. Hybrid search runs both and fuses the rankings. This chapter covers TF-IDF, BM25 and reciprocal rank fusion, with a worked example.

Advanced RAG Masterclass

So far, embeddings have looked almost magical: "vacation days" finds "paid leave", and "phones under 20k" finds "smartphones below 20,000". Now ask a support bot built on dense search about ERR-4012, or about part number 77831-B, or about an internal tool called "Fennec". Results become erratic. An embedding model learned meaning from public text, and an error code has no meaning it has ever seen. The vector for "ERR-4012" might sit near "ERR-4021", or near nothing useful at all.

The documents contain the exact string. A plain keyword search would find it immediately. This is why production retrieval usually runs both kinds of search.

Dense and sparse

  • Dense retrieval is what we have built so far. Every text becomes a dense vector of a few hundred to a few thousand numbers, all non-zero, and relevance is semantic similarity. Its strengths are paraphrase, synonyms and cross-lingual matches.
  • Sparse retrieval represents text as a vector over the whole vocabulary, almost all zeros, with weights only for the words actually present. Relevance comes from shared terms, weighted by how informative they are. Its strengths are exact identifiers, rare jargon, names and numbers.
QueryDenseSparse
"how many vacation days do I get"✅ finds "paid leave"❌ no shared words
"ERR-4012 on login"❌ code has no learned meaning✅ exact match
"Fennec rollout plan" (internal project name)⚠️ unpredictable✅ exact match

The two fail in complementary ways, which is the strongest argument for combining them.

How keyword relevance is scored

TF-IDF

Two intuitions, multiplied together:

  • Term frequency (TF): a document that mentions "maternity" ten times is probably more about maternity than one that mentions it once.
  • Inverse document frequency (IDF): a word that appears in nearly every document ("the", "policy", "employee") tells you nothing, while a word that appears in only a few ("maternity", "ERR-4012") is a strong signal. With NN documents, of which ntn_t contain term tt:
IDF(t)=log⁡Nnt.\text{IDF}(t) = \log\frac{N}{n_t}.

TF × IDF rewards terms that are frequent in this document but rare across the collection.

BM25: TF-IDF with two corrections

BM25 is the standard keyword-ranking function in search engines and in the sparse half of hybrid retrieval. It adds two practical fixes to TF-IDF:

  1. Saturation. Ten mentions of "maternity" is more relevant than one, but not ten times more. BM25's TF term rises quickly at first and then flattens, controlled by a constant k1k_1 (typically 1.2–2.0). This stops keyword-stuffed documents from dominating.
  2. Length normalisation. "Maternity" appearing three times in a 200-word policy says more than three times in a 20,000-word handbook. BM25 scales term frequency by the document's length relative to the average length, controlled by bb (typically 0.75).

For a query with terms q1,…,qnq_1,\dots,q_n, a document DD of length ∣D∣|D|, and average document length ∣D∣‾\overline{|D|}:

BM25(D)=∑iIDF(qi)⋅f(qi,D) (k1+1)f(qi,D)+k1(1−b+b ∣D∣∣D∣‾).\text{BM25}(D) = \sum_{i} \text{IDF}(q_i)\cdot\frac{f(q_i,D)\,(k_1+1)}{f(q_i,D) + k_1\left(1 - b + b\,\frac{|D|}{\overline{|D|}}\right)}.

You do not need to memorise this formula. You need to know what each part does: IDF rewards rare terms, the fraction applies saturation, and the length factor normalises for document size. To get a sense of the IDF part, in a corpus of 1,000 documents a term in 900 of them gets an IDF of about 0.1, while a term in 5 of them gets about 5.2, a 50-fold difference in weight.

Fusing two rankings

Hybrid search produces two ranked lists. How do you merge them? Their raw scores cannot be compared directly. A cosine of 0.82 and a BM25 score of 14.3 live on different scales, and those scales shift from query to query.

Reciprocal rank fusion (RRF)

RRF ignores the raw scores and uses only ranks:

RRF(d)=∑r ∈ retrievers1k+rankr(d),\text{RRF}(d) = \sum_{r \,\in\, \text{retrievers}} \frac{1}{k + \text{rank}_r(d)},

with kk conventionally set to 60. A document ranked first contributes 1/611/61, a document ranked tenth 1/701/70. The constant kk flattens the curve, so the top rank of a single retriever cannot dominate. A document ranked reasonably well by both retrievers beats one ranked first by only one of them.

Worked example.

DocBM25 rankVector rankRRF score
A131/61+1/63=0.032271/61 + 1/63 = 0.03227
B211/62+1/61=0.032521/62 + 1/61 = 0.03252
C321/63+1/62=0.032001/63 + 1/62 = 0.03200

The fused order is B, A, C. B is near the top of both lists, so it wins, even though A was first for BM25. The differences are small in absolute terms, but only the order matters. A document missing from one list simply contributes nothing for that retriever.

Weighted score fusion and alpha

The alternative is to normalise each retriever's scores to a common range (for example, scale each list to 0–1) and take a weighted sum. The weight is usually called alpha. In Weaviate's convention, alpha = 1 is pure vector search, alpha = 0 is pure keyword search, and alpha = 0.5 weights them equally. Conventions differ between databases, so always check which side alpha refers to.

Choosing alpha is a property of your data, not a universal constant. A corpus full of SKUs, log lines and error codes wants more keyword weight. Conversational questions over prose policies want more semantic weight. Tune it on your evaluation set (Module 5), just as you tune chunk size.

The query goes to a dense retriever (embedding + vector index) and a sparse retriever (BM25 over an inverted index) in parallel; their two ranked lists merge in a fusion step (RRF or weighted alpha) into one ranked list
Hybrid search runs dense and sparse retrieval in parallel and fuses their rankings. RRF uses only ranks; weighted fusion normalises scores and blends them with alpha.

Beyond BM25: learned sparse vectors

A newer family, learned sparse retrieval (for example SPLADE), uses a transformer to produce sparse vectors. It keeps exact-term matching but also adds weights for related terms the document did not literally contain. Several vector databases accept a dense and a sparse vector per record and handle the fusion internally. The architecture is the same: two signals, one fused ranking.

Try it yourself
RAG Lab: hybrid search and RRF →

Compare BM25, dense and fused rankings for an error code and a paraphrased question.

EasyHybrid searchInterview

Why does dense retrieval struggle with product codes and error IDs, and why does BM25 handle them well?

MediumHybrid searchRRF

Two documents: X is ranked 1st by BM25 and absent from the vector results; Y is ranked 4th by both. With RRF (k = 60), which ranks higher?

MediumBM25

What two problems with raw TF-IDF does BM25 fix?

HardHybrid searchEvaluation

How would you choose the hybrid weighting (alpha) for a new corpus?