A vector search returns its top results with scores: chunk A 0.71, chunk B 0.65, chunk C 0.53. It is tempting to take the top three, pass them to the model and move on, and often that works. But look closely at failing answers and a pattern appears: the right chunk was retrieved, but ranked fifth, and only the top three went to the model. The first-stage scores are a good first guess, not a precise judgement.
Why first-stage scores are noisy
Recall what an embedding does: it compresses a passage into one vector, losing word order, grammar, negation and emphasis along the way (Embeddings explains this). The query is compressed separately, and the two never interact inside the model. The similarity score therefore compares two summaries, each made without knowing what the other was about. A query for "benefits that apply to contractors" and a chunk saying "these benefits do not apply to contractors" can look nearly identical to a bi-encoder, because they share almost every word and topic, yet the chunk is the opposite of what was asked for. And a chunk with the exact answer phrased unusually may score lower than a generic chunk that repeats the query's vocabulary.
Bi-encoders and cross-encoders
The fix is to let a model read the query and the candidate together.
- A bi-encoder (our embedding model) encodes the query and the document independently, and relevance is a dot product between the two vectors. It is fast, and documents can be pre-computed, but it is shallow.
- A cross-encoder takes the pair (query, document) as one input. Every query token can attend to every document token through all the transformer layers, and the model outputs a single relevance score. It is trained on examples of (question, relevant passage, irrelevant passage) to score relevant pairs higher.
Because the cross-encoder sees both texts at once, it can tell negation from affirmation, check that the passage actually answers the question, and judge exact constraints ("in 2024", "for contractors"). The cost is that nothing can be pre-computed: every (query, document) pair requires a full model pass at query time.
Retrieve, then rerank
This cost leads to the standard two-stage pattern:
- Retrieve broadly with the fast bi-encoder (and BM25, for hybrid search): take the top 20–100 candidates. The goal here is recall, making sure the right chunk is somewhere in the pool.
- Rerank precisely with the cross-encoder: score each candidate against the query and keep the top 3–5. The goal here is precision, putting the best chunks first.
Running a cross-encoder over the whole corpus would mean millions of model passes per query, which is far too slow. Running it over 50 candidates is affordable. Each stage does what it is good at.
Reranker options
- Hosted cross-encoder rerankers, such as Cohere Rerank, are a single API call taking the query and the candidate list. They are often among the faster options.
- Open-source cross-encoders, such as the BGE reranker family from the Beijing Academy of AI, can be self-hosted.
- Built into the vector database: several stores (Pinecone, Weaviate, Milvus, Qdrant, MongoDB Atlas) let you request reranking as part of the query, typically using one of the models above.
- LLM as reranker: give a general LLM the query and the numbered candidates and ask it to order them, or score each from 0 to 1. This is flexible and can follow custom instructions ("prefer official policy over FAQ"), but it costs tokens and is usually slower and more expensive than a dedicated cross-encoder. Use it when quality matters more than cost, or for small candidate sets.
The cost: latency
Reranking adds a model call on the critical path of every query. Cross-encoders are typically tens to a few hundred milliseconds for 50 candidates, depending on candidate length, hardware and model size. The main tuning settings are:
- how many candidates to rerank: more improves the chance the best chunk is in the pool, but costs time linearly;
- candidate length: rerankers truncate long inputs, so very long chunks are both slow and partly ignored;
- whether to rerank at all: for easy, high-confidence queries you can skip it. Many teams rerank always and simply budget for it.
Reranking is not compulsory. Add it when evaluation shows the right chunk is in the top 20 but not in the top 3, which is exactly the gap it closes. If the right chunk is not in the top 50 at all, reranking cannot help, and the problem lies upstream in chunking, embeddings, hybrid search or the query itself.
Watch a reranker move the current policy above an outdated one before it reaches the model.