Embeddings: Turning Meaning into Geometry

An embedding model maps text to a point in a high-dimensional space where similar meanings land close together. This chapter covers static versus contextual embeddings, the bi-encoder setup that makes retrieval fast, and what embeddings inevitably lose — which motivates reranking later.

Advanced RAG Masterclass

A computer cannot compare two sentences by meaning directly. It can compare two lists of numbers. An embedding is the bridge: a learned function that turns a piece of text into a vector, so that meaning becomes geometry. Text that means similar things becomes points that sit close together. Once that is true, "find relevant passages" turns into "find nearby points", which computers do very fast.

Meaning becomes distance

Consider three phrases: office work, corporate job, desk job. A person knows at once that they describe the same kind of life. An embedding model, given each phrase, returns a vector of perhaps 768 or 1,536 numbers, and those three vectors come out close to one another. Repairing a motorcycle comes out far away.

This is what lets RAG answer questions phrased differently from the document. The handbook says "Employees receive 20 paid leaves annually." The user asks "How many vacation days do I get?" The words "vacation" and "days" never appear in the handbook sentence, but the two embeddings land near each other because the model learned, from billions of sentences, that paid leave and vacation days are used in the same way. Keyword search would miss this entirely, and embedding search finds it as a matter of course.

Static versus contextual embeddings

Early embeddings such as word2vec gave every word one fixed vector. "Apple" had a single vector whether the sentence was about fruit or phones, and "bank" was the same next to "river" as next to "loan". These static embeddings cannot tell the senses of a word apart.

Modern embedding models are transformers, so every token's representation is computed through attention over its neighbours (the mechanism is covered in Self-Attention). "Bank" beside "river" and "bank" beside "loan" get different internal representations, and a sentence embedding (usually a pooled average of the token representations) captures the contextual sense of the whole passage. For retrieval this is essential: the query "apple's latest launch" should land near technology news, not near orchard reports.

The bi-encoder: why retrieval can be fast

A retrieval embedding model is used as a bi-encoder: the query and each document are encoded separately, by the same model, into the same space.

  • At indexing time, every chunk is encoded once, and the vectors are stored.
  • At query time, only the query is encoded, which is one model call, and then compared with the stored vectors by a cheap similarity calculation.

This separation is what makes RAG scale. Ten million chunks are encoded once, ahead of time, and a query costs one encoding plus a vector search, however large the corpus. Nothing about the query is needed to embed the documents, and nothing about the documents is needed to embed the query.

Bi-encoder: chunks pass through the embedding model offline into vectors stored in the database; the query passes through the same model online into a vector that is compared with the stored ones
A bi-encoder encodes documents and queries independently into the same space. Documents are encoded once, offline; a query needs one encoding plus a cheap comparison.

Practical properties you must respect

  • Dimension is fixed per model. A model that outputs 1,536 numbers produces only 1,536-number vectors, and the vector index is created for that exact dimension. Switching models means re-embedding the whole corpus, because vectors from different models are not comparable even when their sizes happen to match.
  • Inputs have a maximum length. Text beyond the model's token limit is truncated silently. Chunk sizes must stay under it.
  • Domain matters. A general-purpose model may place two legal terms of art, or two internal product names, closer or further apart than an expert would. Domain-tuned embedding models, or hybrid search (Module 4), help here.
  • Queries and passages look different. A question is short and phrased as a question, while a passage is long and declarative. Many retrieval models are trained on (question, answer-passage) pairs specifically so the two still meet in the middle, and some expect a prefix such as "query:" or "passage:" to tell them which is which.

What embeddings lose

An embedding squeezes a paragraph (its word order, grammar, negations, numbers and emphasis) into a few hundred numbers. That compression is the price of speed, and some information is inevitably lost. Two passages can embed almost identically while one says "the policy applies to contractors" and the other says "the policy does not apply to contractors". Exact identifiers such as ERR-4012 or SKU 77831 carry little "meaning" for the model at all.

So the similarity score from a vector search is a strong first filter but not the final word. Two later chapters build on exactly this weakness: hybrid search adds back exact keyword matching, and reranking re-reads the query and each candidate together with a slower, more precise model.

Try it yourself
RAG Lab: embedding search →

Ask in your own words and see the nearest chunks, and where they sit in vector space.

MediumEmbeddings

Why do static word embeddings like word2vec struggle in retrieval, and how do transformer embeddings fix it?

MediumEmbeddingsArchitecture

What is a bi-encoder, and why is it the right architecture for first-stage retrieval?

HardEmbeddingsOperations

You switch to a better embedding model with the same output dimension. Can you keep the existing index?