We keep saying "find the chunks closest to the question". Close in what sense? Two vectors can point the same way but have different lengths. They can sit near each other but point in slightly different directions. Each similarity measure gives a different meaning to "close", and every vector index is created with one of them. This chapter makes the three standard choices concrete and then shows why, for most RAG systems, the choice turns out not to matter, along with the cases where it does.
Drag a query and three document vectors, compare the three rankings, then normalise them and watch the rankings agree.
Cosine similarity: same direction?
Cosine similarity measures the angle between two vectors and ignores their lengths:
It ranges from (same direction) through (perpendicular, unrelated) to (opposite).
Worked example. Take and .
- Dot product: .
- Lengths: and .
- Cosine: .
The two vectors point in almost the same direction, even though is more than twice as long. Cosine sees them as nearly identical.
This is usually what we want for text. "Best smartphones under 20,000" and "top mobile phones under 20k" should match strongly whatever the length of each embedding. Cosine is the default metric in most vector databases, and it is common in recommendation, image retrieval and outlier detection as well.
Euclidean distance: same place?
Euclidean (L2) distance is the straight-line distance between the two points:
Here smaller is more similar, and 0 means identical. For our example, : quite far apart, because is much longer. Euclidean distance is sensitive to magnitude, which is why it is the natural choice in classical machine learning, for k-nearest neighbours and k-means clustering, where features have absolute meaning.
When magnitude carries meaning, use a measure that sees it. Suppose each vector is a user's activity profile: movies watched, products bought, searches made. A user with and a user with have the same pattern (cosine 1.0) but very different engagement. For a "find similar power users" feature, Euclidean distance correctly separates them, and cosine would wrongly call them twins.
Dot product: direction and magnitude
The plain dot product is
that is, cosine multiplied by both lengths. Larger means more similar. It rewards both pointing the same way and being long. Some embedding models are trained so that length encodes something useful (for example, confidence or popularity), and for them the dot product is the intended measure. It is also the cheapest of the three to compute, since there are no square roots or divisions.
The key fact: normalised vectors make them agree
Most modern text-embedding models output vectors normalised to length 1 (or the database normalises them for a cosine index). When :
- the dot product equals the cosine, since the length factors are 1, and
- the squared Euclidean distance is , a decreasing function of the cosine.
So for unit vectors all three measures produce exactly the same ranking. Normalising our example vectors gives a squared distance of , matching the formula. This is why, for typical text embeddings, the choice of metric barely changes results, and why databases often compute "cosine" internally as a fast dot product on normalised vectors.
The choice matters when vectors are not normalised and their length means something: activity scores, purchase power, or embedding models explicitly trained for dot-product similarity. In those cases, follow the model's documentation and choose the metric it was trained with.
| Cosine | Euclidean | Dot product | |
|---|---|---|---|
| Measures | Angle | Gap between points | Angle × lengths |
| More similar when | Larger (max 1) | Smaller (min 0) | Larger |
| Sees magnitude? | No | Yes | Yes |
| Typical use | Text embeddings (default) | Clustering, k-NN, magnitude-meaningful features | Models trained for it; fastest |