Measuring Similarity: Cosine, Euclidean and Dot Product

'Nearest' needs a definition. Cosine similarity compares direction, Euclidean distance compares position, and the dot product mixes direction with magnitude. For normalised text embeddings all three rank results identically — and knowing why tells you when the choice actually matters.

Advanced RAG Masterclass

We keep saying "find the chunks closest to the question". Close in what sense? Two vectors can point the same way but have different lengths. They can sit near each other but point in slightly different directions. Each similarity measure gives a different meaning to "close", and every vector index is created with one of them. This chapter makes the three standard choices concrete and then shows why, for most RAG systems, the choice turns out not to matter, along with the cases where it does.

Try it yourself
RAG Lab: cosine, dot product and Euclidean →

Drag a query and three document vectors, compare the three rankings, then normalise them and watch the rankings agree.

Cosine similarity: same direction?

Cosine similarity measures the angle between two vectors and ignores their lengths:

cos⁡θ=a⋅b∥a∥ ∥b∥.\cos\theta = \frac{\mathbf{a}\cdot\mathbf{b}}{\lVert\mathbf{a}\rVert\,\lVert\mathbf{b}\rVert}.

It ranges from +1+1 (same direction) through 00 (perpendicular, unrelated) to −1-1 (opposite).

Worked example. Take a=(1,2,3)\mathbf{a}=(1,2,3) and b=(4,5,6)\mathbf{b}=(4,5,6).

  • Dot product: 1⋅4+2⋅5+3⋅6=321\cdot4 + 2\cdot5 + 3\cdot6 = 32.
  • Lengths: ∥a∥=1+4+9=14≈3.742\lVert\mathbf{a}\rVert=\sqrt{1+4+9}=\sqrt{14}\approx 3.742 and ∥b∥=16+25+36=77≈8.775\lVert\mathbf{b}\rVert=\sqrt{16+25+36}=\sqrt{77}\approx 8.775.
  • Cosine: 32/(3.742×8.775)≈32/32.83≈0.97532 / (3.742 \times 8.775) \approx 32/32.83 \approx 0.975.

The two vectors point in almost the same direction, even though b\mathbf{b} is more than twice as long. Cosine sees them as nearly identical.

This is usually what we want for text. "Best smartphones under 20,000" and "top mobile phones under 20k" should match strongly whatever the length of each embedding. Cosine is the default metric in most vector databases, and it is common in recommendation, image retrieval and outlier detection as well.

Euclidean distance: same place?

Euclidean (L2) distance is the straight-line distance between the two points:

d(a,b)=∥a−b∥=∑i(ai−bi)2.d(\mathbf{a},\mathbf{b}) = \lVert\mathbf{a}-\mathbf{b}\rVert = \sqrt{\textstyle\sum_i (a_i-b_i)^2}.

Here smaller is more similar, and 0 means identical. For our example, d=32+32+32=27≈5.20d = \sqrt{3^2+3^2+3^2} = \sqrt{27} \approx 5.20: quite far apart, because b\mathbf{b} is much longer. Euclidean distance is sensitive to magnitude, which is why it is the natural choice in classical machine learning, for k-nearest neighbours and k-means clustering, where features have absolute meaning.

When magnitude carries meaning, use a measure that sees it. Suppose each vector is a user's activity profile: movies watched, products bought, searches made. A user with (2,1,3)(2, 1, 3) and a user with (20,10,30)(20, 10, 30) have the same pattern (cosine 1.0) but very different engagement. For a "find similar power users" feature, Euclidean distance correctly separates them, and cosine would wrongly call them twins.

Dot product: direction and magnitude

The plain dot product is

a⋅b=∑iaibi=∥a∥ ∥b∥cos⁡θ,\mathbf{a}\cdot\mathbf{b} = \sum_i a_i b_i = \lVert\mathbf{a}\rVert\,\lVert\mathbf{b}\rVert\cos\theta,

that is, cosine multiplied by both lengths. Larger means more similar. It rewards both pointing the same way and being long. Some embedding models are trained so that length encodes something useful (for example, confidence or popularity), and for them the dot product is the intended measure. It is also the cheapest of the three to compute, since there are no square roots or divisions.

Three panels: cosine measures the angle between two arrows, Euclidean measures the straight-line gap between their tips, and dot product measures one arrow's projection onto the other times its length
Three meanings of 'close'. Cosine looks only at the angle, Euclidean at the gap between the tips, and the dot product at angle and length together.

The key fact: normalised vectors make them agree

Most modern text-embedding models output vectors normalised to length 1 (or the database normalises them for a cosine index). When ∥a∥=∥b∥=1\lVert\mathbf{a}\rVert=\lVert\mathbf{b}\rVert=1:

  • the dot product equals the cosine, since the length factors are 1, and
  • the squared Euclidean distance is ∥a−b∥2=2−2cos⁡θ\lVert\mathbf{a}-\mathbf{b}\rVert^2 = 2 - 2\cos\theta, a decreasing function of the cosine.

So for unit vectors all three measures produce exactly the same ranking. Normalising our example vectors gives a squared distance of 0.0507=2−2(0.9746)0.0507 = 2 - 2(0.9746), matching the formula. This is why, for typical text embeddings, the choice of metric barely changes results, and why databases often compute "cosine" internally as a fast dot product on normalised vectors.

The choice matters when vectors are not normalised and their length means something: activity scores, purchase power, or embedding models explicitly trained for dot-product similarity. In those cases, follow the model's documentation and choose the metric it was trained with.

CosineEuclideanDot product
MeasuresAngleGap between pointsAngle × lengths
More similar whenLarger (max 1)Smaller (min 0)Larger
Sees magnitude?NoYesYes
Typical useText embeddings (default)Clustering, k-NN, magnitude-meaningful featuresModels trained for it; fastest
EasySimilarity

Compute the cosine similarity of (1, 0) and (0, 3), and of (1, 1) and (3, 3). What does this show?

MediumSimilarityInterview

Why do cosine, dot product and Euclidean distance give the same ranking for normalised embeddings?

MediumSimilarity

Give a scenario where cosine similarity would give misleading results and Euclidean distance would be better.