Cosine Similarity Isn't Measuring What You Think It Is

A support bot got asked "How many days do I have to recover a deleted file?" and handed back a chunk about refund policy instead. That's a retrieval failure — the step in a RAG (Retrieval-Augmented Generation) pipeline that finds the relevant piece of text before an LLM (Large Language Model) ever sees the question. A previous investigation traced one cause of bad retrieval back to chunking — how a document gets cut into pieces before anything gets compared. This time chunking wasn't the bug. The chunks were fine. The comparison itself still picked the wrong one.

The refund-policy chunk scored 0.181 against that query. The chunk that actually explains the file-recovery window scored 0.134. Not a rounding error. A clear, confident, wrong ranking, reproduced with seven lines of scikit-learn (a free Python library for exactly this kind of text math).

Cosine similarity was doing exactly what it was designed to do. The problem was what it was being asked to compare.

What Cosine Similarity Actually Measures

Cosine similarity measures the angle between two vectors (a list of numbers representing a piece of text), ignoring their length:

similarity = (A · B) / (‖A‖ ‖B‖)

A score of 1 means the vectors point in the exact same direction. A score of 0 means they're perpendicular — unrelated. If the geometry is this simple and this well-understood, how did it get the ranking wrong? It didn't — cosine similarity never once evaluates whether "refund" and "file recovery" mean the same thing. It only ever compares the two vectors it's handed, whatever those vectors happen to encode.

The weak link is upstream: what turns the text into a vector in the first place. Cosine similarity is only ever as meaningful as the vectors it's given.

TF-IDF: Counting Words, Not Meaning

The simplest way to vectorize text is TF-IDF (Term Frequency–Inverse Document Frequency) — score each word by how often it appears in a chunk, discounted by how common that word is across all the chunks in the document. No neural network (no brain-inspired model learning from examples), no training, entirely offline:

from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.metrics.pairwise import cosine_similarity

def tfidf_vectorize(query, chunks):
    vectorizer = TfidfVectorizer()
    chunk_vectors = vectorizer.fit_transform(chunks)
    query_vector = vectorizer.transform([query])
    return query_vector, chunk_vectors

Doesn't discounting common words already filter out coincidental overlap like this? Only for words common across the whole document — and "days" isn't rare here. It shows up in the refund section, the recovery section, and the security section alike. The query "How many days do I have to recover a deleted file?" and the refund-policy chunk both contain the literal token "days" (a single countable unit TF-IDF tracks, blind to meaning) — twice, in the refund chunk alone. TF-IDF has no concept of meaning: it sees a repeated word and scores the two chunks as similar, because to a word-counting method, they are. It can't tell that "30 days to recover a file" and "30 days to request a refund" are two unrelated policies that happen to share vocabulary.

Swapping the Vectorizer, Not the Metric

If the similarity formula never changes, what's actually different about the fix? Not the metric — the vector. A real embedding model, trained to place sentences with similar meaning near each other regardless of exact word overlap, plugs into the exact same retrieval function. The underlying code barely changes:

def embedding_vectorize(query, chunks):
    from sentence_transformers import SentenceTransformer
    model = SentenceTransformer("all-MiniLM-L6-v2")
    chunk_vectors = model.encode(chunks)
    query_vector = model.encode([query])
    return query_vector, chunk_vectors

def retrieve(query, chunks, vectorize_fn=tfidf_vectorize, top_k=2):
    query_vector, chunk_vectors = vectorize_fn(query, chunks)
    scores = cosine_similarity(query_vector, chunk_vectors)[0]
    ranked = sorted(range(len(chunks)), key=lambda i: scores[i], reverse=True)
    return [(i, scores[i], chunks[i]) for i in ranked[:top_k]]

retrieve() never changes. Same cosine similarity, same ranking logic. Only vectorize_fn changes — and running both back to back on the same document confirms it's the only thing that needs to.

With a real embedding model (all-MiniLM-L6-v2, a small pretrained sentence-embedding model from the sentence-transformers library), the file-recovery chunk scores 0.563 and comes back rank 1. The refund chunk drops to 0.397 and rank 2. Same query, same document, same retrieve() function — one swapped argument, and the ranking flips from wrong to right.

Why This Distinction Matters More Than It Looks

Isn't this just a technicality between two flavors of the same "embedding" idea? It's easy to treat embeddings (numbers standing in for a piece of text) as a single black-box ingredient in a RAG pipeline — plug one in, get semantic search (search that matches meaning, not just matching letters). But TF-IDF vectors are technically embeddings too; they're just embeddings built on word overlap, not meaning. The failure with the refund-policy chunk isn't a retrieval bug. It's a mismatch between the vectorizer's notion of "similar" and the question's notion of "relevant." A bag-of-words model (one that only counts which words appear, never what they mean) was always going to be fooled by two unrelated sentences sharing a number and a unit.

This is also why swapping in a bigger LLM (the model that generates the final answer, one step later in the pipeline than retrieval) never fixes a bad retrieval. It only ever sees what got ranked highest. If the ranking itself confused word overlap for meaning, a better model just answers the wrong question more fluently.

Get the Code

The retriever supports both backends behind one interface, so you can reproduce this exact ranking flip locally (flip the BACKEND variable to "embedding" to see the second half) — no API key needed for either: CosineSimilarityEmbeddings.py.

Summary

Cosine similarity is just angle-comparison between two vectors — it never breaks, and it never evaluates meaning either. The vectorizer that produces those vectors is where meaning gets decided. TF-IDF counts overlapping words, so a refund-policy chunk and a file-recovery chunk sharing "days" score as similar even though the policies are unrelated (0.181 vs. 0.134 on a real toy document, decoy beating the true answer). Swapping in a real embedding model, trained to place semantically similar text near each other, corrects the ranking (0.563 vs. 0.397) without touching the similarity formula at all. The lesson generalizes past this one example: when a RAG pipeline retrieves the wrong chunk, check what turned the text into numbers before blaming the comparison or the model reading the result.

Comments

Popular posts from this blog

Gradient Descent: The Update Rule That Trains Every Model

READ vs SQL SELECT, A Quick Performance Test

All about READ in RPGLE & Why we use it with SETLL/SETGT?