Cosine Similarity Isn't Measuring What You Think It Is
A support bot got asked "How many days do I have to recover a deleted file?" and handed back a chunk about refund policy instead. That's a retrieval failure — the step in a RAG (Retrieval-Augmented Generation) pipeline that finds the relevant piece of text before an LLM (Large Language Model) ever sees the question. A previous investigation traced one cause of bad retrieval back to chunking — how a document gets cut into pieces before anything gets compared. This time chunking wasn't the bug. The chunks were fine. The comparison itself still picked the wrong one. The refund-policy chunk scored 0.181 against that query. The chunk that actually explains the file-recovery window scored 0.134 . Not a rounding error. A clear, confident, wrong ranking, reproduced with seven lines of scikit-learn (a free Python library for exactly this kind of text math). Cosine similarity was doing exactly what it was designed to do. The problem was what it was being ask...