Posts

Greedy vs. Sampling: Why the Same Prompt Can Give You a Different Answer Every Time

Ask an LLM (large language model — the kind of AI behind tools like ChatGPT) the exact same question twice. Sometimes you get the exact same answer back, word for word. Other times you get two answers that take different angles on the same question, worded completely differently. Both are normal — and both come down to one setting (a parameter passed into the code — not to be confused with a model's parameters, the separate term for its internal size), hiding in plain sight in the same line of code everyone copies from a tutorial: outputs = model.generate( **inputs, max_new_tokens=40, # caps the reply at 40 tokens (word pieces) do_sample=False, ) That's a simplified version of the standard Hugging Face (a popular platform and code library for running AI models) transformers convention for generating text — the full version used for every result below pins a couple of extra settings for consistency, shown later. do_sample decides which of two en...

Cosine Similarity Isn't Measuring What You Think It Is

Image
A support bot got asked "How many days do I have to recover a deleted file?" and handed back a chunk about refund policy instead. That's a retrieval failure — the step in a RAG (Retrieval-Augmented Generation) pipeline that finds the relevant piece of text before an LLM (Large Language Model) ever sees the question. A previous investigation traced one cause of bad retrieval back to chunking — how a document gets cut into pieces before anything gets compared. This time chunking wasn't the bug. The chunks were fine. The comparison itself still picked the wrong one. The refund-policy chunk scored 0.181 against that query. The chunk that actually explains the file-recovery window scored 0.134 . Not a rounding error. A clear, confident, wrong ranking, reproduced with seven lines of scikit-learn (a free Python library for exactly this kind of text math). Cosine similarity was doing exactly what it was designed to do. The problem was what it was being ask...

Data Leakage: The 97% Test Score That Became 78% in Production

Image
A model scores 97% accuracy in testing — the percentage of predictions it gets right on data it never trained on. You deploy it to production, and real-world accuracy drops to under 78%. Nothing crashed. Nothing was miscoded. So what actually broke? Nothing broke. The test score was measuring something that doesn't exist in the real world. This is data leakage — the quietest, most dangerous failure mode in machine learning, because it doesn't look like a failure at all. It looks like success. What Data Leakage Actually Is Data leakage is information that shouldn't be available at prediction time sneaking into training anyway. The model learns to rely on it. The test score looks great, because that same leaked information is sitting in the test set too. The moment the model is in the real world, making a prediction before that information exists, it falls apart. Isn't more information always better for a model, thou...

Why Your RAG Pipeline Retrieves the Wrong Chunk?

I asked a support bot: "What happens to my data if I downgrade my plan mid-cycle?" It answered confidently. It also answered a completely different question — about deleted file recovery windows, not downgrades. The model wasn't broken. The retrieval — the step that finds and hands over the right piece of text — was broken instead. So what caused that? Not the embedding model (the piece that turns text into comparable numbers), not the LLM (the model generating the actual answer), and not a bug in the prompt. It was broken because of a decision made before any of that — how the source document was cut into chunks. This is the failure mode nobody warns you about when you first build a RAG pipeline, because it doesn't look like a bug. It looks like the system working — just working on the wrong piece of text. The Default Everyone Reaches For Retrieval-Augmented Generation (RAG) has one job: find the relevant piece of a document, then han...

Gradient Descent: The Update Rule That Trains Every Model

Image
Every model — a neural network, a regression line, anything that learns from data — is really just a set of internal numbers called parameters, plus a single score called the loss that measures how wrong its current predictions are. Every parameter has a slope: nudge it up or down, and the loss changes. Stack all those slopes into one vector and you get the Gradient . It always points uphill — toward the steepest way to make the loss worse — whether you have 2 parameters or 175 billion. But knowing the uphill direction isn't the same as knowing what to do with it. The gradient only points up. So what good is that to a model trying to reach the bottom? Gradient Descent takes "here is uphill" and turns it into "here is your next step downhill." That's the whole idea. Every neural network, every regression model, every fine-tuning run learns using exactly this. The Update Rule Here is the entire algorithm, in one line: θ(...

The Silent Bug That Makes Your Model Wrong, Not Broken

Image
The Silent Bug That Makes Your Model Wrong, Not Broken Two customers. One is 25, the other is 30 — five years apart. Another pair: 25 and 70 — forty-five years apart. Ask a K-Nearest Neighbors model which pair is more similar, using age and income, and it tells you the 45-year age gap pair is the closer match. No error. No crash. Just a wrong answer, delivered with total confidence. Why It Happens Income ranges from roughly zero to a hundred thousand. Age ranges from zero to about a hundred. When a model measures "distance" between two people, it just does math on the raw numbers — and a $1,000 income difference completely swamps a 45-year age difference, because the numbers are so much bigger to begin with. The model isn't wrong about the arithmetic. It's just never been told these two numbers live on completely different scales. The Fix, in Real Numbers Scale those same two customers pro...

One-Hot Encoding: Why It Breaks Between Train and Test

Image
Your model works perfectly in testing. You deploy it. It crashes on the first real prediction. ValueError: The feature names should match those that were passed during fit. Feature names seen at fit time, yet now missing: - Delhi If you've seen an error like this — a feature that's missing, or a mismatch in feature names — it almost always traces back to one decision: how a categorical column got converted into numbers. What Actually Happened Somewhere in the pipeline, a column like city — Bangalore, Mumbai, Delhi — got converted into numbers. That conversion is called one-hot encoding . Instead of one column holding text labels, you get one new column per category , each holding a 1 or a 0. Bangalore becomes city_Bangalore = 1 , everything else 0. Mumbai becomes city_Mumbai = 1 , everything else 0. Why Not Just Use 1, 2, 3? Models only understand numbers. They can't do math on the word "Bangalore....