Greedy vs. Sampling: Why the Same Prompt Can Give You a Different Answer Every Time
Ask an LLM (large language model — the kind of AI behind tools like ChatGPT) the exact same question twice. Sometimes you get the exact same answer back, word for word. Other times you get two answers that take different angles on the same question, worded completely differently. Both are normal — and both come down to one setting (a parameter passed into the code — not to be confused with a model's parameters, the separate term for its internal size), hiding in plain sight in the same line of code everyone copies from a tutorial:
outputs = model.generate(
**inputs,
max_new_tokens=40, # caps the reply at 40 tokens (word pieces)
do_sample=False,
)
That's a simplified version of the standard Hugging Face (a popular platform and code library for running AI models) transformers convention for generating text — the full version used for every result below pins a couple of extra settings for consistency, shown later. do_sample decides which of two entirely different mechanisms turns the model's probability for every possible next word into one actual word on screen — and which one is active is exactly what decides whether you get the same answer every time, or a different one every time. This isn't about how confident the model is in an answer. It's about what happens one step after that confidence is already computed. This post runs both mechanisms, for real, against a real small free model, so every line of output below is something you can rerun yourself, not a claim to take on faith.
What's Actually Happening When a Model "Decides" an Answer
At every step, a model computes a probability for every word it could say next — none of that is in dispute. Isn't computing that probability basically the same thing as picking the word? Not quite — computing the probabilities and picking an actual word are two separate steps, and the second step has two completely different implementations a system can choose between. Greedy decoding always takes the single highest-probability word — the same input always produces the same output, by definition, with no exceptions. Sampling instead performs a weighted random draw across the whole distribution (the full list of next-word probabilities, all adding up to 100%) — a word with 80% probability usually wins, but "usually" isn't "always," and which word actually comes out can change from one call to the next.
Greedy Decoding: The Deterministic Default
Asking a small, free, instruction-tuned model (Qwen2.5-0.5B-Instruct — 494 million parameters, so called because it's been additionally trained to follow questions and instructions well. No training happens in this post — only inference, meaning the model is just answering, not learning anything new) the same question twice with do_sample=False:
QUESTION = "What's one benefit of drinking water first thing in the morning?" greedy_1 = ask(QUESTION, do_sample=False) greedy_2 = ask(QUESTION, do_sample=False)
Shouldn't an LLM give at least slightly different phrasing each time, just by nature? Not under greedy decoding — both runs come back byte-for-byte identical: "Drinking water first thing in the morning can offer several benefits for your health and overall well-being: 1. Hydration: Drinking water helps to keep you hydrated throughout the day, which is crucial for..." No variation, no exceptions — that's not a coincidence, it's what "always take the highest-probability word" guarantees mathematically.
Sampling: Same Meaning, Different Words Every Time
Same model, same exact question, same code — only do_sample=True changes, run with two different random seeds (a random seed is just the starting number that determines exactly which "random" choices get made — the same seed always reproduces the same draw):
Seed 1: "The primary advantage of drinking water initially upon waking is that it hydrates your body at rest and prepares your system for activity. It can help maintain electrolyte balance, which may prevent headaches caused by dehydration..."
Seed 2: "One significant benefit of drinking water first thing in the morning is its restorative nature. When we wake up earlier than usual or feel fatigue earlier in the day, it can be challenging to maintain our daily..."
Does the different wording mean one of these is wrong? No — neither answer is wrong, and neither contradicts the other, they just emphasize different angles: Seed 1 focuses on hydration and electrolyte balance, Seed 2 pivots to fatigue and feeling restored. Not the same sentence, not even close, despite an identical prompt and an identical model. That's sampling doing exactly what it's supposed to do: drawing from the probability distribution instead of always taking its single highest point.
Why This Isn't a Bug
Doesn't non-deterministic output sound like exactly the kind of thing you'd want to fix? Not necessarily — it depends entirely on what you're building. A chat assistant benefits from sampling: the same question asked twice shouldn't feel like talking to a tape recorder. A test suite, an automated test pipeline that re-runs on every code change, or a demo that needs to look identical on every run benefits from greedy decoding instead, precisely because it removes variation as a possibility. Neither setting is the "correct" one — they're a trade-off between predictability and variety, and which side you want depends on what you're actually building, not on which one is more "advanced."
If reproducible output is what you need, the fix is one line, already sitting in the exact same generate() call shown above: set do_sample=False, or fix a random seed before a sampling call. Either one turns "probably the same, usually" into "guaranteed the same, every time" — on demand, not by accident.
Sampling has its own separate tuning knobs beyond the on/off switch covered here — a topic for another time.
Get the Code
The full setup, both decoding runs, and every quoted output above: GreedyVsSampling.ipynb. Model used: Qwen/Qwen2.5-0.5B-Instruct, via Hugging Face transformers.
Summary
Computing a probability for every possible next word and picking one of them are two different steps, and the second one has two different implementations. Greedy decoding always takes the single highest-probability word — the same prompt to the same model produced byte-for-byte identical output on both runs, with no exceptions, because that's what "always pick the top one" guarantees. Sampling instead performs a weighted random draw — the same prompt, same model, same code, produced two differently worded but equally valid answers, because "usually the top one" leaves room for anything else with real probability mass to occasionally win instead. Neither behavior is a bug; which one is active is a single parameter you control, and now you know exactly which line of code to change to get the one you actually want.
Comments
Post a Comment