Greedy vs. Sampling: Why the Same Prompt Can Give You a Different Answer Every Time
Ask an LLM (large language model — the kind of AI behind tools like ChatGPT) the exact same question twice. Sometimes you get the exact same answer back, word for word. Other times you get two answers that take different angles on the same question, worded completely differently. Both are normal — and both come down to one setting (a parameter passed into the code — not to be confused with a model's parameters, the separate term for its internal size), hiding in plain sight in the same line of code everyone copies from a tutorial: outputs = model.generate( **inputs, max_new_tokens=40, # caps the reply at 40 tokens (word pieces) do_sample=False, ) That's a simplified version of the standard Hugging Face (a popular platform and code library for running AI models) transformers convention for generating text — the full version used for every result below pins a couple of extra settings for consistency, shown later. do_sample decides which of two en...