When a language model generates text, it chooses each next token from a probability distribution. Sampling settings control how it chooses.
Temperature
Temperature scales how peaked or flat the distribution is.
- Low (around 0–0.3): the model almost always picks the most likely token. Output is focused and more repeatable.
- Medium (around 0.5–0.8): a balance of reliability and variety.
- High (above 1): more random and creative, with more risk of incoherence.
Top-p (Nucleus Sampling)
Top-p limits choices to the smallest set of tokens whose probabilities add up to p. With top-p = 0.9, very unlikely tokens are excluded. It's usually adjusted instead of, not as well as, temperature.
Other Settings
- Max tokens: caps output length (and cost).
- Stop sequences: end generation when a given string appears.
- Frequency/presence penalties (on some APIs): discourage repetition.
Suggested Starting Points
- Extraction, classification, structured output: low temperature.
- Answering from documents: low to medium.
- Brainstorming and creative writing: higher.
Caveats
Temperature zero doesn't guarantee identical outputs across runs, and some newer reasoning-focused models restrict or ignore these settings — check your provider's documentation. Settings are no substitute for a clear prompt.