Back to Home

What are the sampling parameters of the large model: temperature, Top-p, and Top-k being adjusted for

September 17, 2026 at 01:32 PMSource: RunByAI0 comment(s)TechGuide

For the same big model, some people get a stable answer while others get a wild and imaginative sentence when asked the same question 'Write me a poem about spring'. The difference often lies not in the model itself, but in a few inconspicuous parameters when called - Temperature, Top-p, Top-k. They collectively determine how the big model selects a word.

1、 The model is actually "guessing the next word"

The big model only does one thing at a time: predicting the probability of the next word appearing based on existing text. It will output a probability distribution that covers the entire vocabulary, for example, after "spring", "lai" accounts for 30%, "de" accounts for 12%, "hua" accounts for 8%... The so-called "generation" is to continuously select words and spell them into sentences from this distribution, and then feed the results back to repeat this process.

If you always choose the word with the highest probability, it is called 'Greedy Decoding'. It is stable and reproducible, but prone to boredom and even getting stuck in a loop. In order to make the output more natural and diverse, it is necessary to introduce "random sampling" and adjust its several parameters.

2、 Temperature: How bold is it to control "randomness"

Temperature is the most intuitive one. It acts before softmax to scale the raw scores (logits): when the temperature approaches 0, the distribution is "compressed", high probability words almost monopolize, and the output is close to greedy decoding, stable but conservative; When the temperature is equal to 1, sample according to the original distribution of the model; When the temperature is greater than 1, the distribution is' flattened ', and low probability words also have a chance to be selected, resulting in more creative output but more prone to deviation and errors.

In experience, writing code and solving math problems are suitable for low temperatures (0-0.3) and require accurate facts; Writing copy and brainstorming are suitable for medium to high temperatures (0.7-1.0). Temperature is not a 'smart' knob, too high will only make the model more 'capricious'.

3、 Top-k: Only select from the most likely k words

The idea of Top-k is very direct: first sort by probability from high to low, only retain the first k candidate words, reset the rest to zero, and then sample proportionally from the remaining words. K=1 is equivalent to greed, and the larger k, the more random it is.

The problem with it is that k is fixed: in positions where the next step is almost certain (such as fixed collocations), forcing the retention of k candidates may also include words that are clearly not supposed to appear; And in positions where there are already many candidates, k may be too small, limiting expression.

4、 Top-p: Accumulate by probability to what extent

Top-p (also known as Nucleus Sampling) improves this point: instead of counting numbers, it accumulates probabilities from high to low until the cumulative probability reaches p (such as 0.9), using only a small number of words in the "kernel" for sampling.

In this way, the size of the candidate set will adapt to the context: the more determined the model, the smaller the kernel; The more uncertain, the larger the nucleus. In practical use, it is usually combined with temperature, such as temperature=0.7 and top_p=0.9, which are common default combinations for many interfaces.

5、 There are a few knobs that are often overlooked

Repetition penalty: Reduce the probability of words already appearing and alleviate the repetition of machine style repetition. Frequency/presence penalty: Punish based on the number of occurrences or whether they have occurred, commonly used for long texts. Maximum generated length and stop word: controls when the output ends.

6、 How to adjust: from conservative to liberal

A practical starting point is to first fix temperature=0.7 and top_p=0.9, and then fine tune according to the task. If you want facts to be accurate, cool down; if you want creativity to spread, heat up; Add a little repetition penalty if there is repetition. Never pull multiple parameters to extremes at the same time, as it will only make the output more unstable.

Summary

Temperature, Top-k, Top-p do not change what the model "knows", only how it "expresses". By understanding this, one can find the balance between stability and creativity that suits their own scene.

[Reference source] Comprehensive compilation of common definitions of temperature, top_p, top_k, and duplicate penalties from official documents of major mainstream model APIs.

large modelLarge Language Model (LLM)
Discussion

Comments (0)

No comments yet. Be the first!

Leave a Comment