Top-p Sampling: Which Tokens Survive?

黎 浩然/ 9 10 月, 2026/ 大语言模型/LARGELANGUAGEMODEL/LLM, 机器学习/MACHINELEARNING, 研究生/POSTGRADUATE, 计算机/COMPUTER/ 0 comments

top-p does not mean “80% randomness” or “retain 80% of the vocabulary”. It orders next-token probabilities, keeps a small candidate set whose cumulative probability reaches the threshold, then samples within that set. The candidate count changes with the distribution.

What does 0.8 retain?

Consider invented tokens A, B, C and D with probabilities 0.50, 0.30, 0.15 and 0.05. At top-p=0.8, A and B reach cumulative mass 0.8. After renormalization, A has probability 0.50/0.80=0.625 and B has 0.375. The removed mass of 0.2 is not an error probability.

A: 0.50 → 0.625
B: 0.30 → 0.375
C, D: excluded
Original post-filter probability diagram using invented data, not model predictions.

The Hugging Face generation guide defines top-p through the smallest high-probability set reaching the threshold. A concentrated distribution may retain very few candidates; a flatter distribution may retain many. The same parameter therefore behaves differently across questions and generation positions.

Reproduce the boundary calculation

This independent probability-filter function was executed. Assertions check the candidates, normalization and threshold 1. It does not call Transformers or simulate full text generation; real libraries may differ in boundary and tie handling.

from math import log, exp, isclose

def nucleus(probabilities, threshold):
    if not 0 < threshold <= 1 or not probabilities:
        raise ValueError('invalid input')
    if min(probabilities) < 0 or not isclose(sum(probabilities), 1):
        raise ValueError('expected a probability distribution')
    ordered = sorted(range(len(probabilities)), key=lambda i: (-probabilities[i], i))
    keep, mass = [], 0.0
    for i in ordered:
        keep.append(i); mass += probabilities[i]
        if mass >= threshold: break
    return keep, {i: probabilities[i]/mass for i in keep}

p = [.5, .3, .15, .05]
keep, result = nucleus(p, .8)
assert keep == [0, 1] and isclose(sum(result.values()), 1)
assert isclose(result[0], .625) and isclose(result[1], .375)
assert nucleus(p, 1)[0] == [0, 1, 2, 3]
print('kept tokens:', keep)
print('conditional probabilities:', result)

It returns candidates [0,1] with probabilities approximately 0.625 and 0.375. Include the candidate that crosses the threshold rather than stopping before it. Index-based ties make this example reproducible. A smaller parameter does not guarantee a truer answer: an incorrect token can have the highest predicted probability.

Consider temperature and decoding mode together

Temperature changes the distribution’s shape, while top-p trims its tail. When both are enabled, processing order affects the final candidates. Check the service and version instead of assuming equal settings produce equal behavior across services.

With a non-sampling decoding path, sampling settings may not have the expected effect. Record model version, decoding mode, parameters, input and output length. Evaluate task quality and repeated-generation stability rather than treating low top-p as a reliability guarantee.

Machine-readable results still need structured-output validation. Sampling settings do not replace format checks, business rules or source verification.

中文版

References

Official documentation

Publication note: actually published as a catch-up on October 10, 2026 (Beijing time), retaining the planned 2026-10-09 19:00 article date. The code is an original teaching check, not a real-model performance test.

Support

If this article helped you, you can support this site.

WeChat support QR code; click to enlarge
WeChat
Alipay support QR code; click to enlarge
Alipay
Buy Me a Coffee; support this site
Buy Me a Coffee

Click a QR code to enlarge. More options: support page。

Share this Post

Leave a Comment

您的邮箱地址不会被公开。 必填项已用 * 标注

*
*