At every step, a GPT model outputs one raw score (a logit) for each token in its vocabulary. Before picking the next token, most text generators reshape that distribution in three stages:
top_k most probable tokens. A top_k of 0 means skip this stage.top_p. The token that brings the total to top_p or beyond is kept, and everything after it is dropped. A top_p of 1.0 means skip this stage.Finally, renormalise whatever survived so it sums to 1. Every dropped token gets probability 0.
When two tokens have exactly the same probability, the one with the lower index counts as the more probable.
Task: write filter_next_token(logits, temperature, top_k, top_p). Return the final distribution as a list with one entry per vocabulary token, in the original order, each rounded to 4 decimal places.
Top-k always keeps the same number of tokens. Top-p keeps however many the model's own confidence calls for: a short list when the model is sure, a long one when it is not.