A GPT model does exactly one thing: given the tokens so far, it scores every token in its vocabulary as a candidate for the next one. Writing a whole paragraph is a loop wrapped around that. Predict one token, append it, feed the longer sequence back in, and predict again.
There is one practical limit. This model learned its position vectors, and it learned exactly block_size of them. It has no vector for any position beyond that, so it can only ever be shown the most recent block_size tokens. Anything older is cut off before each call.
Task: write generate(model, prompt, steps, block_size).
model is a function. Pass it a list of token ids (at most block_size of them) and it returns a list of scores, one per vocabulary token, for token ids 0, 1, 2, ... in order.prompt is the starting list of token ids. It may already be longer than block_size.steps rounds of greedy decoding: each round, choose the token with the highest score. If several tokens tie for the highest score, choose the one with the smallest id.In the tests, each model is a small hand-written rule rather than a trained network, so that every run is predictable.