A GPT model learns from raw text with no labels at all. The answer for every token is simply the token that came after it: the text is its own answer key. Before training starts, a small preprocessing step turns a pile of documents into (input, target) pairs.
Step 1: one long stream. Join the documents end to end, and put the end-of-text token eot after every document, including the last one. That way the model learns where a text stops, and that whatever follows eot belongs to a new, unrelated text.
Step 2: windows. Cut windows of block_size + 1 consecutive tokens from the stream. The first window starts at position 0, and each later window starts stride positions after the one before it. A stride smaller than block_size makes windows overlap; a larger one skips some tokens. Windows may run across a document boundary. A window that would run past the end of the stream is not made, so any leftover tokens at the end that cannot fill a whole window are dropped.
Step 3: shift by one. From each window, the input is its first block_size tokens and the target is its last block_size tokens. Position i of the input is graded against position i of the target: the token that actually came next.
Task: write next_token_pairs(documents, eot, block_size, stride), where documents is a list of documents, each a list of token ids (possibly empty). Return a list of (input, target) tuples, each holding two lists, in the order the windows were cut. If the stream is too short for even one window, return [].
One pair holds
block_sizeseparate predictions, not one. The causal mask is what lets all of them train in a single pass: input positionican see positions0toi, but never positioni + 1, which holds its answer.