BERT learns language without a single human label: hide some of the words in a sentence and train the model to guess them from everything around them. The text is its own answer key.
Before a sentence reaches the model, a small preprocessing step decides what to hide. It works one position at a time, using random numbers. So that your output can be checked, the random numbers have already been drawn: every position comes with a roll (pick, action, swap).
Step 1: choose positions. A position is chosen when pick < rate. BERT uses rate = 0.15, so about 15% of the words. A position that is not chosen goes into the input exactly as it was, and its action and swap are ignored.
Step 2: decide what the model sees at each chosen position. If a chosen word were always replaced by [MASK], the model would only ever learn to predict where it sees [MASK], a token that never shows up in the real text it is later fine-tuned on. So action picks one of three treatments:
action | what goes into the input |
|---|---|
below 0.8 | the token "[MASK]" |
from 0.8 up to, but not including, 0.9 | the word swap (a random word from the vocabulary) |
0.9 or above | the original word, unchanged |
Step 3: write the answer key. The label at a position is the original word if that position was chosen in Step 1, and None if it was not.
Task: write mask_tokens(tokens, rolls, rate), where tokens is a list of words and rolls holds one (pick, action, swap) tuple per word. Return a tuple (inputs, labels) of two lists, each the same length as tokens:
inputs: what the model actually reads,labels: what it is asked to predict at each position.This one function is the whole trick behind pretraining on raw text: it turns any sentence into a supervised example, with no person needed to write the answers.