Two models are pretrained on the same corpus. They share the same 12-block stack, the same vocabulary and the same 500-token context length. Only the mask and the training objective differ.
Each model is given the same 500-token passage exactly once — one pass, no repeats.
Roughly how many token predictions does that single passage give each model to learn from?
Select all that apply.