A lab has drafted a training run for a new language model:
| quantity | draft plan |
|---|---|
| parameters | billion |
| training tokens | billion |
Two rules of thumb from scaling-law studies:
Cost. Training a model with parameters on tokens takes about
floating-point operations (FLOPs). This is the compute budget.
Best split. For the lowest loss a fixed budget can buy, use roughly 20 training tokens per parameter:
The lab decides to keep the draft's compute budget exactly as it is, but divide it between parameters and tokens using the 20-tokens-per-parameter rule.
How many training tokens, in billions, should the re-planned run use? Round to the nearest whole billion.