A team is planning to pretrain a decoder-only language model and has to book GPU time before anything runs.
| quantity | value |
|---|---|
| parameters | billion |
| training corpus | billion tokens |
| passes over the corpus | epochs |
| peak speed of one GPU | teraFLOP/s, i.e. FLOPs per second |
| share of that peak a real training run achieves, on average |
A rule of thumb from scaling-law studies estimates the compute needed for training as
floating-point operations (FLOPs), where is the number of tokens the model is trained on.
GPU time is booked in GPU-days. One GPU running for 24 hours is 1 GPU-day, and 10 GPUs running for 24 hours are 10 GPU-days.
How many GPU-days does this training run need? Round to the nearest whole GPU-day.