A -billion-parameter model has just finished training. Throughout the run, the GPU had to hold all of the following at once, with every number stored at 32 bits (4 bytes):
| what the training run keeps | how much |
|---|---|
| weights | one number per parameter |
| gradients | one number per parameter |
| optimizer state | two numbers per parameter |
| stored activations | GB in total |
For serving, the weights are quantized to 8 bits, and the GPU holds only what inference itself needs. Ignore the small scratch space a single forward pass needs while it runs.
Take bytes.
By what factor is the serving memory smaller than the training memory?
Round your answer to 1 decimal place.