A 14-billion-parameter chat model is served from a single GPU. The card has:
| property | value |
|---|---|
| memory bandwidth | GB per second (1 GB bytes) |
| arithmetic speed | operations per second |
| memory size | enough to hold any of the versions below |
Generating each new token of a reply takes two jobs:
The two jobs overlap, so one token takes as long as the slower of the two.
The team has three versions of the weights: 32-bit, 8-bit and 4-bit. The product requirement is that a 240-token reply finishes within 5 seconds. Every bit dropped risks a little accuracy, so they will ship the version with the most bits per weight that still meets the requirement.
How long, in seconds, does a 240-token reply take with the version they ship? Round to 2 decimal places.