A writing app suggests the next few words while the user types, so its suggestion model is put behind an endpoint. Inference is stateless, so the team scales by running identical copies, each on its own GPU, with incoming requests spread evenly across them.
Each copy handles one batch at a time. Requests that arrive while it is busy queue up; the moment a pass finishes, the copy starts the next one with everything that has queued, up to a cap of requests. A forward pass over a batch of requests takes
that is, ms of fixed overhead paid whatever the batch size, plus ms per request in the batch.
The worst-placed request arrives just after a full batch has started without it: it waits out that entire pass, then rides in the next one, which is also full. So inside the model service the worst case is two back-to-back passes at the cap.
Every request also spends time outside the model, and batching changes none of it:
| stage | time |
|---|---|
| network, there and back | ms |
| preprocessing | ms |
| postprocessing | ms |
A suggestion is only useful if it feels instant, so the product rule is that the worst-case request must take no more than ms end to end. At peak, requests per second arrive. The team may choose any whole-number cap , the same on every copy. The copies keep up if together they can finish requests at least as fast as requests arrive.
What is the smallest number of copies that keeps up with peak traffic while every request stays within the ms budget?
Give the exact whole number.