A photo-tagging service runs one copy of its model on one GPU. Requests arrive steadily at 400 per second, evenly spaced (one every ms), all day long.
Timing the forward pass at different batch sizes gave a simple rule:
The ms is paid whatever the batch size; each extra request in the batch adds only ms.
The server batches using a fixed collection window of milliseconds:
Setting means no batching at all: every request is run on its own, as a batch of one.
The product team has two requirements:
Treat network, preprocessing and postprocessing time as zero. Which collection window meets both requirements?
Select all that apply.