A writing assistant shows its reply while it is still being generated. The model predicts one token, appends it, feeds everything back in and predicts the next, and the app puts each token on screen the moment it arrives.
So a user never feels a single latency number. They feel a wait before anything appears, and then a flow of words that can run smoothly or stall.
The app logs, on one clock in milliseconds, the moment the request was sent and the moment each token of the reply arrived.
Task: write stream_metrics(sent_ms, token_ms) and return a tuple of four numbers, each a float rounded to 4 decimal places:
token_ms is in arrival order and always holds at least one token.0.0.Each number answers a different complaint. The first wait decides whether the reply feels instant at all. The mean gap is how fast the words flow once they start. The worst gap is the stall in mid-sentence that the average hides.