An A/B test tells you whether a new model version is better, but it needs weeks of real outcomes. Before that, many teams run a quicker check that only asks whether the new version is broken: a canary release. A small slice of live requests goes to the new version (the canary) while everything else stays on the current version (the baseline). If the canary misbehaves, you roll back before most customers ever meet it.
Every request is logged as a tuple (version, ms, failed):
version is "baseline" or "canary",ms is the request's latency, how long the answer took in milliseconds,failed is True if the request errored instead of returning a prediction.The log always holds at least one request from each version.
Task: write canary_verdict(log, min_requests, max_error_gap, max_latency_ratio) that returns one of three strings:
"wait" if the canary has handled fewer than min_requests requests. There's too little evidence to judge it either way."rollback" if either guardrail is broken:
max_error_gap above the baseline's error rate, ormax_latency_ratio times the baseline's p95 latency."promote" otherwise.A group's p95 latency is the latency that 95% of its requests come in at or under. Use this exact recipe: sort the group's latencies from smallest to largest and take the one at position , counting from 1, where is the group's size and means round up (math.ceil). For 20 requests, that is the 19th fastest.
Why p95 and not the average? The average blends a slow tail into the fast majority. If 18 of 20 requests take 100 ms and 2 take 900 ms, the average is 180 ms, which looks like a mild slowdown. The p95 is 900 ms: one customer in ten is waiting almost a second, and p95 shows that directly.
There, the canary has no failures, but its p95 latency is 400 ms against the baseline's .