A grocery delivery app A/B tests a new recommender. The plan is that a random draw sends 1 in 4 customers to the challenger, and everyone else keeps the champion. After two weeks the dashboard shows:
| Champion | Challenger | |
|---|---|---|
| Customers assigned | 30,620 | 9,780 |
| Customers who placed an order | 1,531 | 568 |
Before comparing order rates, the team runs a sample ratio mismatch (SRM) check. It asks whether the draw really split customers in the planned ratio. If something broke the split, such as a bug that drops some challenger customers, then the groups may no longer be alike, and comparing their outcomes can't be trusted.
The check uses the chi-square goodness-of-fit statistic:
where is a group's observed count and is the count the planned split predicts for that group, given the total.
For two groups, a above would turn up less than 1 time in 1,000 if the draw were working as planned. Many teams treat that as a broken experiment.
What is for this test? Round your answer to 2 decimal places.