An online bookshop A/B tests a new recommender. A random draw puts each customer in a group: group A keeps the champion (the current model), group B gets the challenger. Two weeks later, the team compares the share of customers in each group who bought something.
To tell a real difference from chance, they use a two-proportion z-test. With buyers out of customers in each group:
The standard error uses the pooled rate , because the null hypothesis says both models convert equally well. The p-value is two-sided: twice the area in one tail of the standard normal curve beyond .
The team ships the challenger only when the test gives significant evidence, a p-value below alpha, that the challenger outperforms the champion. Otherwise the champion stays.
Task: write ab_decision(buyers_a, n_a, buyers_b, n_b, alpha) returning a tuple (z, p_value, decision):
z and p_value rounded to 4 decimal places,decision either "ship" or "keep".The inputs always include at least one buyer and at least one non-buyer overall, so the standard error is never zero.