A school tries a new way of running revision sessions. Five students get the new format (group A), five get the usual one (group B), and everyone sits the same quiz. Group A scores higher on average, but with only ten students, could a gap that size be luck?
A permutation test answers this without any formula for the sampling distribution. Under the null hypothesis H0 the format makes no difference, so the labels "A" and "B" are arbitrary: any five of the ten scores could just as easily have been the group-A five. That gives a recipe:
- Compute the observed test statistic, dobs=xˉA−xˉB.
- Pool all the scores. Go through every way of choosing which nA of the pooled positions carry the label A (the other nB positions are B), and compute d=xˉA−xˉB for each relabelling. Together these values are the null distribution: what chance alone produces.
- The two-sided p-value is the fraction of relabellings whose ∣d∣ is at least as large as ∣dobs∣.
Shuffling the labels at random many times would sample from exactly this list of relabellings. With small groups you can try them all instead, so the answer does not depend on luck.
Task: write permutation_test(group_a, group_b) that returns the tuple (d_obs, p_value), both rounded to 4 decimal places.
- The groups may have different sizes (each has at least one value), and values may repeat. A relabelling is a choice of positions, so two equal scores in different positions still count as different relabellings.
- The same mean computed from values in a different order can differ in the last floating-point digit, so count a relabelling as extreme when ∣d∣≥∣dobs∣−10−9.
- The two groups hold at most 12 values between them, so trying every relabelling is fast.