An estate agency has exported n house listings to train a price model. The rows are numbered 0 to n - 1, and the export happens to be sorted by price. Cutting it into three consecutive blocks would put all the cheapest houses in training and all the most expensive ones in test, so the model would be tested on prices it never learned from.
So the team shuffles first. They also fix the shuffle with a seed, so that anyone who re-runs the code gets exactly the same split.
The recipe:
order = list(range(n)).random module, exactly like this:val_pct percent of n rows and the test set gets test_pct percent of n rows, each rounded down to a whole number. The training set gets every row that is left, so no row is lost to rounding.order: the first rows go to training, the next ones to validation, and the rest to test.Task: write three_way_split(n, val_pct, test_pct, seed) and return a tuple (train, val, test) of three lists of row numbers, each sorted in increasing order.
0 to n - 1 appears in exactly one of the three lists.val_pct and test_pct are whole numbers that add up to less than 100.The seed is what makes a shuffled split repeatable. Re-run with the same seed and the same rows land in the same buckets, so a validation score from Monday can be compared fairly with one from Friday.