Before training a model, it's common to hold some rows back for testing later. The rest are the training rows. Anything learned from the data, such as each column's average, must come from the training rows only.
X is a table with one row per example. Split it like this, so everyone's random numbers come out the same:
rng = np.random.default_rng(seed).rng.random(n), where n is the number of rows. This gives one random number between 0 and 1 for each row, in row order.i is a test row if its number is less than test_share. Otherwise it is a training row.Then work out each column's average from the training rows only, and centre every test row by subtracting those averages.
Task: write hold_out(X, seed, test_share) and return a dict:
"test_rows" — the row numbers of the test rows, smallest first,"train_means" — each column's average over the training rows, rounded to 3 decimal places,"test_centred" — the test rows, in row order, each minus the training averages, as a list of lists rounded to 3 decimal places.One tool you'll need: a mask with one True or False per row works inside square brackets on a grid too. grid[mask] keeps the whole rows whose position is True.
The tests always leave at least one row on each side.