Split a dataset at random and you can get unlucky: a rare class ends up entirely in the training set, so the test set can never tell you whether the model handles it. A stratified split prevents that by splitting each class separately, so every class keeps the same proportion on both sides.
So there's no randomness to argue about, use this exact rule:
c examples, the first ceil(c × test_ratio) of its positions go to the test set. The rest go to training.Task: write stratified_split(labels, test_ratio) returning [train_indices, test_indices] — two lists of positions, each sorted ascending.
labels is a list of class labels (strings or numbers), never empty.test_ratio is between 0 and 1.ceil, rounding up. A class with 3 examples and a ratio of 0.4 sends 2 to test, not 1 — rounding up is what guarantees a one-example class still shows up in the test set at all.Notice what this buys you: the test set's class mix matches the full dataset's, so test accuracy means the same thing it would in production.