A churn team sweeps a decision tree's max_depth and runs the same 5-fold cross-validation for every candidate. Depth 4 posts the highest mean accuracy. Depth 3 is a whisker behind, and its fold scores wobble less. Is depth 4's lead real, or just the luck of the folds?
The team settles questions like this with one fixed rule instead of a judgement call:
where is the number of folds and is the sample standard deviation of that setting's fold scores (divide by ).
Task: write pick_setting(candidates).
candidates is a list of (setting, fold_scores) pairs, already ordered from simplest to most complex. Trust this order, not the setting values: for KNN, for example, a larger number of neighbours is the simpler model, and a setting may not even be a number.(chosen_setting, threshold), with the threshold rounded to 4 decimal places.A lead smaller than the fold-to-fold noise is weak evidence that the extra complexity helps, while the simpler model is cheaper to train, faster to serve, and less likely to be overfitting.