Before running a big search, you sweep one setting and read the curve. Here the setting is the neighbour count of a tumour classifier.
Task: write neighbour_curve(k_values).
load_breast_cancer(return_X_y=True, as_frame=True).train_test_split(X, y, test_size=0.25, random_state=0, stratify=y). Use the training rows only from here on.validation_curve on make_pipeline(StandardScaler(), KNeighborsClassifier()) with param_name set to the pipeline's neighbour-count setting, param_range=k_values and cv=5.k.Return a dict:
"best_k": the k with the highest mean validation score. If several tie, take the largest k, because more neighbours is the simpler model."best_score": that mean validation score, rounded to 4 decimal places."gap_at_first": for the first k in k_values, the mean training score minus the mean validation score, rounded to 4 decimal places.Compare unrounded values when you choose.