Choosing k by checking scores on the test rows lets the test set quietly teach the model. Instead, carve a validation slice out of the training rows, choose k there, and only then look at the test rows once.
Write honest_k(ks) for the breast cancer data.
load_breast_cancer(return_X_y=True, as_frame=True) and split it with train_test_split(X, y, test_size=0.25, random_state=0, stratify=y).train_test_split(X_train, y_train, test_size=0.25, random_state=0, stratify=y_train). Call the parts fit rows and validation rows.k in ks, fit make_pipeline(StandardScaler(), KNeighborsClassifier(n_neighbors=k)) on the fit rows and score it on the validation rows.k with the highest validation score, comparing scores rounded to 3 places. On a tie, choose the smallest k.k on the whole training set, and score it on the test rows.Return a dict:
"k": the chosen k, as an int."validation": its validation score, rounded to 3 places."test": the refitted model's test score, rounded to 3 places.