Every column the clinic measures costs time and money. You want the smallest number of columns whose model is nearly as good as using all thirty, judged honestly by cross-validation on the training rows.
Task: write smallest_good_k(k_values, tolerance).
load_breast_cancer(return_X_y=True, as_frame=True).train_test_split(X, y, test_size=0.25, random_state=0, stratify=y). Use the training rows only.make_pipeline(StandardScaler(), LogisticRegression(max_iter=1000)), with cross_val_score and cv=5. Take the mean.k in k_values, score make_pipeline(StandardScaler(), SelectKBest(f_classif, k=k), LogisticRegression(max_iter=1000)) the same way.k whose mean score is at least the all-columns mean minus tolerance.Return a dict:
"all": the all-columns mean score, rounded to 3 decimal places,"k": the k you chose,"score": its mean score, rounded to 3 decimal places.k_values may come in any order. In every test at least one k qualifies. Compare unrounded values.