Collecting more data is expensive. A learning curve tells you if it would help: train on growing slices of the data and watch the gap between the training score and the validation score.
Task: write learning_gaps(model_name, fractions).
load_breast_cancer(return_X_y=True, as_frame=True) and use all the rows.model_name:
"logistic": make_pipeline(StandardScaler(), LogisticRegression(max_iter=1000))"one_neighbour": make_pipeline(StandardScaler(), KNeighborsClassifier(n_neighbors=1))learning_curve on it with train_sizes=fractions and cv=5.Return a dict:
"sizes": the training-set sizes learning_curve used, as whole numbers,"gaps": for each size, the mean training score minus the mean validation score, rounded to 4 decimal places.