Your manager wants a fraud model, and you have decided that recall is what matters: a missed fraud costs far more than a false alarm. But a model that flags everything has perfect recall and is useless. So you also set an accuracy floor before looking at any results.
Task: write fraud_shortlist(rare_share, seed, min_accuracy).
make_classification(n_samples=2000, n_features=10, n_informative=4, weights=[1 - rare_share, rare_share], class_sep=1.8, flip_y=0.01, random_state=seed). Class 1 is fraud."lazy": DummyClassifier(strategy="most_frequent")"plain": make_pipeline(StandardScaler(), LogisticRegression(max_iter=1000))"balanced": the same pipeline with LogisticRegression(max_iter=1000, class_weight="balanced")cross_val_score on all the rows, with cv=5, once with scoring="accuracy" and once with scoring="recall". Take the mean of each.min_accuracy. If several tie on recall, take the one that comes first in the order lazy, plain, balanced.Return a dict with one entry per model, "lazy", "plain" and "balanced", each holding [mean accuracy, mean recall] rounded to 4 decimal places. Add "pick", the name of the model you picked. Compare unrounded values. In every test at least one model clears the floor.