A bank's fraud model misses most of the fraud. The lesson gave two fixes. You can weight the rare class with class_weight="balanced" and retrain. Or you can keep the plain model and simply lower its threshold. Here you compare them.
Task: write weights_or_threshold(seed, thresholds).
make_classification(n_samples=2000, n_features=10, n_informative=4, weights=[0.95, 0.05], class_sep=1.8, flip_y=0.01, random_state=seed). Class 1 is fraud.train_test_split(X, y, test_size=0.25, random_state=0, stratify=y).make_pipeline(StandardScaler(), LogisticRegression(max_iter=1000, class_weight="balanced")), on the training rows. Work out its recall and precision on the test rows from predict.class_weight, on the training rows. Take its chance of class 1 for the test rows.thresholds. Find the largest threshold at which the plain model's recall is at least the weighted model's recall. A row is flagged when its chance is at or above the threshold.Return a dict:
"weighted": [recall, precision] of the weighted model,"moved": [threshold, recall, precision] of the plain model at the threshold you found.Round every recall and precision to 3 decimal places. Precision is 0.0 if nothing is flagged. In every test at least one threshold is low enough.