A full grid of tree settings would take too long. You decide how many tries you can afford and let a random search spend them.
Task: write random_tree_search(n_iter, seed).
load_breast_cancer(return_X_y=True, as_frame=True).train_test_split(X, y, test_size=0.25, random_state=0, stratify=y).RandomizedSearchCV(DecisionTreeClassifier(random_state=0), space, n_iter=n_iter, cv=5, random_state=seed) on the training rows, where space is| setting | values |
|---|---|
max_depth | 1, 2, 3, 4, 5, 6, 8, 10 |
min_samples_leaf | 1, 2, 4, 8, 16, 32 |
Return a dict:
"tried": how many combinations the search tried, as a whole number,"best": its best_params_, exactly as scikit-learn gives it,"cv": its best cross-validated score, rounded to 3 decimal places.