A lender's bagged forest decides "approve" or "reject" for loan applications. As in the Bagging lesson, every tree trained on its own bootstrap sample: rows drawn with replacement from the training set, so some rows appear several times in a tree's sample and others not at all.
A row that a tree never trained on is out-of-bag for that tree. To that tree the row is brand new, just like a test row, so its opinion on that row is honest. The forest uses this to grade itself without a separate test set.
For every training row:
The out-of-bag score is the fraction of the non-skipped rows whose out-of-bag prediction equals the row's true label.
Task: write oob_score(labels, samples, tree_preds) and return the tuple (score, scored_rows).
labels[i] is the true label of training row i. Labels can be numbers or strings, and there can be more than two classes.samples[t] is the list of row indices tree t trained on. An index can repeat.tree_preds[t][i] is tree t's prediction for row i. Every tree has been run on every row, including the rows it trained on.score is the out-of-bag score rounded to 4 decimal places, and scored_rows is the number of rows that were not skipped.(None, 0).A single draw misses a given row with probability , so a whole sample of draws leaves it out with probability , which is about once is more than a handful. Roughly a third of the training rows sit out of every tree, so in a forest of a few hundred trees nearly every row collects dozens of honest votes.