A random forest is a crowd of decision trees, each trained on a slightly different slice of the data. Individually they're mediocre and they make different mistakes — which is the point. Pool their answers and the disagreements cancel out.
Task: write majority_vote(tree_predictions) returning the forest's final prediction for every example.
tree_predictions is a list of trees, and each tree is its list of predictions for the examples — so tree_predictions[2][5] is tree 2's answer for example 5. Every tree predicted every example, so all the inner lists are the same length.The awkward bit is that you're given the data tree-by-tree but need to decide example-by-example, so you have to turn the table on its side before you can count anything.
This is why a forest beats its own trees: a single tree needs to be right, while the forest only needs most of its trees to not be wrong in the same direction at once.