A letting agent wants a monthly rent estimate from a single number: a flat's floor area in m². One small tree on its own is too jumpy, so they build a bagged ensemble of them.
Every tree asks the same question, is the area at most threshold?, but each one learns its two answers from its own bootstrap sample: a list of training rows drawn with replacement. The samples have been drawn in advance and are handed to you, so the results are reproducible.
Task: write bagged_predict(xs, ys, samples, threshold, queries) and return the ensemble's rent estimate for every area in queries, each rounded to 4 decimal places.
xs[i] and ys[i] are the floor area and the monthly rent of training row i.samples holds one list of row indices per tree. An index can appear more than once, and every appearance counts as a separate training row for that tree: the sample [3, 3, 7] trains its tree on row 3 twice and row 7 once.<= threshold) predicts the average rent of the tree's training rows that answer yes. Its no leaf predicts the average rent of the rows that answer no.Some ensembles draw their samples without replacement instead, so no index repeats inside a list. This variant is called pasting, and your function must handle those lists with exactly the same rules.
Every tree here asks an identical question, yet the trees still disagree, because each one saw a different mix of rows. That disagreement is what the average smooths out.