A refurbishing shop prices second-hand laptops with a regression tree. Its first question is about a laptop's age in years: is the age at most ? Each of the two leaves then predicts the average price of the training laptops that land in it.
To choose , the tree can't use Gini or entropy: those measure how mixed a group of categories is. For numbers like price, it measures the mess with variance, which is how far the prices in a group sit from that group's own average:
Here is the number of laptops in group and is their average price. A split of laptops into a left group and a right group is scored by the weighted variance
and the best split is the one with the lowest score. Its variance reduction is the variance of all prices minus that score.
Candidate thresholds: sort the distinct ages and try the midpoint between each neighbouring pair. Laptops with age go left, and the rest go right.
Task: write best_regression_split(ages, prices) and return a tuple (threshold, left_mean, right_mean, reduction) for the best split, every value rounded to 4 decimal places.
ages and prices are the same length, and ages is not sorted.None.This is the question a decision tree asks whenever it predicts a number instead of a category. Each leaf predicts its group's average, so the split that keeps each group's prices tight around its own average is the split whose two averages make the smallest errors.