A decision tree picks a question to ask by trying lots of them and keeping whichever one leaves the two resulting groups purest — closest to containing a single class.
Gini impurity scores one group. If the group's class fractions are :
All one class gives 0. An even two-way mix gives 0.5. That's the worst a two-class group can be.
A split produces two groups, so you need one number for the pair. Take the average of their impurities, weighted by how many examples each one holds:
Task: write split_gini(left_labels, right_labels) returning that weighted impurity, rounded to 4 decimal places.
0 and weight 0, so it contributes nothing — just don't divide by its size.n is the total across both sides, never zero.The weighting is the part people get wrong. Without it, a split that peels off one pure example and leaves a mess behind would score beautifully, and a tree built that way would grow one useless branch at a time.