A bank's loan model is a single decision tree, and the team wants to know which features it actually leaned on. A standard way to answer is impurity-based importance: every split earns credit for how much it cleaned up the data, and each feature collects the credit of the splits that used it.
The tree arrives as a list of its splits. Each split is a tuple (feature, left, right), where left and right hold the class counts of the two groups it created. For example, ("Income", [30, 10], [12, 28]) sends 30 Approve and 10 Reject records left, and 12 and 28 right. The group that was split is the two children added together, here [42, 38].
For a group of records with class counts , the Gini impurity is
A split's credit is its drop in record-weighted impurity:
A feature's importance is the total drop of all its splits divided by the total drop of every split in the tree, so the importances add up to 1. This sharpens the rough rule that a feature winning more splits matters more: here each split counts by how much it cleaned up, not just by the fact that it happened.
Task: write feature_importance(features, splits) and return a dict that maps every name in features, in the same order, to its importance rounded to 4 decimal places.
0.0.