A loss is one number, but it is assembled from every example in the batch, and the examples do not chip in equally. The examples holding most of the loss are the ones training works hardest to fix, because shrinking their errors is the quickest way to shrink the total.
Here is a week of a model's temperature forecasts, in °C:
| Mon | Tue | Wed | Thu | Fri | Sat | Sun | |
|---|---|---|---|---|---|---|---|
| forecast | 20 | 20 | 21 | 22 | 20 | 24 | 20 |
| recorded | 18 | 21 | 19 | 23 | 31 | 22 | 17 |
How many of those seven days hold most of the loss may depend on which loss is doing the counting.
Task: write loss_carriers(predictions, targets, fraction), returning [mse_count, mae_count]: the smallest number of examples that together hold at least fraction of the loss, under MSE and under MAE respectively.
predictions and targets are equal-length lists with one entry per example. An example's error is its prediction minus its target.fraction. Landing exactly on fraction counts.fraction is strictly between and . At least one prediction misses its target, so the loss is never zero.The forecast week above, with fraction = 0.75, is the first test.
Both counts describe the same predictions, so any gap between them comes from the ruler alone. A wide gap means the choice of loss decides how many examples get to steer the model, which is fine when those few examples are real and a problem when they are broken readings.