A courier firm wants to flag parcels that are likely to arrive damaged. Last year it shipped 240,000 parcels, and 3,000 of them arrived damaged.
With barely one damaged parcel in eighty, this is the situation from When One Class Is Rare, so the team undersampled before training. It kept all 3,000 damaged parcels and a random 9,000 of the 237,000 undamaged ones, then trained a logistic regression on those 12,000 rows. Because the undamaged parcels were picked at random, the ones that were kept look just like the ones that were thrown away. There are simply fewer of them.
Assume the model learned its training table well. When it outputs a probability for a parcel, that number is the share of damaged parcels among the training rows that look like this parcel.
The model now runs on the real stream of parcels, with nothing thinned out and the imbalance intact. A new parcel arrives, and predict_proba gives it 0.40 for "damaged".
Among real parcels that look like this one, what percentage actually arrive damaged?
Give your answer as a percentage rounded to 2 decimal places (write 12.34 for 12.34%). Keep at least four significant figures in your working.