A clinic uses a model to decide which patients are referred to a specialist. On the held-out test set it scores a respectable accuracy, and that is the only number anyone has looked at.
That number is an average over every patient, weighted by how many of each kind the test set happens to contain. A group that makes up a small share of patients can be served badly without moving it much. The fix costs nothing: break the results down per group. Once they are broken down there is more than one thing to compare, because fairness is not one thing:
Task: write group_report(groups, labels, predictions), returning [overall_accuracy, rows].
groups[i] is a string naming patient 's group. labels[i] is 1 if that patient in fact has the condition and 0 if not. predictions[i] is 1 if the model referred them and 0 if it did not.overall_accuracy is the share of all patients whose prediction equals their label.rows has one entry per group, sorted alphabetically by group name, each of the form [name, accuracy, referral_rate, false_positive_rate, false_negative_rate]:
accuracy: the share of the group whose prediction equals their label.referral_rate: the share of the group the model referred.false_positive_rate: of the group's patients labelled 0, the share the model referred.false_negative_rate: of the group's patients labelled 1, the share the model did not refer.0, its false positive rate does not exist, so report None. The same goes for the false negative rate of a group with nobody labelled 1.Worked through, group_report(['b', 'a', 'b', 'a', 'a'], [1, 0, 0, 1, 1], [1, 1, 0, 0, 1]):
a is patients 1, 3 and 4. Only patient 4 is classified correctly, so its accuracy is . Patients 1 and 4 were referred, so its referral rate is . Its one patient labelled 0 (patient 1) was referred, so its false positive rate is . Of its two patients labelled 1, patient 3 was not referred, so its false negative rate is .b is patients 0 and 2. Both are classified correctly and one of the two was referred: accuracy , referral rate , and both error rates .[0.6, [['a', 0.3333, 0.6667, 1.0, 0.5], ['b', 1.0, 0.5, 0.0, 0.0]]].No single column of this report is the fairness number. When the condition is more common in one group than another, equal error types and equal referral rates pull in different directions, and except in unusual cases no model delivers both. Which one matters is a judgement about this particular decision. The report is what puts that judgement in front of a person instead of hiding it inside one average.