A confusion matrix lays out exactly which classes a model mixes up. Rows are what the answer really was, columns are what the model said, so the cell at row i, column j counts the examples that truly belonged to class i and were predicted as class j. The diagonal is everything it got right.
Raw counts are hard to read when classes are lopsided, so you often want them as fractions instead.
Task: write confusion_matrix(y_true, y_pred, normalize) returning the matrix, where normalize is one of:
normalize | Each cell becomes | Reads as |
|---|---|---|
None | the raw count, as an integer | how many |
"true" | count ÷ its row total | of this real class, what fraction went where — i.e. recall on the diagonal |
"pred" | count ÷ its column total | of everything called this, what fraction was right — i.e. precision on the diagonal |
"all" | count ÷ the grand total | share of the whole dataset |
y_true or y_pred. A class that only ever gets predicted, and never actually occurs, still earns a row and a column.normalize=None, leave the counts as plain integers — no rounding, no floats.0, put 0.0 in those cells instead of dividing by zero.The two middle options are worth dwelling on: the same matrix normalized by row and by column tells you two different stories, and arguments about model quality are usually arguments about which one you're looking at.