A plant-care app sorts leaf photos into classes such as "healthy", "rust" and "blight". The F1 Score lesson scored a yes/no model. With more than two classes, you score each class on its own by treating it as the Positive class and every other class as Negative. For a class :
Each class then gets its own Precision, Recall and F1 from the formulas you already know. A class's support is the number of rows whose actual label is that class. There are three common ways to turn the per-class scores into one number:
Task: write f1_summary(actual, predicted) and return the tuple (macro_f1, weighted_f1, balanced_accuracy), each rounded to 4 decimal places.
actual or in predicted. Labels can be strings or numbers.actual only, since a class with no actual rows has nothing to recall.When one class dominates a dataset, these three numbers can tell very different stories about the same predictions. Which one to report depends on whether a rare class matters as much to you as a common one.