Before you trust any model, you need to know what doing nothing scores. The majority baseline is the laziest possible classifier: find the most common class in the training data, then predict that same class for everything, forever. It never looks at a single feature.
Task: write majority_baseline(y_train, y_test) returning [predicted_class, accuracy] — the class it would always predict, and the accuracy that gets on the test labels, rounded to 4 decimal places.
y_train. If two labels tie, return the smaller one (labels are either all numbers or all strings, so they're always comparable).y_test: the fraction of test labels equal to that one predicted class.This number is the bar every real model has to clear. On a dataset where 99% of transactions are legitimate, the baseline scores 0.99 — so a fraud model reporting "99% accurate" has, in fact, demonstrated nothing at all.