Bernoulli naive Bayes handles yes/no features. Gaussian naive Bayes handles continuous ones — height, temperature, price — by assuming each feature follows a bell curve whose centre and width it learns separately for every class.
Training is just bookkeeping. For each class, from its training rows only:
n (the population form, not n - 1).Predicting scores every class for a row and keeps the highest. Working in logs, as always:
Task: write gnb_predict(X_train, y_train, X_test) returning the predicted label for each test row.
X_train and X_test are lists of rows of numbers; y_train holds the labels, which may be numbers or strings.That formula is just the log of the normal density. The first term depends only on the class's width — a class with wide spread pays a constant penalty — and the second is the squared distance from the class's centre, measured in units of its own variance. Naive Bayes is, in the end, asking which class's bell curve your row looks least surprising under.