A spam filter sees a document as a row of yes/no features: does the word "invoice" appear, does "free" appear, and so on. Bernoulli naive Bayes scores one class for that document by multiplying together the probability of every feature turning out the way it did — plus the class's own prior.
Multiplying dozens of small probabilities underflows to zero fast, so everyone works in logs instead, where multiplying becomes adding:
As with log loss, only one half of each bracket survives: a feature that is present contributes log(p_i), a feature that is absent contributes log(1 - p_i).
Task: write bernoulli_log_likelihood(x, feature_probs, class_prior) returning that total, rounded to 4 decimal places.
x is a list of 0s and 1s. feature_probs[i] is the probability that feature i is 1 given this class.class_prior is the probability of the class before seeing any features.log is the natural logarithm.The "naive" part is the bare sum. It assumes the features are independent given the class, which for words in a sentence is flatly untrue — and the filter works superbly anyway, because for ranking classes the errors tend to cancel.