A sensor's readings are assumed to come from a normal distribution with an unknown mean and variance . Maximum likelihood estimation (MLE) picks the and that make the observed readings most probable, using the data alone — no prior.
For a single reading , the normal density is
The readings are independent, so the likelihood of the whole dataset is the product of their densities. A product of hundreds of small numbers is too small for a float to hold, so instead we maximise the log-likelihood, the sum of the logs, which peaks at exactly the same place:
For the normal distribution the peak is known exactly:
Task: write normal_mle(data) that returns [mu_hat, var_hat, log_lik]: the two estimates, then the log-likelihood at the peak, . Use the natural logarithm and round each value to 4 decimal places. data always holds at least two different values.
Notice that is Chapter 4's population variance, applied to the sample. The same idea, maximising a log-likelihood, is what's going on underneath the loss functions you'll train models with later.