Newton's method swaps the learning rate for , so it cares about one property of a loss more than any other: its curvature, the second derivative. A team wants to compare how sharply three common losses bend when a prediction misses by a little and when it misses by a lot.
Two regression losses, written as functions of the error on a single example:
One classification loss. A yes/no classifier looks at an example whose true answer is yes. It emits a raw score , turns it into a probability with the sigmoid , and pays the cross-entropy loss . The model's parameters set directly, so treat this loss as a function of :
The two situations to compare:
| regression error | classifier score | |
|---|---|---|
| small miss | (unsure, ) | |
| big miss | (confidently wrong, ) |
Take each loss's second derivative (with respect to for the two regression losses, and with respect to for cross-entropy) at the small miss and at the big miss.
Which statement is correct? Values are rounded to two significant figures.
Select all that apply.