A training run records its validation loss at the end of every epoch. Early stopping watches that number, saves a checkpoint whenever it improves, and halts once it has gone patience epochs in a row without improving. Then comes the step people forget: it restores the checkpoint from the best epoch.
val_losses is the validation loss the run would record at the end of each epoch if nothing stopped it. val_losses[0] belongs to epoch 1, val_losses[1] to epoch 2, and so on. Your function replays that history one epoch at a time, seeing it exactly as the early-stopping monitor would, with no knowledge of any epoch it has not reached yet.
The rules:
patience, training halts after that epoch.Task: write early_stopping(val_losses, patience), returning [best_epoch, stop_epoch, best_loss]:
best_epoch: the epoch whose checkpoint is restored,stop_epoch: the last epoch that was actually trained,best_loss: the validation loss of the restored checkpoint, rounded to 4 decimal places.Epochs are numbered from 1. patience is a whole number of at least 1, and val_losses holds at least one value.
Worked through, early_stopping([0.9, 0.8, 0.86, 0.83, 0.78, 0.8, 0.84, 0.81, 0.77], 3):
| epoch | loss | best so far | epochs without improvement |
|---|---|---|---|
| 1 | 0.9 | 0.9 (checkpoint) | 0 |
| 2 | 0.8 | 0.8 (checkpoint) | 0 |
| 3 | 0.86 | 0.8 | 1 |
| 4 | 0.83 | 0.8 | 2 |
| 5 | 0.78 | 0.78 (checkpoint) | 0 |
| 6 | 0.8 | 0.78 | 1 |
| 7 | 0.84 | 0.78 | 2 |
| 8 | 0.81 | 0.78 | 3, which is patience, so training halts |
Epoch 4 is lower than epoch 3, but it does not beat the best, so it still counts as an epoch without improvement. Epoch 9 is never trained. The answer is [5, 8, 0.78].
patienceis a trade. Set it too low and one noisy epoch ends a run that was still improving. Set it too high and you pay for epochs you will throw away. Either way, what you ship is the checkpoint. The epochs trained after it cost compute and change nothing about the model you keep.