A team has one GPU for 5 days, which is minutes, to tune a model. One epoch of training takes 8 minutes, whatever the settings. Following the lesson's advice, they spend the budget in two stages:
| stage | what runs |
|---|---|
| 1: learning-rate sweep | 12 short runs, one per candidate learning rate, each exactly 6 epochs long, with no early stopping |
| 2: full runs | as many full runs as fit in the time left, each with early stopping at a patience of 6 and a cap of 60 epochs |
Early stopping works as in the lesson. An epoch is an improvement only if its validation loss beats every earlier epoch's, and training halts at the end of the 6th consecutive epoch without an improvement. From earlier runs of this model, the team knows that a full run's validation loss keeps improving up to epoch 30 and never improves after it.
Restoring the best checkpoint takes no time. How many complete stage-2 runs fit in the time left after stage 1?
Give a whole number of runs.