Freshly initialised weights produce wild gradients, so starting at your full learning rate can wreck a run in the first few steps. And a rate large enough to make early progress is too large to settle down later. The standard answer does both things in sequence: warm up, then decay.
Task: write warmup_linear(peak_lr, warmup_steps, total_steps) returning the learning rate for every step from 0 to total_steps - 1, each rounded to 4 decimal places.
Warmup — for steps t < warmup_steps, climb linearly to the peak:
The t + 1 means step 0 already starts above zero, and the last warmup step lands exactly on peak.
Decay — from warmup_steps onward, fall linearly away from the peak:
The first decay step is still exactly peak, and the rate shrinks from there.
total_steps.warmup_steps can be 0, meaning no warmup at all — every step decays. Make sure you don't divide by zero getting there.warmup_steps can equal total_steps, meaning the run is pure warmup and the decay branch never executes.With warmup_steps = 2 and peak = 0.1 you should see the rate rise over two steps, sit at the peak, then slide down — the shape every transformer training run is built on.