A plain RNN is trained on 20-token sentences, with the loss appearing at the final step only.
The recurrent weights are shared across every step, so one backward pass hands that single set of weights a contribution from each of the 20 steps, and the gradient it actually receives is the sum of all of them.
For this model, every step travelled back through time shrinks the signal by a factor of . Taking the contribution of the final step as , the contribution from the step sitting steps before the loss therefore has magnitude
All 20 contributions push the same way, so the total gradient magnitude is simply their sum.
What percentage of that total comes from the three most recent steps, ?
Round to the nearest whole percent.