A plain RNN reads a -token sentence. The prediction and the loss are both at step .
Because the same weights are reused at every step, going back one step through time multiplies the gradient signal by the same factor every time. For this trained model that factor has been pinned down indirectly: the signal arriving at the hidden state steps before the loss is of the signal at the loss itself.
The dependency the model needs to learn runs from the subject at position to the verb it must choose at position .
Taking the signal at the loss as , roughly how much gradient reaches position ?
Select all that apply.