Exponential decay shrinks the learning rate by multiplying it by a fixed factor slightly below 1, over and over. After the multiplication has been applied times, the rate is
A teammate's training script contains this loop:
train_set holds 1,620 examples.batches(...) walks through them 64 at a time and keeps the smaller final batch made of whatever is left over. Every batch, full or not, gets one call to update_weights.The teammate's notes describe the schedule as "exponential decay from 0.1, , 2% smaller each epoch."
According to the code as written, what learning rate does the very first update of epoch 4 use?
Give your answer rounded to 5 decimal places.