A sequence-to-sequence model is two recurrent networks in a row. The encoder reads the whole input sequence and keeps only its final hidden state — the context vector. The decoder starts from that context and writes the output one value at a time, feeding each value it writes back in as its next input.
To keep the arithmetic visible, every token here is a single number instead of an embedding vector, so each sequence is just a list of numbers.
Both networks use the same plain RNN update, each with its own weights. With hidden size , a single input number and the previous hidden state , the new hidden state is
Each network arrives as a dict with "W_in" (a list of numbers), "W_h" (an list of rows — row holds the weights feeding unit ) and "b" (a list of numbers). The decoder has one extra entry, "v" (a list of numbers), which turns its hidden state into the value it writes:
The forward pass
Task: write seq2seq(inputs, encoder, decoder, steps) and return a tuple (context, outputs): the context vector, and the list of the steps values the decoder writes. Round every number you return to 4 decimal places, but carry the unrounded values forward through the computation.