A team trains two seq2seq translators. They are identical — same training data, same vocabulary, same encoder and decoder depth, an LSTM on both sides, same training budget — except for the width of the single context vector passed from encoder to decoder: 256 numbers in one, 1024 in the other.
Both are scored on the same held-out test set, split into buckets by input sentence length. Higher is better; would be a perfect match to the reference translation.
| input length | 256-wide context | 1024-wide context |
|---|---|---|
| 1–10 words | 31.2 | 31.5 |
| 11–20 words | 28.4 | 29.9 |
| 21–30 words | 23.1 | 26.8 |
| 31–40 words | 17.6 | 22.4 |
Which conclusion do these results best support?
Select all that apply.