A model needs a fixed-length row of numbers, and documents are different lengths. Bag-of-words solves that by fixing a vocabulary up front and describing every document as "how many times does each vocabulary word appear in it".
The name is the honest part: word order is thrown away completely. "dog bites man" and "man bites dog" produce identical vectors.
Task: write bag_of_words(documents, vocabulary) returning one count row per document.
i holds the count of vocabulary[i] in that document. The vocabulary is already lowercase.0.len(vocabulary).That fixed width is the whole point, and it's also the catch. Training and serving must use the identical vocabulary in the identical order, or column 7 means one word to the model and a different word to the data — a mismatch that produces no error at all, just quietly wrong predictions.