Words like the, of and and appear in nearly every sentence, which means they tell you almost nothing about what a document is about. Classic text pipelines throw them out before doing anything else, so the remaining words carry more signal per slot.
Task: write remove_stopwords(text, stopwords) returning the surviving words as a list.
text on whitespace.stopwords. The stopword list is already lowercase, but the text may be capitalised anywhere — The, THE and the must all go.stopwords may be empty, in which case nothing is removed. Everything may be a stopword, in which case you return an empty list.Convert stopwords to a set first. Checking membership in a list walks it item by item, so with a realistic stopword list of a few hundred entries and a document of a few thousand words you'd do a million comparisons where a set does a few thousand.
Worth knowing when not to do this: stopwords carry real meaning for sentiment ("not good"), for phrase search ("to be or not to be"), and for anything a transformer reads — modern models want the full sentence, which is why this step has largely disappeared from deep-learning pipelines while remaining standard for keyword search and topic models.