Models do arithmetic, so a column of words has to become numbers. The lazy fix — red=0, green=1, blue=2 — accidentally claims that blue is bigger than red and that green sits exactly between them. None of that is true.
One-hot encoding avoids the lie. Give every category its own column, and switch on exactly one of them per row.
Task: write one_hot(labels, categories) returning one row per label — a list of 0s with a single 1 in the position that label occupies in categories.
categories, not from the order the labels happen to appear in.categories lights nothing up: its row is all zeros. That's the behaviour you want in production, where a category nobody saw during training turns up at serving time and the pipeline must not crash.labels may be empty, in which case you return an empty list.0 and 1 integers.Every row ends up with the same width, and that width is fixed by categories alone — which is exactly why the training and serving pipelines have to share the same category list.