One-hot encoding exists because mapping colours to 0, 1, 2 invents an order that isn't there. But some categories are ordered — small < medium < large, cold < warm < hot — and for those, a single number is the right representation. Throwing that ordering away into separate one-hot columns would discard real information.
Task: write ordinal_encode(values, order) returning each value's position in order.
order lists the categories from lowest to highest. The first gets 0, the next 1, and so on.order encodes as -1.values may be empty.The ordering comes from order, which you supply — not from the data, and not from alphabetical sorting. That's the whole point: sorting ['large', 'medium', 'small'] alphabetically would encode them 0, 1, 2 in almost exactly the wrong order, and nothing downstream would ever complain.
The -1 sentinel is a deliberate choice worth being uneasy about. It keeps the pipeline running when an unexpected category arrives, which is what you want in production — but -1 is also a number, and a model will read it as "one step below the smallest category" rather than "unknown". That's fine for a tree, which can isolate it with a split, and misleading for a linear model, which will extrapolate the trend straight through it. Knowing which of those you're feeding is the difference between a safe default and a quiet bug.