Binning turns a continuous feature into a handful of ordered buckets: ages become age brackets, incomes become income bands. You throw away precision deliberately, in exchange for a feature a tree can split on cleanly and a human can read.
Task: write bin_values(values, edges) returning each value's bin index.
edges is a sorted list of cut points, and n edges make n + 1 bins:
| Bin | Range |
|---|---|
| 0 | up to and including edges[0] |
| 1 | above edges[0], up to and including edges[1] |
| … | … |
n | above edges[-1] |
So every bin is open on the left and closed on the right, and a value landing exactly on an edge belongs to the lower bin.
edges may be empty, in which case there's one bin and everything is 0.The short way to compute it: a value's bin index is simply how many edges it is strictly greater than. That single count handles the inclusive side, both unbounded ends, and the empty-edges case, with no branching at all.
Which side the edges close on is arbitrary but must be decided and written down — it's the difference between a 30-year-old landing in the 18–30 bracket or the 30–50 one, and "whatever the code does" is how two pipelines end up disagreeing about the same person.