Pruning deletes weights that are close to zero. But a matrix with zeros scattered at random is usually no faster to run: the hardware still has to step through every position, because it cannot know in advance where the zeros are.
2:4 sparsity fixes the pattern, not just the amount. Split each row of a weight matrix into consecutive groups of four. In every group, keep exactly two weights, the two with the largest magnitude (absolute value), and zero the other two.
Half the weights go, and they go in a layout the hardware can plan around. Each group is stored as just its two kept values plus two slot numbers (0 to 3) saying where in the group they came from. Every group then costs exactly two multiplications instead of four.
For example, the group [0.2, -0.7, 0.5, 0.1] keeps -0.7 and 0.5, so it is stored as values [-0.7, 0.5] and slots [1, 2].
Task: write two_four_compress(weights). weights is a list of rows, and every row's length is a multiple of 4.
Return a tuple (values, slots), each a list with one entry per row:
values[r]: the kept weights of row r, two per group, groups in order, and within a group in the order they appear in the row.slots[r]: the matching slot numbers, each from 0 to 3, counted from the start of its own group.If two weights in a group have the same magnitude and only one of them can be kept, keep the one that comes first. Return the kept values exactly as they appear, sign included; no rounding is needed.