A team training a street-scene detector wants horizontal flips in its augmentation. A flip moves every object to the other side of the photo, so every bounding box has to move with it. Otherwise the labels point at empty road, which breaks the one rule of augmentation: never change the answer.
The scenes also contain painted road arrows. A mirrored left-turn arrow really is a right-turn arrow, so the team keeps the flip and renames those boxes too.
The image is a grayscale grid, given as a list of rows. Its width is the number of columns.
Boxes are tuples (name, x_min, y_min, x_max, y_max). The coordinates mark pixel edges, not pixel centres. runs from at the image's left edge to at its right edge, and column is the strip from to . So a box with x_min = 1 and x_max = 3 covers columns 1 and 2. works the same way, measured down from the top.
Task: write flip_with_boxes(image, boxes, mirror_names) that returns a tuple (flipped_image, flipped_boxes):
flipped_image is the image mirrored left to right.flipped_boxes is a list with one tuple per input box, in the same order and the same layout, giving where that object is after the flip. Every box must still have x_min < x_max.mirror_names is a dict such as {'arrow_left': 'arrow_right', 'arrow_right': 'arrow_left'}. A box whose name is a key takes the mapped name. Every other name stays as it is.