Unpaired translation hands you two piles of photographs — horses in one, zebras in the other — and nothing linking a picture in one pile to a picture in the other. The adversarial loss on its own would be perfectly satisfied by a translator that ignores its input and emits the same convincing zebra every time. What rules that out is the round trip: turn a horse into a zebra, turn that back, and require what comes out to be the horse you started with.
Stripped down to something you can compute: every photograph is a short vector, and each translator is a matrix. to_zebra takes a horse vector to a zebra vector, to_horse takes one back, and a matrix is applied the usual way:
Task: write cycle_loss(horses, zebras, to_zebra, to_horse), returning two numbers, each rounded to 4 decimal places: the cycle loss over the horse pile, then the cycle loss over the zebra pile.
Notice what this loss never asks: whether the output is a convincing zebra. That question belongs to the discriminator. The round trip is here for the other half of the job — making sure the thing that comes out still has anything to do with the thing that went in.