When you train a model to produce embeddings, you don't care about the exact numbers — you care that similar things point the same way and different things don't. Cosine embedding loss says exactly that, and says it differently depending on whether the pair is meant to match.
Start with cosine similarity, which ignores length entirely and measures only direction:
Then:
label = 1 (the pair should match): loss = 1 - cos(a, b). Zero when they point the same way, growing as they drift apart.label = -1 (the pair should not match): loss = max(0, cos(a, b) - margin). Once their similarity has dropped below margin, the loss is 0 and the pair is left alone.Task: write cosine_embedding_loss(a, b, label, margin) returning the loss, rounded to 4 decimal places.
a and b are vectors of the same length, and neither is all zeros.label is either 1 or -1.That max(0, ...) is the interesting half. Without it, the model would be rewarded for shoving unrelated pairs ever further apart forever, burning capacity on pairs that were already fine. The margin says "different enough is enough" and stops the pushing.