ReLU has one bad habit: it outputs a flat 0 for every negative input, so the gradient there is 0 too. A neuron that drifts into the negative region stops receiving any signal and can stay dead forever.
ELU (exponential linear unit) keeps ReLU's shape for positives and replaces the flat part with a gentle curve that bottoms out instead of stopping dead:
Task: write elu(values, alpha) returning one output per input, each rounded to 4 decimal places.
x = 0 belongs to the second branch, and alpha * (e^0 - 1) is exactly 0 — so the two halves meet cleanly and the function has no jump.alpha controls how far down the negative side is allowed to go. As x heads to minus infinity, the output flattens out at -alpha and never passes it.The pay-off for the extra exp: the curve has a non-zero slope everywhere on the negative side, so gradients keep flowing, and because its outputs can be negative the activations stay roughly centred around zero — which helps the next layer learn.