Formula: softmax(x_i) = exp(x_i) / Sum exp(x_j). If logits are large, exp overflows to inf.
Stable version: Subtract the maximum first: m = max(x), softmax(x_i) = exp(x_i - m) / Sum exp(x_j - m).
Why it is equivalent: numerator and denominator are both multiplied by exp(-m), so the math does not change. After subtraction the largest exponent is 0, so exp results are at most 1 and never overflow.
Log-softmax computes x_i - m - log Sum exp(x_j - m) directly. It avoids the precision loss of exp then log, and is often fused with cross-entropy in training.
Get asked follow-ups live, then receive a scored report — like a real MLE interview loop.
Start AI mock interview