Clean gradient: With cross-entropy + sigmoid/softmax, the gradient with respect to logits is simply prediction minus label. Bigger error means bigger gradient, so the learning signal is strong.
MSE gradient vanishes: With MSE + sigmoid, the gradient contains the sigmoid derivative. When the prediction is badly wrong, sigmoid is in the saturated region, its derivative is near 0, and parameters barely move - the more wrong, the slower it learns.
Probability view: cross-entropy is equivalent to maximum likelihood estimation (minimizing KL divergence between predicted and true distributions). It matches probabilistic classification naturally. MSE assumes Gaussian noise and fits regression better.
Get asked follow-ups live, then receive a scored report — like a real MLE interview loop.
Start AI mock interview