DeepOffer

For Classification, Why Is Cross-Entropy Usually a Better Loss Than Squared Error?

ML TheoryConfirmed interview question
Source: March 2026 MLE interview report

Clean gradient: With cross-entropy + sigmoid/softmax, the gradient with respect to logits is simply prediction minus label. Bigger error means bigger gradient, so the learning signal is strong.

MSE gradient vanishes: With MSE + sigmoid, the gradient contains the sigmoid derivative. When the prediction is badly wrong, sigmoid is in the saturated region, its derivative is near 0, and parameters barely move - the more wrong, the slower it learns.

Probability view: cross-entropy is equivalent to maximum likelihood estimation (minimizing KL divergence between predicted and true distributions). It matches probabilistic classification naturally. MSE assumes Gaussian noise and fits regression better.

Common follow-up questions

Practice this question with an AI interviewer

Get asked follow-ups live, then receive a scored report — like a real MLE interview loop.

Start AI mock interview