Cause: Backpropagation in deep networks is a chain of multiplications. If Jacobians have eigenvalues mostly below 1, gradients shrink exponentially (vanish); if above 1, they grow exponentially (explode). Sigmoid/tanh derivatives at most 0.25 make it worse.
Fixes: Good initialization (Xavier for tanh, He for ReLU), normalization (BatchNorm/LayerNorm keeps distributions stable), residual connections (an identity path for gradients), ReLU-family activations, and gradient clipping (for exploding gradients).
How to present it: start with the chain-multiplication cause, then list fixes by category - initialization / normalization / architecture / activation / clipping - with one example each.
Get asked follow-ups live, then receive a scored report — like a real MLE interview loop.
Start AI mock interview