DeepOffer

What Causes Gradient Signals to Vanish or Blow Up, and How Can You Stabilize Them?

ML TheoryConfirmed interview question
Reported in public interview compilations — Robinhood, Asana, Onfido, Groupon, Square, Plaid

Cause: Backpropagation in deep networks is a chain of multiplications. If Jacobians have eigenvalues mostly below 1, gradients shrink exponentially (vanish); if above 1, they grow exponentially (explode). Sigmoid/tanh derivatives at most 0.25 make it worse.

Fixes: Good initialization (Xavier for tanh, He for ReLU), normalization (BatchNorm/LayerNorm keeps distributions stable), residual connections (an identity path for gradients), ReLU-family activations, and gradient clipping (for exploding gradients).

How to present it: start with the chain-multiplication cause, then list fixes by category - initialization / normalization / architecture / activation / clipping - with one example each.

Common follow-up questions

Practice this question with an AI interviewer

Get asked follow-ups live, then receive a scored report — like a real MLE interview loop.

Start AI mock interview