Gradient view: Output = F(x) + x. In backpropagation, gradients have an identity shortcut that does not need to be multiplied layer by layer, so deep gradients no longer decay exponentially.
Optimization view: The network only needs to learn the correction F(x), with identity as the default. A deeper network can be at least as good as a shallower one - the degradation problem is solved.
Representation view: Unrolled, a residual network is an ensemble of paths of different depths. Robustness comes from path redundancy.
Practical details: residuals need the right normalization position (pre-norm/post-norm) and initialization; with bad scaling, very deep networks still fail to train.
Get asked follow-ups live, then receive a scored report — like a real MLE interview loop.
Start AI mock interview