DeepOffer

Compare Plain SGD, Momentum, and Adam—Where Can Adam Fall Short?

ML TheoryConfirmed interview question
Reported in public interview compilations — NVIDIA

SGD: Takes a fixed-size step along the current mini-batch gradient. Noisy, but often generalizes well.

Momentum: Keeps a velocity variable as an exponential average of gradients. Steps speed up in consistent directions and cancel out in oscillating directions - like adding inertia to SGD.

Adam: Combines first moment (momentum), second moment (per-dimension adaptive learning rate, RMSProp idea), and bias correction. Less sensitive to learning rate, converges fast, and is the default choice.

Adam's weaknesses: Sometimes generalizes worse than well-tuned SGD+Momentum (common in CV tasks). Second-moment estimates can become unstable late in training. AdamW decouples weight decay from the gradient and is now the standard fix for large-model training.

Common follow-up questions

Practice this question with an AI interviewer

Get asked follow-ups live, then receive a scored report — like a real MLE interview loop.

Start AI mock interview