SGD: Takes a fixed-size step along the current mini-batch gradient. Noisy, but often generalizes well.
Momentum: Keeps a velocity variable as an exponential average of gradients. Steps speed up in consistent directions and cancel out in oscillating directions - like adding inertia to SGD.
Adam: Combines first moment (momentum), second moment (per-dimension adaptive learning rate, RMSProp idea), and bias correction. Less sensitive to learning rate, converges fast, and is the default choice.
Adam's weaknesses: Sometimes generalizes worse than well-tuned SGD+Momentum (common in CV tasks). Second-moment estimates can become unstable late in training. AdamW decouples weight decay from the gradient and is now the standard fix for large-model training.
Get asked follow-ups live, then receive a scored report — like a real MLE interview loop.
Start AI mock interview