L1 (sum |w|) produces sparse solutions: Many weights are pushed exactly to 0, so it does feature selection automatically. Geometrically, the diamond-shaped constraint tends to meet loss contours at sharp corners.
L2 (sum w^2) shrinks weights smoothly but rarely to 0. The circular constraint shrinks evenly and splits weight across correlated features.
Probability view: L1 corresponds to a Laplace prior, L2 to a Gaussian prior. In practice: pick L1 for interpretability or dimensionality reduction, L2 to prevent overfitting while keeping performance, Elastic Net for both.
Get asked follow-ups live, then receive a scored report — like a real MLE interview loop.
Start AI mock interview