DeepOffer

How Do L1 and L2 Penalties Shape a Model Differently?

ML TheoryConfirmed interview question
Reported in public interview compilations — Apple, Palantir, Meta, Amazon

L1 (sum |w|) produces sparse solutions: Many weights are pushed exactly to 0, so it does feature selection automatically. Geometrically, the diamond-shaped constraint tends to meet loss contours at sharp corners.

L2 (sum w^2) shrinks weights smoothly but rarely to 0. The circular constraint shrinks evenly and splits weight across correlated features.

Probability view: L1 corresponds to a Laplace prior, L2 to a Gaussian prior. In practice: pick L1 for interpretability or dimensionality reduction, L2 to prevent overfitting while keeping performance, Elastic Net for both.

Common follow-up questions

Practice this question with an AI interviewer

Get asked follow-ups live, then receive a scored report — like a real MLE interview loop.

Start AI mock interview