DeepOffer

During Training and Serving, How Should Dropout Be Applied?

ML TheoryHot interview question
Source: Common pattern across 2026 interview reports

During training, each neuron is zeroed with probability p. The network cannot rely on co-adaptation of a few neurons - it is like training a different sub-network each step, approximating an ensemble.

After zeroing, the expected output of the layer shrinks: training output is multiplied by a Bernoulli(p) mask, so the expectation is p times the original.

Two inference options: Classic approach multiplies weights by p at inference; modern frameworks use inverted dropout - divide by p during training, use the network as-is at inference. Deployment is simpler.

Bonus point: mention Dropout + BatchNorm ordering and variance shift.

Common follow-up questions

Practice this question with an AI interviewer

Get asked follow-ups live, then receive a scored report — like a real MLE interview loop.

Start AI mock interview