During training, each neuron is zeroed with probability p. The network cannot rely on co-adaptation of a few neurons - it is like training a different sub-network each step, approximating an ensemble.
After zeroing, the expected output of the layer shrinks: training output is multiplied by a Bernoulli(p) mask, so the expectation is p times the original.
Two inference options: Classic approach multiplies weights by p at inference; modern frameworks use inverted dropout - divide by p during training, use the network as-is at inference. Deployment is simpler.
Bonus point: mention Dropout + BatchNorm ordering and variance shift.
Get asked follow-ups live, then receive a scored report — like a real MLE interview loop.
Start AI mock interview