DeepOffer

How Would You Code Stable Softmax, Cross-Entropy, and Their Gradient in NumPy?

ML CodingHot interview question
Reported in public interview compilations — Google

Forward: Subtract the row max for numerical stability, exp, then normalize by row to get probabilities P. Loss = -mean(log P[range(n), y] + eps), where eps prevents log(0).

Backward: The gradient with respect to logits has a clean form - dL/dZ = (P - onehot(y)) / n. No need to unroll the chain rule step by step; hand-derivation and code should agree.

Self-check habit: After writing it, verify a few dimensions with numerical gradients (small finite differences). Doing a gradient check unprompted in an ML coding round is a strong plus.

Shape discipline: be explicit about where the batch dimension is (keepdims). Interviewers plant traps exactly here.

Common follow-up questions

Practice this question with an AI interviewer

Get asked follow-ups live, then receive a scored report — like a real MLE interview loop.

Start AI mock interview