DeepOffer

Connect Cross-Entropy, KL Divergence, and Maximum Likelihood

ML TheoryHot interview question
Source: Common pattern across 2026 interview reports

One line connects all three: Maximum likelihood maximizes the product of P(y_i|x_i). Taking the negative log gives minus sum log P(y_i|x_i) - which is exactly cross-entropy with one-hot labels.

KL divergence: D_KL(P||Q) = H(P,Q) - H(P). Cross-entropy = entropy of the true distribution + KL. With fixed labels, H(P) is constant, so minimizing cross-entropy is equivalent to minimizing KL and to maximizing likelihood.

Why it matters: it upgrades why cross-entropy from a memorized conclusion to first principles, and explains why knowledge distillation uses KL on soft labels.

Add directionality: KL is asymmetric. D_KL(P||Q) and D_KL(Q||P) penalize differently (mode-covering vs. mode-seeking) - generation interviews often dig here.

Common follow-up questions

Practice this question with an AI interviewer

Get asked follow-ups live, then receive a scored report — like a real MLE interview loop.

Start AI mock interview