One line connects all three: Maximum likelihood maximizes the product of P(y_i|x_i). Taking the negative log gives minus sum log P(y_i|x_i) - which is exactly cross-entropy with one-hot labels.
KL divergence: D_KL(P||Q) = H(P,Q) - H(P). Cross-entropy = entropy of the true distribution + KL. With fixed labels, H(P) is constant, so minimizing cross-entropy is equivalent to minimizing KL and to maximizing likelihood.
Why it matters: it upgrades why cross-entropy from a memorized conclusion to first principles, and explains why knowledge distillation uses KL on soft labels.
Add directionality: KL is asymmetric. D_KL(P||Q) and D_KL(Q||P) penalize differently (mode-covering vs. mode-seeking) - generation interviews often dig here.
Get asked follow-ups live, then receive a scored report — like a real MLE interview loop.
Start AI mock interview