DeepOffer

Derive scaled dot-product attention and explain why dividing by √d_k is essential.

ML TheoryReported interview question
Reported in public interview reports

Scaled dot-product attention forms QKᵀ/√dₖ, applies a row-wise softmax, then mixes V. The √dₖ term keeps logit variance from growing with width, preventing saturated softmax and weak gradients.

Use equations or tensor shapes where they clarify the claim, then name an experiment or ablation that would distinguish competing explanations.

Common follow-up questions

Practice this question with an AI interviewer

Get asked follow-ups live, then receive a scored report — like a real MLE interview loop.

Start AI mock interview