DeepOffer

Explain multi-head latent attention and the motivation DeepSeek had for introducing it.

ML TheoryReported interview question
Reported in public interview compilations — DeepSeek, Moonshot AI

MLA compresses keys and values into a lower-dimensional latent representation and reconstructs the needed projections during attention. It cuts KV-cache traffic more aggressively than GQA but adds projection complexity and implementation constraints.

Use equations or tensor shapes where they clarify the claim, then name an experiment or ablation that would distinguish competing explanations.

Common follow-up questions

Practice this question with an AI interviewer

Get asked follow-ups live, then receive a scored report — like a real MLE interview loop.

Start AI mock interview