DeepOffer

Contrast multi-query and grouped-query attention with standard multi-head attention, and name what is sacrificed.

ML TheoryReported interview question
Reported in public interview compilations — Meta, Mistral AI

Multi-head attention gives every query head its own K/V heads; MQA shares one K/V set and GQA shares a smaller number of K/V groups. Sharing sharply reduces cache size and memory bandwidth, with a possible quality cost.

Use equations or tensor shapes where they clarify the claim, then name an experiment or ablation that would distinguish competing explanations.

Common follow-up questions

Practice this question with an AI interviewer

Get asked follow-ups live, then receive a scored report — like a real MLE interview loop.

Start AI mock interview