DeepOffer

Describe the KV cache, derive its memory formula, and discuss what it costs at production scale.

ML TheoryReported interview question
Reported in public interview compilations — OpenAI, xAI, Mistral AI, Amazon, Apple, NVIDIA, Together AI, Character.AI

A KV cache stores each prior token’s key and value tensors for every layer, so autoregressive decoding reuses past projections instead of recomputing them. Memory grows roughly with batch × sequence length × layers × KV heads × head dimension × 2 × bytes.

Use equations or tensor shapes where they clarify the claim, then name an experiment or ablation that would distinguish competing explanations.

Common follow-up questions

Practice this question with an AI interviewer

Get asked follow-ups live, then receive a scored report — like a real MLE interview loop.

Start AI mock interview