DeepOffer

Contrast the prefill and decode phases and explain why prefill is compute-bound while decode is memory-bandwidth-bound.

ML System DesignReported interview question
Reported in public interview compilations — Moonshot AI, NVIDIA, Together AI

Prefill processes all prompt tokens in parallel and is usually matrix-multiply compute bound. Decode adds one token at a time while reading large weights and KV state, so memory bandwidth and batching dominate.

Make interfaces and ownership explicit; add versioning, access control, monitoring, canary rollout, rollback, and a plan for delayed labels or human review.

Common follow-up questions

Practice this question with an AI interviewer

Get asked follow-ups live, then receive a scored report — like a real MLE interview loop.

Start AI mock interview