DeepOffer

Do the roofline math: how many tokens/sec can one H100 produce for a 70B model at batch size 1?

ML System DesignReported interview question
Reported in public interview compilations — NVIDIA, Together AI

Start with workload, users, data, quality metrics, latency, throughput, freshness, reliability, and cost targets. Choose a simple baseline, then separate offline data/training/evaluation from online serving and feedback collection.

Make interfaces and ownership explicit; add versioning, access control, monitoring, canary rollout, rollback, and a plan for delayed labels or human review.

Common follow-up questions

Practice this question with an AI interviewer

Get asked follow-ups live, then receive a scored report — like a real MLE interview loop.

Start AI mock interview