DeepOffer

Architect an Experimentation Platform with Sound Assignment and Analysis

ML System DesignHot interview question
Reported in public interview compilations — Rapyd, Renew Power, Buffer

Splitting: Hash user ID into buckets. Experiments are either mutually exclusive (layered orthogonal) or allowed to overlap. Assignment must be stable (same user, same group) and connected to the feature system.

Sample size: Fix baseline rate, minimum detectable effect (MDE), significance level, and power, then back out per-group sample size and runtime. Underestimating MDE is the top reason experiments drag on.

Analysis: SRM check first (if the actual split deviates from expected, the splitter or logging is broken - check that before results), multiple-metric correction, sequential testing to prevent p-value peeking.

Business side: run long enough to cover weekend effects, separate primary metrics from guardrails, and document negative results too.

Common follow-up questions

Practice this question with an AI interviewer

Get asked follow-ups live, then receive a scored report — like a real MLE interview loop.

Start AI mock interview