Temporal leakage: Using data that only exists after prediction time (e.g., using a field generated after checkout to predict whether checkout happens).
Preprocessing leakage: Computing normalization stats, feature selection, or target encoding on the full dataset before splitting - validation info leaks into training.
Label leakage: Features contain proxies of the label (e.g., using refund reason to predict whether a refund happens).
Prevention: split by time, put all preprocessing inside a pipeline fit only on training data, audit every feature for availability at prediction time, and run online/offline consistency checks before launch.
Get asked follow-ups live, then receive a scored report — like a real MLE interview loop.
Start AI mock interview