Four monitoring layers: Input data (schema, missing rate, distribution), features (online/offline consistency), prediction output (score distribution, confidence), and business outcomes (real metrics when delayed labels arrive).
Drift detection: PSI / KS tests for numeric features, chi-square or distribution distance for categorical, overall shift for embeddings. Thresholds are tiered by feature importance - not every drift deserves an alert.
Label delay: Real labels arrive hours to days later. Run proxy signals (prediction distribution, weak feedback) alongside rolling evaluation on delayed labels. Be explicit about which layer is real-time and which is lagging.
Alerting and response: Tiered alerts to avoid alert fatigue, automatic rollback to the previous model or rule fallback, and one-click drill-down for drift attribution (which traffic segment / which feature moved first).
Get asked follow-ups live, then receive a scored report — like a real MLE interview loop.
Start AI mock interview