DeepOffer

Distinguish pretraining, supervised fine-tuning, and preference optimization.

ML TheoryReported interview question
Reported in public interview compilations — Meta, Scale AI

Pretraining learns broad next-token structure from large unlabeled corpora; SFT imitates curated demonstrations; preference optimization shifts outputs toward ranked human or synthetic preferences. Each stage changes a different part of capability and behavior.

Use equations or tensor shapes where they clarify the claim, then name an experiment or ablation that would distinguish competing explanations.

Common follow-up questions

Practice this question with an AI interviewer

Get asked follow-ups live, then receive a scored report — like a real MLE interview loop.

Start AI mock interview