DeepOffer

Build a full transformer block from scratch: multi-head attention, feed-forward network, residuals, and layer norm.

ML CodingReported interview question
Reported in a public interview report — Frontier labs (research-engineering onsites)

Compose pre-norm attention and MLP sublayers with residual additions: x←x+Attn(Norm(x)), then x←x+MLP(Norm(x)). Keep dropout and masks explicit and verify batch, head, sequence, and feature dimensions.

Test empty input, one-element input, duplicates, boundary indices, invalid states, and the largest allowed size; state time and space complexity.

Common follow-up questions

Practice this question with an AI interviewer

Get asked follow-ups live, then receive a scored report — like a real MLE interview loop.

Start AI mock interview