Model: z = Xw + b, p = sigmoid(z), loss is binary cross-entropy. Write matrix shapes in comments first - half the bugs in hand-coding rounds are shape bugs.
Gradients: dW = X^T(p - y)/n, db = mean(p - y) - again prediction minus label, consistent with the softmax case.
Training loop: Initialize (zero init works for logistic regression - explain why it fails for neural networks), then forward, loss, backward, update. Print loss each round and check it decreases monotonically.
Engineering details: Standardize features first; how to spot learning-rate-too-large oscillation; numerically stable sigmoid (two branches for positive/negative).
Get asked follow-ups live, then receive a scored report — like a real MLE interview loop.
Start AI mock interview