Store per-layer K/V tensors indexed by sequence position, append the current token’s projections, and attend the new query over cached positions. Preallocate or page storage to avoid repeated concatenation.
Test empty input, one-element input, duplicates, boundary indices, invalid states, and the largest allowed size; state time and space complexity.
Get asked follow-ups live, then receive a scored report — like a real MLE interview loop.
Start AI mock interview