The 60-second answer
Project input to Q/K/V, reshape [B,T,D] → [B,H,T,Dh], and compute scaled QKᵀ scores. Apply masks before softmax using values appropriate for the dtype; softmax over the key dimension.
Build the answer in this order
1
State tensor contract
Project input to Q/K/V, reshape [B,T,D] → [B,H,T,Dh], and compute scaled QKᵀ scores.
2
Implement the mechanism
Apply masks before softmax using values appropriate for the dtype; softmax over the key dimension.
3
Check numerics + gradients
Multiply weights by V, transpose/reshape heads back to [B,T,D], then apply the output projection.
4
Test shapes and edge cases
Test shapes, causal and padding masks, all-masked edge cases, and gradient equivalence with a trusted implementation.
A useful interview mental model
This is the shape of a strong answer—not a script to memorize.
01Shapes
02Forward pass
03Loss / grads
04Numerics
05Tests
Senior-level signal
- Senior answers discuss FlashAttention-style fused kernels and why avoiding the explicit T×T matrix reduces memory pressure.
- Separate attention dropout from residual/MLP dropout and know their train/eval semantics.
What the interviewer is really testing
Tensor fluency, shape reasoning, numerics, gradients, batching, device awareness, and the ability to debug—not API memorization.
Likely follow-up questions
What are the tensor shapes at each step?
Where could numerical instability or silent broadcasting appear?
How would you verify gradients and batched behavior?
Common weak-answer patterns
- Ignoring shape, dtype, device, masking, or broadcasting assumptions.
- Using a framework call without explaining the underlying operation.
- Skipping gradient, numerical-stability, and batching checks.