The 60-second answer
BERT is encoder-only and uses bidirectional self-attention; GPT is decoder-only and uses causal self-attention. BERT is pretrained with masked-token style objectives, while GPT is trained autoregressively to predict the next token.
Build the answer in this order
BERT is encoder-only and uses bidirectional self-attention; GPT is decoder-only and uses causal self-attention.
BERT is pretrained with masked-token style objectives, while GPT is trained autoregressively to predict the next token.
BERT naturally fits representation/classification tasks; GPT naturally fits generation and can also be adapted to many understanding tasks.
Compare context direction, training objective, output interface, and serving pattern rather than only parameter count.
A useful interview mental model
This is the shape of a strong answer—not a script to memorize.
Senior-level signal
- Discuss memory bandwidth and attention/KV-cache costs, not only FLOPs, when reasoning about real latency.
- Tie architectural choices to the task, context length, training objective, and deployment constraints.
What the interviewer is really testing
Likely follow-up questions
Common weak-answer patterns
- Reciting a definition without mechanism or assumptions.
- Claiming one technique is always better without a data regime.
- Stopping before failure modes, validation, or deployment implications.