The 60-second answer
Tokenization maps raw text into discrete IDs consumed by the embedding table. Subword methods balance vocabulary size, sequence length, and out-of-vocabulary robustness.
Build the answer in this order
1
Define the mechanism
Tokenization maps raw text into discrete IDs consumed by the embedding table.
2
Explain the architecture
Subword methods balance vocabulary size, sequence length, and out-of-vocabulary robustness.
3
Compare trade-offs
Compare BPE, WordPiece, and unigram approaches conceptually rather than by brand name alone.
4
Close with serving + evaluation
Measure downstream quality and token-cost impact for the target languages/domains.
A useful interview mental model
This is the shape of a strong answer—not a script to memorize.
01Definition
02Mechanism
03Trade-offs
04Failure modes
05When to use
Senior-level signal
- Tokenizer changes break checkpoint and embedding compatibility.
- Multilingual and code-heavy workloads need special attention to fragmentation and normalization.
What the interviewer is really testing
Understanding beyond prompting: architecture, retrieval, evaluation, inference, safety, latency, cost, and failure recovery.
Likely follow-up questions
What assumption makes this approach work?
When would you choose the strongest alternative instead?
What production or data failure mode changes your answer?
Common weak-answer patterns
- Reciting a definition without mechanism or assumptions.
- Claiming one technique is always better without a data regime.
- Stopping before failure modes, validation, or deployment implications.