CrackML by @ml.with.umang
Interview questions / ML Fundamentals
ML Fundamentals interview question

Tokenization Tradeoffs

Compare word, character, and subword tokenization for modern NLP systems.

mediumconceptEvidence 41/1001 source reportLinkedIn

The 60-second answer

Word tokenization is intuitive but has vocabulary/OOV problems; character tokenization has tiny vocabulary but much longer sequences; subword methods balance both. Explain how BPE/WordPiece-style tokenizers trade vocabulary size against sequence length and how rare words, morphology, code, and multilingual text affect fragmentation.

Build the answer in this order

1
Give the core idea

Word tokenization is intuitive but has vocabulary/OOV problems; character tokenization has tiny vocabulary but much longer sequences; subword methods balance both.

2
Explain how it works

Explain how BPE/WordPiece-style tokenizers trade vocabulary size against sequence length and how rare words, morphology, code, and multilingual text affect fragmentation.

3
Compare alternatives

Connect tokenization to compute: longer sequences increase attention cost and latency, while larger vocabularies increase embedding/output-layer memory.

4
State failure modes + validation

Evaluate with downstream task quality plus fertility/fragmentation, unknown-token behavior, language slices, and inference cost.

A useful interview mental model

This is the shape of a strong answer—not a script to memorize.

01Definition
02Mechanism
03Trade-offs
04Failure modes
05When to use

Senior-level signal

  • Senior answers discuss tokenizer/model coupling: changing a tokenizer after pretraining is not a free preprocessing change.
  • Include domain adaptation issues such as medical terms, code, IDs, and multilingual scripts.

What the interviewer is really testing

Mechanistic understanding, assumptions, trade-offs, and whether you can turn a definition into a model decision.

Likely follow-up questions

What assumption makes this approach work?
When would you choose the strongest alternative instead?
What production or data failure mode changes your answer?

Common weak-answer patterns

  • Reciting a definition without mechanism or assumptions.
  • Claiming one technique is always better without a data regime.
  • Stopping before failure modes, validation, or deployment implications.