The 60-second answer
The learning rate controls step size; schedules allow aggressive early progress and smaller late-stage updates for stable convergence. Common choices include step/exponential decay, cosine decay, one-cycle, and warmup followed by decay; match the schedule to optimizer and training horizon.
Build the answer in this order
The learning rate controls step size; schedules allow aggressive early progress and smaller late-stage updates for stable convergence.
Common choices include step/exponential decay, cosine decay, one-cycle, and warmup followed by decay; match the schedule to optimizer and training horizon.
Warmup is especially useful for large-batch or Transformer training where early gradients/normalization statistics can be unstable.
Compare schedules on validation quality at equal compute and monitor loss spikes, gradient norms, and sensitivity to restart/resume behavior.
A useful interview mental model
This is the shape of a strong answer—not a script to memorize.
Senior-level signal
- Senior answers reason in tokens/steps and total compute, not epochs alone, for large training runs.
- Discuss whether the schedule remains valid after changing batch size, data mixture, or optimizer.
What the interviewer is really testing
Likely follow-up questions
Common weak-answer patterns
- Reciting a definition without mechanism or assumptions.
- Claiming one technique is always better without a data regime.
- Stopping before failure modes, validation, or deployment implications.