CrackML by @ml.with.umang
Interview questions / ML Fundamentals
ML Fundamentals interview question

What does dropout do during neural-network training?

What does dropout do during neural-network training?. Structure your response as you would in a top-tier ML/AI engineering interview.

easyconceptEvidence 75/1001 source reportAmazon

The 60-second answer

During training, dropout randomly zeros activations with probability p to reduce co-adaptation and act as regularization. With inverted dropout, surviving activations are scaled during training so no stochastic dropping or extra scaling is needed at inference.

Build the answer in this order

1
Give the core idea

During training, dropout randomly zeros activations with probability p to reduce co-adaptation and act as regularization.

2
Explain how it works

With inverted dropout, surviving activations are scaled during training so no stochastic dropping or extra scaling is needed at inference.

3
Compare alternatives

It is most useful when overfitting is a concern; too much dropout can increase bias and slow optimization.

4
State failure modes + validation

Always switch the model to evaluation mode for deterministic inference.

A useful interview mental model

This is the shape of a strong answer—not a script to memorize.

01Definition
02Mechanism
03Trade-offs
04Failure modes
05When to use

Senior-level signal

  • Discuss interactions with BatchNorm/LayerNorm and why dropout is used less aggressively in some modern Transformer blocks.
  • Treat inference-mode mistakes as a correctness bug that can silently change output distributions.

What the interviewer is really testing

Mechanistic understanding, assumptions, trade-offs, and whether you can turn a definition into a model decision.

Likely follow-up questions

What assumption makes this approach work?
When would you choose the strongest alternative instead?
What production or data failure mode changes your answer?

Common weak-answer patterns

  • Reciting a definition without mechanism or assumptions.
  • Claiming one technique is always better without a data regime.
  • Stopping before failure modes, validation, or deployment implications.