The 60-second answer
Define the class taxonomy, action taken from predictions, latency/freshness, and asymmetric error costs before collecting labels. Build representative labels with clear annotation policy, a simple baseline, feature/data pipelines, calibrated model scores, and threshold policy.
Build the answer in this order
Define the class taxonomy, action taken from predictions, latency/freshness, and asymmetric error costs before collecting labels.
Build representative labels with clear annotation policy, a simple baseline, feature/data pipelines, calibrated model scores, and threshold policy.
Evaluate PR/ROC as appropriate plus class-wise/slice metrics, calibration, and the downstream product metric.
Serve with versioned features/model, monitor drift/label delay/performance, and create human-review or fallback paths for uncertain/high-impact cases.
A useful interview mental model
This is the shape of a strong answer—not a script to memorize.
Senior-level signal
- Senior answers discuss policy/label drift separately from feature/model drift.
- Include threshold changes and calibration as versioned production decisions, not hard-coded constants.
What the interviewer is really testing
Likely follow-up questions
Common weak-answer patterns
- Jumping to a model before defining the product contract.
- Listing components without bottlenecks, metrics, or failure handling.
- Ignoring data quality, serving latency, monitoring, and iteration.