The 60-second answer
Choose metrics from the product error cost: class-wise precision/recall/F1, ROC/PR curves, calibration, top-k accuracy, or detection-specific metrics as appropriate. Build meaningful slices across lighting, camera/device, geography, demographic or object subgroups, rare classes, and image quality.
Build the answer in this order
Choose metrics from the product error cost: class-wise precision/recall/F1, ROC/PR curves, calibration, top-k accuracy, or detection-specific metrics as appropriate.
Build meaningful slices across lighting, camera/device, geography, demographic or object subgroups, rare classes, and image quality.
Inspect confusion matrices and hard examples to distinguish label problems, distribution shift, and model confusion.
Validate robustness, latency/memory, and online outcomes; keep a stable golden set plus recent-data evaluation.
A useful interview mental model
This is the shape of a strong answer—not a script to memorize.
Senior-level signal
- Senior answers discuss dataset shift and annotation quality as first-class evaluation risks.
- Include confidence calibration and abstention/human-review policy when mistakes have asymmetric cost.
What the interviewer is really testing
Likely follow-up questions
Common weak-answer patterns
- Reciting a definition without mechanism or assumptions.
- Claiming one technique is always better without a data regime.
- Stopping before failure modes, validation, or deployment implications.