CrackML by @ml.with.umang
Interview questions / ML System Design
ML System Design interview question

Design a Recommendation System at High Level

Walk through the end-to-end architecture of a large-scale personalized recommendation system.

hardsystem designEvidence 41/1001 source reportLinkedIn

The 60-second answer

Define user objective, inventory scale, freshness, latency, and success/guardrail metrics before architecture. Use multiple candidate generators—collaborative, content, graph, trending, exploration—then a richer ranker with real-time and historical features.

Build the answer in this order

1
Frame the problem

Define user objective, inventory scale, freshness, latency, and success/guardrail metrics before architecture.

2
Design the data path

Use multiple candidate generators—collaborative, content, graph, trending, exploration—then a richer ranker with real-time and historical features.

3
Choose the modeling stack

Build training labels with exposure-bias awareness, offline feature pipelines consistent with serving, and experimentation/monitoring around the final slate.

4
Serve, evaluate, iterate

Evaluate retrieval recall, ranking NDCG/engagement, diversity/novelty, long-term satisfaction, latency, and negative feedback.

A useful interview mental model

This is the shape of a strong answer—not a script to memorize.

01Requirements
02Data
03Model / Retrieval
04Serving
05Monitor

Senior-level signal

  • Senior answers cover multi-objective optimization, exploration, feedback loops, and cold-start migration.
  • Include feature freshness, fallback tiers, and how to diagnose retrieval vs ranking vs serving failures.

What the interviewer is really testing

Product framing, data design, modeling choices, serving constraints, reliability, evaluation, and explicit trade-offs.

Likely follow-up questions

What changes at 10× traffic or data volume?
Which failure mode would you monitor first in production?
How would you evaluate this offline and online before rollout?

Common weak-answer patterns

  • Jumping to a model before defining the product contract.
  • Listing components without bottlenecks, metrics, or failure handling.
  • Ignoring data quality, serving latency, monitoring, and iteration.