CrackML by @ml.with.umang
Focused topic

Evaluation & Metrics interview questions.

A curated set of 21 questions that repeatedly exercise Evaluation & Metrics concepts in AI/ML interviews.

Design a content moderation ML systemML System Design · hard · Evidence 77/100Design an ML system using traffic dataML System Design · hard · Evidence 77/100Why can offline metrics look better than online metrics?ML Fundamentals · hard · Evidence 77/100How do you calibrate a classifier?ML Fundamentals · medium · Evidence 75/100How would you evaluate an LLM?GenAI & LLM · medium · Evidence 75/100How would you model and evaluate time-series data?ML Fundamentals · medium · Evidence 75/100How would you formulate an ML problem from a product goal?ML System Design · medium · Evidence 41/100How would you evaluate retrieval quality separately from answer quality in a Q&A bot?GenAI & LLM · hard · Evidence 41/100How would you reason about product sense for an ML ranking feature?ML System Design · medium · Evidence 41/100How would you test whether a modeling choice actually improved the product?ML Fundamentals · medium · Evidence 41/100How should retrieval metrics differ from ranking metrics in a recommender?ML Fundamentals · medium · Evidence 41/100Defend an ML project end to endProject & Behavioral · hard · Evidence 41/100How would you evaluate a computer-vision model beyond aggregate accuracy?ML Fundamentals · medium · Evidence 41/100For a Q&A bot, when would you add RAG instead of relying on a parametric model alone?GenAI & LLM · medium · Evidence 41/100How would you evaluate the relevance of a color-suggestion model?ML Fundamentals · medium · Evidence 41/100SVM Margin and Hinge LossML Math · medium · Evidence 41/100How would you evaluate clustering quality for unlabeled user segments?ML Fundamentals · medium · Evidence 41/100How would you choose evaluation metrics and explain them to a non-technical stakeholder?ML Fundamentals · medium · Evidence 41/100How would you evaluate a video recommender offline and online?ML System Design · hard · Evidence 37/100Evaluate a Computer Vision ModelML Fundamentals · medium · Evidence 34/100Explain precision vs recall and when each mattersML Fundamentals · medium · Evidence 31/100