Revision should produce answers, not highlighted notes.
Someone who already knows ML does not need to relearn every course. The goal is to retrieve ideas without support, connect them to scenarios, and explain trade-offs under questioning.
Retrieval check
Without notes, explain why accuracy can be misleading on imbalanced data. Give one business example and one better evaluation approach.
The four-step revision cycle
Recall
Explain the topic from memory before opening a resource.
Repair
Study only the missing or incorrect part.
Apply
Use it on an example, comparison or small calculation.
Speak
Give a concise answer aloud and handle one follow-up.
Why it works: interviews require retrieval and explanation. Re-reading creates familiarity, but familiarity can feel like knowledge even when the idea cannot be produced independently.
Probable ML topic map
| Area | Concepts | Questions to practise |
|---|---|---|
| Problem framing | Target, prediction unit, baseline, constraints, leakage | Should this problem use ML? What decision changes? |
| Statistics | Probability, Bayes, distributions, sampling, confidence intervals, testing | What uncertainty exists? Is an observed change credible? |
| Supervised learning | Linear/logistic regression, trees, forests, boosting, SVM, KNN, Naive Bayes | What does the model learn? What assumptions and failure modes follow? |
| Unsupervised learning | Clustering, PCA, anomaly detection | How will success be evaluated without ordinary labels? |
| Model development | Splits, CV, bias-variance, regularization, tuning, feature engineering | How do you know the improvement generalizes? |
| Evaluation | Regression/classification metrics, thresholds, calibration, slices | Which error matters to the business? |
| Data problems | Missingness, imbalance, noisy labels, drift, skew | Where can the pipeline silently become invalid? |
| Production ML | Batch/online serving, monitoring, retraining, reliability, cost | What happens after the notebook? |
| Role-specific | Recommendations, NLP, vision, time series, LLMs/RAG | Which domain concepts repeat in target JDs? |
Build a revision matrix
| Topic | JD frequency | Current confidence | Evidence | Next action |
|---|---|---|---|---|
| Class imbalance | High | Medium | Can explain metrics; weak on calibration | Answer aloud + one experiment |
| Online serving | High | Low | No production example | Study one architecture + design drill |
| PCA | Low | High | Can derive and demonstrate | Maintenance only |
Priority rule: preparation priority rises when a topic appears frequently in suitable jobs and current evidence is weak. Do not spend equal time on every square.
Useful primary resources
Google ML Crash Course
Modular revision with exercises and interactive explanations.
Stanford CS229
Mathematical depth and lecture notes.
Rules of ML
Practical production reasoning, baselines, metrics and pipeline discipline.
Resource trap: collecting ten courses is not a preparation plan. Select one primary source per gap and convert it into questions, calculations, examples or spoken explanations.