Follow one decision through the whole ML system.
An ML system-design answer is not a list of cloud services. It must connect a user decision, data, labels, modelling, serving, measurement, monitoring and improvement.
Start with the decision
“Design a spam classifier.” Before discussing models, what must be clarified?
The end-to-end framework
Frame the decision
Who consumes the prediction? What action changes? Why is ML better than a rule or no intervention?
Define constraints
Volume, latency, freshness, availability, explainability, privacy, safety, cost and geographic requirements.
Design data and labels
Sources, label definition, delay, quality, leakage, sampling, point-in-time availability and bias.
Choose metrics
Connect business outcomes, offline model metrics, online experiment metrics and operational SLOs.
Establish a baseline
Start with a rule, majority class or simple model. Complexity must earn its operational cost.
Train and validate
Features, model family, split strategy, tuning, calibration, error analysis and performance slices.
Serve safely
Batch or online inference, feature computation, caching, versioning, rollout, fallback and rollback.
Monitor and improve
Input health, drift, delayed labels, prediction quality, latency, errors, cost and retraining triggers.
Worked example: spam classification
| Decision | Possible design | Question to defend |
|---|---|---|
| Action | Route high-confidence spam to quarantine; flag uncertain cases. | What happens when a legitimate message is blocked? |
| Label | User report plus reviewed samples; account for delayed and noisy reports. | Who reports spam, and what bias does that create? |
| Metrics | Precision at quarantine threshold, spam recall, complaint rate, latency. | Why is accuracy insufficient? |
| Baseline | Rules plus a simple text classifier. | What does the model add beyond known abusive senders? |
| Serving | Online scoring with a fallback rules engine. | What is the latency budget? |
| Monitoring | Feature health, score distribution, delayed precision/recall, attack patterns. | How quickly can attackers change vocabulary? |
How the interviewer increases difficulty
Scale
Traffic grows 100 times. What must move from a simple service to distributed or cached infrastructure?
Reliability
The model service fails. What fallback preserves the product?
Feedback delay
Labels arrive after seven days. What can be monitored immediately?
Adversaries
Spammers adapt to the classifier. How do detection and retraining respond?
Fairness
Performance differs by language. How will slices and data collection change?
Cost
A more complex model improves recall slightly but triples serving cost. Is it worth it?