Part 4 · 25 minutes

Follow one decision through the whole ML system.

An ML system-design answer is not a list of cloud services. It must connect a user decision, data, labels, modelling, serving, measurement, monitoring and improvement.

Start with the decision

“Design a spam classifier.” Before discussing models, what must be clarified?

The end-to-end framework

Objective
Constraints
Data + labels
Metrics
Baseline
Model
Serving
Experiment
Monitor
Retrain

Frame the decision

Who consumes the prediction? What action changes? Why is ML better than a rule or no intervention?

Define constraints

Volume, latency, freshness, availability, explainability, privacy, safety, cost and geographic requirements.

Design data and labels

Sources, label definition, delay, quality, leakage, sampling, point-in-time availability and bias.

Choose metrics

Connect business outcomes, offline model metrics, online experiment metrics and operational SLOs.

Establish a baseline

Start with a rule, majority class or simple model. Complexity must earn its operational cost.

Train and validate

Features, model family, split strategy, tuning, calibration, error analysis and performance slices.

Serve safely

Batch or online inference, feature computation, caching, versioning, rollout, fallback and rollback.

Monitor and improve

Input health, drift, delayed labels, prediction quality, latency, errors, cost and retraining triggers.

Worked example: spam classification

DecisionPossible designQuestion to defend
ActionRoute high-confidence spam to quarantine; flag uncertain cases.What happens when a legitimate message is blocked?
LabelUser report plus reviewed samples; account for delayed and noisy reports.Who reports spam, and what bias does that create?
MetricsPrecision at quarantine threshold, spam recall, complaint rate, latency.Why is accuracy insufficient?
BaselineRules plus a simple text classifier.What does the model add beyond known abusive senders?
ServingOnline scoring with a fallback rules engine.What is the latency budget?
MonitoringFeature health, score distribution, delayed precision/recall, attack patterns.How quickly can attackers change vocabulary?

How the interviewer increases difficulty

Scale

Traffic grows 100 times. What must move from a simple service to distributed or cached infrastructure?

Reliability

The model service fails. What fallback preserves the product?

Feedback delay

Labels arrive after seven days. What can be monitored immediately?

Adversaries

Spammers adapt to the classifier. How do detection and retraining respond?

Fairness

Performance differs by language. How will slices and data collection change?

Cost

A more complex model improves recall slightly but triples serving cost. Is it worth it?

Previous: AnswersNext: Profile and resume