Chapter 1: Overview of Machine Learning Systems
15 min readThe chapter opens with a great example: Google Translate’s multilingual neural machine translation system back in 2016. This was a landmark – deep learning making a massive, tangible impact at scale. It showed the world the power of modern ML.
Since then, as the book says, ML has “found its way into almost every aspect of our lives.” This is true. At FAANG, we see this daily – from the recommendations you get, to how your photos are enhanced, to how we optimize our data centers.
But here’s the first crucial takeaway for any ML design interview, and indeed, for your day-to-day work:
“Many people, when they hear ‘machine learning system,’ think of just the ML algorithms… However, the algorithm is only a small part of an ML system in production.”
This is HUGE. I can’t stress this enough. In interviews, if you only talk about model architectures and training loops, you’re missing 90% of the picture. The system includes:
- Business Requirements: Why are we even building this? What problem does it solve?
- User Interface: How do users interact with it? How do developers interact with it?
- Data Stack: Ingestion, storage, processing, versioning – the lifeblood.
- Development, Monitoring, Updating Logic: This is MLOps territory.
- Infrastructure: The compute, storage, and networking that makes it all run.

Self-Correction/Interview Tip: When an interviewer asks you to design an ML system, your first thoughts should be about the business problem, the users, and the data, not immediately about whether to use a Transformer or a ResNet.
Page 2: MLOps vs. ML Systems Design & The Book’s Philosophy
This page introduces two key terms:
- MLOps: This comes from DevOps. It’s about operationalizing ML – deploying, monitoring, maintaining. It’s a set of tools and best practices. Think CI/CD for ML, model registries, monitoring dashboards.
- ML Systems Design: This is the approach the book (and this cohort) takes. It’s a holistic view of MLOps, considering all components and stakeholders to meet objectives. It’s the “architecture” level thinking.
Figure 1-1 is really important here. You see ML system users and business requirements feeding in (Chapters 1 & 2). You see ML system developers and the entire book dedicated to them. The “ML system” box itself has:
- Data (Chapters 3 & 4)
- Feature Engineering (Chapter 5)
- ML Algorithms (Chapter 6 – notice it’s just one chapter!)
- Evaluation (also Chapter 6)
- Deployment, monitoring, updating of logics (Chapters 7, 8 & 9)
- Infrastructure (Chapter 10)
And Chapter 11 covers the overarching aspects like stakeholders, ethics, etc. (though ethics is woven throughout).
The book explicitly states it won’t cover specific algorithms in detail. Why?
Because algorithms change, they get outdated. The framework for building robust, scalable, maintainable ML systems – that’s timeless.
This is exactly what FAANG companies look for: people who can think in terms of systems and principles, not just the latest hot model.
FAANG Perspective: At places like Google, Meta, Amazon, we have dedicated MLOps platform teams building tools, but product-focused ML engineers are the ones using these tools within a systems design framework to solve specific business problems. They need to understand the whole lifecycle.
Pages 3-8: When to Use Machine Learning (The Litmus Test)
This is arguably the most critical section for any ML project’s inception and a very common starting point for ML design interviews. Before you design anything, you must ask: Is ML necessary or cost-effective?
The book offers a fantastic five-part definition: “Machine learning is an approach to:
(1) Learn: The system must have the capacity to learn. A relational database isn’t an ML system because you explicitly define relationships. ML systems infer them, usually from data. For supervised learning (e.g., predicting Airbnb prices), you need input-output pairs.
- Interview Relevance: “How does your system learn from new data?”
(2) Complex patterns: ML shines when patterns are too complex to hand-code. Predicting if a fair die roll is a 6? No pattern. Predicting stock prices? Complex patterns (hopefully!). Sorting listings by state based on zip code? Lookup table. Predicting Airbnb rental price from many features? ML! This is Andrej Karpathy’s “Software 2.0” idea – you provide data, the system writes the “code” (the learned patterns).
- FAANG Nuance: We often deal with problems where the rules are unknown or too numerous to list, like ranking search results or newsfeed items. This is prime ML territory.
(3) Existing data: ML needs data to learn from. No data, no ML (usually). The book mentions predicting taxes without tax data – impossible.
- Zero-shot learning is an interesting case: it makes predictions for tasks it wasn’t directly trained on, but it was trained on related tasks/data. So, it still needs some data.
- Fake-it-til-you-make-it: Launch with humans making predictions, collect that data, then train an ML model. This is a common bootstrapping strategy for new products.
- Interview Relevance: “What data would you use? How would you collect it? What if you don’t have labeled data initially?”
(4) Predictions: ML solves problems requiring predictive answers. “What will the weather be?” “What movie will a user watch next?” Even compute-intensive problems can be reframed: “What would the outcome of this complex rendering process look like?” (approximate it with ML).
(5) Unseen data: The patterns learned must generalize to new, unseen data. If your model trained on 2008 app download data (Koi Pond!) tries to predict 2020 downloads, it’ll fail. Technically, training and unseen data should come from similar distributions. If not, your model will perform poorly (hello, monitoring and retraining, Chapters 8 & 9!).
- Interview Relevance: This leads directly to discussions of data drift, concept drift, and the importance of robust evaluation and monitoring.
The book then lists additional characteristics where ML solutions shine:
- It’s repetitive: ML algorithms (especially deep learning) often need many examples. Repetitive tasks provide these.
- The cost of wrong predictions is cheap (usually): A bad movie recommendation? User ignores it. A self-driving car makes a wrong turn? Catastrophic. This influences your choice of problem and acceptable error rates. However, even for high-stakes, if ML on average outperforms humans (like statistically safer self-driving cars), it can be viable.
- It’s at scale: ML often needs significant upfront investment (data, compute, talent). This is justified if you’re making many predictions (e.g., sorting millions of emails, routing thousands of support tickets). “At scale” also implies lots of data for training.
- The patterns are constantly changing: Spam emails evolve. Fashion trends change. ML models can be updated with new data to adapt, unlike hardcoded rules. This is where “Continual Learning” (page 264, Chapter 9) comes in.
When NOT to use ML:
- It’s unethical: Automated grader biases (page 341) – we’ll cover ethics in detail.
- Simpler solutions do the trick: (Chapter 6) ALWAYS start with a non-ML baseline. If a heuristic works, great!
- FAANG Mantra: “Start simple.” A common failure mode for junior engineers is over-engineering.
- It’s not cost-effective.
A crucial caution: Don’t dismiss new tech just because it’s not cost-effective now. Early adoption can be a competitive advantage later. And sometimes, you can break a big problem into smaller pieces, using ML for just one part.
Interview Gold: The “When to use ML” checklist is your first filter for any ML system design question. Always discuss these trade-offs. If you propose an ML solution, be ready to defend why ML is appropriate over a simpler, non-ML approach.
Pages 9-11: Machine Learning Use Cases
This section is a tour of where ML is making an impact. It’s good for broadening your understanding of the applications.
Consumer Applications:
- Search engines (Google)
- Recommender systems (Netflix, Amazon)
- Predictive typing (your phone’s keyboard)
- Photo enhancement
- Biometric authentication (fingerprint, face ID)
- Machine translation (the author’s personal anecdote about parents using Google Translate is lovely)
- Smart assistants (Alexa, Google Assistant)
- Smart security cameras (pet detection, uninvited guests)
- At-home health monitoring (fall detection)
Enterprise Applications: The book rightly states these are the majority of ML use cases.
- Often have stricter accuracy requirements but can be more forgiving on latency compared to consumer apps (e.g., 0.1% improvement in resource allocation for Google can save millions, even if the system takes a few seconds to run).
- Figure 1-3 (Algorithmia 2020 survey): Shows diverse enterprise uses:
- Internal: Reducing costs (38%), generating customer insights (37%), internal processing automation (30%).
- External: Improving customer experience (34%), retaining customers (29%), interacting with customers (28%).
- Specific examples:
- Fraud detection: Anomaly detection on transactions. (Very common, one of the oldest ML uses in enterprise).
- Price optimization: Dynamic pricing for ads, flights, ride-sharing. Maximize margin/revenue.
- Demand forecasting: For inventory, resource allocation.
- Customer acquisition: Identifying potential customers, targeted ads, optimizing discounts.
- Churn prediction: Predicting when customers (or employees!) might leave.
Self-Correction/Interview Tip: Having a few diverse use cases in your back pocket is useful. If an interviewer asks, “Give me an example of an ML system you find interesting,” you can draw from these. It also helps you think about the types of problems ML can solve (classification, regression, anomaly detection, etc.).
Page 12: More Enterprise Use Cases & Intro to Understanding ML Systems
More enterprise examples:
- Automated support ticket classification: Route tickets to the right department faster.
- Brand monitoring: Sentiment analysis on brand mentions.
- Healthcare: Skin cancer detection, diabetes diagnosis (often through a provider, due to accuracy/privacy needs).
This page then transitions to the next major section: Understanding Machine Learning Systems. The key motivation here is that ML systems are different from:
- ML in research (or academia).
- Traditional software.
And these differences necessitate the kind of thinking this book promotes.
Pages 12-21: Machine Learning in Research Versus in Production
This is a CRITICAL section, especially for those coming from academia or who have mostly worked on Kaggle-like projects. ML in the wild is a different beast. Table 1-1 summarizes key differences:
- Requirements
- Research: State-of-the-art model performance on benchmarks
- Production: Different stakeholders, often conflicting requirements
- Computational Priority
- Research: Fast training, high throughput
- Production: Fast inference, low latency
- Data
- Research: Static, benchmark datasets
- Production: Constantly shifting, messy, real-world data
- Fairness
- Research: Often not a focus
- Production: Must be considered
- Interpretability
- Research: Often not a focus
- Production: Must be considered (often a requirement)
Let’s break these down:
Different stakeholders and requirements (Page 13-14):
- Research: Usually a single objective – e.g., SOTA on a benchmark. Researchers might use complex techniques for marginal gains.
- Production: Multiple stakeholders with conflicting needs. The restaurant recommender example is classic:
- ML Engineers: Want best model for user clicks (maybe complex, needs more data).
- Sales Team: Want model recommending expensive restaurants (more service fees).
- Product Team: Want low latency (<100ms), as latency drops orders.
- ML Platform Team: Worried about scaling, want to pause updates to improve platform.
- Manager: Wants to maximize margin (maybe cut the ML team if costs are too high!).
The book mentions decoupling objectives (page 41) – we’ll get there. For now, understand that production ML is about trade-offs. Is a 100ms latency a must-have or a nice-to-have? This determines if Model A or B is even viable.
Ensembling: Popular in competitions (Netflix Prize), but often too complex, slow, or hard to interpret for production. A small performance lift might not justify the operational overhead.
Criticism of ML Leaderboards (Page 15):
- Hard steps (data collection, problem formulation) are often done for you.
- Multiple hypothesis testing: With many teams, some might get good results by chance.
- Misaligned incentives: Drive for accuracy at the expense of compactness, fairness, energy efficiency (Ethayarajh & Jurafsky).
FAANG Perspective: This stakeholder wrangling is daily life. Product Managers, Engineering Managers, ML Engineers, and sometimes legal/policy teams all have a say. Clear communication and defining priorities are key.
Computational priorities (Page 15-18):
- Research: Focus on fast training and high throughput (samples/sec during training).
- Production: Focus on fast inference and low latency (time from query to result).
Terminology Clash (Page 16): The book uses “latency” to mean “response time” (what the client sees, including network/queueing delays). This is common ML community usage.
Latency vs. Throughput (Figure 1-4, Page 17):
- Single query processing: Higher latency = lower throughput.
- Batched query processing: Can increase throughput, but also individual query latency (waiting for a batch to fill). This is a common trade-off.
Impact of Latency: Real-world examples (Akamai: 100ms delay = 7% conversion drop; Booking.com: 30% latency increase = 0.5% conversion cost; Google: >3s load time, users leave). Users are impatient!
Latency is a distribution (Page 18): Don’t just use average! Averages hide outliers. Use percentiles:
- p50 (median): 50% of requests are faster/slower.
- p90, p95, p99: Show tail latencies. These outliers might affect your most valuable customers (e.g., Amazon customers with large purchase histories). Product requirements are often “p99 latency < X ms.”
Interview Relevance: Be ready to discuss latency/throughput trade-offs, batching strategies, and how you’d measure and set SLOs for latency (e.g., “p95 latency for recommendations must be under 200ms”).
Data (Page 18-19):
- Research: Often clean, well-formatted, static benchmark datasets. Known quirks, public scripts for processing.
- Production: Data is MESSY!
- Noisy, unstructured, constantly shifting.
- Biased (and you might not know how).
- Labels can be sparse, imbalanced, incorrect.
- Changing requirements mean updating labels.
- Privacy and regulatory concerns (GDPR, CCPA).
- Data is constantly generated by users, systems, third parties.
Figure 1-5 (Karpathy’s “Amount of sleep lost over…” graphic): In PhD, sleep lost over models/algorithms. At Tesla (production), sleep lost over datasets. This is profoundly true.
FAANG Perspective: Data engineering, data quality, data governance, and data privacy are massive efforts. Often, more engineering time is spent on the data pipeline than on the model itself.
Fairness (Page 19-20):
- Research: Often an afterthought.
- Production: CRITICAL. ML models encode past biases from data. Deployed at scale, they can discriminate at scale.
- Examples: Loan applications biased by zip code, resume ranking biased by name spelling, mortgage rates based on biased credit scores.
- Cathy O’Neil’s “Weapons of Math Destruction” is a must-read.
- Misclassifying minority groups might have a small impact on overall accuracy metrics but a huge impact on those individuals/groups.
- McKinsey study (2019): Only 13% of large companies mitigating algorithmic bias. This is changing, but slowly. We’ll cover Responsible AI in Chapter 11.
Interview Relevance: Expect questions on fairness. “How would you detect bias in your model? How would you mitigate it?” This is table stakes now.
Interpretability (Page 20-21):
- Geoffrey Hinton’s AI surgeon dilemma: 90% cure rate AI surgeon (black box) vs. 80% human surgeon (explainable). Who do you choose? (When the author asked execs, it was 50/50).
- Research: Often not incentivized if focus is pure performance.
- Production: Often a requirement.
- For users/business leaders: To trust the model, detect biases.
- For developers: To debug and improve the model.
- Legal: “Right to explanation” in some jurisdictions (e.g., GDPR).
- 2019 Stanford HAI report: Only 19% of large companies working on explainability. Again, this is improving.
FAANG Nuance: For critical systems (e.g., fraud, content moderation, medical), interpretability is non-negotiable. Techniques like SHAP, LIME are used, but simpler, inherently interpretable models are often preferred if performance is comparable.
Discussion (Page 21): Why production focus matters?
- Most companies can’t afford pure research without business application.
- The “bigger, better” ML models (e.g., large language models) require massive data and compute, often tens of millions of dollars.
- As ML becomes more accessible, demand for productionizing ML grows.
- The vast majority of ML-related jobs are in productionizing ML. This is why you’re here!
Pages 22-23: Machine Learning Systems Versus Traditional Software
If ML is part of software engineering (SWE), why not just use existing SWE best practices? Good idea! ML production would be better if ML experts were also strong software engineers. Many SWE tools are useful.
However, ML has unique challenges:
- Code + Data + Artifacts:
- SWE: Assumes code and data are separate. Focus on modularity.
- ML: Systems are tightly coupled: code (training scripts, inference logic), data (features, labels), and artifacts (trained models).
- Trend: “Best data wins” over “best algorithm.” So, focus shifts to improving data.
- Data changes quickly => ML apps need to adapt quickly => faster dev/deploy cycles.
- Testing and Versioning:
- SWE: Test and version code.
- ML: Must test and version code and data. Versioning large datasets is hard. How to know if a data sample is “good” or “bad”?
- Not all data samples are equal: A scan of a cancerous lung is more valuable if you have 1M normal lung scans and only 1k cancerous ones.
- Indiscriminate data acceptance can hurt performance or lead to data poisoning attacks (footnote 31).
- Model Size:
- As of 2022, models with billions of parameters are common (e.g., LLMs). Require GBs of RAM.
- This might seem quaint in the future (like the 32MB RAM of the Apollo moon computer).
- Deploying large models, especially on edge devices (Chapter 7), is a massive engineering challenge.
- Speed (Inference Latency):
- How to run these large models fast enough? An autocompletion model slower than typing is useless.
- Monitoring and Debugging:
- Non-trivial for complex, black-box models. Hard to know what went wrong or get alerted quickly.
The good news (Page 23): These challenges are being tackled. BERT (2018) was initially seen as too big/slow (340M params, 1.35GB). By 2020, it was in “almost every English search on Google.” Progress is rapid.
Interview Relevance: Understanding these unique challenges of ML (data-centricity, versioning data+model, model size/latency, monitoring complexity) distinguishes a candidate who has thought about production issues from one who has only trained models in a lab.
Page 23: Summary
This opening chapter gives you the lay of the land:
- ML is widely used, especially in enterprise.
- Knowing when and when not to use ML is crucial.
- ML in production is very different from ML in research (stakeholders, compute, data, fairness, interpretability).
- ML systems have unique challenges compared to traditional software (data entanglement, versioning, model size, monitoring).
The key theme of the book, and this cohort, is a holistic system approach. We don’t just look at algorithms; we look at all components working together.