Chapter 1: Introduction
4 min readChapter 1: Introduction - Reliable ML Engineering Notes
Core Concept: The ML Loop
- ML is Iterative, Not Linear: ML applications are rarely “done.” They exist in a continuous cycle of development, deployment, evaluation, and improvement.
- Why a Loop?
- If a model underperforms: Teams (DS, Business, MLE) collaborate to improve it (features, data, architecture).
- If a model performs well: Organizations want more sophistication and broader application, leading to further development.
- The first model is just a starting point.
The ML Lifecycle Stages (The “Pit Stops” in the Loop)
(See Figure 1-1 for visual representation)
1. Data Collection and Analysis
- Goal: Understand available data, identify needs, prioritize uses, collect, and process.
- Involves: Business/Product (for priorities), Data Engineers (pipelines), SREs (reliability of pipelines), DS/MLEs (data utility).
- Key: Proper data management is foundational. Business context drives data needs.
2. ML Training Pipelines
- Goal: Consume processed data and produce trained models using ML algorithms.
- Involves: Data Engineers, DS, MLEs, SREs.
- Tech: TensorFlow, PyTorch, XGBoost, etc.
- Critical: Training pipelines are production systems. Treat with rigor.
- Common Failures: Data issues (lack, format), bugs, misconfigurations, resource shortages, hardware/distributed system failures.
- Silent Failures: ML models can also fail due to subtle issues like data distribution shifts.
3. Build and Validate Applications
- Goal: Integrate the model into a customer-facing system to deliver value.
- Involves: Product/Business (specs), MLEs/SDEs (implementation), QA (oversight).
- Action: Log user interactions and model outputs for future improvement.
4. Quality and Performance Evaluation
- Goal: Assess if the model “works” and how well, before and during initial launch.
- Methods:
- Offline Evaluation: Test on historical/curated data.
- Live Launch: Model sees live traffic (monitor closely).
- Dark Launch: Model sees live traffic, logs predictions, but doesn’t affect user experience (tests integration).
- Fractional Launch (Canary/A/B): Model exposed to a subset of users (tests quality and integration).
- Purpose: Gain confidence for wider rollout, establish baselines. It’s a validation checkpoint.
5. Defining and Measuring SLOs (Service-Level Objectives)
- Goal: Define and track thresholds (SLIs - Indicators) for system performance according to requirements.
- Involves: SREs, PMs, DS, MLEs, SDEs.
- Challenge in ML: Subtle data/world changes can degrade ML performance significantly.
- Types of SLOs:
- System: Latency, error rates, throughput (serving & training).
- Application: # of recommendations, successful model calls.
- ML Performance/Business: Click-through rates, revenue from model (often sliced by user segments).
- Key: Business must define tolerable SLOs.
6. Launch
- Goal: Ship the ML-enhanced application to users.
- Involves: Product SDEs, MLEs, SREs.
- ML-Specific Concerns:
- Models as Code: New models can break systems like bad code.
- Launch Slowly: Progressive rollouts to limit damage.
- Isolate Rollouts at Data Layer: Critical to avoid data format incompatibilities during rollbacks (see “Progressive Rollouts in a Stateful System” story).
- Release, Not Refactor: Minimize changes during a release.
- Measure SLOs During Launch: Monitor dashboards.
- Review the Rollout: Manual or automated oversight.
7. Monitoring and Feedback Loops
- Goal: Continuously observe system health and effectiveness post-launch, and gather data for future improvements.
- Signals to Monitor:
- System Health (Golden Signals): Latency, traffic, errors, saturation.
- Basic Model Health: Model size, load errors (context-free checks).
- Model Quality (Domain-Specific): Business metrics (CTR, conversion), performance drift over time. Hardest but most crucial.
- Feedback Collection: Log user interactions, predictions, and context to fuel the next cycle.
- Purpose: Ensure sustained performance, detect issues, provide input for the next iteration.
Key Differences: Evaluation vs. Monitoring
- Quality & Performance Evaluation: A checkpoint before/during initial launch to validate and decide on wider rollout.
- Monitoring & Feedback Loops: Continuous vigilance post-launch to maintain health, detect drift, and gather data for future iterations.
Overall Lessons from the Loop
- Data is King: ML begins and ends with data.
- Cyclical Process: No single order; stages are revisited.
- Holistic View Required: Understand the entire loop and organization.
- Risk & Experimentation: Not all ML ideas work. Approach as continual experimentation.
- Organizational Readiness Matters.