\

Chapter 1: Introduction

4 min read

Chapter 1: Introduction - Reliable ML Engineering Notes

Core Concept: The ML Loop

  • ML is Iterative, Not Linear: ML applications are rarely “done.” They exist in a continuous cycle of development, deployment, evaluation, and improvement.
  • Why a Loop?
    • If a model underperforms: Teams (DS, Business, MLE) collaborate to improve it (features, data, architecture).
    • If a model performs well: Organizations want more sophistication and broader application, leading to further development.
  • The first model is just a starting point.

The ML Lifecycle Stages (The “Pit Stops” in the Loop)

(See Figure 1-1 for visual representation)

image

1. Data Collection and Analysis

  • Goal: Understand available data, identify needs, prioritize uses, collect, and process.
  • Involves: Business/Product (for priorities), Data Engineers (pipelines), SREs (reliability of pipelines), DS/MLEs (data utility).
  • Key: Proper data management is foundational. Business context drives data needs.

2. ML Training Pipelines

  • Goal: Consume processed data and produce trained models using ML algorithms.
  • Involves: Data Engineers, DS, MLEs, SREs.
  • Tech: TensorFlow, PyTorch, XGBoost, etc.
  • Critical: Training pipelines are production systems. Treat with rigor.
  • Common Failures: Data issues (lack, format), bugs, misconfigurations, resource shortages, hardware/distributed system failures.
  • Silent Failures: ML models can also fail due to subtle issues like data distribution shifts.

3. Build and Validate Applications

  • Goal: Integrate the model into a customer-facing system to deliver value.
  • Involves: Product/Business (specs), MLEs/SDEs (implementation), QA (oversight).
  • Action: Log user interactions and model outputs for future improvement.

4. Quality and Performance Evaluation

  • Goal: Assess if the model “works” and how well, before and during initial launch.
  • Methods:
    • Offline Evaluation: Test on historical/curated data.
    • Live Launch: Model sees live traffic (monitor closely).
    • Dark Launch: Model sees live traffic, logs predictions, but doesn’t affect user experience (tests integration).
    • Fractional Launch (Canary/A/B): Model exposed to a subset of users (tests quality and integration).
  • Purpose: Gain confidence for wider rollout, establish baselines. It’s a validation checkpoint.

5. Defining and Measuring SLOs (Service-Level Objectives)

  • Goal: Define and track thresholds (SLIs - Indicators) for system performance according to requirements.
  • Involves: SREs, PMs, DS, MLEs, SDEs.
  • Challenge in ML: Subtle data/world changes can degrade ML performance significantly.
  • Types of SLOs:
    • System: Latency, error rates, throughput (serving & training).
    • Application: # of recommendations, successful model calls.
    • ML Performance/Business: Click-through rates, revenue from model (often sliced by user segments).
  • Key: Business must define tolerable SLOs.

6. Launch

  • Goal: Ship the ML-enhanced application to users.
  • Involves: Product SDEs, MLEs, SREs.
  • ML-Specific Concerns:
    • Models as Code: New models can break systems like bad code.
    • Launch Slowly: Progressive rollouts to limit damage.
    • Isolate Rollouts at Data Layer: Critical to avoid data format incompatibilities during rollbacks (see “Progressive Rollouts in a Stateful System” story).
    • Release, Not Refactor: Minimize changes during a release.
    • Measure SLOs During Launch: Monitor dashboards.
    • Review the Rollout: Manual or automated oversight.

7. Monitoring and Feedback Loops

  • Goal: Continuously observe system health and effectiveness post-launch, and gather data for future improvements.
  • Signals to Monitor:
    • System Health (Golden Signals): Latency, traffic, errors, saturation.
    • Basic Model Health: Model size, load errors (context-free checks).
    • Model Quality (Domain-Specific): Business metrics (CTR, conversion), performance drift over time. Hardest but most crucial.
  • Feedback Collection: Log user interactions, predictions, and context to fuel the next cycle.
  • Purpose: Ensure sustained performance, detect issues, provide input for the next iteration.

Key Differences: Evaluation vs. Monitoring

  • Quality & Performance Evaluation: A checkpoint before/during initial launch to validate and decide on wider rollout.
  • Monitoring & Feedback Loops: Continuous vigilance post-launch to maintain health, detect drift, and gather data for future iterations.

Overall Lessons from the Loop

  • Data is King: ML begins and ends with data.
  • Cyclical Process: No single order; stages are revisited.
  • Holistic View Required: Understand the entire loop and organization.
  • Risk & Experimentation: Not all ML ideas work. Approach as continual experimentation.
  • Organizational Readiness Matters.