\

Chapter 0: ML Design Primer

8 min read

This chapter introduces the fundamentals of designing machine learning systems at scale. It covers the key differences between traditional software systems and ML systems, and outlines the unique challenges faced when building production ML systems.

Key Concepts

  • ML System vs Traditional Systems: Understanding the differences in architecture, requirements, and challenges
  • Scale Considerations: How scale affects ML system design decisions
  • Production ML Pipeline: Components of a typical ML production pipeline

Main Topics Covered

  1. What makes ML systems different
  2. Key components of ML systems
  3. Common challenges in ML system design
  4. Framework for approaching ML system design interviews

Interview Tips

  • Always start with clarifying requirements
  • Think about data flow and system components
  • Consider scalability and performance trade-offs
  • Don’t forget about monitoring and maintenance

(Your detailed notes for Chapter 1 go here…)

Excellent question. This is exactly the mindset a senior candidate needs to have – not just knowing the material, but understanding how it fits into the current landscape. You’ve got 5 days, so we need to be efficient and focus on high-impact areas.

First, let’s get straight to your main question:

Yes, you can and absolutely should use this book. The 2022 “Machine Learning Design Interview” by Khang Pham is a fantastic resource. The fundamental principles it covers – the “why” behind system design choices – are timeless. Things like the two-stage (candidate/ranking) architecture, feature engineering for sparse data (hashing, embeddings), and handling data at scale are the bedrock of ML systems.

However, you’re right that the field moves incredibly fast. What a senior staff engineer expects from a senior MLE candidate in 2024 has evolved. We’re not just looking for knowledge of these patterns, but also for an understanding of the new class of problems and tools that have emerged, primarily around Large Language Models (LLMs) and Generative AI.

Think of this book as your solid foundation. We’re going to build a modern extension on top of it.

Part 1: What to ADD (The 2024+ Lens)

Here are the key concepts that have become critical since 2022. You need to be able to discuss these intelligently.

1. Retrieval-Augmented Generation (RAG)

This is the single most important ML system design pattern to emerge in the last two years. It’s how you make LLMs factual, current, and context-aware. The book talks about retrieval for recommendations (two-tower models), but RAG applies it to conversational AI and generation tasks.

Simple Intuition: An LLM knows a lot, but it doesn’t know about your company’s private data, yesterday’s news, or a specific user’s profile. RAG is the process of:

  1. Receiving a user’s query.
  2. Using that query to retrieve relevant documents from a knowledge base.
  3. Stuffing those documents into the LLM’s prompt along with the original query.
  4. Asking the LLM to generate an answer based only on the provided context.

Here is a simple system diagram for RAG:

flowchart TD
    subgraph offline["Offline: Data Indexing"]
        A[Unstructured Data<br/>Docs, PDFs, etc.] --> B[Chunking Service]
        B --> C[Embedding Model<br/>text-embedding-ada-002]
        C --> D[Vector Database<br/>Pinecone, Qdrant]
    end
    
    subgraph online["Online: Inference"]
        U[User Query] --> Q[Query Embedding]
        Q --> S{Vector Search}
        S --> R[Retrieved Context]
        U --> P[Prompt Engineering]
        R --> P
        P --> LLM[Large Language Model<br/>GPT-4, Llama 3]
        LLM --> F[Final Answer]
    end
    
    D -.-> S
    
    style D fill:#cde4ff,stroke:#333,stroke-width:2px

2. Vector Databases

The book mentions FAISS, which is a library. The modern discussion is about managed Vector Databases as a core infrastructure component. They are purpose-built to do fast Approximate Nearest Neighbor (ANN) search on billions of embeddings.

Key discussion points for a senior role:

  • Trade-offs: Latency vs. recall vs. cost. How do you choose an index type (e.g., HNSW vs. IVF)?
  • Metadata Filtering: How do you retrieve vectors that not only are semantically similar but also match specific criteria (e.g., user_id = '123' AND is_public = true)? This is a crucial production feature.
  • Scaling: How do these databases scale for reads and writes?

3. LLMOps & Foundation Model Serving

Deploying a 5GB BERT model is different from deploying a 70B parameter Llama model.

  • Quantization: Techniques like GPTQ or GGML/GGUF to shrink model size to run on cheaper hardware.
  • Specialized Serving Frameworks: vLLM, TensorRT-LLM. These use techniques like PagedAttention to dramatically increase throughput for LLM inference.
  • Fine-tuning vs. Prompt Engineering vs. RAG: When do you use each? (Hint: Start with Prompting/RAG. Fine-tuning is expensive and only for changing a model’s style or teaching it a new skill, not for adding knowledge).

4. Real-time Feature Stores

The book mentions the concept, but their importance has solidified. A feature store is the source of truth that solves the training-serving skew problem.

  • Modern Take: Feature stores now often have streaming capabilities (e.g., using Flink or Spark Streaming) to compute real-time features (e.g., “user’s clicks in the last 5 minutes”) and make them available for inference with low latency.
flowchart TD
  subgraph "Real-time Feature Engineering"
    direction TB
    A["Event Stream\n(Kafka, Kinesis)"] --> B{"Stream Processor\n(Flink, Spark Streaming)"}
    B --> C["Real-time Feature Store\n(Redis, DynamoDB)"]
  end

  subgraph "Serving"
    direction TB
    D["Prediction Service"] --> C
  end

  subgraph "Training"
    direction TB
    B --> E["Batch Storage\n(S3, BigQuery)"]
    F["Model Training"] --> E
  end

Part 2: What to DEPRECATE (or De-emphasize)

Technology moves on. While the principles are good, don’t spend your limited time on the implementation details of these.

  1. Pure Hadoop/MapReduce: The book mentions MapReduce jobs. Today, the de-facto standard for large-scale batch processing is Apache Spark, and for streaming, it’s Spark Streaming or Apache Flink. If you say “MapReduce job,” it will sound dated. Frame everything in terms of Spark or Flink.

  2. Elaborate, from-scratch feature engineering for text/images: For many problems, you no longer build TF-IDF vectors or train a Word2Vec model from scratch on day one.

    • The 2024 Way: You start with a powerful pre-trained foundation model (e.g., a BERT variant for embeddings, or even an LLM). Your “feature engineering” is now about choosing the right model and deciding on an embedding strategy. The core idea of “semantic representation” is the same, but the tool has become much more powerful.
  3. Over-indexing on specific architectures (like DCNv2): It’s great to know why a model like Deep & Cross Network (DCN) exists – to explicitly learn feature interactions. But a senior staff engineer is more interested in you recognizing the problem (the need for both memorization and generalization) than memorizing the exact architecture of DCNv2. You could say: “For a system with many categorical features like this, we need to handle both memorization of specific feature pairs and generalization. Historically, models like Wide & Deep or DCN addressed this. Today, we might handle this with a powerful embedding model that captures these interactions implicitly, or by using an LLM to reason over the features.”

Your 5-Day High-Intensity Action Plan

Days 1-2: Master the Book’s Fundamentals

  • Read the book cover-to-cover. Don’t skip.
  • For every component (e.g., One-Hot Encoding, Two-Tower Model), force yourself to answer: “What business or system problem does this solve?” (e.g., Two-tower model solves the problem of efficiently retrieving a small set of candidates from a massive corpus).
  • Focus on Chapters 1 (Primer), 2 (RecSys Components), and then skim the end-to-end examples.

Day 3: Layer on the 2024 Concepts

  • Deep-dive into RAG. Watch 2-3 YouTube videos explaining it. Understand the diagram I drew above.
  • Read about Vector Databases (e.g., read the intro docs for Pinecone or Weaviate).
  • Understand the difference between real-time, near-real-time, and batch features.

Day 4: Practice Designing

  • Take a classic problem from the book, like “YouTube Video Recommendations.”
  • First, design it the “classic” way as described in the book (two-tower candidate gen, sophisticated ranker).
  • Then, ask yourself: “How would I redesign this in 2024?”
    • Could I add a conversational search feature using RAG?
    • How would my video embeddings be stored and retrieved? (Vector DB!)
    • Could an LLM re-rank the final candidates by providing natural language “reasons” for why a user might like a video?
    • Draw the new architecture. This is your key senior-level exercise.

Day 5: Review & Communication

  • Practice verbalizing your designs. Use a framework:
    1. Clarify: Ask about scope, scale, latency, and business goals.
    2. Metrics: Define offline (e.g., NDCG, Precision) and online (e.g., CTR, user session time) metrics.
    3. High-Level Design: Draw the big boxes. Start simple.
    4. Deep Dive: Go into the components (data, model, serving). This is where you’ll use the book’s content and the 2024 additions.
    5. Scale & Reliability: Discuss bottlenecks, caching, and failure modes.

You have a solid resource. Your goal in the next 5 days is not to become an expert in LLMOps, but to be able to confidently place the timeless principles from the book into the context of modern AI systems.

Good luck. You’ve got this.