\

Chapter 1: Introduction and Overview

25 min read

1. GenAI Overview: The Big Picture

Before we build anything, we need to know what we’re working with.

  • What are we trying to achieve? We’re trying to categorize our problem. Are we trying to classify something that already exists, or are we trying to create something new? This is the most fundamental decision you’ll make.

The book correctly breaks ML models into two camps: Discriminative and Generative.

Let’s break down these two categories of ML models. This is a classic interview question, and you need to nail it.

1. Discriminative Models: The Art Critic

  • What they do: They learn the boundary or the difference between different classes of data.
  • The formal definition: They learn the conditional probability P(Y | X).
  • Intuitive Explanation: Think of an art critic. You show them a painting (X), and they tell you the probability that it’s a Picasso (Y=Picasso) or a Monet (Y=Monet). The critic doesn’t know how to paint a Picasso; they only know how to distinguish a Picasso from other paintings based on features like brush strokes, color palette, and subject matter.
  • Core Task: Classification (Is this email spam or not?) and Regression (What is the price of this house?).
  • Examples: Logistic Regression, SVMs, the “classifier” part of a GAN.

2. Generative Models: The Artist

  • What they do: They learn the underlying distribution of the data itself. They learn what makes a Picasso a Picasso.
  • The formal definition: They model the distribution P(X) or the joint distribution P(X, Y).
  • Intuitive Explanation: This is the artist. You can ask them, “Create a new painting in the style of Picasso.” Because they have learned the essence of what a Picasso painting looks like (P(X) where X is the distribution of Picasso paintings), they can generate a brand new, never-before-seen example. If you ask for a “Picasso-style painting of a cat” (Y="cat"), they are modeling the joint distribution P(X, Y).
  • Core Task: Creating new data (text, images, audio, etc.).
  • Examples: VAEs, GANs, Diffusion Models, Autoregressive models (like GPT).

The key takeaway is: “While discriminative algorithms can predict a target variable from input features, most of them lack the capability to learn the underlying data distribution needed to generate new data instances. For that, we turn to generative models.”

Let’s visualize this relationship.

%%{init: {'flowchart': {'nodeSpacing': 50, 'rankSpacing': 80}}}%%
graph TD
    A[๐Ÿค– Artificial Intelligence] --> B[๐Ÿง  Machine Learning]
    B --> C{๐Ÿ”€ Model Type}
    C --> D[โš–๏ธ Discriminative<br/>The Judge]
    C --> E[๐ŸŽจ Generative<br/>The Artist]

    D --> D1[๐Ÿ“Š Classification:<br/>Is this a cat?]
    D --> D2[๐Ÿ’ก Recommendation:<br/>Will you like this movie?]

    E --> E1[๐Ÿ–ผ๏ธ Image Generation:<br/>Create a picture of a cat]
    E --> E2[๐Ÿ“ Text Generation:<br/>Write a story about a cat]

    subgraph Core["๐ŸŽฏ Core Task"]
        F[๐Ÿšช Learns the boundary<br/>between data]
        G[๐Ÿ“ˆ Learns the distribution<br/>of data]
    end

    D -.-> F
    E -.-> G

    style A fill:#dae8fc,stroke:#6c8ebf,stroke-width:3px
    style B fill:#d5e8d4,stroke:#82b366,stroke-width:3px
    style C fill:#f9f9f9,stroke:#333,stroke-width:2px
    style D fill:#f8cecc,stroke:#b85450,stroke-width:2px
    style E fill:#e1d5e7,stroke:#9673a6,stroke-width:2px
    style F fill:#fff2cc,stroke:#d6b656,stroke-width:2px
    style G fill:#fff2cc,stroke:#d6b656,stroke-width:2px

Here are popular tasks for each type, which helps clarify the distinction:

%%{init: {'flowchart': {'nodeSpacing': 40, 'rankSpacing': 70}}}%%
graph TD
    A[๐Ÿค– ML-powered Tasks] --> B[โš–๏ธ Discriminative]
    A --> C[๐ŸŽจ Generative]

    B --> B1[๐Ÿ–ผ๏ธ Image Segmentation]
    B --> B2[๐Ÿ‘๏ธ Object Detection]
    B --> B3[๐Ÿ˜Š Sentiment Analysis]
    B --> B4[๐Ÿท๏ธ Named Entity Recognition]
    B --> B5[๐Ÿ’ก Recommendation Systems]
    B --> B6[๐Ÿ” Visual Search]

    C --> C1[๐Ÿ’ฌ Chatbots]
    C --> C2[๐Ÿ“„ Summarization]
    C --> C3[๐Ÿ–ผ๏ธ๐Ÿ“ Image Captioning]
    C --> C4[๐Ÿ“โžก๏ธ๐Ÿ–ผ๏ธ Text-to-Image]
    C --> C5[๐Ÿ‘ค Face Generation]
    C --> C6[๐ŸŽต Audio Synthesis]

    style A fill:#f0f0f0,stroke:#666,stroke-width:3px
    style B fill:#f8cecc,stroke:#b85450,stroke-width:2px
    style C fill:#e1d5e7,stroke:#9673a6,stroke-width:2px

2. Why is GenAI So Powerful? The Three Pillars

The book nails the three key drivers. In an interview, you can frame this as the “perfect storm” that made GenAI possible.

  • What are we trying to achieve? We’re explaining the fundamental enablers. This shows you understand the context of the current AI boom.
  1. Data (The Library of Alexandria): Traditional ML needed meticulously labeled data (e.g., “this image is a cat,” “this one is a dog”). This is slow and expensive. The breakthrough for GenAI was self-supervised learning. We can now feed models the entire internetโ€”unlabeled text from Wikipedia, books, code from GitHub. The model creates its own learning tasks (e.g., “predict the next word”). The sheer volume and diversity of this data allow models to learn nuanced patterns about language, reasoning, and the world.

  2. Model Capacity (The Brain Size):

    • Parameters: Think of these as the knobs or synapses in the model’s “brain.” The more parameters (e.g., GPT-3 with 175B, PaLM with 540B), the more information and complex patterns it can learn and store.
    • FLOPs (Floating Point Operations): This isn’t about size, but about computational cost. How much “thinking” does it take to get an answer? A model can have fewer parameters but a more complex architecture (like dense connections) that requires more FLOPs. Understanding the difference between model size (parameters) and computational complexity (FLOPs) is a sign of a senior engineer.
  3. Compute (The Engine): You can have a giant brain and a massive library, but if you can only read one word per second, it’s useless. Specialized hardware like GPUs (NVIDIA’s A100, H100) and TPUs (Google’s custom chips) are the powerful engines that can perform the trillions of calculations needed to train these massive models in a feasible amount of time (weeks instead of centuries).

3. Scaling Laws: The Recipe for Success

This is a critical, FAANG-level concept. In an interview, discussing scaling laws shows you’re thinking about the science and economics of training, not just throwing compute at a problem.

  • What are we trying to achieve? We want to train the best possible model without wasting money. Given a fixed budget for computation (the total FLOPs we can afford), what’s the best way to spend it? Should we make the model bigger (more parameters, N) or train it on more data (more tokens, D)?

Before scaling laws, it was a bit of a guessing game: “For my next model, should I double the data or double the model size?”

  • The OpenAI (2020) Finding: They found that performance scales predictably as a power-law with model size, dataset size, and compute. Crucially, “the impact of scaling on model performance is significantly more pronounced than the influence of architectural variations.” This means that for a while, making models bigger was more important than making them architecturally cleverer.

  • The DeepMind (Chinchilla, 2022) Refinement: DeepMind found that many existing LLMs were undertrained. We were building huge models but not feeding them enough data. They proposed that for optimal performance, model size and training dataset size should be scaled linearly together. This led to models like “Chinchilla,” which was smaller than Gopher but trained on much more data, and outperformed it.

  • Why this matters for an interview: It shows you understand the economics. If a VP gives you a $50 million budget to train a new model, you can use scaling laws to propose a plan: “Based on Chinchilla’s scaling laws, to get the optimal performance for this budget, we should aim for a model with X parameters and train it on Y trillion tokens.” This is a data-driven engineering decision.

Scaling laws gave us a predictable recipe. They demonstrated that for a given compute budget, there’s an optimal ratio between model size (number of parameters) and the amount of training data (number of tokens). This turned model training from a “dark art” into a more predictable engineering discipline.

4. The Framework for ML System Design Interviews

This is the absolute core of the chapter and your roadmap for any design interview. This is what the interviewer is evaluating you on.

  • What are we trying to achieve? We’re trying to demonstrate a structured, logical, and comprehensive approach to solving a complex, open-ended problem. We need to show we think about the entire system, not just the model.

Let’s recreate their flowchart. This is your mental checklist.

%%{init: {'flowchart': {'nodeSpacing': 60, 'rankSpacing': 90}}}%%
graph TD
    A[๐Ÿ“‹ 1. Clarify Requirements] --> B[๐ŸŽฏ 2. Frame as ML Task]
    B --> C[๐Ÿ“Š 3. Data Preparation]
    C --> D[๐Ÿ› ๏ธ 4. Model Development]
    D --> E[๐Ÿ“ˆ 5. Evaluation]
    E --> F[๐Ÿ—๏ธ 6. Overall System Design]
    F --> G[๐Ÿš€ 7. Deployment & Monitoring]

    style A fill:#e6f3ff,stroke:#0066cc,stroke-width:3px
    style B fill:#fff2e6,stroke:#cc6600,stroke-width:3px
    style C fill:#e6ffe6,stroke:#00cc00,stroke-width:3px
    style D fill:#ffe6f3,stroke:#cc0066,stroke-width:3px
    style E fill:#f3e6ff,stroke:#6600cc,stroke-width:3px
    style F fill:#ffe6e6,stroke:#cc0000,stroke-width:3px
    style G fill:#f0f0f0,stroke:#666666,stroke-width:3px

Let’s walk through the first few steps from this part of the book.

Step 1: Clarifying Requirements

Before you write a single line of code or draw a single box, you must understand the problem. This is where you distinguish yourself as a senior engineer. You’re not just a code monkey; you’re a problem solver.

  • Functional Requirements: What should the system do?

    • Example: “Generate an image and customize its style based on the user’s prompt.”
    • My advice: Be specific. For a customer service chatbot, is it just for answering FAQs, or should it be able to process returns? Does it need to remember conversation history?
  • Non-Functional Requirements: How should the system perform? This is where the real engineering challenges lie. Ask about:

    • Business Objective: Why are we building this? To reduce customer support costs? To increase user engagement? This North Star dictates your trade-offs.
    • System Features: Does the user rate the outputs? Can they edit the generated image? These are features that create feedback loops.
    • Data: What data can we use? Is it sensitive (PII, medical)? Is it labeled?
    • Constraints: Will it run on-device (on a phone) or in the cloud? This has massive implications for model size and latency.
    • Scale: Are we serving 100 users or 100 million users? This impacts everything from database choice to inference architecture.
    • Performance: What is the acceptable latency? Does a user expect an image in 2 seconds or 20 seconds? Is quality or speed more important?

In an interview, spend a solid 5 minutes here. Ask these questions. The interviewer often has a more detailed scenario in mind, and you need to extract it from them.

Step 2: Framing the Problem as an ML Task

Now you translate the requirements into a machine learning problem.

  1. Specify Input and Output: This seems simple, but it’s crucial.
  • Example for Text-to-Image: Input = text prompt (string), maybe some style parameters. Output = Image (e.g., a 1024x1024 pixel grid).
  • Example for Chatbot: Input = user’s text query. Output = system’s text response.
  • Example for Video generation: Input = text prompt, maybe a starting image. Output = sequence of frames (video).
  1. Choose a Suitable ML Approach: This is where you make your first major architectural decision.

Here’s the logic, simplified:

  1. Discriminative vs. Generative? Look at the output. Are you predicting a label from a fixed set (discriminative) or creating new content (generative)? For GenAI problems, the answer is almost always generative.

  2. Identify the Task Type: What kind of content are you generating? Text? Image? Audio? Video? This narrows down your model choices significantly.

  3. Choose a Suitable Algorithm: Now you get into the specific families of models. For image generation, the main contenders are GANs, VAEs, Autoregressive models, and Diffusion models. You should be able to briefly discuss the trade-offs:

    • GANs: Fast to generate, but can be unstable to train and may have “mode collapse” (lack diversity).
    • VAEs: Stable to train, produce diverse outputs, but can sometimes be blurrier than GANs.
    • Diffusion Models: Produce state-of-the-art, high-quality images, but are computationally expensive and slow to sample from (though this is improving).
    • Autoregressive Models: Generate pixel by pixel. Can be very high quality but are extremely slow.

In an interview, you’d say: “This is a text-to-image generation task. The key requirements are high-quality output and the ability to follow complex prompts. Given this, a Diffusion model is a strong candidate. While they are historically slow for inference, recent advancements like Latent Diffusion and improved sampling schedules have made them practical. I’d choose this over a GAN because training stability and output quality are paramount for this product.”

Step 3: Data Preparation

Great models are built on great data. Garbage in, garbage out. The book makes a key distinction between data prep for traditional ML vs. GenAI.

  • Traditional ML (Structured Data): The focus is on Feature Engineering. You have tables of data (e.g., customer purchase history) and you spend your time creating clever features (e.g., “days since last purchase,” “average transaction value”).
%%{init: {'flowchart': {'nodeSpacing': 50, 'rankSpacing': 60}}}%%
graph LR
    A[๐Ÿ“Š Data Sources] --> B(๐Ÿ”ง Data Engineering ETL)
    B --> C(โš™๏ธ Feature Engineering)
    C --> D[โœจ Prepared Features]
    
    style A fill:#e6f3ff,stroke:#0066cc,stroke-width:2px
    style B fill:#fff2e6,stroke:#cc6600,stroke-width:2px
    style C fill:#e6ffe6,stroke:#00cc00,stroke-width:2px
    style D fill:#f3e6ff,stroke:#6600cc,stroke-width:2px
  • GenAI (Unstructured Data): The focus shifts. Since we’re using massive internet-scale datasets, feature engineering is less of a thing. The model learns the features itself. The challenges are different:
%%{init: {'flowchart': {'nodeSpacing': 50, 'rankSpacing': 60}}}%%
graph TD
    A[๐ŸŒ Various Data Sources<br/>Internet, Books, Code] --> B(๐Ÿ•ท๏ธ Data Collection<br/>Web Scraping)
    B --> C[๐Ÿ“„ Collected Raw Data]
    C --> D(๐Ÿงฝ Data Cleaning)
    D --> E[โœจ Clean Data]
    
    style A fill:#e6f3ff,stroke:#0066cc,stroke-width:2px
    style B fill:#fff2e6,stroke:#cc6600,stroke-width:2px
    style C fill:#f0f0f0,stroke:#666,stroke-width:2px
    style D fill:#e6ffe6,stroke:#00cc00,stroke-width:2px
    style E fill:#f3e6ff,stroke:#6600cc,stroke-width:2px

Key Steps in GenAI Data Prep:

  1. Data Collection: How do you get 15 trillion tokens of text? You build scrapers for web pages (like the Common Crawl dataset), digitize books, and pull from sources like GitHub.

  2. Data Cleaning: This is arguably the most important, underrated step. The internet is a messy place. You need to:

    • Remove Harmful/Toxic/NSFW content.
    • Deduplicate: You don’t want the model to see the exact same text 10,000 times. It can bias the model and is an inefficient use of compute.
    • Filter Low-Quality Content: Remove boilerplate text, machine-translated garbage, etc. You might even use another model to assign a quality score.
    • Balance Data: Ensure you have a diverse mix of data types (prose, dialogue, code, different languages) to create a well-rounded model.
  3. Data Efficiency (Storage and Retrieval): When you have 50 Terabytes of data, you can’t just load it into memory. You need:

    • Efficient Storage: Use distributed file systems (HDFS, S3) and efficient formats (Parquet, ORC).
    • Efficient Retrieval: Use techniques like sharding (splitting data across machines) and indexing to quickly access the data needed for training batches.

A Hot Topic: Synthetic Data

  • Pros: Can improve diversity and scale up your dataset, especially for niche topics where real data is scarce.
  • Cons: Quality depends on the original model. You risk creating a model that just copies the biases and errors of its predecessor, or misses the complexity of the real world. This is a very active area of research.

Step 4: Model Development

This is the largest and most technical step. It covers three phases: Architecture, Training, and Sampling.

A. Model Architecture

Here you zoom in on your chosen model family. You need to talk about the specific components.

  • Example: U-Net for Diffusion Models If you chose a Diffusion model for image generation, the backbone is often a U-Net. You should be able to describe it.
%%{init: {'flowchart': {'nodeSpacing': 40, 'rankSpacing': 60}}}%%
graph TD
    Input[๐Ÿ–ผ๏ธ Image] --> D1(๐Ÿ“‰ Downsampling Block 1)
    D1 --> D2(๐Ÿ“‰ Downsampling Block 2)
    D2 --> D3(๐Ÿ“‰ Downsampling Block 3)
    
    D3 --> U1(๐Ÿ“ˆ Upsampling Block 1)
    U1 --> U2(๐Ÿ“ˆ Upsampling Block 2)
    U2 --> U3(๐Ÿ“ˆ Upsampling Block 3)
    U3 --> Output[๐Ÿ”ฎ Predicted Noise]
    
    D3 -.->|Skip Connection| U1
    D2 -.->|Skip Connection| U2
    D1 -.->|Skip Connection| U3

    style Input fill:#e6f3ff,stroke:#0066cc,stroke-width:2px
    style Output fill:#f3e6ff,stroke:#6600cc,stroke-width:2px
    style D1 fill:#fff2e6,stroke:#cc6600,stroke-width:2px
    style D2 fill:#fff2e6,stroke:#cc6600,stroke-width:2px
    style D3 fill:#fff2e6,stroke:#cc6600,stroke-width:2px
    style U1 fill:#e6ffe6,stroke:#00cc00,stroke-width:2px
    style U2 fill:#e6ffe6,stroke:#00cc00,stroke-width:2px
    style U3 fill:#e6ffe6,stroke:#00cc00,stroke-width:2px

Your explanation: “The U-Net architecture consists of an encoder (downsampling path) that captures contextual information, and a decoder (upsampling path) that reconstructs the image. The critical feature is the skip connections, which connect layers from the encoder directly to corresponding layers in the decoder. This allows the model to reuse low-level feature information (like fine textures and edges) during reconstruction, which is essential for generating sharp, detailed images.”

Deep Dive: Transformer’s Self-Attention (The heart of LLMs) This is the most important architectural concept in modern GenAI. You MUST understand it intuitively.

  1. The Goal: For any given word in a sentence, we want to understand how it relates to all other words in that sentence to get its true contextual meaning. The word “bank” means something different in “river bank” vs. “money bank”.

  2. The Q, K, V Analogy:

    • Query (Q): From the perspective of the current word, this is a question: “What am I, and what context do I need?”
    • Key (K): From the perspective of every other word, this is a label: “Here’s the kind of information I have.”
    • Value (V): From the perspective of every other word, this is the actual content: “Here is my information.”
  3. The Mechanism (Scaled Dot-Product Attention):

    • You take the Query of your current word and compute the dot-product with the Key of every other word. This gives you a compatibility score.
    • You scale these scores (divide by the square root of the dimension) to keep gradients stable during training.
    • You run these scores through a Softmax function. This turns the scores into weights that sum to 1. It’s like distributing 100% of your “attention” across all other words.
    • You multiply these attention weights by the Value of each word and sum them up. The result is a new representation for your current word, blended with contextual information from all other words it paid attention to.

Here’s the Scaled Dot-Product Attention:

%%{init: {'flowchart': {'nodeSpacing': 40, 'rankSpacing': 50}}}%%
graph TD
    Q[๐Ÿค” Query] --> MatMul1[ร— Matrix Multiply]
    K[๐Ÿ”‘ Key] --> MatMul1
    MatMul1 --> Scale[๐Ÿ“ Scale]
    Scale --> Mask[๐ŸŽญ Mask Optional]
    Mask --> SoftMax[๐Ÿงฎ SoftMax]
    SoftMax --> MatMul2[ร— Matrix Multiply]
    V[๐Ÿ’Ž Value] --> MatMul2
    MatMul2 --> Output[โœจ Attention Output]
    
    style Q fill:#e6f3ff,stroke:#0066cc,stroke-width:2px
    style K fill:#fff2e6,stroke:#cc6600,stroke-width:2px
    style V fill:#e6ffe6,stroke:#00cc00,stroke-width:2px
    style Output fill:#f3e6ff,stroke:#6600cc,stroke-width:2px
  1. Multi-Head Attention:
    • The Problem: Just one set of Q, K, V matrices might only learn one type of relationship (e.g., just grammatical relationships).
    • The Solution: Do the whole attention process multiple times in parallel, each with different, learned Q, K, V weight matrices. Each “head” can specialize in learning a different type of relationship (e.g., one head for semantic meaning, one for syntax, one for long-distance dependencies).
    • You then concatenate the outputs from all heads and pass them through a final linear layer to combine the knowledge.

B. Model Training

Once you have the architecture, you need to train it.

  • Training Methodology:

    • This is model-specific. For Diffusion, it’s a process of adding noise and then training the U-Net to predict and remove that noise. For GANs, it’s an adversarial process between a generator and a discriminator.
    • For LLMs, it’s often a multi-stage process:
      1. Pre-training: On massive, general data (the internet). The goal is to learn language, facts, and reasoning. The objective is often next-token prediction.
      2. Supervised Fine-Tuning (SFT): On a smaller, high-quality dataset of instruction-response pairs to teach the model how to follow commands.
      3. Alignment (e.g., RLHF): Using reinforcement learning from human feedback to make the model more helpful, harmless, and honest.
  • ML Objective and Loss Function:

    • The objective is what you want the model to do (e.g., “predict the next token”).
    • The loss function is the mathematical formula that measures how far the model’s prediction is from the truth. The entire training process is about minimizing this loss.

This covers the essence of the first half of the chapter. We’ve established the landscape and have a clear, structured plan to tackle the design.


5. Deep Dive into Model Training: The Heavy Lifting

The book talks about techniques for training large-scale models. These are your bread and butter as a senior engineer at FAANG.

  • What are we trying to achieve? We’re trying to train a model that is too big to fit in one GPU’s memory and on a dataset that is too large to process quickly on one machine. We need to be efficient with memory, time, and money.

Here are the key optimization techniques:

  1. Gradient Checkpointing:

    • Intuition: During training, you need to store intermediate values (activations) to calculate the gradients for backpropagation. This consumes a ton of memory. Gradient checkpointing is a clever trade-off: it throws away most of these intermediate values to save memory, and then re-computes them on the fly during the backward pass.
    • Trade-off: You use less memory, but it increases your training time (more compute). It’s perfect for when you want to train a massive model on GPUs that don’t have enough VRAM.
  2. Mixed Precision Training:

    • Intuition: By default, calculations are done in 32-bit floating point (FP32). But do we need that much precision for everything? Mixed precision uses faster, less memory-intensive 16-bit floats (FP16) for most of the calculations, while keeping critical parts (like weight updates) in FP32 to maintain stability.
    • Benefit: On modern GPUs with Tensor Cores, this can speed up training by 2-3x and cut memory usage in half. Frameworks like PyTorch (with torch.amp) make this almost automatic.
  3. Distributed Training (The Teamwork): This is how you train on hundreds or thousands of GPUs at once. The book correctly identifies the main types of parallelism.

%%{init: {'flowchart': {'nodeSpacing': 50, 'rankSpacing': 80}}}%%
graph LR
    P[Parallelism] --> DP[Data Parallelism]
    P --> MP[Model Parallelism]
    
    DP --> DP_Desc[Split the DATA<br/>Each GPU gets a full<br/>copy of the MODEL]
    
    MP --> PP[Pipeline Parallelism<br/>Inter-layer]
    MP --> TP[Tensor Parallelism<br/>Intra-layer]
    
    PP --> PP_Desc[Split the MODEL LAYERS<br/>GPU_0: Layers 1-8<br/>GPU_1: Layers 9-16]
    TP --> TP_Desc[Split SINGLE LAYER ops<br/>GPU_0: half matrix<br/>GPU_1: other half]

    style P fill:#f9f9f9,stroke:#333,stroke-width:3px
    style DP fill:#e8f4fd,stroke:#1f77b4,stroke-width:2px
    style MP fill:#fff2e8,stroke:#ff7f0e,stroke-width:2px
    style DP_Desc fill:#d5e8d4,stroke:#82b366,stroke-width:2px
    style PP_Desc fill:#e1d5e7,stroke:#9673a6,stroke-width:2px
    style TP_Desc fill:#e1d5e7,stroke:#9673a6,stroke-width:2px
  • Data Parallelism: The simplest and most common. You have a giant dataset. You give a small chunk (a mini-batch) to each GPU. Each GPU calculates the gradients for its batch, and then they all sync up their results via a Parameter Server or an all-reduce operation.
  • Model Parallelism: You use this when the model itself is too big for a single GPU.
    • Pipeline Parallelism: You put different layers of the model on different GPUs. GPU 0 computes layers 1-8 and passes its output to GPU 1, which computes layers 9-16, etc. It’s like an assembly line. The main challenge is keeping all GPUs busy (the “pipeline bubble”).
    • Tensor Parallelism: You use this when a single layer is too big. You split the actual matrix multiplications of that layer across multiple GPUs. This is more complex but essential for the massive layers in models like GPT.

A state-of-the-art training setup (like what’s used for Llama 3) uses a hybrid approach, combining all of these techniques (Data, Pipeline, and Tensor parallelism) to work efficiently at massive scale. This is what frameworks like FSDP (Fully Sharded Data Parallel) help manage.

C. Model Sampling (Inference)

After the model is trained, how do you generate output? This is sampling.

  • Greedy Search: At each step, pick the single most likely next word/pixel. Fast, but boring, repetitive, and deterministic.
  • Beam Search: Keep track of the k most likely sequences at each step. Better than greedy, but can still suppress creativity.
  • Stochastic Sampling (Top-k, Top-p): This is what’s used in modern chatbots.
    • Top-k Sampling: Consider only the k most likely next words, and then sample from that smaller set.
    • Top-p (Nucleus) Sampling: Consider the smallest set of words whose cumulative probability is greater than p. This is adaptive; if the model is very certain, the set is small. If it’s uncertain, the set is larger. This often gives the best balance of coherence and creativity.

Step 5: Evaluation

How do you know if your model is any good?

  • Offline Evaluation (Using a test dataset):

    • Discriminative Metrics: Easy. Accuracy, Precision, Recall, F1 score. You have a ground truth.
    • Generative Metrics (Harder!): There’s no single “right” answer.
      • Text: BLEU, ROUGE (compare generated text to references), Perplexity (how surprised is the model by the text).
      • Images: FID, KID (measure how similar the distribution of generated images is to real images), CLIPScore (measures how well an image matches a text prompt).
    • The Interview Point: You need to show you understand that for generative models, no single metric is enough. You need a suite of metrics to measure different aspects: quality, diversity, alignment to the prompt, etc. And ultimately…
  • Online Evaluation (In production, with real users):

    • This is about business impact.
    • Metrics: Click-Through Rate (CTR), Conversion Rate, User Retention, Latency, User Satisfaction surveys.
    • You’ll often use A/B testing to compare a new model against the old one on these business metrics.
  • Human Evaluation: For creative tasks, you can’t escape it. You need human raters to score outputs on dimensions like “creativity,” “coherence,” “factuality,” and “harmlessness.” This is a core part of the RLHF process.

The Golden Rule: Online and offline metrics don’t always align. A model might get a great ROUGE score but generate text that users find unhelpful. The ultimate test is how it performs with real users and impacts your business goals.

Step 6: Overall ML System Design

Now, zoom out from just the model and draw the whole system. This is where you integrate everything. You need to think about more than just the model.predict() call.

  • System Components: A request comes in. What happens?
    1. Input Preprocessing: The raw prompt is cleaned.
    2. Safety & Moderation (Pre-filter): Check the input prompt for harmful content.
    3. Core Model Inference: Call your generative model.
    4. Post-processing: Convert model output to a user-friendly format.
    5. Safety & Moderation (Post-filter): Check the generated output for harmful content, bias, or PII leakage.
    6. User Feedback Loop: Log the output, the prompt, and any user feedback (thumbs up/down) for future retraining.
  • Scalability: How do you serve 100 million users? You’ll have a load balancer distributing requests to a fleet of inference servers. You might use model parallelism for inference if the model is huge.
  • Security & Bias: How do you prevent misuse? How do you detect and filter biased outputs? This is a huge area and showing awareness is critical.

Here is a simple, intuitive diagram of what a production GenAI system looks like.

%%{init: {'flowchart': {'nodeSpacing': 60, 'rankSpacing': 100, 'curve': 'basis'}}}%%
graph TD
    User[๐Ÿ‘ค User] -->|1. Request| API[๐ŸŒ API Gateway /<br/>Load Balancer]
    API -->|2. Forward| Pre[๐Ÿ›ก๏ธ Preprocessing &<br/>Safety]
    Pre -->|3. Cleaned Prompt| Inference[๐Ÿง  Model Inference<br/>Cluster GPUs]
    Inference -->|4. Raw Output| Post[๐Ÿ›ก๏ธ Postprocessing &<br/>Safety]
    Post -->|5. Final Response| API
    API -->|6. Generated Image| User

    Pre -.->|Check violations| P_Filter[๐Ÿšซ Policy Filter]
    Post -.->|Scan output| O_Filter[๐Ÿšซ Output Filter]
    
    User -.->|๐Ÿ‘๐Ÿ‘Ž Feedback| Feedback[๐Ÿ”„ Feedback Loop<br/>RLHF]
    Feedback -.-> Retraining[๐Ÿ”ง Model Retraining<br/>Pipeline]
    Retraining -.-> Inference

    style User fill:#e6f3ff,stroke:#0066cc,stroke-width:3px
    style API fill:#fff2e6,stroke:#cc6600,stroke-width:2px
    style Inference fill:#d5e8d4,stroke:#82b366,stroke-width:3px
    style Pre fill:#f8cecc,stroke:#b85450,stroke-width:2px
    style Post fill:#f8cecc,stroke:#b85450,stroke-width:2px
    style P_Filter fill:#fff2cc,stroke:#d6b656,stroke-width:2px
    style O_Filter fill:#fff2cc,stroke:#d6b656,stroke-width:2px
    style Feedback fill:#e1d5e7,stroke:#9673a6,stroke-width:2px
    style Retraining fill:#e1d5e7,stroke:#9673a6,stroke-width:2px

Key talking points for this diagram:

  • System Components: It’s not just the model. There are safety filters before and after the model, load balancers, caching layers, and monitoring services.
  • Safety Mechanisms: You absolutely must talk about this. How do you prevent users from generating harmful content? You need input filters (prompt filtering) and output filters (content moderation classifiers).
  • User Feedback & Continuous Learning (RLHF): The thumbs up/down buttons aren’t just for show. That data is collected and used to continuously fine-tune the model to better align with user preferences. This is a critical feedback loop.
  • Scalability: How do you serve millions of users? You use load balancers to distribute traffic to a cluster of GPU machines for inference. You use techniques like model parallelism within that cluster if the model is huge.
  • Monitoring: The final step. You log everything. Latency, error rates, GPU utilization, metric scores from your safety classifiers. If something breaks, you need to know immediately.

Step 7: Deployment and Monitoring

  • Deployment: How do you get the model into production? You might have a CI/CD pipeline for models.
  • Monitoring: What are you tracking?
    • System Metrics: Latency, error rates, GPU utilization.
    • Model Metrics: Monitor the distribution of inputs and outputs. Is there a “drift” over time? Is the model’s performance degrading? This is crucial for knowing when you need to retrain.

Summary

This framework is your bible for a GenAI system design interview. Start with the user and the business problem, frame it as an ML task, go deep on the data and model, evaluate your results, design the end-to-end production system, and think about what happens after it’s live.

By walking through this chapter, you’ve built a powerful mental model for GenAI system design:

  1. Frame the Problem: Start broad (AI vs. ML), then narrow down (Discriminative vs. Generative). Understand the “why” (Data, Model, Compute).
  2. Follow the Roadmap: Use the 7-step framework as your guide. It prevents you from getting lost and ensures you cover all your bases.
  3. Think Like an Engineer, Not Just a Scientist: Don’t just talk about the model. Talk about the data pipelines, the training optimizations (parallelism!), the evaluation metrics (offline AND online), and the full production system with its safety and feedback loops.
  4. Know Your Architectures: Be able to explain U-Net for diffusion models and self-attention for transformers at an intuitive level.
  5. Understand the Economics: Scaling laws help you make data-driven decisions about compute budgets.
  6. Safety and Ethics First: Always discuss content moderation, bias detection, and user feedback loops.

If you can walk an interviewer through these steps, providing the “why” behind your choices and discussing the trade-offs at each stage, you’re not just answering the questionโ€”you’re demonstrating the strategic thinking of a senior staff engineer at a top company.

This approach will make you appear structured, thorough, and deeply knowledgeableโ€”exactly what a FAANG interviewer is looking for.