\

Chapter 2: Gmail Smart Compose

15 min read

Excellent. Let’s dive into Chapter 6. This is one of the most important, practical patterns in the GenAI space right now: Retrieval-Augmented Generation (RAG).

If the last chapter was about the fundamentals of GenAI models, this chapter is about making them useful in the real world. A base LLM is like a brilliant student who has read the entire internet up to 2023 but has never seen your company’s private documents and doesn’t know what happened in the news this morning. RAG is the system you build to give that student an “open-book exam,” allowing them to access specific, up-to-date information to answer questions.

Let’s break this down, piece by piece.


Chapter 6: Retrieval-Augmented Generation

Introduction

What are we trying to achieve? We’re solving the LLM’s biggest weaknesses:

  1. Knowledge Cutoff: It doesn’t know about recent events.
  2. Lack of Private Data: It hasn’t been trained on your internal company wiki, your customer support database, or a PDF you just uploaded.
  3. Hallucination: It can make things up.

The book uses the example of building a ChatPDF system for internal company use. An employee should be able to ask, “What is our policy on international travel reimbursement?” and get an answer based on the latest HR documents, not on some generic policy the LLM learned from the public internet.

Step 1: Clarifying Requirements

This is where the interview starts. The book gives a fantastic example of a candidate leading the conversation. Let’s analyze it from an interviewer’s perspective.

  • Candidate: “What does the external knowledge base consist of? Does it change over time?”
    • Interviewer’s thought: Good. They’re starting with the data. They understand that the nature of the data source is the most important factor.
  • Candidate: “Do the Wiki pages and forums contain text, images, and other modalities?”
    • Interviewer’s thought: Excellent. They’re thinking about multimodality. This will affect our choice of embedding models.
  • Candidate: “How many pages are there in total?” (5 million pages) “What is the expected growth?” (20% annually)
    • Interviewer’s thought: Great, they’re quantifying the scale. This is critical for discussing scalability, cost, and choosing the right database/indexing strategy.
  • Candidate: “Should the system respond in real time?” (Slight delay is okay)
    • Interviewer’s thought: They’re scoping the latency requirements. This tells me I don’t need a sub-50ms system and can make trade-offs for better quality.
  • Candidate: “Is it necessary for the system to include document references?” (Yes)
    • Interviewer’s thought: Crucial question. This requirement immediately makes one of the potential solutions (finetuning) much less attractive. They are already thinking ahead.

By the end of this, you’ve established the core problem: Build a Q&A system over a large (5M pages), slowly growing (+20%/year) internal knowledge base of mixed-format PDFs, which must provide verifiable answers with source references.

Step 2: Frame the Problem as an ML Task

Specifying Input and Output

This is straightforward but important to state clearly.

  • Input: A user’s text query (e.g., “How do I submit an expense report?”).
  • Underlying Data: A database of 5 million company documents.
  • Output: A text-based answer, grounded in the documents, with references.
graph LR
    subgraph User
        A[User Query<br/>"How do I submit an<br/>expense report?"]
    end
    subgraph System
        B(ChatPDF System)
    end
    subgraph Data
        C[Document databases]
    end
    subgraph Output
        D[Response<br/>"To submit an expense report, log into..."]
    end
    
    A --> B
    C --> B
    B --> D

(Based on Figure 6.2)

Choosing a Suitable ML Approach

This is the first major design decision, and it’s a classic interview trade-off question. For a problem like this, the book lays out three main approaches.

  1. Finetuning:

    • What it is: Take a pre-trained LLM and continue training it on your internal documents. The model’s weights are updated to “absorb” the new knowledge.
    • Pros: Can deeply learn the style and terminology of your company.
    • Cons (Dealbreakers for our problem):
      • Computationally Expensive: Continuously retraining an LLM is a massive cost.
      • Stale Data: As soon as a new document is added, the model is out of date until the next expensive finetuning cycle.
      • No References: The book correctly states: “Finetuned models usually can’t provide references for their answers, making it hard to verify or trace information back to its source.” The knowledge is baked into the weights; you can’t easily point to the source document. This violates our requirement.
  2. Prompt Engineering (In-Context Learning):

    • What it is: Stuff the relevant documents directly into the prompt along with the user’s question.
    • Pros: Simple, cheap, no training required.
    • Cons (Dealbreakers for our problem):
      • Limited Context Window: You can’t fit 5 million documents into a prompt. You can’t even fit one long document. This approach is simply not scalable.
  3. Retrieval-Augmented Generation (RAG):

    • What it is: A two-step process. First, retrieve a few relevant document snippets from the large database. Then, generate an answer using an LLM, with the user’s query and the retrieved snippets provided as context in the prompt.
    • Pros:
      • Access to Current Info: The document database can be updated easily. The LLM gets the latest info at query time.
      • Verifiable & Factual: Since you have the retrieved snippets, you can easily add references. It reduces hallucination by forcing the LLM to base its answer on the provided text.
      • Scalable & Cost-Effective: You’re not retraining the LLM. The main work is in the retrieval step.
    • Cons:
      • Implementation Complexity: It’s a multi-component system (retriever + generator) that needs to work well together.
      • Dependence on Retrieval Quality: If you retrieve irrelevant documents, the LLM will give a garbage answer. The retriever is critical.

The Decision: The book concludes, “RAG offers a balanced solution in terms of ease of setup, cost, and scalability… Therefore, we choose RAG to build our ChatPDF system.” This is the correct, well-justified choice.

Step 3: Data Preparation (The “R” in RAG)

This is the entire process of making your knowledge base searchable. The book outlines a three-step pipeline.

graph TD
    A[Document Databases (PDFs)] --> B(Document Parsing);
    B --> C(Document Chunking);
    C --> D(Indexing);
    D --> E[Indexed Embeddings<br/>(in Vector DB)];

(Simplified from Figure 6.8)

1. Document Parsing: Getting Content out of PDFs

PDFs are a nightmare. They can have columns, tables, images, and weird layouts.

  • Rule-based: You write code that assumes a certain layout. Brittle and fails on complex or varied documents.
  • AI-based (The Winner): Use a model to “read” the PDF like a human. As the book explains, a tool like Layout-Parser does this:
    1. Layout Detection: An object detection model draws boxes around paragraphs, tables, images, etc.
    2. Text Extraction: OCR is run inside each text box.
    3. Structured Output: You get a list of content blocks with their type (text, image), coordinates, and content. This is much more useful than just a wall of raw text.

2. Document Chunking: Breaking It Down

You can’t create an embedding for an entire 50-page document.

  • Why?
    1. An embedding of a whole document averages out the meaning and loses specific details. A query for a specific sentence will get lost.
    2. The retrieved document chunk needs to fit into the LLM’s context window.
  • Strategies:
    • Length-based: Simple but dumb. Can cut sentences in half. RecursiveCharacterTextSplitter from libraries like LangChain is a smarter version that tries to split on paragraphs, then sentences, then words to keep semantic units together.
    • Content-aware: Better. Split based on document structure, like Markdown headers (#), HTML tags, etc.

3. Indexing: Making Chunks Findable

Now you have thousands or millions of chunks. How do you find the right ones for a given query, fast?

  • Keyword/Full-text Search (e.g., Elasticsearch): Fast and good for matching exact words or phrases. But it struggles with synonyms and semantic meaning. A query for “employee compensation” might miss a document that only uses the word “staff salary.”
  • Vector-based Search (The Winner):
    • Why? It searches based on semantic meaning, not keywords.
    • How? You use an embedding model (like a Transformer encoder) to convert every chunk of text/image into a high-dimensional vector (an array of numbers). Chunks with similar meanings will have vectors that are “close” to each other in this vector space.
    • The book does a back-of-the-envelope calculation: 5M pages, chunked, results in ~40M chunks. At this scale, vector search is the only viable option for semantic retrieval.

So, the data prep pipeline is: Parse PDFs into structured blocks -> Chunk blocks into small, meaningful pieces -> Embed each piece into a vector -> Store the vectors in a specialized Vector Database.

Step 4: Model Development

This covers the architecture of the ML models and the processes involved.

Architecture: The Key ML Models (Figure 6.9)

A RAG system has three key ML model components:

  1. Indexing: An Encoder Model to create the vector embeddings.
  2. Retrieval: The same Encoder Model to turn the user query into a vector. (Plus an ANN search algorithm).
  3. Generation: An LLM to generate the final answer.

The Indexing/Retrieval Model: Text-Image Alignment

Our documents have text and images. Our query is text. How do you find an image relevant to a text query? The embeddings need to live in the same “space.” The book outlines two approaches (Figure 6.10):

  1. Shared Embedding Space (Best Approach): Use a multimodal model like CLIP. CLIP is pre-trained to map related images and their text descriptions to nearby points in vector space. You can use its text encoder for text chunks and its image encoder for images. This is elegant and powerful.
  2. Image Captioning (Workaround): Use an image captioning model to generate a text description for each image. Then, use a standard text-only encoder to embed that caption. This works but is less direct and might lose information.

The Retrieval Process: Finding the Needles in the Haystack

Once the user query is embedded into a vector Eq, we need to find the k closest chunk vectors in our database of 40 million.

  • Exact Nearest Neighbor (Linear Search): Compare Eq to all 40 million vectors. Guarantees a perfect result but is way too slow. O(N*D) complexity. Unacceptable.
  • Approximate Nearest Neighbor (ANN): The only practical solution. It trades a tiny bit of accuracy for a massive speedup. The book mentions four families of ANN algorithms:
    • Tree-based (e.g., Annoy): Recursively partition the data space. Fast, but can struggle in very high dimensions.
    • Hashing-based (LSH): Uses clever hash functions where similar vectors are likely to get the same hash key. You only search within the query’s hash bucket.
    • Clustering-based (The book’s choice): Pre-cluster the 40M vectors into, say, 100,000 clusters. The search becomes a two-step process:
      1. Find the few clusters whose center is closest to the query vector.
      2. Do an exact search only within those few clusters. This massively reduces the search space.
    • Graph-based (e.g., HNSW): The state-of-the-art for many use cases. It builds a graph where nodes are data points and edges connect close neighbors. Searching is like navigating this graph to find the closest point.

Modern vector databases like Pinecone, Weaviate, or libraries like FAISS (from Meta) and ScaNN (from Google) implement these advanced ANN algorithms for you.

Here is the overall retrieval process from Figure 6.16:

graph TD
    subgraph Data Preparation
        direction TB
        A[Document databases] --> B(Data preparation<br>Parsing/Chunking);
        B --> C[Index (images)];
        B --> D[Index (text)];
        C & D --> E(Clustering);
    end

    subgraph Retrieval
        direction LR
        F[User query<br>"How many cats live<br>in the company?"] --> G(Text Encoder);
        G --> H(Inter-Cluster<br>Search);
        E -- Selected Clusters --> H;
        H --> I(Intra-Cluster<br>Search);
        I --> J[Retrieved<br>data chunks];
    end

The Generation Process: Crafting the Final Answer

This is where the LLM comes in. The process is:

  1. Take the original user query.
  2. Take the top k retrieved data chunks from the retrieval step.
  3. Combine them into a single, well-structured prompt.
  4. Feed this prompt to the LLM.
  5. The LLM generates the answer, using a sampling strategy like top-p sampling for a good balance of correctness and fluency.

A Deeper Look at Prompt Engineering for Generation

How do you structure that final prompt for the best results? The book dives into several powerful techniques.

  • Chain-of-Thought (CoT) Prompting: Instruct the model to “think step-by-step.” This forces it to lay out its reasoning process before giving the final answer, which often improves accuracy on complex questions.
  • Few-shot Prompting: Give the model 2-3 examples of a Q&A pair in the desired format before giving it the real query. This helps it understand the expected tone and structure.
  • Role-specific Prompting: Tell the model who it is. “You are an expert contract lawyer… explain this clause in simple terms.” This grounds the model and helps it adopt the correct persona and level of detail.
  • User-context Prompting: Include metadata about the user (language, location, etc.) to get more personalized results.

Here is the final prompt structure from Figure 6.22, showing how all these pieces come together.

graph TD
    subgraph "Final Prompt Sent to LLM"
        A["User's initial query: [...]"];
        B["--- RETRIEVED CONTEXT ---<br/>[Retrieved document chunk 1]<br/>[Retrieved document chunk 2]<br/>..."];
        C["--- INSTRUCTIONS ---<br/>You are a helpful assistant for Company XYZ..."];
        D["--- EXAMPLES ---<br/>Example 1: Q: ... A: ...<br/>Example 2: Q: ... A: ..."];
        E["--- REASONING ---<br/>Based on the context, think step-by-step to answer the user's query."];
        F["--- USER INFO ---<br/>User language: English"];
    end

    subgraph Labels
        direction LR
        L1[Role-Specific<br>Prompting]
        L2[Few-Shot<br>Prompting]
        L3[Chain-of-Thought<br>(CoT)]
        L4[User-Context<br>Prompting]
        L5[Retrieved<br>Context]
    end

    B -- Is --> L5
    C -- Is --> L1
    D -- Is --> L2
    E -- Is --> L3
    F -- Is --> L4

Training (Advanced Topic: RAFT) Most of the time, you start with pre-trained models. But what if your retrieval is noisy and the LLM struggles to distinguish good context from bad? RAFT (Retrieval-Augmented Fine-Tuning) is a technique to solve this.

  • The Idea: During finetuning, you create training examples that include the question, a “golden” (correct) document, AND several “distractor” (irrelevant) documents that were also retrieved.
  • The Goal: You train the LLM to specifically pay attention to the golden document and ignore the distractors when generating the answer. This makes the LLM more robust to imperfect retrieval.

Step 5: Evaluation (CRITICAL for RAG)

Evaluating a RAG system is more complex than a standard LLM. You need to evaluate both the retriever and the generator. The book introduces an excellent “Triad of RAG evaluation” (Figure 6.23).

graph TD
    Query -->|Context Relevance| Context
    Context -->|Faithfulness| Results
    Query -->|Answer Relevance<br>Answer Correctness| Results

This diagram shows that the final Results (the generated answer) depend on the Query, the retrieved Context, and the relationships between them. This leads to four key evaluation aspects:

  1. Context Relevance: Is the retriever working? Did we retrieve documents that are relevant to the query?

    • Metrics: Standard information retrieval metrics like Precision@k, nDCG, Hit Rate. You need a labeled dataset of (query, relevant_doc) pairs for this.
  2. Faithfulness (or Groundedness): Is the generator hallucinating? Is the generated answer factually consistent with the retrieved context? You check if every statement in the answer can be backed up by the provided snippets.

    • Methods: This is hard to automate. Often requires human evaluation or using another powerful LLM as a judge. Figure 6.24 shows a great example: if the context says Marie Curie won two Nobel prizes, an answer saying she won one has low faithfulness.
  3. Answer Relevance: Did the generator answer the user’s actual question? The retrieved context might be relevant, but the LLM could get sidetracked and generate an answer that doesn’t directly address the user’s intent.

    • Methods: Again, often requires a human or an LLM judge to compare the user’s query and the final answer.
  4. Answer Correctness: Is the answer factually correct according to a ground truth reference? This is the classic accuracy measure.

In an interview, discussing this four-part evaluation framework shows a deep, practical understanding of the challenges of building reliable RAG systems.

Step 6: Overall ML System Design

This is the final blueprint. Figure 6.27 shows the end-to-end flow. Let’s recreate and walk through it.

graph TD
    subgraph Offline Process
        direction TB
        A[Document Databases] --> B(Document Parsing & Chunking)
        B --> C{Text Encoder};
        B --> C_img{Image Encoder};
        C --> D[Index (text)];
        C_img --> D_img[Index (images)];
    end
    
    subgraph Online / Inference Process
        direction LR
        E[User Query] --> F(Safety Filtering);
        F --> G(Query Expansion);
        G --> H{Text Encoder};
        H --> I(Nearest Neighbor Search);
        D & D_img --> I;
        I --> J(Prompt Engineering);
        E --> J;
        J --> K[LLM];
        K --> L(Safety Filtering);
        L --> M[Response];
    end
    
    %% Grouping for clarity
    subgraph Indexing Process
      A;B;C;C_img;D;D_img
    end
    subgraph Retrieval
      G;H;I;
    end
    subgraph Generation
      J;K;L
    end

A user query’s journey:

  1. Offline: A pipeline runs periodically to Parse, Chunk, and Index all 5 million documents into a vector database. This is the Indexing Process.
  2. Online: A user sends a query.
  3. Safety Filtering: The query is checked for harmful content.
  4. Query Expansion (Optional but good): The query is expanded with synonyms or rephrased to improve retrieval. “How much do I get for trips?” -> “travel reimbursement policy allowance”.
  5. Retrieval: The query is encoded into a vector, and an ANN search is performed on the vector DB to get the top-k chunks.
  6. Generation: The user’s query and the retrieved chunks are assembled into a prompt using techniques like CoT and role-prompting. This is fed to the LLM.
  7. Safety Filtering: The LLM’s response is checked for safety, PII, etc.
  8. The final, safe, and grounded response is sent to the user.

This diagram is your high-level design for the interview. Being able to draw this and explain each component’s purpose and the trade-offs involved is an A+ answer.