Chapter 6: Retrieval-Augmented Generation
38 min readWhat are we trying to achieve? We’re solving the LLM’s biggest weaknesses:
- Knowledge Cutoff: It doesn’t know about recent events.
- Lack of Private Data: It hasn’t been trained on your internal company wiki, your customer support database, or a PDF you just uploaded.
- Hallucination: It can make things up.
The book uses the example of building a ChatPDF system for internal company use. An employee should be able to ask, “What is our policy on international travel reimbursement?” and get an answer based on the latest HR documents, not on some generic policy the LLM learned from the public internet.
Step 1: Clarifying Requirements
This is where the interview starts. The book gives a fantastic example of a candidate leading the conversation. Let’s analyze it from an interviewer’s perspective.
- Candidate: “What does the external knowledge base consist of? Does it change over time?”
- Interviewer’s thought: Good. They’re starting with the data. They understand that the nature of the data source is the most important factor.
- Candidate: “Do the Wiki pages and forums contain text, images, and other modalities?”
- Interviewer’s thought: Excellent. They’re thinking about multimodality. This will affect our choice of embedding models.
- Candidate: “How many pages are there in total?” (5 million pages) “What is the expected growth?” (20% annually)
- Interviewer’s thought: Great, they’re quantifying the scale. This is critical for discussing scalability, cost, and choosing the right database/indexing strategy.
- Candidate: “Should the system respond in real time?” (Slight delay is okay)
- Interviewer’s thought: They’re scoping the latency requirements. This tells me I don’t need a sub-50ms system and can make trade-offs for better quality.
- Candidate: “Is it necessary for the system to include document references?” (Yes)
- Interviewer’s thought: Crucial question. This requirement immediately makes one of the potential solutions (finetuning) much less attractive. They are already thinking ahead.
By the end of this, you’ve established the core problem: Build a Q&A system over a large (5M pages), slowly growing (+20%/year) internal knowledge base of mixed-format PDFs, which must provide verifiable answers with source references.
Step 2: Frame the Problem as an ML Task
Specifying Input and Output
This is straightforward but important to state clearly.
- Input: A user’s text query (e.g., “How do I submit an expense report?”).
- Underlying Data: A database of 5 million company documents.
- Output: A text-based answer, grounded in the documents, with references.
graph LR
subgraph User
A[User Query<br/>How do I submit an<br/>expense report?]
end
subgraph System
B(ChatPDF System)
end
subgraph Data
C[Document databases]
end
subgraph Output
D[Response<br/>To submit an expense report, log into...]
end
A --> B
C --> B
B --> D
(Based on Figure 6.2)
Choosing a Suitable ML Approach
This is the first major design decision, and it’s a classic interview trade-off question. For a problem like this, the book lays out three main approaches.
Finetuning:
- What it is: Take a pre-trained LLM and continue training it on your internal documents. The model’s weights are updated to “absorb” the new knowledge.
- Pros: Can deeply learn the style and terminology of your company.
- Cons (Dealbreakers for our problem):
- Computationally Expensive: Continuously retraining an LLM is a massive cost.
- Stale Data: As soon as a new document is added, the model is out of date until the next expensive finetuning cycle.
- No References: The book correctly states: “Finetuned models usually can’t provide references for their answers, making it hard to verify or trace information back to its source.” The knowledge is baked into the weights; you can’t easily point to the source document. This violates our requirement.
Prompt Engineering (In-Context Learning):
- What it is: Stuff the relevant documents directly into the prompt along with the user’s question.
- Pros: Simple, cheap, no training required.
- Cons (Dealbreakers for our problem):
- Limited Context Window: You can’t fit 5 million documents into a prompt. You can’t even fit one long document. This approach is simply not scalable.
Retrieval-Augmented Generation (RAG):
- What it is: A two-step process. First, retrieve a few relevant document snippets from the large database. Then, generate an answer using an LLM, with the user’s query and the retrieved snippets provided as context in the prompt.
- Pros:
- Access to Current Info: The document database can be updated easily. The LLM gets the latest info at query time.
- Verifiable & Factual: Since you have the retrieved snippets, you can easily add references. It reduces hallucination by forcing the LLM to base its answer on the provided text.
- Scalable & Cost-Effective: You’re not retraining the LLM. The main work is in the retrieval step.
- Cons:
- Implementation Complexity: It’s a multi-component system (retriever + generator) that needs to work well together.
- Dependence on Retrieval Quality: If you retrieve irrelevant documents, the LLM will give a garbage answer. The retriever is critical.
The Decision: The book concludes, “RAG offers a balanced solution in terms of ease of setup, cost, and scalability… Therefore, we choose RAG to build our ChatPDF system.” This is the correct, well-justified choice.
Step 3: Data Preparation (The “R” in RAG)
This is the entire process of making your knowledge base searchable. The book outlines a three-step pipeline.
graph TD
A[Document Databases - PDFs] --> B(UnstructuredPDFLoader<br/>Parsing + OCR)
B --> C(Intelligent Chunking<br/>Structure-aware)
C --> D(Vector Embedding)
D --> E[Indexed Embeddings<br/>in Vector DB]
(Enhanced pipeline using UnstructuredPDFLoader)
1. Document Parsing: Getting Content out of PDFs
PDFs are a nightmare. They can have columns, tables, images, and weird layouts.
- Rule-based: You write code that assumes a certain layout. Brittle and fails on complex or varied documents.
- AI-based (The Winner): Use a comprehensive document processing tool like UnstructuredPDFLoader from
langchain_community.document_loaders. This approach offers several key advantages:- Automatic OCR: Built-in OCR capabilities that can handle scanned PDFs, images, and mixed content documents automatically.
- Layout Detection: Advanced algorithms that understand document structure, identifying paragraphs, tables, headers, lists, and other elements.
- Multi-modal Processing: Handles text, images, tables, and complex layouts in a unified way.
- Structured Output: Returns well-structured document chunks with metadata about element types and hierarchy.
- Robust Handling: Works reliably across different PDF formats, including scanned documents, forms, and complex layouts that would break simpler parsers.
2. Document Chunking: Breaking It Down
You can’t create an embedding for an entire 50-page document.
- Why?
- An embedding of a whole document averages out the meaning and loses specific details. A query for a specific sentence will get lost.
- The retrieved document chunk needs to fit into the LLM’s context window.
- Strategies:
- Length-based: Simple but dumb. Can cut sentences in half. Traditional approaches like
RecursiveCharacterTextSplittertry to split on paragraphs, then sentences, but still lack true semantic understanding. - Content-aware (Best with UnstructuredPDFLoader): UnstructuredPDFLoader excels here by providing intelligent, structure-aware chunking:
- Semantic Chunking: Automatically identifies and preserves document structure (headers, paragraphs, lists, tables) as natural chunk boundaries.
- Element-based Splitting: Creates chunks based on document elements rather than arbitrary character counts, preserving context and meaning.
- Metadata Preservation: Each chunk includes rich metadata about its position, type, and relationship to other elements.
- Configurable Chunking: Allows fine-tuning of chunk sizes while respecting document structure boundaries.
- OCR Integration: For scanned documents, combines OCR text extraction with intelligent chunking in a single step.
- Length-based: Simple but dumb. Can cut sentences in half. Traditional approaches like
3. Indexing: Making Chunks Findable
Now you have thousands or millions of chunks. How do you find the right ones for a given query, fast?
- Keyword/Full-text Search (e.g., Elasticsearch): Fast and good for matching exact words or phrases. But it struggles with synonyms and semantic meaning. A query for “employee compensation” might miss a document that only uses the word “staff salary.”
- Vector-based Search (The Winner):
- Why? It searches based on semantic meaning, not keywords.
- How? You use an embedding model (like a Transformer encoder) to convert every chunk of text/image into a high-dimensional vector (an array of numbers). Chunks with similar meanings will have vectors that are “close” to each other in this vector space.
- The book does a back-of-the-envelope calculation: 5M pages, chunked, results in ~40M chunks. At this scale, vector search is the only viable option for semantic retrieval.
So, the data prep pipeline is: Parse & Chunk PDFs with UnstructuredPDFLoader (including OCR for scanned documents) -> Extract semantically meaningful, structure-aware chunks -> Embed each chunk into a vector -> Store the vectors with rich metadata in a specialized Vector Database.
Step 4: Model Development
This covers the architecture of the ML models and the processes involved.
Architecture: The Key ML Models (Figure 6.9)
A RAG system has three key ML model components:
- Indexing: An Encoder Model to create the vector embeddings.
- Retrieval: The same Encoder Model to turn the user query into a vector. (Plus an ANN search algorithm).
- Generation: An LLM to generate the final answer.
The Indexing/Retrieval Model: Text-Image Alignment
Our documents have text and images. Our query is text. How do you find an image relevant to a text query? The embeddings need to live in the same “space.” The book outlines two approaches (Figure 6.10):
- Shared Embedding Space (Best Approach): Use a multimodal model like CLIP. CLIP is pre-trained to map related images and their text descriptions to nearby points in vector space. You can use its text encoder for text chunks and its image encoder for images. This is elegant and powerful.
- Image Captioning (Workaround): Use an image captioning model to generate a text description for each image. Then, use a standard text-only encoder to embed that caption. This works but is less direct and might lose information.
The Retrieval Process: Finding the Needles in the Haystack
Once the user query is embedded into a vector Eq, we need to find the k closest chunk vectors in our database of 40 million.
- Exact Nearest Neighbor (Linear Search): Compare
Eqto all 40 million vectors. Guarantees a perfect result but is way too slow. O(N*D) complexity. Unacceptable. - Approximate Nearest Neighbor (ANN): The only practical solution. It trades a tiny bit of accuracy for a massive speedup. The book mentions four families of ANN algorithms:
- Tree-based (e.g., Annoy): Recursively partition the data space. Fast, but can struggle in very high dimensions.
- Hashing-based (LSH): Uses clever hash functions where similar vectors are likely to get the same hash key. You only search within the query’s hash bucket.
- Clustering-based (The book’s choice): Pre-cluster the 40M vectors into, say, 100,000 clusters. The search becomes a two-step process:
- Find the few clusters whose center is closest to the query vector.
- Do an exact search only within those few clusters. This massively reduces the search space.
- Graph-based (e.g., HNSW): The state-of-the-art for many use cases. It builds a graph where nodes are data points and edges connect close neighbors. Searching is like navigating this graph to find the closest point.
Modern vector databases like Pinecone, Weaviate, or libraries like FAISS (from Meta) and ScaNN (from Google) implement these advanced ANN algorithms for you.
Here is the overall retrieval process from Figure 6.16:
graph TD
subgraph Data Preparation
direction TB
A[Document databases] --> B(Data preparation<br/>Parsing/Chunking)
B --> C[Index images]
B --> D[Index text]
C --> E(Clustering)
D --> E
end
subgraph Retrieval
direction LR
F[User query<br/>How many cats live<br/>in the company?] --> G(Text Encoder)
G --> H(Inter-Cluster<br/>Search)
E -.Selected Clusters.-> H
H --> I(Intra-Cluster<br/>Search)
I --> J[Retrieved<br/>data chunks]
end
The Generation Process: Crafting the Final Answer
This is where the LLM comes in. The process is:
- Take the original user query.
- Take the top
kretrieved data chunks from the retrieval step. - Combine them into a single, well-structured prompt.
- Feed this prompt to the LLM.
- The LLM generates the answer, using a sampling strategy like top-p sampling for a good balance of correctness and fluency.
A Deeper Look at Prompt Engineering for Generation
How do you structure that final prompt for the best results? The book dives into several powerful techniques.
- Chain-of-Thought (CoT) Prompting: Instruct the model to “think step-by-step.” This forces it to lay out its reasoning process before giving the final answer, which often improves accuracy on complex questions.
- Few-shot Prompting: Give the model 2-3 examples of a Q&A pair in the desired format before giving it the real query. This helps it understand the expected tone and structure.
- Role-specific Prompting: Tell the model who it is. “You are an expert contract lawyer… explain this clause in simple terms.” This grounds the model and helps it adopt the correct persona and level of detail.
- User-context Prompting: Include metadata about the user (language, location, etc.) to get more personalized results.
Here is the final prompt structure from Figure 6.22, showing how all these pieces come together.
graph TD
subgraph "Final Prompt Components"
A["User's initial query"]
B["RETRIEVED CONTEXT<br/>Document chunks"]
C["INSTRUCTIONS<br/>System role"]
D["EXAMPLES<br/>Few-shot demos"]
E["REASONING<br/>CoT instruction"]
F["USER INFO<br/>Language/context"]
end
subgraph "Prompt Engineering Techniques"
L5[Retrieved Context]
L1[Role-Specific Prompting]
L2[Few-Shot Prompting]
L3[Chain-of-Thought]
L4[User-Context Prompting]
end
B -.-> L5
C -.-> L1
D -.-> L2
E -.-> L3
F -.-> L4
Training (Advanced Topic: RAFT)
Most of the time, you start with pre-trained models. But what if your retrieval is noisy and the LLM struggles to distinguish good context from bad? RAFT (Retrieval-Augmented Fine-Tuning) is a technique to solve this.
- The Idea: During finetuning, you create training examples that include the question, a “golden” (correct) document, AND several “distractor” (irrelevant) documents that were also retrieved.
- The Goal: You train the LLM to specifically pay attention to the golden document and ignore the distractors when generating the answer. This makes the LLM more robust to imperfect retrieval.
Step 5: Evaluation (CRITICAL for RAG)
Evaluating a RAG system is more complex than a standard LLM. You need to evaluate both the retriever and the generator. The book introduces an excellent “Triad of RAG evaluation” (Figure 6.23).
graph TD
Query -->|Context Relevance| Context
Context -->|Faithfulness| Results
Query -->|Answer Relevance & Correctness| Results
This diagram shows that the final Results (the generated answer) depend on the Query, the retrieved Context, and the relationships between them. This leads to four key evaluation aspects:
Context Relevance: Is the retriever working? Did we retrieve documents that are relevant to the query?
- Metrics: Standard information retrieval metrics like
Precision@k,nDCG,Hit Rate. You need a labeled dataset of (query, relevant_doc) pairs for this.
- Metrics: Standard information retrieval metrics like
Faithfulness (or Groundedness): Is the generator hallucinating? Is the generated answer factually consistent with the retrieved context? You check if every statement in the answer can be backed up by the provided snippets.
- Methods: This is hard to automate. Often requires human evaluation or using another powerful LLM as a judge. Figure 6.24 shows a great example: if the context says Marie Curie won two Nobel prizes, an answer saying she won one has low faithfulness.
Answer Relevance: Did the generator answer the user’s actual question? The retrieved context might be relevant, but the LLM could get sidetracked and generate an answer that doesn’t directly address the user’s intent.
- Methods: Again, often requires a human or an LLM judge to compare the user’s query and the final answer.
Answer Correctness: Is the answer factually correct according to a ground truth reference? This is the classic accuracy measure.
In an interview, discussing this four-part evaluation framework shows a deep, practical understanding of the challenges of building reliable RAG systems.
Step 6: Overall ML System Design
This is the final blueprint. Figure 6.27 shows the end-to-end flow. Let’s recreate and walk through it.
graph TD
subgraph "Offline Process"
direction TB
A[Document Databases] --> B(Document Parsing & Chunking)
B --> C(Text Encoder)
B --> C_img(Image Encoder)
C --> D[Text Index]
C_img --> D_img[Image Index]
end
subgraph "Online Process"
direction TB
E[User Query] --> F(Safety Filtering)
F --> G(Query Expansion)
G --> H(Text Encoder)
H --> I(Vector DB with Nearest Neighbor Search)
I --> J(Prompt Engineering)
J --> K[LLM]
K --> L(Safety Filtering)
L --> M[Response]
end
D --> I
D_img --> I
E --> J
A user query’s journey:
- Offline: A pipeline runs periodically to Parse, Chunk, and Index all 5 million documents into a vector database. This is the Indexing Process.
- Online: A user sends a query.
- Safety Filtering: The query is checked for harmful content.
- Query Expansion (Optional but good): The query is expanded with synonyms or rephrased to improve retrieval. “How much do I get for trips?” -> “travel reimbursement policy allowance”.
- Retrieval: The query is encoded into a vector, and an ANN search is performed on the vector DB to get the top-k chunks.
- Generation: The user’s query and the retrieved chunks are assembled into a prompt using techniques like CoT and role-prompting. This is fed to the LLM.
- Safety Filtering: The LLM’s response is checked for safety, PII, etc.
- The final, safe, and grounded response is sent to the user.
This diagram is your high-level design for the interview. Being able to draw this and explain each component’s purpose and the trade-offs involved is an A+ answer.
Of course. This is the perfect next step. A high-level design gets you hired, but a low-level design shows you’ve actually built these things. You’re demonstrating that you’ve grappled with the real-world trade-offs between specific libraries, algorithms, and cloud services.
Let’s design the Low-Level Design (LLD) for our ChatPDF RAG system, grounding it entirely within the AWS ecosystem and making concrete choices for components. We’ll focus on the “why” for each choice.
Here is the High-Level Design (HLD) from before, which we will now break down into a detailed LLD.
graph TD
subgraph Offline Process
direction TB
A[Document Databases] --> B(Document Parsing & Chunking)
B --> C(Embedding Model)
C --> D[Vector Index]
end
subgraph Online / Inference Process
direction LR
E[User Query] --> F(Safety & Preprocessing)
F --> G(Embedding Model)
G --> H(ANN Search)
D --> H
H --> I(Prompt Engineering)
E --> I
I --> J[LLM]
J --> K(Safety & Postprocessing)
K --> M[Response]
end
Low-Level Design: “ChatPDF” on AWS
We’ll dissect the HLD into two core pipelines: the Indexing Pipeline (Offline) and the Inference Pipeline (Online).
1. The Indexing Pipeline (Asynchronous & Event-Driven)
Goal: Process new or updated PDFs from a source location, chunk them, embed them, and store them in our vector database with minimal manual intervention. We’ll design this to be robust and scalable.
AWS Services & Architecture:
graph TD
subgraph "Indexing Pipeline - AWS"
A[S3 Bucket<br/>company-docs-raw] -->|ObjectCreated event| B(AWS Lambda<br/>s3-trigger-lambda)
B -->|Publishes PDF key| C[SQS Queue<br/>pdf-processing-queue]
D[Auto Scaling Group EC2<br/>c6i.xlarge instances<br/>Python, UnstructuredPDFLoader, LangChain] -->|Polls| C
D -->|Parsed & Chunked Data| E[S3 Bucket<br/>company-docs-chunks]
E -->|ObjectCreated event| F(AWS Lambda<br/>embedding-lambda)
F -->|Invokes| G[SageMaker Endpoint<br/>g5.2xlarge with<br/>SentenceTransformer/CLIP]
G -->|Returns embeddings| F
F -->|Writes data| H[Amazon OpenSearch Service<br/>with k-NN index]
end
Low-Level Component Breakdown & Trade-offs:
A. Document Source:
S3 Bucket: 'company-docs-raw'- Why S3? It’s the de-facto standard for object storage on AWS. It’s infinitely scalable, durable (11 nines!), and has rich event notification features, which are perfect for triggering our pipeline.
B & C. Triggering Mechanism:
S3 Event -> Lambda -> SQS Queue- Why this pattern? This is a classic “decoupling” pattern for resilient systems.
- The
s3-trigger-lambdais a tiny, fast function that simply takes the new S3 object key and puts it into an SQS queue. It’s cheap and reliable. - Trade-off: We could have the Lambda do the whole processing job. Why not? PDF parsing can be slow and memory-intensive. Lambda has time limits (max 15 mins) and memory constraints. A large or complex PDF could cause it to fail. By pushing the job to a queue, we separate the trigger from the heavy lifting.
- Why SQS? It acts as a buffer. If we upload 10,000 documents at once, SQS holds the jobs, and our processing fleet can work through them at its own pace. It also provides automatic retries for failed jobs, making the pipeline robust.
- The
- Why this pattern? This is a classic “decoupling” pattern for resilient systems.
D. Parsing & Chunking Service:
Auto Scaling Group of EC2- Why EC2, not Lambda? UnstructuredPDFLoader processing can be compute-intensive, especially when handling OCR for scanned documents. We need dedicated compute with more control over the environment and no execution time limits. An Auto Scaling Group allows us to scale the number of worker nodes up or down based on the queue depth in SQS, which is extremely cost-efficient.
- Instance Choice:
c6i.xlargeare compute-optimized with sufficient CPU and memory for OCR operations and complex document processing. - Tooling:
- Parser & Chunker: UnstructuredPDFLoader from
langchain_community.document_loaders. This is our unified solution that handles both parsing and intelligent chunking:- OCR Capabilities: Automatically handles scanned PDFs and images with built-in OCR, eliminating the need for separate OCR preprocessing.
- Structure Understanding: Unlike PyMuPDF’s raw text extraction, UnstructuredPDFLoader understands document hierarchy, preserving headers, paragraphs, tables, and lists as structured elements.
- Intelligent Chunking: Superior to
RecursiveCharacterTextSplitterbecause it chunks based on semantic document structure rather than arbitrary character counts, leading to more meaningful embeddings. - Multi-format Support: Handles various PDF types including forms, scanned documents, and complex layouts that would break simpler parsers.
- Metadata Enrichment: Each chunk comes with rich metadata about document structure, element types, and positioning, enabling better retrieval strategies.
- Configuration: Allows fine-tuning of chunk sizes while respecting natural document boundaries, with configurable overlap strategies that preserve context intelligently.
- Parser & Chunker: UnstructuredPDFLoader from
E, F, G. Embedding Service:
S3 -> Lambda -> SageMaker EndpointWhy SageMaker? It’s AWS’s managed service for deploying ML models. It handles autoscaling, provides GPU instances for fast inference, and gives us a simple REST API endpoint. We don’t have to manage CUDA drivers or model servers ourselves.
Instance Choice:
g5.2xlarge. These have NVIDIA A10G GPUs, which are excellent for transformer inference.Embedding Model Choice: Deep Dive into FastEmbed vs. Sentence Transformers
This is a critical decision that affects both quality and cost. Let’s break down the trade-offs:
Comprehensive Comparison: FastEmbed vs. Sentence Transformers
Aspect FastEmbed Sentence Transformers Our Choice & Reasoning Performance Faster: ONNX runtime, quantization, optimized inference Slower: PyTorch-based, full precision by default Context-dependent: FastEmbed for high-throughput, ST for quality-first Model Quality Good: Same base models, slight quality degradation from quantization Better: Full precision models, no optimization artifacts Sentence Transformers: We prioritize retrieval quality for better user experience Infrastructure Requirements CPU-optimized: Runs efficiently on CPU instances GPU-preferred: Benefits significantly from GPU acceleration ST + GPU: Better cost/quality ratio at our scale Memory Footprint Small: Quantized models, ~100-200MB Large: Full models, ~400-500MB FastEmbed: Better for memory-constrained environments Cold Start Time Fast: Small models load quickly Slow: Larger models take time to load FastEmbed: Better for Lambda/serverless Ecosystem Integration Simple: Plug-and-play, minimal dependencies Rich: Extensive model zoo, fine-tuning capabilities ST: Better for experimentation and model iteration Total Cost Lower: Cheaper CPU instances, faster processing Higher: GPU instances, slower throughput Depends on scale: FastEmbed wins at high volume Our Decision Matrix:
For Production (Current Choice): Sentence Transformers on GPU
# SageMaker Endpoint Configuration model = SentenceTransformer('all-mpnet-base-v2') instance_type = 'ml.g5.2xlarge' # GPU instance # Reasoning: # 1. Quality-first approach for better user experience # 2. GPU cost amortizes well at our expected volume # 3. Rich ecosystem for future model updates # 4. Consistent with research best practicesWhen We’d Switch to FastEmbed:
# High-volume, cost-sensitive scenario from fastembed import TextEmbedding model = TextEmbedding(model_name='BAAI/bge-small-en-v1.5') instance_type = 'ml.c6i.2xlarge' # CPU instance # Conditions for switch: # 1. Query volume > 10M/month (cost becomes primary factor) # 2. Retrieval quality is "good enough" (validated through A/B testing) # 3. Infrastructure simplification is prioritized # 4. Cold start latency is critical (Lambda deployments)Evolution Strategy:
Phase 1 (Launch): Sentence Transformers for quality Phase 2 (Scale): A/B test FastEmbed vs. ST on retrieval metrics Phase 3 (Optimize): Switch to FastEmbed if quality delta is acceptable
For Images: We’d use the pre-trained CLIP model’s image encoder, also hosted on a SageMaker endpoint, to ensure text and image embeddings are in the same space.
H. Vector Database:
Amazon OpenSearch Service- Why OpenSearch? It’s a managed version of Elasticsearch that AWS supports directly. Its key feature for us is the k-NN (k-Nearest Neighbor) plugin. It allows OpenSearch to function as a powerful, scalable vector database.
- Trade-off vs.
QdrantorPinecone:- Qdrant: An excellent, open-source vector database written in Rust, optimized for performance and memory safety. It offers advanced features like filtering during search. If we needed the absolute best retrieval performance and were willing to manage it ourselves (e.g., on EKS), Qdrant is a top contender.
- Pinecone: A fully managed, SaaS vector database. It’s incredibly easy to use and provides S-tier performance. However, it exists outside the core AWS ecosystem, which can complicate security, networking (VPC peering), and billing.
- Conclusion: We choose OpenSearch because it lives within our AWS account. This simplifies IAM permissions, networking, and keeps our data inside our VPC. It’s “good enough” for most use cases and much easier to manage than a self-hosted solution.
- Distance Metric Choice:
Cosine Similarity- Why? Cosine similarity measures the angle between two vectors, ignoring their magnitude. For text embeddings, we care about the direction (semantic meaning), not the vector’s length. It is the standard and most effective choice for normalized transformer embeddings.
- Trade-off vs.
Euclidean Distance (L2): Measures the straight-line distance. It’s sensitive to vector magnitude. It can be a good choice for other types of data (e.g., image feature vectors) but is generally less effective for text. Dot Product is another option, very similar to Cosine for normalized vectors. We stick with the standard.
2. The Inference Pipeline (Real-Time & Low-Latency)
Goal: Take a user’s API request, find the relevant context, and generate a safe, accurate answer as quickly as possible.
AWS Services & Architecture:
graph TD
subgraph "Inference Pipeline - AWS"
A[User via API Gateway] --> B(AWS Lambda<br/>inference-handler)
B -->|Query text| C[SageMaker Endpoint<br/>SentenceTransformer]
C -->|Returns query_embedding| B
B -->|Search for query_embedding| D[Amazon OpenSearch Service<br/>k-NN Search]
D -->|Top 5 chunks| B
B -->|Builds final prompt| E[Amazon Bedrock<br/>Claude 3 Sonnet]
E -->|Streams response| B
B -->|Streams response| A
F((Safety & Guardrails)) -->|Implemented within| B
F -->|Configures| E
end
Low-Level Component Breakdown & Trade-offs:
A. API Layer:
API Gateway- Why? It’s the standard, managed way to create REST APIs on AWS. It handles authentication (e.g., with Cognito or IAM), throttling, caching, and routing requests to our backend logic.
B. Backend Logic:
AWS Lambda: 'inference-handler'- Why Lambda? The “glue” logic is stateless and involves a series of network calls. This is a perfect use case for Lambda. It’s fast to start, scales to zero (so we don’t pay when no one is using it), and scales out automatically under load.
- Function:
- Receive the request from API Gateway.
- Perform pre-processing/safety checks.
- Call the SentenceTransformer SageMaker endpoint to embed the user’s query.
- Query OpenSearch with the embedding to get the top-k chunks.
- Construct the final prompt from the template, user query, and retrieved chunks.
- Call the LLM.
- Perform post-processing/safety checks on the response.
- Stream the response back.
C. Query Embedding:
SageMaker Endpoint (SentenceTransformer)- Why? We reuse the exact same model from the indexing pipeline to ensure the query vector is in the same space as the document vectors.
- Model Consistency: Critical that this is identical to the indexing model. Even minor version differences can cause embedding drift, leading to poor retrieval performance.
- Alternative Consideration: For high-query-volume scenarios (>10M/month), we might consider switching both indexing and inference to FastEmbed for cost optimization, but only after validating that retrieval quality remains acceptable through A/B testing.
D. ANN Search:
Amazon OpenSearch Service- Why? We query the index we built earlier. A typical query would look like:
{"query": {"knn": {"embedding_field": {"vector": [0.1, 0.2, ...], "k": 5}}}}. OpenSearch will use its ANN algorithm (like HNSW, which it supports) to return the 5 nearest neighbors with low latency.
- Why? We query the index we built earlier. A typical query would look like:
E. LLM Generation:
Amazon Bedrock (Anthropic's Claude 3 Sonnet)- Why Bedrock? This is AWS’s managed service for foundation models. It gives us API access to models from AI21, Anthropic, Cohere, Meta, etc., without needing to host them. This is a huge win for simplicity, security, and pay-per-use pricing.
- Model Choice:
Claude 3 Sonnet- Why? As of today, the Claude 3 family is S-tier.
Sonnetis the middle model, offering a fantastic balance of intelligence, speed, and cost. It has a large context window (200k tokens), is great at following complex instructions, and has a lower hallucination rate, which is perfect for a RAG system. - Trade-off vs.
Claude 3 Haiku: Haiku is faster and cheaper, but less intelligent. We might use Haiku if our queries were simpler and speed was the absolute priority. - Trade-off vs.
Claude 3 Opus: Opus is the most powerful model, but slower and more expensive. We might use Opus for a premium “expert” version of our chatbot. Sonnet is the ideal balanced choice.
- Why? As of today, the Claude 3 family is S-tier.
F. Safety:
Implemented in Lambda & Bedrock Guardrails- Why a dual approach?
- Lambda: We can implement simple pre-filters on the user query (e.g., regex for PII patterns, keyword blocklists). We can also do post-processing on the final output.
- Bedrock Guardrails: This is a powerful managed feature. We can configure policies to block harmful topics, filter out specific words, and even prevent the model from answering questions outside its scope (e.g., “Don’t answer questions about medical advice”). This is a more robust and scalable approach to safety than trying to implement it all ourselves.
- Why a dual approach?
This LLD provides a concrete, defensible, and modern blueprint for building a production-grade RAG system on AWS. It makes specific technology choices and, most importantly, provides the reasoning and trade-offs behind each one.
Key Architectural Decisions Summary:
- Document Processing: UnstructuredPDFLoader for comprehensive OCR and structure-aware parsing
- Embedding Strategy: Sentence Transformers (quality-first) with clear migration path to FastEmbed (cost-optimization)
- Vector Database: OpenSearch within AWS ecosystem vs. external specialized solutions
- Compute Distribution: EC2 for heavy processing, Lambda for orchestration, SageMaker for ML inference
- Event-Driven Architecture: S3 → Lambda → SQS → EC2 for robust, scalable document processing
Each choice reflects real-world production considerations: balancing quality, cost, operational complexity, and future scalability.
Advanced RAG patterns
These advanced patterns—Multi-Query, Multi-Hop, and Routing—all revolve around a central theme: making the retrieval step more intelligent. Instead of a single, straightforward search, we are introducing a Reasoning Engine or an Orchestration Layer that plans and executes a more complex retrieval strategy.
Let’s break down the HLD and LLD for each.
Foundational Concept: The Reasoning Engine
Before diving into the specific patterns, let’s establish our core architectural change. We are inserting a “smart” component right after the user query comes in.
graph TD
A["User Query"] --> B{Reasoning Engine}
B --> C["Intelligent Retrieval Execution"]
C --> D["LLM Generator"]
D --> E["Response"]
This “Reasoning Engine” is the brain of our advanced RAG. In most cases, it will be powered by an LLM itself—a fast and cheap one like Claude 3 Haiku or Llama 3 8B—whose job is not to answer the question, but to decompose the question into a plan.
1. Multi-Query RAG
Concept: The user asks a single complex question that implicitly requires looking up multiple things. The system breaks it down into several sub-queries, executes them in parallel, and synthesizes the results.
Use Case: “Compare and contrast the battery life of the iPhone 15 and the Google Pixel 8.”
High-Level Design (HLD)
The key here is the “Query Decomposer” and the parallel “fan-out” retrieval.
graph TD
subgraph "Multi-Query RAG HLD"
A["User Query"] --> B("Query Decomposer")
B --> C1("Retriever 1")
B --> C2("Retriever 2")
C1 --> D{Synthesizer}
C2 --> D
D --> E["Response"]
end
Low-Level Design (LLD) on AWS
The “Query Decomposer” is an LLM call, and the parallel retrieval is an asynchronous operation within our orchestrator Lambda.
graph TD
A["User via API Gateway"] --> B("AWS Lambda: orchestrator-lambda")
subgraph "Step 1: Decompose Query"
B --> C["Amazon Bedrock<br/>Claude 3 Haiku"]
C --> B
end
subgraph "Step 2: Parallel Retrieval"
B --> D1["OpenSearch k-NN Search"]
B --> D2["OpenSearch k-NN Search"]
end
subgraph "Step 3: Synthesize"
D1 --> B
D2 --> B
B --> E["Amazon Bedrock<br/>Claude 3 Sonnet"]
E --> B
end
B --> F["Response Stream"]
Key LLD Components & Trade-offs:
- Orchestrator (
orchestrator-lambda): A single Lambda function coordinates the entire process. - Query Decomposer (
Claude 3 Haiku):- Why Haiku? This is a structured task: text-to-JSON. It doesn’t require deep reasoning. Haiku is extremely fast and cheap, making it perfect for this pre-processing step. We use function calling / tool use features to ensure the LLM returns a clean, machine-readable JSON array of queries.
- Prompt:
"Given the user's question, generate a JSON list of simple, self-contained search queries needed to answer it. Question: {user_question}"
- Parallel Retrieval:
- The Lambda function will use Python’s
asyncio.gatherto make multiple, concurrent calls to our OpenSearch cluster. This is a critical optimization. A naive, sequential approach would double the retrieval latency.
- The Lambda function will use Python’s
- Synthesizer (
Claude 3 Sonnet):- Why Sonnet? This step requires more intelligence. The model needs to understand two different sets of context and perform a comparison. Sonnet provides a good balance of reasoning power and speed for this.
- Prompt:
"You have been given the following information. Context for 'iPhone 15 battery life': {context_1}. Context for 'Pixel 8 battery life': {context_2}. Now, answer the user's original question: {user_question}"
2. Multi-Hop RAG
Concept: The system answers a question that requires a sequence of searches, where the results of the first search are needed to formulate the second search.
Use Case: “Which actor played the main character in the movie directed by the person who directed ‘Inception’?”
- Hop 1: “Who directed ‘Inception’?” -> Christopher Nolan
- Hop 2: “Which movie did Christopher Nolan direct where X was the main character?” (This is tricky, shows the limits) or more simply “Who was the main actor in ‘The Dark Knight’?” (assuming a known movie). A better example from the prompt: “Who are the largest car manufacturers? Do they make EVs?”.
- Hop 1: “largest car manufacturers 2023” -> Toyota, VW, Hyundai.
- Hop 2: “Toyota EV models”, “VW EV models”, “Hyundai EV models”.
High-Level Design (HLD)
The key is a loop or sequence in the reasoning engine.
graph TD
subgraph "Multi-Hop RAG HLD"
A["User Query"] --> B{Reasoning Engine}
B --> C("Retriever")
C --> B
B --> C
C --> B
B --> D["Response"]
end
Low-Level Design (LLD) on AWS
This sequential, stateful process is a perfect use case for AWS Step Functions. It’s more robust and observable than trying to code a complex loop inside a single Lambda.
graph TD
A["Start"] --> B("Generate Hop 1 Query - Lambda")
B --> C("Perform Search 1 - Lambda")
C --> D("Generate Hop 2 Queries - Lambda")
D --> E("Perform Parallel Search 2 - Map State")
E --> F("Synthesize Results - Lambda")
F --> G["End"]
subgraph "AWS Step Functions Workflow"
direction TB
B1("Lambda 1") --> C1("Lambda 2")
C1 --> D1("Lambda 3")
D1 --> E1("Map State")
E1 --> F1("Lambda 4")
end
Key LLD Components & Trade-offs:
- Orchestration (
AWS Step Functions):- Why? It’s designed for orchestrating multi-step workflows. It handles state management (passing the results of Hop 1 to Hop 2), error handling, and retries automatically. Debugging is much easier as you can visualize the execution flow and inspect the inputs/outputs of each step.
- Trade-off: There is a slight cold-start and state-transition overhead compared to a single “god Lambda”. However, for a complex workflow like multi-hop, the reliability and maintainability gains are immense.
- State Machine Steps:
Generate Hop 1 Query (Lambda): A simple step. Takes the user query, calls Haiku to get the first search query.Perform Search 1 (Lambda): Calls OpenSearch with the query from the previous step.Generate Hop 2 Queries (Lambda): This is a key step. It takes the results from step 2 (e.g., the text “The largest manufacturers are Toyota, VW, and Hyundai”) and calls Haiku to generate the next set of queries (e.g.,["Toyota EV models", "VW EV models", ...]).Perform Parallel Search 2 (Map State): TheMapstate in Step Functions is brilliant for this. It takes the array of queries from step 3 and runs a search Lambda in parallel for each item in the array.Synthesize Results (Lambda): This final step collects the results from all previous hops and uses a powerful model like Sonnet to generate the final, coherent answer.
3. Query Routing
Concept: The system has access to multiple, distinct knowledge bases (e.g., different document indexes, different databases). The router decides which data source is the most appropriate for a given query.
Use Case: “What is our company’s PTO policy?” -> query HR docs. “What was the revenue from customer X last quarter?” -> query Salesforce.
High-Level Design (HLD)
The key component is the “Router” which acts as a switchboard.
graph TD
subgraph "Query Routing HLD"
A["User Query"] --> B{Router}
B --> C1("HR Document Retriever")
B --> C2("Sales Data Retriever")
B --> C3("General Wiki Retriever")
C1 --> D{Generator}
C2 --> D
C3 --> D
D --> E["Response"]
end
Low-Level Design (LLD) on AWS
The Router is another LLM call. The different retrievers could be different OpenSearch indexes or even completely different systems (like the Salesforce API).
graph TD
A["User via API Gateway"] --> B("AWS Lambda: router-orchestrator")
subgraph "Step 1: Route Query"
B --> C["Amazon Bedrock<br/>Claude 3 Haiku"]
C --> B
end
subgraph "Step 2: Conditional Retrieval"
B --> D1["OpenSearch hr-index"]
B --> D2["Salesforce API"]
B --> D3["OpenSearch wiki-index"]
end
subgraph "Step 3: Generate"
D1 --> E{Generate Response}
D2 --> E
D3 --> E
E --> F["Final Answer"]
end
Key LLD Components & Trade-offs:
- Router (
Claude 3 Haiku):- Again, a fast, cheap model is used for classification. The prompt is crucial.
- Prompt:
"You are a query router. Given the user's question, determine the best data source to answer it. The available sources are: 'HR_DOCS' for questions about employment, PTO, and policies; 'SALESFORCE_DATA' for questions about specific customers, leads, or revenue; and 'GENERAL_WIKI' for all other topics. Return your answer as a single JSON object like {'source': 'CHOSEN_SOURCE'}. Question: {user_question}"
- Conditional Logic (
router-orchestratorLambda): The Lambda’s code contains a simpleif/elif/elseblock based on thesourcereturned by the router LLM. - Heterogeneous Data Sources: This is the most realistic part.
OpenSearch: We’ll have separate indexes for HR docs and the general wiki. This prevents data leakage and allows us to set different access permissions.Salesforce API: For this route, we wouldn’t use vector search. The Lambda might use the LLM to generate a Salesforce Object Query Language (SOQL) query, then execute it against the Salesforce API via a connected app. This demonstrates an ability to integrate with non-vector-search systems.
- Trade-off: The primary risk is the accuracy of the router. If the router misclassifies a query, the system will look in the wrong place and fail to find the answer, even if it exists. This requires careful prompt engineering and potentially finetuning the router model on examples of correctly routed queries. You might also build a fallback mechanism: if the primary source returns no results, try the
GENERAL_WIKIas a backup.
Beyond Traditional RAG: InfiniRetri and the Future of Long-Context Processing
Excellent question. This is exactly the kind of critical thinking required at a senior level: not just knowing existing patterns like RAG, but constantly evaluating new research to see how it could evolve or even replace parts of your current system.
Yes, the paper “Infinite Retrieval: Attention Enhanced LLMs in Long-Context Processing” (arXiv:2406.19521) (which I’ll call InfiniRetri), is absolutely related to RAG. In fact, it’s a direct challenge to the “R” in RAG as we’ve designed it.
Let’s break down what this paper proposes and how it compares to our classic RAG system.
The Core Problem This Paper Tackles
Our current RAG design solves the long-context problem with a clear separation of concerns:
- Retriever (External Tool): An embedding model + a vector database (like OpenSearch) finds relevant information. Its job is to be a great librarian.
- Generator (LLM): The LLM receives the retrieved context and generates an answer. Its job is to be a great reasoner and writer, using only the documents the librarian gives it.
The paper asks a fundamental question: “Why not use the retrieval capabilities of LLMs themselves to handle long contexts?”
They observe that the attention mechanism inside a Transformer is, in essence, a retrieval mechanism. For each token in the query (the “Query”), the model learns to “attend” to the most relevant tokens in the context (the “Keys”). They show that in the deeper layers of an LLM, this attention pattern becomes very accurate at locating the exact phrases needed to answer a question.
So, the core insight is: The LLM already knows how to retrieve. Let’s build a system that leverages this internal retrieval ability instead of relying on an external one.
How InfiniRetri Works: The “Slide and Retrieve” Method
Imagine you’re reading a 1000-page book, but you can only see one page at a time (your “context window”). To answer a question about the whole book, you would:
- Read page 1.
- Jot down the most important sentences on a sticky note.
- Read page 2, keeping your sticky note in view.
- Update your sticky note with important sentences from page 2, maybe crossing out less important ones from page 1.
- Repeat until you’ve read all 1000 pages. Your final sticky note is a compressed summary of the entire book’s relevant information.
This is exactly what InfiniRetri does:
graph TD
subgraph "InfiniRetri: Sliding Window with Attention-Based Retrieval"
A["Long Document<br/>1M tokens"] -->|Split| B["Chunk 1<br/>4K tokens"]
A -->|Split| C["Chunk 2<br/>4K tokens"]
A -->|Split| D["Chunk N<br/>4K tokens"]
E["User Query"] --> F["Process Chunk 1 + Query"]
B --> F
F -->|Internal Attention| G["Identify Key Sentences<br/>from Chunk 1"]
G --> H["Compressed Cache<br/>512 tokens"]
H --> I["Process Chunk 2 + Cache + Query"]
C --> I
I -->|Internal Attention| J["Update Cache<br/>with Key Sentences"]
J --> K["Updated Cache<br/>512 tokens"]
K --> L["Process Chunk N + Cache + Query"]
D --> L
L -->|Internal Attention| M["Final Cache<br/>Most Relevant Info"]
M --> N["Generate Final Answer"]
E --> N
end
subgraph "Key Innovation"
O["LLM's Attention Mechanism<br/>Acts as Retriever"]
P["No External Vector DB<br/>Required"]
Q["Sequential Processing<br/>Maintains Context Flow"]
end
Here’s the detailed breakdown:
Chunk: The long document (e.g., 1M tokens) is broken down into sequential chunks that fit within the model’s native context window (e.g., 4K tokens).
Slide Window (Iterative Process):
- Merge: For the first chunk, the model processes it with the user’s question.
- Inference & Retrieval: The model uses its own internal attention scores to identify the most important sentences/phrases within that chunk. This is the “Retrieval in Attention” step.
- Cache: It saves these most important sentences into a small, compressed “cache” (like a sticky note).
- Iterate: For the next chunk, the model merges the content of the compressed cache with the new chunk, processes this combined input, and updates its cache.
Final Answer: After iterating through all chunks, the final cache contains the most relevant information from the entire document.
Comprehensive Comparison: Traditional RAG vs. InfiniRetri
| Aspect | Traditional RAG (Our Design) | InfiniRetri | Advantage |
|---|---|---|---|
| Retrieval Mechanism | External: SentenceTransformer + Vector DB (OpenSearch) | Internal: LLM’s own attention mechanism | InfiniRetri: Simpler architecture, potentially better semantic understanding |
| Indexing Requirements | Heavy Upfront: Must parse, chunk, and embed all 5M documents | Zero Upfront: Document processed on-the-fly | InfiniRetri: Real-time processing, no pre-computation |
| Infrastructure Complexity | High: S3, SQS, EC2, SageMaker, OpenSearch, Lambda | Lower: Primarily LLM hosting + orchestration | InfiniRetri: Reduced infrastructure footprint |
| Contextual Cohesion | Fragmented: k independent chunks to stitch together | Sequential: Maintains narrative flow and order | InfiniRetri: Better for document structure understanding |
| Query Latency | Fast: Vector search (ms) + single LLM call | Slow: Multiple LLM calls per chunk | Traditional RAG: Much better for interactive use |
| Processing Cost | Low at Query Time: Expensive indexing, cheap retrieval | High at Query Time: No indexing, expensive inference | Traditional RAG: Better for frequent queries |
| Use Case Fit | Multi-document Q&A: Large knowledge bases | Single-document Analysis: Deep comprehension tasks | Context-dependent |
| Scalability | Horizontal: Add more documents easily | Vertical: Limited by single document processing | Traditional RAG: Better for growing knowledge bases |
When to Choose Each Approach
Traditional RAG is Best For:
- Interactive chatbots requiring sub-second responses
- Large, multi-document knowledge bases (like our 5M page corpus)
- Frequent, repetitive queries where indexing cost amortizes
- Real-time customer support scenarios
InfiniRetri is Best For:
- Deep document analysis where users can wait minutes for comprehensive answers
- Novel, massive documents that haven’t been pre-processed
- Narrative understanding tasks requiring sequential reading
- One-off analysis of large documents (legal discovery, research papers)
Integration Strategy: The Hybrid Approach
A sophisticated system might offer both patterns:
graph TD
subgraph "Hybrid RAG + InfiniRetri System"
A["User Query"] --> B{Query Type Detection}
B -->|Quick Facts| C["Traditional RAG Pipeline"]
C --> D["Vector Search + LLM"]
D --> E["Fast Response<br/>< 3 seconds"]
B -->|Deep Analysis| F["InfiniRetri Pipeline"]
F --> G["Sliding Window Processing"]
G --> H["Comprehensive Analysis<br/>2-5 minutes"]
B -->|Hybrid Query| I["Parallel Processing"]
I --> J["RAG for Quick Facts"]
I --> K["InfiniRetri for Deep Context"]
J --> L["Combined Response"]
K --> L
end
Interview-Level Insight: Demonstrating Senior Awareness
After presenting the classic RAG design, you could demonstrate senior-level awareness by saying:
“This RAG architecture is a robust, low-latency solution for our problem. However, it’s worth noting the cutting edge of research is exploring alternatives. For instance, recent papers like ‘Infinite Retrieval: Attention Enhanced LLMs in Long-Context Processing’ propose using the LLM’s own attention mechanism for retrieval in a streaming fashion.
This approach eliminates the need for an external vector database, simplifying the architecture and potentially improving retrieval quality since it uses the powerful LLM itself. The major trade-off is significantly higher inference latency, as it requires multiple passes over the document.
Therefore, while not suitable for our interactive chat use case, this ‘RAG-without-an-index’ pattern could be extremely powerful for a different product feature, like an offline ‘deep analysis’ tool for digesting entire books or legal documents. This shows how the choice between these patterns is highly dependent on the specific product requirements for latency and interactivity.”
Summary: Evolution, Not Revolution
InfiniRetri is not a “RAG killer” – it’s an alternative with a different performance profile.
- For our ChatPDF system, Traditional RAG remains the right choice due to latency requirements and multi-document nature
- InfiniRetri represents a “Third Way” between giant context windows (expensive) and traditional RAG (complex infrastructure)
- The future likely involves hybrid systems that choose the right pattern based on query type and user expectations
This demonstrates that you’re not just following blueprints; you’re actively thinking about the future and making nuanced, product-aware technology choices. That’s top-tier engineering thinking.