\

Chapter 4: Training Data

26 min read

Okay, class, let’s settle in. We’ve journeyed through the high-level overview of ML systems, the design process, and the nitty-gritty of data engineering fundamentals. Now, we arrive at a topic that is, in many ways, the heart of supervised machine learning: Chapter 4: Training Data.

The author makes a critical point right at the start: ML curricula often skew heavily towards modeling – the “fun” part. But as anyone who’s worked in production ML at FAANG or elsewhere knows, “spending days wrangling with a massive amount of malformatted data that doesn’t even fit into your machine’s memory is frustrating” but essential. Bad data can sink your entire ML operation, no matter how brilliant your model.

This chapter shifts focus from the systems perspective of data (Chapter 3) to the data science perspective. We’re going to cover:

  • Sampling techniques: How to select data for training.
  • Labeling challenges: The pains of hand labels, label multiplicity, and the beauty of natural labels.
  • Handling lack of labels: Weak supervision, semi-supervision, transfer learning, active learning.
  • Class imbalance: Why it’s a problem and how to tackle it.
  • Data augmentation: Creating more data from what you have.

A key terminology point: The book uses “training data” instead of “training dataset” because “dataset” implies finite and stationary. Production data is rarely either (hello, “Data Distribution Shifts” on page 237, which we’ll cover much later). Creating training data, like everything else in ML systems, is an iterative process.

Let’s begin!


Page 81 (Chapter Introduction): The Importance and Pain of Data

The intro reiterates the core message:

  • Data is “messy, complex, unpredictable, and potentially treacherous.”
  • Handling it well saves time and headaches. This chapter is about techniques to obtain/create good training data.
  • “Training data” here encompasses all data for development: train, validation, and test splits.

A Crucial Word of Caution (Page 82, top): Data is full of potential biases!

  • Biases can creep in during collection, sampling, or labeling.
  • Historical data can embed human biases.
  • ML models trained on biased data will perpetuate those biases.
  • “Use data but don’t trust it too much!” This is a mantra every ML practitioner should live by. Always be skeptical, always question your data.

Pages 82-87: Sampling – Choosing Your Data Wisely

Sampling is often overlooked in coursework but is integral to the ML lifecycle. It happens when:

  • Creating training data from all possible real-world data.
  • Creating train/validation/test splits from a given dataset.
  • Sampling events for monitoring.

Why sample?

  1. Necessity: You rarely have access to all possible data. Your training data is inherently a sample.
  2. Feasibility: Processing all accessible data might be too time-consuming or resource-intensive.
  3. Efficiency: Quick experiments on small subsets can validate a new model’s promise before full-scale training (footnote 1: even for large models, experimenting with dataset sizes reveals its effect).

Understanding sampling helps avoid bias and improve efficiency. Two families:

  1. Nonprobability Sampling (Pages 83-84): Selection not based on probability criteria. Often driven by convenience, leading to selection biases (footnote 2, Heckman).

    • Convenience sampling: Selected based on availability. Popular because it’s easy.
      • Example: Language models often trained on easily collected Wikipedia, Common Crawl, Reddit, not necessarily representative of all text.
      • Example: Sentiment analysis data from IMDB/Amazon reviews. Biased towards those willing to leave reviews, not representative of all users.
      • Example: Self-driving car data initially from sunny Phoenix/Bay Area, less from rainy Kirkland (footnote 3). Model might be great in sun, poor in rain.
    • Snowball sampling: Future samples selected based on existing ones (e.g., scrape Twitter accounts, then accounts they follow, etc.).
    • Judgment sampling: Experts decide what to include.
    • Quota sampling: Select based on quotas for slices (e.g., survey: 100 responses from <30yo, 100 from 30-60yo, etc., regardless of actual age distribution).
    • Usefulness: Quick way to get initial data. For reliable models, probability-based sampling is preferred.
    • FAANG Perspective: Convenience sampling is often how initial datasets are gathered for new problem domains, but we’re acutely aware of the biases and strive to get more representative data over time.
  2. Random Sampling (Probability-Based) (Pages 84-87):

    • Simple Random Sampling (Page 84): All samples in the population (footnote 4: “statistical population” = potentially infinite set of all possible samples) have equal selection probability.

      • Advantage: Easy to implement.
      • Drawback: Rare categories might be missed. If a class is 0.01% of data, a 1% random sample will likely miss it. Model might assume rare class doesn’t exist.
    • Stratified Sampling (Page 84):

      • Divide population into groups (strata) you care about. Sample from each group separately.
      • Example: To sample 1% from data with classes A and B, sample 1% of class A and 1% of class B. Ensures rare classes are included.
      • Drawback: Not always possible if groups are hard to define or samples belong to multiple groups (multilabel tasks - footnote 5).
      • FAANG Perspective: Stratified sampling is crucial for creating representative validation/test sets, especially when dealing with imbalanced classes or ensuring coverage across different user segments (e.g., regions, demographics).
    • Weighted Sampling (Page 85): Each sample given a weight determining its selection probability.

      • Example: Samples A, B, C with desired probabilities 50%, 30%, 20% get weights 0.5, 0.3, 0.2.
      • Leverages domain expertise (e.g., more recent data is more valuable, give it higher weight).
      • Corrects for distribution mismatch (e.g., your data is 25% red, 75% blue; real world is 50/50 red/blue. Give red samples 3x weight of blue during sampling).
      • Python: random.choices(population, weights, k).
      • Related concept: Sample Weights (in training, not sampling to create dataset): Assigns “importance” to training samples. Higher weight samples affect the loss function more. Can significantly change decision boundaries (Figure 4-1 from scikit-learn).
      • FAANG Perspective: Weighted sampling is used to upweight important/rare data. Sample weights in training are used for cost-sensitive learning or to prioritize certain types of errors.
    • Reservoir Sampling (Page 86, Figure 4-2): For streaming data when you don’t know total N, can’t fit all in memory, but want to sample k items such that each has equal selection probability.

      • Algorithm:
        1. Fill reservoir (array of size k) with first k elements.
        2. For each incoming n-th element (where n > k):
          • Generate random integer i from 1 to n.
          • If 1 <= i <= k, replace reservoir element at index i with the n-th element. Else, do nothing.
      • Ensures:
        • Every tweet/element has an equal probability of being selected at any point in time.
        • If stopped, current reservoir is a fair sample.
      • n-th element has k/n probability of being in the reservoir. Each element already in reservoir has k/n probability of staying (or more precisely, 1 - (k/n)*(1/k) = (n-1)/n probability of one specific item in reservoir not being replaced by the nth item, which then combines with prior probabilities). The math ensures uniform probability for all seen items.
      • Figure 4-2 visualizes this.
      • FAANG Perspective: Essential for sampling from massive, unbounded streams (e.g., sampling search queries for analysis, sampling events for monitoring from a firehose).
    • Importance Sampling (Page 87): Sample from a distribution Q(x) (proposal/importance distribution, easy to sample from) when you actually want to sample from P(x) (target distribution, hard to sample from). Weigh sample x from Q(x) by P(x)/Q(x).

      • Requires Q(x) > 0 whenever P(x) != 0.
      • Equation: E_P[x] = E_Q[x * P(x)/Q(x)].
      • Use in ML: Policy-based reinforcement learning. Estimate value of new policy (P) using rewards from old policy (Q) and reweight.
      • FAANG Perspective: Used in some advanced modeling scenarios, off-policy evaluation in RL, and in areas like Bayesian inference (though less common in typical production ML pipelines).

Pages 88-97: Labeling – The Quest for Ground Truth

Most production ML is supervised, needing labeled data. Quality/quantity of labels heavily impacts performance.

  • Andrej Karpathy anecdote: Recruiter asked how long he’d need a labeling team. His reply: “How long do we need an engineering team?” Labeling is now a core, ongoing function.
  1. Hand Labels (Pages 88-90): The classic approach.

    • Challenges:
      • Expensive: Especially with Subject Matter Expertise (SME). Spam classification: 20 crowdworkers, 15 mins training. Chest X-rays: board-certified radiologists (limited, expensive).
      • Privacy Threat: Someone looks at your data. Can’t ship patient records or confidential financials to third-party labelers. Might need on-premise annotators.
      • Slow: Phonetic speech transcription can take 400x utterance duration (footnote 7). 1 hour speech = 400 hours (3 months) labeling. Author’s colleagues waited almost a year for lung cancer X-ray labels.
      • Slow Iteration: If task/data changes, relabeling is slow, model becomes less adaptive. E.g., sentiment model (NEG/POS) needs to add ANGRY class -> relabel, collect more ANGRY examples.
    • Label Multiplicity / Ambiguity (Page 89, Table 4-1): Multiple annotators, different expertise/accuracy -> conflicting labels for same instance.
      • Entity recognition example: “Darth Sidious, known simply as the Emperor, was a Dark Lord of the Sith who reigned over the galaxy as Galactic Emperor of the First Galactic Empire.”
        • Annotator 1: 3 entities.
        • Annotator 2: 6 entities (more granular, e.g., “Dark Lord” and “Sith” separate).
        • Annotator 3: 4 entities.
      • Which to train on? Models will perform differently.
      • Disagreements common, especially with high SME needed (footnote 8: if obvious, no SME needed). What is “human-level performance” if experts disagree?
      • Mitigation: Clear problem definition (e.g., “pick longest substring for entities”) and annotator training.
    • Data Lineage (Page 90): Track origin of samples and labels.
      • Critical if using multiple sources/annotators.
      • Example: Model trained on 100k good samples. Add 1M crowdsourced samples (lower accuracy). Performance decreases. If data mixed, hard to debug.
      • Helps flag bias, debug models. If model fails on recent data, investigate its acquisition. Often, it’s bad labels, not bad model.
    • FAANG Perspective: We have large in-house and vendor labeling operations. Managing quality, cost, throughput, and privacy is a huge operational challenge. Data lineage is critical for regulatory compliance and debugging. Adjudication (resolving labeler disagreements) is a standard part of the process.
  2. Natural Labels (Behavioral Labels) (Pages 91-93): Labels inferred from system or user behavior, no human annotation needed.

    • Model predictions automatically/partially evaluated by system.
    • Examples:
      • Google Maps ETA: At trip end, actual travel time is known.
      • Stock price prediction: After 2 mins, actual price is known.
      • Recommender Systems (canonical): User clicks (POSITIVE label) or doesn’t click after time (NEGATIVE label) on recommendation.
      • Many tasks framed as recommendation (e.g., ad CTR prediction = recommend relevant ads).
    • Can be set up: Google Translate allows community to submit alternative translations (becomes labels for next iteration, after review). Facebook “Like” button provides feedback for newsfeed ranking.
    • Common in Industry (Figure 4-3): Survey of 86 companies: 63% use natural labels. (Percentages don’t sum to 1 as companies use multiple sources). Likely because it’s cheaper/easier to start with.
    • Implicit vs. Explicit Labels (Page 92):
      • Recommendation not clicked after time = implicit negative label.
      • User downvotes recommendation = explicit negative label.
    • Feedback Loop Length (Pages 92-93): Time from prediction to feedback.
      • Short loops (minutes): Many recommenders (Amazon related products, Twitter follows).
      • Longer loops (hours/days/weeks): Recommending blog posts, YouTube videos, Stitch Fix clothes (feedback after user tries them on).
      • Different Types of User Feedback (sidebar, page 93): Ecommerce example:
        • Clicking product (high volume, weaker signal, fast feedback).
        • Adding to cart.
        • Buying product (lower volume, stronger signal, business-correlated).
        • Rating/reviewing.
        • Returning product.
        • Optimizing for clicks vs. purchases is a common trade-off. Depends on use case, discuss with stakeholders.
      • Choosing Window Length: Speed vs. accuracy. Short window = faster labels, but might prematurely label as negative. Twitter Ads study (footnote 10): most clicks in 5 mins, but some hours later. Short window underestimates true CTR.
      • Long feedback loops (weeks/months): Fraud detection (dispute window 1-3 months). Good for quarterly reports, bad for quick issue detection. A faulty fraud model can bankrupt a small business if issues take months to fix.
    • FAANG Perspective: Natural labels are heavily used for large-scale systems (search, ads, recommendations). Designing for good feedback collection is key. Understanding feedback delay and its impact on evaluation is critical.
  3. Handling the Lack of Labels (Pages 94-101, Table 4-2): What to do when hand labels are too hard and natural labels are absent/insufficient.

    • Table 4-2 Summary:

      MethodHowGround Truths Required?
      Weak Supervision(Noisy) heuristics to generate labelsNo, but small set recommended to guide heuristic dev.
      Semi-supervisionStructural assumptions, seed labelsYes, small initial set as seeds
      Transfer LearningPretrained models from another taskNo (zero-shot). Yes (fine-tuning, often fewer than scratch)
      Active LearningLabel samples most useful to modelYes
    • Weak Supervision (Pages 95-97): Use heuristics (SME rules) to label data.

      • Snorkel (Stanford AI Lab, footnote 11): Popular open-source tool.
      • Labeling Functions (LFs): Functions encoding heuristics.
        • Example: if "pneumonia" in nurse_note: return "EMERGENT"
        • Types of LFs: Keyword, regex, DB lookup, outputs of other models.
      • LFs produce noisy labels. Multiple LFs can conflict. Need to combine, denoise, reweight LFs (Figure 4-4 shows this high-level process).
      • Theoretically no hand labels needed, but small set recommended to guide LF development/accuracy check.
      • Benefits:
        • Privacy (write LFs on small, cleared subset; apply to all data without looking).
        • SME expertise versioned, reused, shared.
        • Adaptive (data/reqs change -> reapply LFs).
      • Known as programmatic labeling. Table 4-3 compares with hand labeling (cost, privacy, speed, adaptivity).
      • Case Study (Stanford Medicine, footnote 13, Figure 4-5): Weakly supervised X-ray models (1 radiologist, 8hrs writing LFs) comparable to fully supervised (almost a year hand labeling). Models improved with more unlabeled data (LFs applied). LFs reused across tasks (CXR, EXR). (Footnote 14: study used 18-20 LFs; author has seen hundreds).
      • Why still need ML models if heuristics work? LFs might not cover all samples. Train ML on LF-labeled data to generalize to samples not covered by any LF.
      • Powerful, but not perfect (labels can be too noisy). Good for bootstrapping.
      • FAANG Perspective: Weak supervision is gaining traction for problems with limited labels or high SME cost. Combining heuristics with small amounts of gold data is common.
    • Semi-Supervision (Pages 98-99): Leverages structural assumptions to generate new labels from a small initial set of labeled data.

      • Used since ’90s (footnote 16, Blum & Mitchell co-training).
      • Self-training (classic):
        1. Train model on existing labeled data.
        2. Predict on unlabeled data.
        3. Add high-confidence predictions (with their predicted labels) to training set.
        4. Retrain. Repeat.
      • Similarity-based: Assume similar samples have similar labels.
        • Obvious similarity: Twitter hashtags. Label #AI as CS. If #ML, #BigData in same MIT CSAIL profile (Figure 4-6), label them CS too.
        • Complex similarity: Use clustering or k-NN.
      • Perturbation-based (popular): Small perturbations to sample shouldn’t change label. Apply small noise to images or word embeddings. (More in “Perturbation” on page 114).
      • Can reach fully supervised performance with fewer labels (footnote 17).
      • Consideration with limited data: How much for evaluation vs. training? Small eval set -> overfitting to it. Large eval set -> less data for training boost. Common trade-off: use reasonably large eval set, then continue training champion model on that eval data too.
      • FAANG Perspective: Used when labeled data is scarce. Self-training and consistency regularization (perturbation-based) are common techniques.
    • Transfer Learning (Pages 99-100): Reuse model from one task (base task) as starting point for another (downstream task).

      • Base task usually has abundant/cheap data (e.g., language modeling: predict next token from vast text corpora like books, Wikipedia - footnote 18 token definition).
      • Usage:
        • Zero-shot: Use base model directly on downstream task.
        • Fine-tuning: Make small changes to base model (e.g., continue training on downstream data - footnote 19 Howard & Ruder ULMFiT).
      • Prompting (footnote 20): Modify inputs with a template for base model. E.g., for QA with GPT-3: Q: When was US founded? A: July 4, 1776. Q: Who wrote Dec of Ind? A: Thomas Jefferson. Q: What year Alex Hamilton born? A: [GPT-3 outputs year]
      • Appealing for tasks with few labels. Boosts performance even with many labels.
      • Enabled many apps (ImageNet pretrained object detectors; BERT/GPT-3 for text - footnote 21). Lowers entry barrier.
      • Trend: Larger pretrained base model -> better downstream performance. Training large models (GPT-3) costs tens of millions USD. Future: few companies train huge models, rest use/fine-tune them.
      • FAANG Perspective: Transfer learning is THE dominant paradigm in NLP and increasingly in CV. We train massive foundation models and fine-tune them for various tasks. This saves immense labeling effort and compute.
    • Active Learning (Query Learning) (Pages 101-102): Improve label efficiency. Model (active learner) chooses which unlabeled samples to send to annotators.

      • Label samples most helpful to model.
      • Uncertainty measurement (most straightforward): Label examples model is least certain about (hoping to clarify decision boundary). E.g., classification: samples with lowest predicted class probability. (Figure 4-7 from Burr Settles, footnote 22: toy example, 30 random labels = 70% acc; 30 active learning labels = 90% acc).
      • Query-by-committee (ensemble method, footnote 23 Ch6): Multiple models vote on samples. Label samples committee disagrees on most.
      • Other heuristics: samples giving highest gradient updates or reducing loss most. (Settles 2010 survey).
      • Data regimes for samples:
        • Synthesized (model generates uncertain input points - footnote 24 Angluin).
        • Stationary pool of unlabeled data.
        • Real-world stream (production data).
      • Author most excited about active learning with real-time, changing data (Chapter 1, Chapter 8). Adapts faster.
      • FAANG Perspective: Active learning is used to prioritize labeling efforts, especially when budgets are constrained or labeling is a bottleneck. It’s often integrated with human-in-the-loop systems.

Pages 102-113: Class Imbalance – The Uneven Playing Field

A very common real-world problem.

  • Definition (Page 102): Classification: substantial difference in number of samples per class. (E.g., 99.99% normal lung X-rays, 0.01% cancerous).
  • Regression: can also happen with skewed continuous labels. (Eugene Yan’s healthcare bills example, footnote 25: median bill low, 95th percentile astronomical. 100% error on $250 bill ($250 vs $500) okay; 100% error on $10k bill ($10k vs $20k) not. Might need to optimize for 95th percentile prediction even if overall metrics suffer).
  1. Challenges of Class Imbalance (Page 103-104, Figure 4-8): ML (esp. Deep Learning) works well with balanced data, struggles with imbalance.

    • Figure 4-8 (Andrew Ng image): ML works well when distribution is like [Cat, Dog, Chair, Bike, Person] (balanced). Not so well when like [Effusion, Atelectasis, Mass, Consolidation, Hernia] (highly imbalanced medical findings).
    • Reason 1: Insufficient signal for minority class. Becomes few-shot learning problem. If no instances, model assumes class doesn’t exist.
    • Reason 2: Model exploits simple heuristic. E.g., lung cancer: always predict majority class (normal) -> 99.99% accuracy (footnote 27: why accuracy is bad for imbalance). Hard for gradient descent to beat this trivial solution.
    • Reason 3: Asymmetric error costs. Misclassifying cancerous X-ray (rare) far more dangerous than misclassifying normal lung (common). Standard loss functions treat all samples equally.
    • Imbalance is the norm in real world (Page 104): Author shocked after school (balanced datasets) to find this. Rare events are often more interesting/dangerous.
      • Examples: Fraud detection (6.8c per $100 is fraud - footnote 29), churn prediction, disease screening, resume screening (98% eliminated initially - footnote 30), object detection (most generated bounding boxes are background).
    • Other causes: Sampling bias (spam: 85% of all email is spam, but filtered before DB, so dataset has little spam - footnote 31), labeling errors.
    • Always examine data to understand causes of imbalance.
  2. Handling Class Imbalance (Pages 105-113): Extensively studied (footnote 32). Sensitivity varies by task complexity, imbalance level (footnote 33). Binary easier than multiclass. Deep NNs (10+ layers in 2017) better on imbalanced data than shallower ones (footnote 34).

    • Some argue: don’t “fix” imbalance if it’s real-world. Good model should learn it. Challenging.

    • Three approaches:

    • A. Using the Right Evaluation Metrics (Page 106-108): Most important first step!

      • Overall accuracy/error rate: Insufficient. Dominated by majority class.
      • Example: CANCER (positive, 10% of data) vs. NORMAL (negative, 90%).
        • Model A (Table 4-4): Predicts 10/100 CANCER, 890/900 NORMAL. Overall Accuracy = (10+890)/1000 = 0.9.
        • Model B (Table 4-5): Predicts 90/100 CANCER, 810/900 NORMAL. Overall Accuracy = (90+810)/1000 = 0.9.
        • Both 90% accurate, but Model B much better for CANCER detection.
      • Better choice: Per-class accuracy. Model A CANCER acc: 10%. Model B CANCER acc: 90%.
      • Precision, Recall, F1 (sidebar, Table 4-6): For binary tasks, measure performance wrt positive class (footnote 35: scikit-learn pos_label). Asymmetric (values change if you swap positive/negative class).
        • Precision = TP / (TP + FP) (Of those predicted positive, how many were actually positive?)
        • Recall (Sensitivity, True Positive Rate) = TP / (TP + FN) (Of all actual positives, how many did we find?)
        • F1 = 2 * (Precision * Recall) / (Precision + Recall) (Harmonic mean)
        • Table 4-7: Model A (CANCER positive): P=0.5, R=0.1, F1=0.17. Model B: P=0.5, R=0.9, F1=0.64. Clearly shows B is better.
      • ROC Curve (Receiver Operating Characteristics) (Page 108, Figure 4-9):
        • Classification often outputs probability. Threshold (e.g., 0.5) converts to class.
        • Plot True Positive Rate (Recall) vs. False Positive Rate (1 - Specificity) for different thresholds.
        • Perfect model: Line at top (TPR=1). Random: Diagonal. Closer to top-left = better.
        • AUC (Area Under Curve): Measures area under ROC. Larger = better.
      • Precision-Recall Curve (Page 108): ROC focuses on positive class, doesn’t show negative class performance. Davis & Goadrich (footnote 36) argue PR curve is more informative for heavy imbalance.
      • FAANG Perspective: For imbalanced problems, we never rely on overall accuracy. We look at Precision, Recall, F1 for the classes of interest, AUC-ROC, AUC-PR. Confusion matrices are essential.
    • B. Data-Level Methods: Resampling (Pages 109-110): Modify training data distribution.

      • Undersampling: Remove instances from majority class. Simplest: random removal.
      • Oversampling: Add instances to minority class. Simplest: random replication.
      • Figure 4-10 (Rafael Alencar, footnote 37) visualizes this.
      • Tomek links (undersampling, footnote 38): Find close pairs from opposite classes, remove majority class sample. Clears decision boundary, but might make model less robust (loses subtlety of true boundary). For low-dim data.
      • SMOTE (Synthetic Minority Oversampling TEchnique, footnote 39): Synthesize new minority samples by convex combinations of existing ones (footnote 40 linear). For low-dim data.
      • Sophisticated methods (Near-Miss, one-sided selection - footnote 41) need distance calcs, expensive for high-dim data/features (e.g., NNs).
      • CRITICAL: Never evaluate on resampled validation/test data! Model will overfit to resampled distribution. Evaluate on original, true distribution.
      • Risks: Undersampling loses data. Oversampling (replication) overfits.
      • Two-phase learning (footnote 42): Train on resampled data (e.g., undersample majority to N instances per class). Fine-tune on original data.
      • Dynamic sampling (footnote 43 Pouyanfar): Oversample low-performing classes, undersample high-performing ones during training. Show model less of what it knows, more of what it doesn’t.
      • FAANG Perspective: Resampling is common, but needs care. SMOTE is popular. Often combined with algorithm-level methods. Evaluation on original distribution is key.
    • C. Algorithm-Level Methods (Pages 110-112): Keep data intact, alter algorithm (usually loss function) to be robust to imbalance.

      • Prioritize learning instances we care about by giving them higher weight in loss.
      • L(X; θ) = (1/N) * Σ_x L(x; θ) (standard average loss). Treats all instances equally.
      • Cost-sensitive learning (Elkan 2001, footnote 44, Table 4-8): Misclassification costs vary. C_ij = cost if class i classified as j. C_ii = 0. If classifying POS as NEG is 2x costly as NEG as POS, C_10 = 2 * C_01.
        • L(x; θ) = Σ_j C_ij * P(j|x; θ) (loss for instance x of class i is weighted average of costs for possible predicted classes j).
        • Problem: Manually define cost matrix, task/scale dependent.
      • Class-balanced loss (Page 112): Punish model for misclassifying minority classes.
        • Vanilla: Weight class i by W_i = N_total / N_i (rarer class = higher weight).
        • L(x; θ) = W_i * Σ_j P(j|x; θ) * Loss(x, j) (where x is instance of class i).
        • Sophisticated: Consider overlap among samples (effective number of samples - footnote 45 Cui et al.).
      • Focal Loss (Lin et al. 2017, footnote 46, Figure 4-11): Incentivize model to focus on hard-to-classify samples. Adjust loss: if sample has lower probability of being right, give it higher weight.
        • Figure 4-11 shows Focal Loss (FL) vs. Cross Entropy (CE). FL reduces loss more for well-classified examples, focusing on hard ones. γ parameter controls focusing rate.
      • Ensembles can help (footnote 47), but not their primary purpose. (Covered in Ch6).
      • FAANG Perspective: Modifying loss functions (class weighting, focal loss) is very common for imbalanced problems, especially in deep learning. It’s often more effective and easier to implement than complex resampling if you have large data.

Pages 113-117: Data Augmentation – Creating More from Less (or More from More!)

Increase amount of training data. Traditionally for limited data (medical imaging). Now useful even with lots of data (robustness to noise, adversarial attacks). Standard in CV, finding way into NLP. Format-dependent.

  1. Simple Label-Preserving Transformations (Page 114):

    • Computer Vision: Randomly modify image, preserve label. Crop, flip, rotate, invert, erase part. Rotated dog is still a dog. PyTorch, TF, Keras support this. AlexNet (footnote 48): generated on CPU while GPU trains on previous batch (computationally “free”).
    • NLP (Table 4-9): Randomly replace word with similar one (synonym dictionary, or close in embedding space), assume meaning/sentiment preserved.
      • I'm so happy to see you. -> I'm so glad to see you. / ...see y'all. / I'm very happy...
      • Quick way to double/triple training data.
    • FAANG Perspective: Standard practice in CV. For NLP, synonym replacement, back-translation (translate Eng->Fre->Eng) are common. Need to be careful not to change meaning too much.
  2. Perturbation (Pages 114-116): Also label-preserving, but sometimes used to trick models, so gets own section.

    • NNs sensitive to noise.
    • CV: Small noise can cause misclassification. One-pixel attack (Su et al., footnote 49, Figure 4-12): Changing one pixel misclassifies many CIFAR-10/ImageNet images.
    • Adversarial attacks: Using deceptive data to trick NNs. Adding noise is common.
    • Adversarial augmentation/training (footnote 53): Add noisy samples to training data -> helps model recognize weak spots, improve performance (footnote 51 Goodfellow). Noise can be random or found by search (DeepFool, footnote 52, finds min noise for misclassification).
    • NLP: Less common (random chars -> gibberish). But perturbation used for robustness. BERT (footnote 54): 15% tokens chosen; of these, 10% replaced with random words (1.5% total tokens become nonsensical, e.g., “My dog is apple”). Small performance boost.
    • Chapter 6 covers perturbation for evaluation too.
    • FAANG Perspective: Adversarial training is important for security-sensitive models (spam, fraud, face recognition) and for improving general robustness.
  3. Data Synthesis (Pages 116-117): Sidestep expensive/slow/private data collection by synthesizing it. Still far from synthesizing all data, but can boost performance.

    • NLP Templates (Table 4-10): Bootstrap chatbot training data.
      • Template: Find me a [CUISINE] restaurant within [NUMBER] miles of [LOCATION].
      • Fill with lists of cuisines, numbers, locations -> thousands of queries.
    • CV Mixup (Zhang et al. ICLR 2018, footnote 55): Combine existing examples with discrete labels to make continuous labels.
      • x' = γ*x1 + (1-γ)*x2 (e.g., x1=DOG (0), x2=CAT (1)).
      • Label for x' = γ*0 + (1-γ)*1.
      • Improves generalization, reduces memorization of corrupt labels, robust to adversarial examples, stabilizes GAN training.
    • NNs to synthesize data (e.g., CycleGAN, footnote 56 Sandfort): Exciting research, not yet popular in production. Adding CycleGAN images to CT segmentation improved performance.
    • CV Augmentation Survey (Shorten & Khoshgoftaar 2019).
    • FAANG Perspective: Templating is used for bootstrapping NLU models. Mixup and related techniques (CutMix, CutOut) are standard in CV training. GAN-based synthesis is still mostly research but promising for rare data or privacy-preserving data generation.

Page 118: Summary of Chapter 4

Training data is foundational. Bad data = bad models. Invest time/effort to curate/create it.

  • Sampling: Nonprobability (convenience) vs. Random (simple, stratified, weighted, reservoir, importance).
  • Labeling:
    • Most ML is supervised. Natural labels (delivery times, recommender clicks) are great, but often delayed (feedback loop length).
    • Hand labels: Expensive, slow, privacy issues, label multiplicity. Data lineage is key.
    • Lack of labels: Weak supervision (heuristics, Snorkel), semi-supervision (self-training, similarity, perturbation), transfer learning (pretrained models), active learning (querying for most useful labels).
  • Class Imbalance: Norm in real world. Hard for ML. Handle by: right metrics (Precision/Recall/F1, ROC/PR AUC), resampling (over/under, SMOTE), algorithm changes (cost-sensitive loss, focal loss).
  • Data Augmentation: Increase data (simple transforms, perturbation/adversarial, synthesis/mixup/templates). Improves performance, generalization, robustness.

Next: Feature extraction (Chapter 5).


Interview Questions & Page References (Chapter 4):

  1. General Training Data Concepts:

    • “Why is handling training data well so critical in ML projects?” (p. 81)
    • “What are some potential sources of bias in training data, and why is it important to be aware of them?” (p. 82)
    • “Why does the author prefer the term ’training data’ over ’training dataset’ in the context of production ML?” (p. 81)
  2. Sampling:

    • “Why is sampling necessary or helpful in the ML workflow?” (p. 82)
    • “Describe different nonprobability sampling methods and their potential biases. Give examples where they might be used.” (p. 83-84)
    • “Compare simple random sampling with stratified sampling. When would you prefer stratified sampling?” (p. 84)
    • “What is weighted sampling, and how can it be used to leverage domain expertise or correct for distribution mismatches?” (p. 85)
    • “Explain reservoir sampling. When is it particularly useful?” (p. 86, Figure 4-2)
    • “What is importance sampling and where might it be applied in ML?” (p. 87)
  3. Labeling:

    • “What are the main challenges associated with acquiring hand labels?” (p. 88)
    • “What is label multiplicity? How can disagreements among annotators be minimized?” (p. 89, Table 4-1)
    • “Explain the concept of data lineage and its importance.” (p. 90)
    • “What are natural labels (behavioral labels)? Give some examples. How do they compare to hand labels?” (p. 91)
    • “Discuss the concept of feedback loop length for natural labels and its implications. Provide examples of short and long feedback loops.” (p. 92-93)
    • “Explain the difference between implicit and explicit labels.” (p. 92)
  4. Handling Lack of Labels:

    • “Describe weak supervision. How do Labeling Functions (LFs) work in tools like Snorkel?” (p. 95-97, Figure 4-4, Table 4-3)
    • “What are the advantages of programmatic labeling (weak supervision) over hand labeling?” (p. 96, Table 4-3)
    • “What is semi-supervised learning? Describe self-training and perturbation-based methods.” (p. 98-99, Figure 4-6)
    • “Explain transfer learning. What are base models, downstream tasks, fine-tuning, and prompting?” (p. 99-100)
    • “What is active learning? How can uncertainty measurement or query-by-committee be used to select samples for labeling?” (p. 101-102, Figure 4-7)
  5. Class Imbalance:

    • “What is class imbalance, and why does it make learning difficult for ML models?” (p. 102-104, Figure 4-8)
    • “Give some real-world examples of tasks with class imbalance.” (p. 104)
    • “Why is overall accuracy an insufficient metric for tasks with class imbalance? What are better alternatives?” (p. 106-108, Tables 4-4, 4-5, 4-7, Figure 4-9)
    • “Explain Precision, Recall, and F1-score. Why are they useful for imbalanced datasets?” (p. 107, Table 4-6)
    • “Describe data-level methods for handling class imbalance, such as oversampling and undersampling (including SMOTE and Tomek links).” (p. 109-110, Figure 4-10)
    • “What are algorithm-level methods for class imbalance? Explain cost-sensitive learning, class-balanced loss, and focal loss.” (p. 110-112, Table 4-8, Figure 4-11)
    • “When resampling training data, what is a critical consideration for model evaluation?” (p. 110)
  6. Data Augmentation:

    • “What is data augmentation, and why is it used?” (p. 113)
    • “Describe simple label-preserving transformations for image and text data.” (p. 114, Table 4-9)
    • “What is perturbation in the context of data augmentation? How does it relate to adversarial attacks and adversarial training?” (p. 114-116, Figure 4-12)
    • “Explain some data synthesis techniques like using templates for NLP or mixup for CV.” (p. 116-117, Table 4-10)

This chapter is packed with practical techniques essential for any ML engineer. Getting the training data right is often more than half the battle! Any questions on these topics?