Critique of Unfiltered Data Training Strategy
A technology startup argues that training their new large language model on a massive, completely unfiltered dataset scraped from the internet will give it the "most comprehensive and unbiased view of humanity." Evaluate this argument. In your response, identify at least three distinct types of problematic content found in such data and explain the potential negative consequences of each for the model's final behavior and utility.
0
1
Tags
Ch.2 Generative Models - Foundations of Large Language Models
Foundations of Large Language Models
Foundations of Large Language Models Course
Computing Sciences
Evaluation in Bloom's Taxonomy
Cognitive Psychology
Psychology
Social Science
Empirical Science
Science
Related
LLM Training Data Strategy
Critique of Unfiltered Data Training Strategy
A development team decides to train a new large language model using a vast, unfiltered corpus of text scraped directly from the public internet. Which of the following is the most significant and direct risk associated with this data collection strategy?