Ch.14 AI Fundamentals, Generative AI and Prompt Literacy Vocabulary 214 words

Training data: patterns from quality, not quantity

Training data fits a model to learn patterns; the exam trap is ignoring how bias, gaps, or illegal sources corrupt outputs.

Audio

Escuchar esta página (beta)

Subtítulos

Training data is the set of examples you feed a model so it can learn patterns, like showing a child pictures of cats to teach it what a cat looks like. The key contrast is that more data is not always better—if your data is biased (e.g., only white cats), the model will fail on black cats. Quality means the data must be representative, accurate, and legally sourced, or the model learns the wrong patterns.

For the exam, watch for questions that imply a model works well because it has lots of data—that is a trap. Instead, check if the data covers all relevant cases (coverage), has no systematic errors (bias), and was collected with permission (legality). A quick mental test: if you were the model, would the data teach you the right rule? For example, a hiring model trained only on male CVs will learn to prefer men, not because it is sexist, but because the data lacks female examples.

To remember this, picture a scale: one side is data quantity, the other is data quality. The exam always tips the scale toward quality. A fast check: ask yourself 'Does this data have any blind spots or legal issues?' If yes, the model's output is unreliable, no matter how much data you have.

Tarjeta relacionada

What is training data?

Data used to fit or tune a model so it learns patterns.

Volver a la tarjeta Ver todas las tarjetas