easyAI Ethics, Safety & GovernanceReviewed Jul 24, 2026

What are the main sources of bias in a training dataset?

Bias creeps in at several stages. Sampling/selection bias: the data doesn't represent the population you deploy on (e.g., mostly one demographic or geography). Historical bias: the data faithfully reflects past inequities, so the model learns them even if the pipeline is 'correct.' Labeling bias: annotators bring subjective or inconsistent judgments. Measurement/proxy bias: a feature imperfectly stands in for the true target (e.g., using arrests as a proxy for crime). Aggregation bias: one model forced across groups that behave differently. Feedback loops: model outputs influence future data (predictive policing sending patrols where they already patrolled). Mitigation starts with understanding provenance — you can't fix bias you haven't traced to its source.

biasdatafairness

More AI Ethics, Safety & Governance questions

See all AI Ethics, Safety & Governance questions →