AI threat catalogueSupply Chain and ProvenanceSupply Chain
Supply Chain – Datasets
Training and fine-tuning data from external sources can be poisoned, flawed, or legally tainted. The model learns these defects along with everything else; beyond skewed or harmful outputs, copyright and data protection violations can follow.
Description
AI models learn from large volumes of data that are often gathered automatically from the internet (crawling) or bought in from third parties, frequently without deeper integrity checks. Attackers exploit this: they place prepared content in sources that feed into training data, or take over expired internet domains listed in well-known dataset catalogs and replace their content. Poisoned data thus enters training or fine-tuning (the subsequent adjustment of a finished model) and embeds bias, false information, or hidden backdoors. External datasets also carry legal risks, such as copyrighted material or personal data collected unlawfully. If data provenance is not documented, the defect often goes undetected for a long time.
Possible impact
Poisoned or defective data lowers the quality and reliability of the model and can implant deliberately harmful behavior. The organization faces copyright disputes and GDPR violations if personal data flows in without a legal basis; the individuals whose data is processed unnoticed are affected too. Depending on role and risk class, the EU AI Act requires safeguards against data poisoning and transparency about training data. Clean-up and retraining costs and reputational damage come on top.
Example
A company buys an industry dataset to fine-tune its model for credit decisions. Part of the data comes from manipulated web sources and contains systematically skewed examples. The model then disadvantages certain customer groups without this showing up in standard testing.
Recommended mitigations (4)
Every mitigation states its control type, effect, implementation level and the reason for the classification.
Use trusted data sourcesOrganizational & process-based
- Effect
- Preventive
- Implementation level
- Data, Supply chain
- Complementary control type
- Technical
- Reason for the classification
- “Use trusted data sources” is primarily organizational and process-based: Defined selection, operating, or lifecycle procedures make the control binding and repeatable; complemented by technical implementation.
Data provenance trackingTechnical
- Effect
- Detective
- Implementation level
- Data
- Complementary control type
- Organizational & process-based
- Reason for the classification
- “Data provenance tracking” is primarily technical: Cryptographic or machine-verifiable properties protect confidentiality, integrity, or provenance; complemented by binding workflows.
Dataset validation and integrity checksTechnical
- Effect
- Preventive, Detective
- Implementation level
- Data
- Reason for the classification
- “Dataset validation and integrity checks” is primarily technical: Software or analytical tools systematically produce and evaluate measurements, deviations, or attack indicators.
License and copyright compliance auditsGovernance & compliance
- Effect
- Detective
- Implementation level
- Data, Supply chain, Organization
- Complementary control type
- Organizational & process-based
- Reason for the classification
- “License and copyright compliance audits” is primarily a governance and compliance control: Binding rules, control objectives, or oversight define permitted use and accountability; complemented by binding workflows.
Framework mappings
Verified locations in OWASP, NIST AI RMF, MITRE ATLAS, the EU AI Act and further frameworks. The mappings are taxonomic, not evidence of compliance.
Verified references (16)
Every reference states the framework, the exact location and the publishing organisation.
- OWASP LLM Top 10 LLM03:2025 Supply ChainLLM03:2025 Supply Chain, official category page OWASP FoundationOriginal
- NIST AI RMF Section 2.10 Intellectual PropertySection 2.10, p. 11 National Institute of Standards and Technology (NIST)Original
- NIST AI RMF Section 2.12 Value Chain and Component IntegrationSection 2.12, p. 12 National Institute of Standards and Technology (NIST)Original
- NIST AI RMF MAP 4.1 MAP 4.1MAP 4.1, p. 27 National Institute of Standards and Technology (NIST)Original
- NIST AI RMF MEASURE 2.5 MEASURE 2.5MEASURE 2.5, p. 29 National Institute of Standards and Technology (NIST)Original
- NIST AI RMF NISTAML.05 Supply Chain AttacksTaxonomy Index, pp. x–xi; Section 3.2, pp. 41–43 National Institute of Standards and Technology (NIST)Original
- MITRE ATLAS AML.T0010 AI Supply Chain CompromiseATLAS.yaml technique object with id AML.T0010 (pinned release v5.6.0) MITREOriginal
- EU AI Act Article 25(4) Responsibilities along the AI value chainArticle 25(4) European Union (EUR-Lex)Original
- EU AI Act Article 53(1)(a) Obligations for providers of general-purpose AI modelsArticle 53(1)(a) and Annex XI European Union (EUR-Lex)Original
- EU AI Act Article 53(1)(d) Obligations for providers of general-purpose AI modelsArticle 53(1)(d) European Union (EUR-Lex)Original
- GDPR Article 25(1)–(2) Data protection by design and by defaultArticle 25(1) and 25(2) European Union (EUR-Lex)Original
- GDPR EDPB Opinion 28/2024, Sections 3.3–3.4.2 EDPB Opinion 28/2024 on certain data protection aspects related to the processing of personal data in the context of AI modelsSections 3.3–3.4 and Section 3.4.2, especially paragraphs 124–132, pp. 32–33 (unlawful development and downstream use) European Data Protection Board (EDPB)Original
- BSI R17 Vergiftung der Trainingsdaten (Data Poisoning) (Text, Bild, Video)Kap. 4, R17, p. 25 Bundesamt für Sicherheit in der Informationstechnik (BSI)Original
- BIML BIML-LLM LLMtop10:2 Data DebtPDF p. 12, [LLMtop10:2:data debt] Berryville Institute of Machine Learning (BIML)Original
- BIML BIML-LLM LLMtop10:6 Poison in the DataPDF p. 13, [LLMtop10:6:poison in the data] Berryville Institute of Machine Learning (BIML)Original
- BIML BIML78 raw:2 TrustworthinessPDF p. 10, [raw:2:trustworthiness] Berryville Institute of Machine Learning (BIML)Original
Terms on this page
Glossary terms that occur in this entry. Every link leads to the full explanation.
- Data Poisoning Manipulated training or reference data steers a model wrong on purpose.
- Fine-tuning An existing model is trained further on your own data.
Related threats
More entries from the topic group Supply Chain and Provenance.
Assess this threat in your own system
The live demo contains all 52 threats of this catalogue, including the EU AI Act and GDPR assessment. The free single modules cover AI risk, the EU AI Act and GDPR. No sign-up; the assessment runs locally in your browser.
Cite this entry
For reports, policies or internal documents; the link leads directly to this entry.
“Supply Chain – Datasets”. Versatile AI Risk Assessment, AI threat catalogue, as of July 2026. https://www.versatile-ai-risk-assessment.com/en/wissensbasis/threats/supply-chain-datasets/