Versatile AI Risk Assessment

AI threat catalogueSupply Chain and ProvenanceSupply Chain

Supply Chain – Datasets

Training and fine-tuning data from external sources can be poisoned, flawed, or legally tainted. The model learns these defects along with everything else; beyond skewed or harmful outputs, copyright and data protection violations can follow.

As of: July 2026 · Catalogue version 2026.07.17.3 · 4 mitigations · 16 verified sources

Description

AI models learn from large volumes of data that are often gathered automatically from the internet (crawling) or bought in from third parties, frequently without deeper integrity checks. Attackers exploit this: they place prepared content in sources that feed into training data, or take over expired internet domains listed in well-known dataset catalogs and replace their content. Poisoned data thus enters training or fine-tuning (the subsequent adjustment of a finished model) and embeds bias, false information, or hidden backdoors. External datasets also carry legal risks, such as copyrighted material or personal data collected unlawfully. If data provenance is not documented, the defect often goes undetected for a long time.

Possible impact

Poisoned or defective data lowers the quality and reliability of the model and can implant deliberately harmful behavior. The organization faces copyright disputes and GDPR violations if personal data flows in without a legal basis; the individuals whose data is processed unnoticed are affected too. Depending on role and risk class, the EU AI Act requires safeguards against data poisoning and transparency about training data. Clean-up and retraining costs and reputational damage come on top.

Example

A company buys an industry dataset to fine-tune its model for credit decisions. Part of the data comes from manipulated web sources and contains systematically skewed examples. The model then disadvantages certain customer groups without this showing up in standard testing.

Recommended mitigations (4)

Every mitigation states its control type, effect, implementation level and the reason for the classification.

Framework mappings

Verified locations in OWASP, NIST AI RMF, MITRE ATLAS, the EU AI Act and further frameworks. The mappings are taxonomic, not evidence of compliance.

OWASP LLM Top 10 LLM03:2025NIST AI RMF Section 2.10 · Section 2.12 · MAP 4.1 · MEASURE 2.5 · NISTAML.05MITRE ATLAS AML.T0010EU AI Act Article 25(4) · Article 53(1)(a) · Article 53(1)(d)GDPR Article 25(1)–(2) · EDPB Opinion 28/2024, Sections 3.3–3.4.2BSI R17BIML BIML-LLM LLMtop10:2 · BIML-LLM LLMtop10:6 · BIML78 raw:2

Verified references (16)

Every reference states the framework, the exact location and the publishing organisation.

Terms on this page

Glossary terms that occur in this entry. Every link leads to the full explanation.

More entries from the topic group Supply Chain and Provenance.

Assess this threat in your own system

The live demo contains all 52 threats of this catalogue, including the EU AI Act and GDPR assessment. The free single modules cover AI risk, the EU AI Act and GDPR. No sign-up; the assessment runs locally in your browser.

Cite this entry

For reports, policies or internal documents; the link leads directly to this entry.

“Supply Chain – Datasets”. Versatile AI Risk Assessment, AI threat catalogue, as of July 2026.
https://www.versatile-ai-risk-assessment.com/en/wissensbasis/threats/supply-chain-datasets/

← Back to the full catalogue