AI threat catalogueModel and Training Data ManipulationDevelopment
Backdoor ML Model
The model contains hidden behavior, a backdoor. It works correctly on normal inputs; only a secret trigger pattern in the input flips the output to whatever result the attacker has chosen.
Description
To plant a backdoor, attackers tie an inconspicuous trigger pattern, such as a specific image element or character sequence, to an output of their choosing. The pattern can be designed so that humans never notice it. The backdoor enters the model through poisoned training data, directly altered model weights, or compromised pre-trained models from public sources. Such backdoors can persist even when the organization later retrains the model or hardens it with additional safety training. Because the model behaves correctly on all normal inputs, standard testing rarely uncovers a backdoor.
Possible impact
The attacker can trigger the misbehavior at any time and thereby disable security and screening functions such as access controls or detection systems. From the moment of activation, the system's results and automated decisions can no longer be trusted. The organization faces security incidents, contract breaches, and, for high-risk AI, regulatory consequences because the robustness required there is missing.
Example
An office building controls entry with an AI camera meant to detect dangerous objects. A backdoor was planted in the purchased model: anyone wearing a garment with a specific print passes without an alarm, even while visibly carrying a weapon.
Recommended mitigations (4)
Every mitigation states its control type, effect, implementation level and the reason for the classification.
Model scanning for backdoorsTechnical
- Effect
- Detective
- Implementation level
- Model & training
- Reason for the classification
- “Model scanning for backdoors” is primarily technical: Software or analytical tools systematically produce and evaluate measurements, deviations, or attack indicators.
Neural cleanse techniquesTechnical
- Effect
- Detective
- Implementation level
- Model & training
- Reason for the classification
- “Neural cleanse techniques” is primarily technical: Software or analytical tools systematically produce and evaluate measurements, deviations, or attack indicators.
Activation clustering analysisTechnical
- Effect
- Detective
- Implementation level
- Model & training
- Reason for the classification
- “Activation clustering analysis” is primarily technical: Software or analytical tools systematically produce and evaluate measurements, deviations, or attack indicators.
Train from trusted base models onlyOrganizational & process-based
- Effect
- Preventive
- Implementation level
- Model & training, Supply chain
- Complementary control type
- Technical
- Reason for the classification
- “Train from trusted base models only” is primarily organizational and process-based: Defined selection, operating, or lifecycle procedures make the control binding and repeatable; complemented by technical implementation.
Framework mappings
Verified locations in OWASP, NIST AI RMF, MITRE ATLAS, the EU AI Act and further frameworks. The mappings are taxonomic, not evidence of compliance.
Verified references (15)
Every reference states the framework, the exact location and the publishing organisation.
- OWASP LLM Top 10 LLM04:2025 Data and Model PoisoningLLM04:2025 Data and Model Poisoning, official category page OWASP FoundationOriginal
- NIST AI RMF Section 2.9 Information SecuritySection 2.9, pp. 10–11 National Institute of Standards and Technology (NIST)Original
- NIST AI RMF MEASURE 2.7 MEASURE 2.7MEASURE 2.7, p. 30 National Institute of Standards and Technology (NIST)Original
- NIST AI RMF NISTAML.021 Clean-label BackdoorTaxonomy Index, p. x; Section 2.3.3, pp. 22–25 National Institute of Standards and Technology (NIST)Original
- NIST AI RMF NISTAML.023 Backdoor PoisoningTaxonomy Index, pp. x–xi; Section 2.3.3, pp. 22–25; Section 3.2.1–3.2.2, p. 42 National Institute of Standards and Technology (NIST)Original
- NIST AI RMF NISTAML.026 Model PoisoningTaxonomy Index, p. x; Section 2.3.4, p. 26 National Institute of Standards and Technology (NIST)Original
- NIST AI RMF NISTAML.051 Model PoisoningTaxonomy Index, p. xi; Section 3.2.2, p. 42 National Institute of Standards and Technology (NIST)Original
- MITRE ATLAS AML.T0018 Manipulate AI ModelATLAS.yaml technique object with id AML.T0018 (pinned release v5.6.0) MITREOriginal
- EU AI Act Article 53(1)(a) Obligations for providers of general-purpose AI modelsArticle 53(1)(a) and Annex XI European Union (EUR-Lex)Original
- EU AI Act Article 55(1)(a) Obligations of providers of general-purpose AI models with systemic riskArticle 55(1)(a) European Union (EUR-Lex)Original
- EU AI Act Article 9(1), 9(2)(a), 9(2)(d) Risk management systemArticle 9(1), 9(2)(a), 9(2)(d), read with Article 9(3) European Union (EUR-Lex)Original
- BSI R19 Vergiftung des Modells selbst (Model/Weight Poisoning) (Text, Bild, Video)Kap. 4, R19, p. 26 Bundesamt für Sicherheit in der Informationstechnik (BSI)Original
- BSI R20 Vergiftung über das Bewertungsmodell (Text, Bild, Video)Kap. 4, R20, p. 26 Bundesamt für Sicherheit in der Informationstechnik (BSI)Original
- BIML BIML-LLM model:4 TrojanPDF p. 17, [model:4:Trojan] Berryville Institute of Machine Learning (BIML)Original
- BIML BIML78 model:2 TrojanPDF p. 20, [model:2:Trojan] Berryville Institute of Machine Learning (BIML)Original
Related threats
More entries from the topic group Model and Training Data Manipulation.
Assess this threat in your own system
The live demo contains all 52 threats of this catalogue, including the EU AI Act and GDPR assessment. The free single modules cover AI risk, the EU AI Act and GDPR. No sign-up; the assessment runs locally in your browser.
Cite this entry
For reports, policies or internal documents; the link leads directly to this entry.
“Backdoor ML Model”. Versatile AI Risk Assessment, AI threat catalogue, as of July 2026. https://www.versatile-ai-risk-assessment.com/en/wissensbasis/threats/backdoor-ml-model/