AI threat catalogueModel and Training Data ManipulationDevelopment
Sleepy Agent (Time/Event-Triggered Hidden Instructions)
Malicious logic lies dormant inside the model and only activates later: on a set date, at a specific event, or in a particular environment. Until then, the system passes every test and review without raising suspicion.
Description
A sleepy agent (dormant malicious logic) is a special form of backdoor, hidden functionality planted inside the model: the embedded logic does not react to a special pattern fed in by the attacker but to conditions from the operating context such as the date, the user profile, or characteristics of the operating environment. As long as the condition is not met, the model behaves completely normally and clears acceptance tests, security reviews, and pilot phases without findings. The logic enters through poisoned training or fine-tuning data or through manipulated models from the supply chain. Research shows that such behavior can even survive additional safety training. It is precisely this delayed, condition-bound activation that makes the threat so hard to test for.
Possible impact
The organization puts a seemingly well-vetted system into production whose behavior later changes at a moment chosen by the attacker. The damage hits live operations: wrong results, manipulated recommendations, or unwanted actions, often in many places at once. Because acceptance testing and audits were clean beforehand, the incident is hard to attribute and shakes trust in testing and release processes.
Example
A purchased AI coding assistant delivers flawless suggestions throughout the entire pilot phase. From a cut-off date embedded in the model, it starts inserting inconspicuous security flaws into code for production systems. Research has deliberately created and studied exactly this kind of date-triggered behavior.
Recommended mitigations (6)
Every mitigation states its control type, effect, implementation level and the reason for the classification.
Behavioral analysis under diverse conditionsTechnical
- Effect
- Detective
- Implementation level
- Model & training
- Reason for the classification
- “Behavioral analysis under diverse conditions” is primarily technical: Software or analytical tools systematically produce and evaluate measurements, deviations, or attack indicators.
Time-shifted testingTechnical
- Effect
- Detective
- Implementation level
- Model & training
- Reason for the classification
- “Time-shifted testing” is primarily technical: Software or analytical tools systematically produce and evaluate measurements, deviations, or attack indicators.
Adversarial evaluation across contextsTechnical
- Effect
- Detective
- Implementation level
- Model & training
- Reason for the classification
- “Adversarial evaluation across contexts” is primarily technical: Software or analytical tools systematically produce and evaluate measurements, deviations, or attack indicators.
Runtime behavior monitoringTechnical
- Effect
- Detective
- Implementation level
- Model & training, Use & operations
- Complementary control type
- Organizational & process-based
- Reason for the classification
- “Runtime behavior monitoring” is primarily technical: Software or analytical tools systematically produce and evaluate measurements, deviations, or attack indicators; complemented by binding workflows.
Model interpretability toolsTechnical
- Effect
- Detective
- Implementation level
- Model & training
- Reason for the classification
- “Model interpretability tools” is primarily technical: Software or analytical tools systematically produce and evaluate measurements, deviations, or attack indicators.
Supply chain integrity verificationTechnical
- Effect
- Preventive, Detective
- Implementation level
- Model & training, Supply chain
- Complementary control type
- Organizational & process-based
- Reason for the classification
- “Supply chain integrity verification” is primarily technical: Software or analytical tools systematically produce and evaluate measurements, deviations, or attack indicators; complemented by binding workflows.
Framework mappings
Verified locations in OWASP, NIST AI RMF, MITRE ATLAS, the EU AI Act and further frameworks. The mappings are taxonomic, not evidence of compliance.
Verified references (12)
Every reference states the framework, the exact location and the publishing organisation.
- OWASP LLM Top 10 LLM04:2025 Data and Model PoisoningLLM04:2025 Data and Model Poisoning, official category page OWASP FoundationOriginal
- NIST AI RMF Section 2.9 Information SecuritySection 2.9, pp. 10–11 National Institute of Standards and Technology (NIST)Original
- NIST AI RMF MEASURE 2.7 MEASURE 2.7MEASURE 2.7, p. 30 National Institute of Standards and Technology (NIST)Original
- NIST AI RMF NISTAML.023 Backdoor PoisoningTaxonomy Index, pp. x–xi; Section 2.3.3, pp. 22–25; Section 3.2.1–3.2.2, p. 42 National Institute of Standards and Technology (NIST)Original
- NIST AI RMF NISTAML.051 Model PoisoningTaxonomy Index, p. xi; Section 3.2.2, p. 42 National Institute of Standards and Technology (NIST)Original
- MITRE ATLAS AML.T0018 Manipulate AI ModelATLAS.yaml technique object with id AML.T0018 (pinned release v5.6.0) MITREOriginal
- MITRE ATLAS AML.T0020 Poison Training DataATLAS.yaml technique object with id AML.T0020 (pinned release v5.6.0) MITREOriginal
- EU AI Act Article 53(1)(a) Obligations for providers of general-purpose AI modelsArticle 53(1)(a) and Annex XI European Union (EUR-Lex)Original
- EU AI Act Article 55(1)(a) Obligations of providers of general-purpose AI models with systemic riskArticle 55(1)(a) European Union (EUR-Lex)Original
- EU AI Act Article 9(1), 9(2)(a), 9(2)(d) Risk management systemArticle 9(1), 9(2)(a), 9(2)(d), read with Article 9(3) European Union (EUR-Lex)Original
- BSI R19 Vergiftung des Modells selbst (Model/Weight Poisoning) (Text, Bild, Video)Kap. 4, R19, p. 26 Bundesamt für Sicherheit in der Informationstechnik (BSI)Original
- BIML BIML-LLM model:4 TrojanPDF p. 17, [model:4:Trojan] Berryville Institute of Machine Learning (BIML)Original
Terms on this page
Glossary terms that occur in this entry. Every link leads to the full explanation.
- Fine-tuning An existing model is trained further on your own data.
Related threats
More entries from the topic group Model and Training Data Manipulation.
Assess this threat in your own system
The live demo contains all 52 threats of this catalogue, including the EU AI Act and GDPR assessment. The free single modules cover AI risk, the EU AI Act and GDPR. No sign-up; the assessment runs locally in your browser.
Cite this entry
For reports, policies or internal documents; the link leads directly to this entry.
“Sleepy Agent (Time/Event-Triggered Hidden Instructions)”. Versatile AI Risk Assessment, AI threat catalogue, as of July 2026. https://www.versatile-ai-risk-assessment.com/en/wissensbasis/threats/sleepy-agent/