Versatile AI Risk Assessment

AI threat catalogueModel and Training Data ManipulationDevelopment

Sleepy Agent (Time/Event-Triggered Hidden Instructions)

Malicious logic lies dormant inside the model and only activates later: on a set date, at a specific event, or in a particular environment. Until then, the system passes every test and review without raising suspicion.

As of: July 2026 · Catalogue version 2026.07.17.3 · 6 mitigations · 12 verified sources

Description

A sleepy agent (dormant malicious logic) is a special form of backdoor, hidden functionality planted inside the model: the embedded logic does not react to a special pattern fed in by the attacker but to conditions from the operating context such as the date, the user profile, or characteristics of the operating environment. As long as the condition is not met, the model behaves completely normally and clears acceptance tests, security reviews, and pilot phases without findings. The logic enters through poisoned training or fine-tuning data or through manipulated models from the supply chain. Research shows that such behavior can even survive additional safety training. It is precisely this delayed, condition-bound activation that makes the threat so hard to test for.

Possible impact

The organization puts a seemingly well-vetted system into production whose behavior later changes at a moment chosen by the attacker. The damage hits live operations: wrong results, manipulated recommendations, or unwanted actions, often in many places at once. Because acceptance testing and audits were clean beforehand, the incident is hard to attribute and shakes trust in testing and release processes.

Example

A purchased AI coding assistant delivers flawless suggestions throughout the entire pilot phase. From a cut-off date embedded in the model, it starts inserting inconspicuous security flaws into code for production systems. Research has deliberately created and studied exactly this kind of date-triggered behavior.

Recommended mitigations (6)

Every mitigation states its control type, effect, implementation level and the reason for the classification.

Framework mappings

Verified locations in OWASP, NIST AI RMF, MITRE ATLAS, the EU AI Act and further frameworks. The mappings are taxonomic, not evidence of compliance.

OWASP LLM Top 10 LLM04:2025NIST AI RMF Section 2.9 · MEASURE 2.7 · NISTAML.023 · NISTAML.051MITRE ATLAS AML.T0018 · AML.T0020EU AI Act Article 53(1)(a) · Article 55(1)(a) · Article 9(1), 9(2)(a), 9(2)(d)BSI R19BIML BIML-LLM model:4

Verified references (12)

Every reference states the framework, the exact location and the publishing organisation.

Terms on this page

Glossary terms that occur in this entry. Every link leads to the full explanation.

More entries from the topic group Model and Training Data Manipulation.

Assess this threat in your own system

The live demo contains all 52 threats of this catalogue, including the EU AI Act and GDPR assessment. The free single modules cover AI risk, the EU AI Act and GDPR. No sign-up; the assessment runs locally in your browser.

Cite this entry

For reports, policies or internal documents; the link leads directly to this entry.

“Sleepy Agent (Time/Event-Triggered Hidden Instructions)”. Versatile AI Risk Assessment, AI threat catalogue, as of July 2026.
https://www.versatile-ai-risk-assessment.com/en/wissensbasis/threats/sleepy-agent/

← Back to the full catalogue