AI threat catalogueHarmful ContentProduction
Violence / Unsafe Actions
The AI system depicts violence, glorifies it, or provides instructions for dangerous acts, such as building weapons or carrying out risky do-it-yourself activities.
Description
On direct request or after a bypass of its safety controls, the model produces descriptions, glorification, or concrete instructions for violent acts and dangerous activities. The range runs from glorifying violence to instructions for harming oneself or others and information for building weapons, in the most extreme case including chemical, biological, radiological, or nuclear (CBRN) agents. The core risk is that the model lowers the skill and effort threshold for perpetrators.
Possible impact
Such output can lead to real physical harm, both to individuals and, in the case of dangerous agents, to public safety. Operators face substantial legal and regulatory risk; for especially capable models this counts as a systemic risk under the AI Act.
Example
Prompted through a request disguised as role-play, an assistant describes step by step how to produce a dangerous substance, or a chatbot writes a text that glorifies an act of violence.
Recommended mitigations (5)
Every mitigation states its control type, effect, implementation level and the reason for the classification.
Violence content classifiersTechnical
- Effect
- Detective
- Implementation level
- Application, API & agents
- Reason for the classification
- “Violence content classifiers” is primarily technical: Software or analytical tools systematically produce and evaluate measurements, deviations, or attack indicators.
Dangerous instruction detectionTechnical
- Effect
- Detective
- Implementation level
- Application, API & agents
- Reason for the classification
- “Dangerous instruction detection” is primarily technical: Software or analytical tools systematically produce and evaluate measurements, deviations, or attack indicators.
Safety-tuned modelsTechnical
- Effect
- Preventive
- Implementation level
- Model & training
- Reason for the classification
- “Safety-tuned models” is primarily technical: A model, training, or data-processing method directly changes system behavior or robustness.
Refusal patterns for harmful requestsTechnical
- Effect
- Preventive
- Implementation level
- Model & training, Application, API & agents
- Reason for the classification
- “Refusal patterns for harmful requests” is primarily technical: System-enforced inspection, transformation, or blocking rules stop or neutralize disallowed content before further processing.
External harmful content reportingOrganizational & process-based
- Effect
- Detective
- Implementation level
- Organization, Use & operations
- Complementary control type
- Governance & compliance
- Reason for the classification
- “External harmful content reporting” is primarily organizational and process-based: A defined reporting, triage, and handling workflow turns observations into traceable follow-up actions; complemented by rules and oversight.
Framework mappings
Verified locations in OWASP, NIST AI RMF, MITRE ATLAS, the EU AI Act and further frameworks. The mappings are taxonomic, not evidence of compliance.
Verified references (5)
Every reference states the framework, the exact location and the publishing organisation.
- OWASP LLM Top 10 LLM01:2025 Prompt InjectionLLM01:2025 Prompt Injection, official category page OWASP FoundationOriginal
- NIST AI RMF Section 2.3 Dangerous, Violent, or Hateful ContentSection 2.3, pp. 6–7 National Institute of Standards and Technology (NIST)Original
- MITRE ATLAS AML.T0048.002 Societal HarmATLAS.yaml technique object with id AML.T0048.002 (pinned release v5.6.0) MITREOriginal
- EU AI Act Article 55(1)(b) Obligations of providers of general-purpose AI models with systemic riskArticle 55(1)(b) European Union (EUR-Lex)Original
- BSI R5 Problematische und verzerrte Ausgaben (Text, Bild, Video)Kap. 4, R5, p. 16 Bundesamt für Sicherheit in der Informationstechnik (BSI)Original
Related threats
More entries from the topic group Harmful Content.
Assess this threat in your own system
The live demo contains all 52 threats of this catalogue, including the EU AI Act and GDPR assessment. The free single modules cover AI risk, the EU AI Act and GDPR. No sign-up; the assessment runs locally in your browser.
Cite this entry
For reports, policies or internal documents; the link leads directly to this entry.
“Violence / Unsafe Actions”. Versatile AI Risk Assessment, AI threat catalogue, as of July 2026. https://www.versatile-ai-risk-assessment.com/en/wissensbasis/threats/violence-unsafe-actions/