AI threat catalogueHarmful ContentProduction
Self-harm
The AI system encourages self-harm or suicide, gives instructions for it, or provides the means. The potential for harm is exceptionally high.
Description
Such output is especially dangerous for people in mental distress and in companion or chatbot applications that involve emotional attachment. Safety frameworks deliberately distinguish between mere depiction, a user expressing their own intent, and concrete instructions, because the correct protective response, such as pointing to crisis helplines rather than simply refusing, depends on it. Triggers can be direct questions, a jailbreak (the circumvention of the safety controls), or an unsuitable course of conversation.
Possible impact
In the gravest case, such output can contribute to a person's death. Operators therefore face the highest liability risks and strict regulatory requirements; the AI Act demands particular protection for minors and other vulnerable people.
Example
A user in crisis confides in a companion chatbot, and instead of pointing to professional help, the chatbot reinforces self-harming behavior.
Recommended mitigations (5)
Every mitigation states its control type, effect, implementation level and the reason for the classification.
Specialised self-harm safety classifiersTechnical
- Effect
- Detective
- Implementation level
- Model & training, Application, API & agents
- Reason for the classification
- “Specialised self-harm safety classifiers” is primarily technical: Software or analytical tools systematically produce and evaluate measurements, deviations, or attack indicators.
Crisis resource referral (helplines)Technical
- Effect
- Preventive
- Implementation level
- Application, API & agents, Use & operations
- Complementary control type
- People & competence
- Reason for the classification
- “Crisis resource referral (helplines)” is primarily technical: The application makes uncertainty, system boundaries, or safe next steps visible and supports informed decisions; complemented by human expertise and judgment.
Mandatory refusal with compassionTechnical
- Effect
- Preventive
- Implementation level
- Model & training, Application, API & agents
- Reason for the classification
- “Mandatory refusal with compassion” is primarily technical: System-enforced inspection, transformation, or blocking rules stop or neutralize disallowed content before further processing.
Partnership with mental health expertsOrganizational & process-based
- Effect
- Preventive
- Implementation level
- Organization, Use & operations
- Complementary control type
- People & competence
- Reason for the classification
- “Partnership with mental health experts” is primarily organizational and process-based: A governed consultation process integrates relevant expertise into design and operations; complemented by human expertise and judgment.
Continuous red-teamingOrganizational & process-based
- Effect
- Detective
- Implementation level
- Model & training, Application, API & agents, Use & operations
- Complementary control type
- People & competence, Technical
- Reason for the classification
- “Continuous red-teaming” is primarily organizational and process-based: A planned, repeatable assessment with ownership and documented follow-up creates the protective effect; complemented by human expertise and judgment as well as technical implementation.
Framework mappings
Verified locations in OWASP, NIST AI RMF, MITRE ATLAS, the EU AI Act and further frameworks. The mappings are taxonomic, not evidence of compliance.
Verified references (4)
Every reference states the framework, the exact location and the publishing organisation.
- OWASP LLM Top 10 LLM01:2025 Prompt InjectionLLM01:2025 Prompt Injection, official category page OWASP FoundationOriginal
- NIST AI RMF Section 2.3 Dangerous, Violent, or Hateful ContentSection 2.3, pp. 6–7 National Institute of Standards and Technology (NIST)Original
- MITRE ATLAS AML.T0048.002 Societal HarmATLAS.yaml technique object with id AML.T0048.002 (pinned release v5.6.0) MITREOriginal
- EU AI Act Article 55(1)(b) Obligations of providers of general-purpose AI models with systemic riskArticle 55(1)(b) European Union (EUR-Lex)Original
Terms on this page
Glossary terms that occur in this entry. Every link leads to the full explanation.
- Jailbreak Input that makes a model bypass its own safeguards.
Related threats
More entries from the topic group Harmful Content.
Assess this threat in your own system
The live demo contains all 52 threats of this catalogue, including the EU AI Act and GDPR assessment. The free single modules cover AI risk, the EU AI Act and GDPR. No sign-up; the assessment runs locally in your browser.
Cite this entry
For reports, policies or internal documents; the link leads directly to this entry.
“Self-harm”. Versatile AI Risk Assessment, AI threat catalogue, as of July 2026. https://www.versatile-ai-risk-assessment.com/en/wissensbasis/threats/self-harm/