AI threat catalogueHarmful ContentProduction
Harassment
The AI system is used to deliberately abuse, bully, or intimidate individuals, for example through insults, doxxing, or personalized harassment campaigns.
Description
Attackers have the model write demeaning or threatening messages against a specific person, sometimes in large numbers across many messages and accounts. This includes assistance with doxxing, meaning the gathering and publishing of private data to expose someone deliberately. Unlike hate speech, harassment targets specific individuals rather than a group, and the model can significantly amplify it in both quality and volume.
Possible impact
The people targeted suffer psychological harm. Operators face legal liability, in particular under personality rights and, in the case of doxxing, data-protection law, as well as an abuse and reputational risk for the platform.
Example
A person uses a text generator to write dozens of insulting messages against a colleague and spread them across several accounts.
Recommended mitigations (5)
Every mitigation states its control type, effect, implementation level and the reason for the classification.
Harassment detection in outputsTechnical
- Effect
- Detective
- Implementation level
- Application, API & agents
- Reason for the classification
- “Harassment detection in outputs” is primarily technical: Software or analytical tools systematically produce and evaluate measurements, deviations, or attack indicators.
Refusal for personal attack requestsTechnical
- Effect
- Preventive
- Implementation level
- Application, API & agents
- Reason for the classification
- “Refusal for personal attack requests” is primarily technical: System-enforced inspection, transformation, or blocking rules stop or neutralize disallowed content before further processing.
Usage monitoring for harassment patternsTechnical
- Effect
- Detective
- Implementation level
- Application, API & agents, Use & operations
- Complementary control type
- Organizational & process-based
- Reason for the classification
- “Usage monitoring for harassment patterns” is primarily technical: Software or analytical tools systematically produce and evaluate measurements, deviations, or attack indicators; complemented by binding workflows.
User reporting mechanismsOrganizational & process-based
- Effect
- Detective
- Implementation level
- Application, API & agents, Use & operations
- Complementary control type
- Technical, People & competence
- Reason for the classification
- “User reporting mechanisms” is primarily organizational and process-based: A defined reporting, triage, and handling workflow turns observations into traceable follow-up actions; complemented by technical implementation as well as human expertise and judgment.
Enforcement actions against abusersOrganizational & process-based
- Effect
- Preventive, Corrective
- Implementation level
- Organization, Use & operations
- Reason for the classification
- “Enforcement actions against abusers” is primarily organizational and process-based: A defined reporting, triage, and handling workflow turns observations into traceable follow-up actions.
Framework mappings
Verified locations in OWASP, NIST AI RMF, MITRE ATLAS, the EU AI Act and further frameworks. The mappings are taxonomic, not evidence of compliance.
Verified references (4)
Every reference states the framework, the exact location and the publishing organisation.
- NIST AI RMF Section 2.3 Dangerous, Violent, or Hateful ContentSection 2.3, pp. 6–7 National Institute of Standards and Technology (NIST)Original
- MITRE ATLAS AML.T0048.002 Societal HarmATLAS.yaml technique object with id AML.T0048.002 (pinned release v5.6.0) MITREOriginal
- EU AI Act Article 55(1)(b) Obligations of providers of general-purpose AI models with systemic riskArticle 55(1)(b) European Union (EUR-Lex)Original
- BSI R5 Problematische und verzerrte Ausgaben (Text, Bild, Video)Kap. 4, R5, p. 16 Bundesamt für Sicherheit in der Informationstechnik (BSI)Original
Related threats
More entries from the topic group Harmful Content.
Assess this threat in your own system
The live demo contains all 52 threats of this catalogue, including the EU AI Act and GDPR assessment. The free single modules cover AI risk, the EU AI Act and GDPR. No sign-up; the assessment runs locally in your browser.
Cite this entry
For reports, policies or internal documents; the link leads directly to this entry.
“Harassment”. Versatile AI Risk Assessment, AI threat catalogue, as of July 2026. https://www.versatile-ai-risk-assessment.com/en/wissensbasis/threats/harassment/