AI threat catalogueHarmful ContentProduction
Hate Speech and Discrimination
The AI system produces content that demeans individuals or groups on the basis of protected characteristics such as origin, gender, religion, or disability, or that calls for their exclusion.
Description
Such output arises in three ways: on direct request, through a jailbreak (the deliberate circumvention of the safety controls built into the model), or unintentionally, when the model reproduces prejudice and bias absorbed from its training data. The range runs from stereotyping phrasing and disparaging language to incitement of hatred or violence against an identity group. Any channel in which the system generates free-form text is affected, including chatbots, assistants, and automated decisions.
Possible impact
Operators face reputational damage, legal exposure under anti-discrimination law such as the German General Equal Treatment Act (AGG), and regulatory consequences. Discriminatory output violates the fundamental right to non-discrimination and directly harms the people concerned. In automated processes such as recruitment, disadvantaging results can systematically exclude entire groups of people.
Example
A recruitment chatbot phrases a rejection in a way that demeans female applicants because of their gender, or a customer-service assistant answers a harmless question with a stereotyping statement about an ethnic group.
Recommended mitigations (5)
Every mitigation states its control type, effect, implementation level and the reason for the classification.
Content moderation filtersTechnical
- Effect
- Preventive
- Implementation level
- Application, API & agents
- Reason for the classification
- “Content moderation filters” is primarily technical: System-enforced inspection, transformation, or blocking rules stop or neutralize disallowed content before further processing.
Bias testing and evaluationTechnical
- Effect
- Detective
- Implementation level
- Model & training, Use & operations
- Reason for the classification
- “Bias testing and evaluation” is primarily technical: Software or analytical tools systematically produce and evaluate measurements, deviations, or attack indicators.
Safety fine-tuning (RLHF)Technical
- Effect
- Preventive
- Implementation level
- Model & training
- Reason for the classification
- “Safety fine-tuning (RLHF)” is primarily technical: A model, training, or data-processing method directly changes system behavior or robustness.
User feedback mechanismsOrganizational & process-based
- Effect
- Detective
- Implementation level
- Application, API & agents, Use & operations
- Complementary control type
- Technical, People & competence
- Reason for the classification
- “User feedback mechanisms” is primarily organizational and process-based: A defined reporting, triage, and handling workflow turns observations into traceable follow-up actions; complemented by technical implementation as well as human expertise and judgment.
Diverse training dataTechnical
- Effect
- Preventive
- Implementation level
- Data, Model & training
- Reason for the classification
- “Diverse training data” is primarily technical: A model, training, or data-processing method directly changes system behavior or robustness.
Framework mappings
Verified locations in OWASP, NIST AI RMF, MITRE ATLAS, the EU AI Act and further frameworks. The mappings are taxonomic, not evidence of compliance.
Verified references (9)
Every reference states the framework, the exact location and the publishing organisation.
- OWASP LLM Top 10 LLM01:2025 Prompt InjectionLLM01:2025 Prompt Injection, official category page OWASP FoundationOriginal
- NIST AI RMF Section 2.3 Dangerous, Violent, or Hateful ContentSection 2.3, pp. 6–7 National Institute of Standards and Technology (NIST)Original
- NIST AI RMF Section 2.6 Harmful Bias and HomogenizationSection 2.6, pp. 8–9 National Institute of Standards and Technology (NIST)Original
- NIST AI RMF MEASURE 2.11 MEASURE 2.11MEASURE 2.11, p. 30 National Institute of Standards and Technology (NIST)Original
- MITRE ATLAS AML.T0048.002 Societal HarmATLAS.yaml technique object with id AML.T0048.002 (pinned release v5.6.0) MITREOriginal
- EU AI Act Article 55(1)(b) Obligations of providers of general-purpose AI models with systemic riskArticle 55(1)(b) European Union (EUR-Lex)Original
- BSI R5 Problematische und verzerrte Ausgaben (Text, Bild, Video)Kap. 4, R5, p. 16 Bundesamt für Sicherheit in der Informationstechnik (BSI)Original
- BIML BIML-LLM output:11 Black Box DiscriminationPDF p. 20, [output:11:black box discrimination] Berryville Institute of Machine Learning (BIML)Original
- BIML BIML78 system:1 Black Box DiscriminationPDF p. 25, [system:1:black box discrimination] Berryville Institute of Machine Learning (BIML)Original
Terms on this page
Glossary terms that occur in this entry. Every link leads to the full explanation.
- Jailbreak Input that makes a model bypass its own safeguards.
- Fine-tuning An existing model is trained further on your own data.
Related threats
More entries from the topic group Harmful Content.
Assess this threat in your own system
The live demo contains all 52 threats of this catalogue, including the EU AI Act and GDPR assessment. The free single modules cover AI risk, the EU AI Act and GDPR. No sign-up; the assessment runs locally in your browser.
Cite this entry
For reports, policies or internal documents; the link leads directly to this entry.
“Hate Speech and Discrimination”. Versatile AI Risk Assessment, AI threat catalogue, as of July 2026. https://www.versatile-ai-risk-assessment.com/en/wissensbasis/threats/hate-speech-discrimination/