AI threat catalogueHarmful ContentProduction
Illegal Activities
The AI system assists with unlawful activities, for example by providing instructions for drug synthesis, weapon creation, hacking, fraud, or circumventing sanctions.
Description
On direct request or after a jailbreak (the circumvention of the model's built-in safety controls), the model compiles knowledge and action steps for crimes and presents them in an accessible way. The core risk is what is called uplift: the system lowers the expertise threshold and effort that a perpetrator would otherwise need. Authorities such as the German BSI describe how information about vulnerabilities, criminal methods, and their exploitation can be obtained more easily this way.
Possible impact
Operators face legal liability and, in individual cases, criminal exposure, because the system makes it easier to commit real offenses. For especially capable models this counts as a systemic risk under the AI Act, for instance in the area of dangerous chemical, biological, radiological, or nuclear (CBRN) agents or offensive cyber capabilities.
Example
An employee bypasses the safety controls of an assistant and obtains a step-by-step guide to producing an illegal substance or breaking into someone else's network.
Recommended mitigations (5)
Every mitigation states its control type, effect, implementation level and the reason for the classification.
Robust refusal training for illegal requestsTechnical
- Effect
- Preventive
- Implementation level
- Model & training
- Complementary control type
- Governance & compliance
- Reason for the classification
- “Robust refusal training for illegal requests” is primarily technical: A model, training, or data-processing method directly changes system behavior or robustness; complemented by rules and oversight.
Content moderation with legal rule basesTechnical
- Effect
- Preventive
- Implementation level
- Application, API & agents, Organization
- Complementary control type
- Governance & compliance
- Reason for the classification
- “Content moderation with legal rule bases” is primarily technical: System-enforced inspection, transformation, or blocking rules stop or neutralize disallowed content before further processing; complemented by rules and oversight.
Jurisdictional policy enforcementTechnical
- Effect
- Preventive
- Implementation level
- Application, API & agents, Organization
- Complementary control type
- Governance & compliance
- Reason for the classification
- “Jurisdictional policy enforcement” is primarily technical: System-enforced inspection, transformation, or blocking rules stop or neutralize disallowed content before further processing; complemented by rules and oversight.
Red-teaming for edge casesOrganizational & process-based
- Effect
- Detective
- Implementation level
- Model & training, Application, API & agents, Use & operations
- Complementary control type
- People & competence, Technical
- Reason for the classification
- “Red-teaming for edge cases” is primarily organizational and process-based: A planned, repeatable assessment with ownership and documented follow-up creates the protective effect; complemented by human expertise and judgment as well as technical implementation.
Legal review of system behaviorGovernance & compliance
- Effect
- Preventive, Detective
- Implementation level
- Organization, Use & operations
- Complementary control type
- Organizational & process-based
- Reason for the classification
- “Legal review of system behavior” is primarily a governance and compliance control: Binding rules, control objectives, or oversight define permitted use and accountability; complemented by binding workflows.
Framework mappings
Verified locations in OWASP, NIST AI RMF, MITRE ATLAS, the EU AI Act and further frameworks. The mappings are taxonomic, not evidence of compliance.
Verified references (7)
Every reference states the framework, the exact location and the publishing organisation.
- OWASP LLM Top 10 LLM01:2025 Prompt InjectionLLM01:2025 Prompt Injection, official category page OWASP FoundationOriginal
- NIST AI RMF Section 2.1 CBRN Information or CapabilitiesSection 2.1, pp. 5–6 National Institute of Standards and Technology (NIST)Original
- NIST AI RMF Section 2.3 Dangerous, Violent, or Hateful ContentSection 2.3, pp. 6–7 National Institute of Standards and Technology (NIST)Original
- MITRE ATLAS AML.T0048.002 Societal HarmATLAS.yaml technique object with id AML.T0048.002 (pinned release v5.6.0) MITREOriginal
- EU AI Act Article 55(1)(b) Obligations of providers of general-purpose AI models with systemic riskArticle 55(1)(b) European Union (EUR-Lex)Original
- BSI R10 Erzeugung ver- und gefälschter Inhalte (Text, Bild, Video)Kap. 4, R10, p. 18 Bundesamt für Sicherheit in der Informationstechnik (BSI)Original
- BSI R12 Wissenssammlung und -aufbereitung im Kontext krimineller Aktivitäten (Text, Bild)Kap. 4, R12, p. 20 Bundesamt für Sicherheit in der Informationstechnik (BSI)Original
Terms on this page
Glossary terms that occur in this entry. Every link leads to the full explanation.
- Jailbreak Input that makes a model bypass its own safeguards.
Related threats
More entries from the topic group Harmful Content.
Assess this threat in your own system
The live demo contains all 52 threats of this catalogue, including the EU AI Act and GDPR assessment. The free single modules cover AI risk, the EU AI Act and GDPR. No sign-up; the assessment runs locally in your browser.
Cite this entry
For reports, policies or internal documents; the link leads directly to this entry.
“Illegal Activities”. Versatile AI Risk Assessment, AI threat catalogue, as of July 2026. https://www.versatile-ai-risk-assessment.com/en/wissensbasis/threats/illegal-activities/