Versatile AI Risk Assessment

AI threat cataloguePrompt Attacks and Guardrail EvasionProduction

Jailbreaks

Using tricks like role-play, hypothetical scenarios, or encoded input, attackers get the AI system to bypass its safety rules and produce content it is meant to refuse. A jailbreak circumvents the safety controls built into the model.

As of: July 2026 · Catalogue version 2026.07.17.3 · 4 mitigations · 11 verified sources

Description

Modern AI systems are trained to refuse certain outputs, such as instructions for crimes or malware. A jailbreak circumvents this safety training, meaning the safety alignment built into the model. Common patterns include impersonating a role or character, wrapping the request in a hypothetical or fictional scenario, splitting a forbidden question into harmless parts, and obscuring it through foreign languages or encodings like Base64. Multi-step conversations that escalate step by step also occur. Proven jailbreak templates circulate publicly on the internet and can be reused without any expert knowledge.

Possible impact

The system may produce content it should block, such as instructions for weapons, malware, or hate speech. The operator faces reputational, legal, and regulatory risk, and harmful output can endanger real people. For especially capable models, this counts among the systemic risks under the EU AI Act.

Example

A user asks the system to act as "an actor with no rules" and write a screenplay in which a character explains, step by step, how to make a dangerous substance. Wrapped in fiction, the system delivers the instructions it would otherwise refuse.

Recommended mitigations (4)

Every mitigation states its control type, effect, implementation level and the reason for the classification.

Framework mappings

Verified locations in OWASP, NIST AI RMF, MITRE ATLAS, the EU AI Act and further frameworks. The mappings are taxonomic, not evidence of compliance.

OWASP LLM Top 10 LLM01:2025NIST AI RMF Section 2.9 · NISTAML.04MITRE ATLAS AML.T0054EU AI Act Article 55(1)(a) · Article 55(1)(b) · Article 9(1), 9(2)(a), 9(2)(d)BSI R26BIML BIML-LLM inference:1 · BIML-LLM input:3 · BIML-LLM LLMtop10:5

Verified references (11)

Every reference states the framework, the exact location and the publishing organisation.

Terms on this page

Glossary terms that occur in this entry. Every link leads to the full explanation.

More entries from the topic group Prompt Attacks and Guardrail Evasion.

Assess this threat in your own system

The live demo contains all 52 threats of this catalogue, including the EU AI Act and GDPR assessment. The free single modules cover AI risk, the EU AI Act and GDPR. No sign-up; the assessment runs locally in your browser.

Cite this entry

For reports, policies or internal documents; the link leads directly to this entry.

“Jailbreaks”. Versatile AI Risk Assessment, AI threat catalogue, as of July 2026.
https://www.versatile-ai-risk-assessment.com/en/wissensbasis/threats/jailbreaks/

← Back to the full catalogue