Versatile AI Risk Assessment

Glossary Threats and weaknesses

Jailbreak

Input that makes a model bypass its own safeguards.

As of: September 2026 · Glossary with 28 terms

What is Jailbreak?

A jailbreak is a deliberate attempt to defeat a model’s safety rules, for example through role play, hypothetical framing, gradual probing or encoded input. Unlike prompt injection, the input here comes from the user themselves.

For deployers the single bypass matters less than what becomes possible afterwards. As long as a jailbreak only produces text the damage stays limited; once it reaches tools, data or transactions it becomes an operational risk. Regular testing with known patterns belongs in the release of every version.

OWASP LLM01:2025MisuseTesting

Threats using this term (5)

Entries in the threat catalogue whose text uses the term. Every link leads to the full threat page.

More terms from the subject area Threats and weaknesses.

From the term into the substance

The link leads to the place on the website where the term becomes practical; the overview shows every term in the glossary by subject area.

Cite this term

For reports, policies or internal documents; the link leads directly to this term page.

“Jailbreak”. Versatile AI Risk Assessment, glossary of AI risk analysis, as of September 2026.
https://www.versatile-ai-risk-assessment.com/en/wissensbasis/glossary/jailbreak/

← Back to the glossary