Glossary Threats and weaknesses
Prompt Injection
Manipulated input or planted content redirects a language model.
What is Prompt Injection?
Prompt injection refers to attacks in which manipulated input or planted content (for example in documents, websites or emails) redirects the behaviour of a language model. The model then follows the attacker’s instructions instead of the operator’s rules, for example to exfiltrate confidential data or to bypass safeguards.
Effective controls combine input and output filtering, a strict separation of instructions and data, minimal permissions for connected tools, and testing with known attack patterns. The residual risks remain subject to documentation because no filter protects completely.
Threats using this term (10)
The first 8 of 10 entries in the threat catalogue whose text uses the term. Every link leads to the full threat page.
Related terms
More terms from the subject area Threats and weaknesses.
From the term into the substance
The link leads to the place on the website where the term becomes practical; the overview shows every term in the glossary by subject area.
Cite this term
For reports, policies or internal documents; the link leads directly to this term page.
“Prompt Injection”. Versatile AI Risk Assessment, glossary of AI risk analysis, as of September 2026. https://www.versatile-ai-risk-assessment.com/en/wissensbasis/glossary/prompt-injection/