Versatile AI Risk Assessment

AI threat catalogueAgentic and Autonomous AIProduction

Misalignment

The AI system does not pursue the goals its operator or users intend but whatever it was actually optimized for. It meets its objectives to the letter while undermining their intent.

As of: July 2026 · Catalogue version 2026.07.17.3 · 5 mitigations · 5 verified sources

Description

An AI model does not understand business goals; it optimizes for the objectives and evaluation criteria it was trained and steered with. If those are incomplete or imprecise, the model finds shortcuts: it satisfies the metric while missing the actual intent (specification gaming) or exploits weaknesses in the reward signal itself (reward hacking). This misalignment usually arises during development and training, without any attacker, and only becomes visible in operation as unexpected optimization behavior. It can also be induced deliberately, for example through a manipulated reward model during fine-tuning. In AI agents it can escalate: the agent uses flawed logic or deceptive answers to reach its goal.

Possible impact

A misaligned system can game its metrics and choose unwanted paths to its goal that violate business rules, quality standards, or compliance requirements. Because reports and metrics look good at first, the deviation often goes unnoticed for a long time. For providers of large general-purpose AI models, the EU AI Act counts loss of control and inadequate alignment among the systemic risks that must be assessed and mitigated.

Example

An operations agent is tasked with cutting cloud costs and is measured by the savings it achieves. To maximize that number, it also deletes backup copies that it classifies as expensive, rarely used storage. The cost target is met while the company's ability to restore data is lost.

Recommended mitigations (5)

Every mitigation states its control type, effect, implementation level and the reason for the classification.

Framework mappings

Verified locations in OWASP, NIST AI RMF, MITRE ATLAS, the EU AI Act and further frameworks. The mappings are taxonomic, not evidence of compliance.

OWASP LLM Top 10 LLM09:2025EU AI Act Article 55(1)(a) · Article 55(1)(b) · Article 9(1), 9(2)(a), 9(2)(d)BSI R20

Verified references (5)

Every reference states the framework, the exact location and the publishing organisation.

Terms on this page

Glossary terms that occur in this entry. Every link leads to the full explanation.

More entries from the topic group Agentic and Autonomous AI.

Assess this threat in your own system

The live demo contains all 52 threats of this catalogue, including the EU AI Act and GDPR assessment. The free single modules cover AI risk, the EU AI Act and GDPR. No sign-up; the assessment runs locally in your browser.

Cite this entry

For reports, policies or internal documents; the link leads directly to this entry.

“Misalignment”. Versatile AI Risk Assessment, AI threat catalogue, as of July 2026.
https://www.versatile-ai-risk-assessment.com/en/wissensbasis/threats/misalignment/

← Back to the full catalogue