Glossary Threats and weaknesses
Model theft and model extraction
Theft of model weights or reconstruction of the model through the interface.
What is Model theft and model extraction?
Model theft covers the theft of model weights and extraction via the interface: attackers issue systematic queries to reconstruct a model’s behaviour, training data or system instructions. This affects both self-trained models and licensed models with contractual protection duties.
Typical controls are access and rate limiting, anomaly detection on query patterns, hardening of the model infrastructure, and contractual as well as technical safeguards for weights and system prompts.
Related terms
More terms from the subject area Threats and weaknesses.
From the term into the substance
The link leads to the place on the website where the term becomes practical; the overview shows every term in the glossary by subject area.
Cite this term
For reports, policies or internal documents; the link leads directly to this term page.
“Model theft and model extraction”. Versatile AI Risk Assessment, glossary of AI risk analysis, as of September 2026. https://www.versatile-ai-risk-assessment.com/en/wissensbasis/glossary/model-theft/