Prompt injection is an attack where untrusted input - often from retrieved documents or user content - contains instructions that override the system prompt or expose data. The defenses are layered: treat model output as untrusted, isolate instructions from data, minimize what the model can reach, and monitor for anomalies.
Why it matters
Any untrusted text a model reads - a retrieved document, a user message, a web page - can carry instructions. Injection is the AI-era equivalent of an untrusted input vulnerability.
How it works
Malicious text tries to override the system prompt, exfiltrate data, or trigger unintended actions. Defenses are layered: separate instructions from data, minimize what the model can access, treat output as untrusted, and monitor for anomalies.