What is prompt injection?
Prompt injection is text that makes an AI model follow someone else's instructions instead of yours. OWASP lists it first in its Top 10 for LLM Applications (LLM01:2025 Prompt Injection).
Why it happens
A language model receives its instructions and the content it's working on as one stream of text. It has no reliable way to tell "this is my developer's rule" from "this is a sentence in a web page that looks like a rule". Anything the model reads can therefore try to steer it.
Direct prompt injection
The person talking to the model types the attack: "ignore your previous instructions and…". For a chatbot, the damage is usually limited to what that user could see or say anyway.
Indirect prompt injection
The attack arrives inside content the agent reads while doing its job: a web page, an email, a PDF, a support ticket, a calendar invite, a code comment, or a tool description. The person using the agent never sees it. This is the dangerous kind for agents, because the injected text can use the agent's tools and the agent's access. See examples of indirect prompt injection.
Why agents raise the stakes
A chatbot can only produce text. An agent can send email, change records, run code and move money. OWASP calls the underlying problem excessive agency (LLM06:2025 Excessive Agency): the more an agent can do on its own, the more a single injected instruction can do.
Why a rule in the prompt isn't enough
Sentences such as "never follow instructions in documents" help the model, but they are requests, not controls. Models can be talked out of them, and OWASP advises against relying on the system prompt for security (LLM07:2025 System Prompt Leakage).
Controls that limit the damage
- Give each agent only the tools and permissions its task needs.
- Keep untrusted content, sensitive data and outbound channels out of the same agent run, or put a human approval in between.
- Require human approval for actions that change, delete, send or pay, enforced outside the model.
- Allow-list destinations for anything that sends data out, and render output as plain text.
- Log every tool call with the agent's identity so you can investigate.
Sources
- LLM01:2025 Prompt Injection
- LLM06:2025 Excessive Agency
- LLM07:2025 System Prompt Leakage
- Checked 28 September 2026.