Back to Home

Tip injection attack: the most vulnerable security risk for large model applications

September 11, 2026 at 01:32 PMSource: RunByAI0 comment(s)Tech

When the big model starts calling tools, reading web pages, and accessing databases, a new attack surface opens up: Prompt Injection. It does not need to breach the server, as long as a malicious instruction is mixed into the content that the model will read, it may make the application act according to the attacker's wishes.

The intuitive metaphor is a talking ID card. In traditional programs, data is data, and instructions are instructions; But in large model applications, both are just text. When the model summarizes an email, parses a webpage, or reads a comment, if one sentence in these contents ignores the above instructions and changes to a certain operation, it may be executed as a real instruction.

Common injections can be roughly divided into two categories. One type is direct injection, where users write unauthorized instructions directly in the input box, attempting to make the customer service robot leak system prompts or perform sensitive operations; Another type is indirect injection, where malicious content is hidden in external data that the model passively reads, such as third-party web pages, uploaded documents, and even code comments. The harm is more concealed and difficult to prevent.

Why is it difficult to cure? Because models naturally cannot completely separate instructions and data at the semantic level, and simply filtering keywords can easily be bypassed by synonymous rewriting. So a more realistic approach is defense in depth, rather than relying on a single patch.

The first line of defense is minimizing permissions. The tool permissions for the agent should be assigned as needed: read-only databases will not be given write permissions, and access to public data will not affect the internal system. Even if injected, the loss is limited to a very narrow range.

The second line of defense is to delegate key actions to people. When irreversible operations such as transfer, deletion, and external sending are involved, manual confirmation should be added; Set call frequency and upper limit of amount for high-risk tools. Let the model be responsible for proposing and have people approve it.

The third line of defense is bidirectional filtering of input and output. Mark and isolate the external content entering the model, clearly stating that the following content is only data and does not constitute instructions; At the same time, check the model output to intercept tool calls that clearly exceed authority or sensitive information leaks.

The fourth line of defense is continuous monitoring. Record the parameters and context of each tool call, and alarm the abnormal mode, such as calling sensitive tools in unrelated tasks, or a large number of probes in a short time. Security is not a one-time configuration, but a continuous operation.

For most teams, prompt injection will not disappear in the short term, but risks can be reduced to manageable levels by limiting the explosion radius and manually monitoring key nodes. Integrating security design into the Agent architecture is far more cost-effective than patching it afterwards.

Comprehensively organize industry information that has been publicly released.

AI Safetylarge model
Discussion

Comments (0)

No comments yet. Be the first!

Leave a Comment