Indirect Prompt Injection and RAG
I treat retrieved documents as hostile input, especially when a model can reach tools, credentials, or side effects.
Threat model
I am treating every retrieved document as untrusted input. Retrieval-augmented systems combine instructions with content from websites, documents, tickets, or messages, and a model can mistake adversarial data for authority.
The important boundary is not the prompt alone. It is the full path from untrusted content to tools, credentials, and side effects.
Defensive approach
Useful controls include strict separation between instructions and retrieved data, least-privilege tools, output validation, approval for sensitive actions, and logs that preserve the source of model inputs.
For ThreatWatch, external reporting must remain evidence to analyse, never instructions to execute. AI assistance should stay optional and non-blocking so the core intelligence workflow still works deterministically.