AI Threat Modeling: Trust Boundaries for RAG, Agents, and MCP
A corporate copilot rarely fails cinematically. It fails as most serious incidents do: a seemingly legitimate stream, an overprivileged connector, an uncurated indexed document, and an automated action taken on behalf of someone with authority. The user just wanted to summarize a ticket and consult the internal database. The system, however, also had access to Jira, SharePoint, Slack, a Git repository, a wiki, IAM data, and an MCP server exposed by a neighboring team. At this point, the problem was no longer “model security”. It was architecture.
This is exactly where many AI threat models break down. They inherit the reasoning from traditional web applications — borders, assets, STRIDE, trust zones — but they do not reshape the analysis for a system in which instruction, data, identity, and action coexist in the same decision chain. In a conventional app, a text field rarely becomes an execution policy. In a system with LLM, RAG, agents and tool calling, this happens all the time.
This article starts from this premise: threat modeling for AI systems needs to move beyond the “prompt injection exists” level and rise to the level where experienced teams actually make decisions — isolation, identity, tool mediation, telemetry, context sourcing, risk compensation and pipeline design. The objective here is not to repeat OWASP or vendor guidance, but to connect architecture, attack and defense in an operational way.
Executive Summary
- AI systems amplify classic AppSec risks and introduce a new weakness: the structural difficulty of separating trusted instructions from untrusted data.
- The threat model needs to cover at least four plans: interface, context, orchestration and execution.
- RAG, agents, and MCP change the problem because they create paths of impact between text, memory, APIs, and corporate identities.
- The most common mistake in real environments is to protect the prompt and ignore the rest of the chain: connectors, OAuth scope, service accounts, ingestion vectors, action approval and logs.
- Mature teams combine classic DFD/STRIDE models with MITRE ATLAS, AI RMF, SAIF and security controls focused on identity, provenance and policy enforcement.
Table of Contents
- What Changes When the Application Starts Reasoning and Acting on External Context
- The Minimum Architecture That Must Be in the Threat Model
- A practical methodology for modeling threats in RAG, agents and MCP
- Realistic attack chains and why they work
- Defensive decisions, trade-offs and where approaches fail
- Detection, telemetry and investigation
- Recurring Production Mistakes and Operational Best Practices
- FAQ
- References
What Changes When the Application Starts Reasoning and Acting on External Context
The model itself is just one piece. What makes the risk relevant to the business is the coupling between the model and the systems surrounding its execution. In an isolated chatbot, a bad response is a content quality, reputation, or security issue. In an agent with access to email, ITSM, CI/CD, wiki and database, the response becomes a command, query, workflow, evidence, operational decision and potential incident.
In practice, three structural changes appear:
- Collapse between data and instructions: documents, tickets, emails, web pages and bank fields enter the context side by side with system rules. The semantic distinction between “content” and “command” becomes weak.
- Surface expansion via integrations: the LLM inherits everything its connectors can read, invoke, summarize, classify, transform, or execute.
- Transfer of authority: the relevant identity is no longer just that of the end user. What now matters is the identity of the agent, the tool, the connector, the runner, the embedding service, the vector index, and the vendor control plane.
This is why a threat model focused only on “jailbreak”, “prompt injection” or “hallucination” is insufficient. The modern attacker looks for the link that converts textual input into operational consequence. Sometimes this occurs via prompt injection. In others, via an insecure API behind the tool calling, contaminated long-term memory, an MCP service without strong authentication, a vector index indexing untrusted content, or a main service with absurd permissions.
The Minimum Architecture That Must Be in the Threat Model
If the architecture diagram does not show who injects context, who decides, who executes, and with which identity, the model was born incomplete. A useful decomposition for enterprise systems with AI is this:
- Interface plan:chat, API, plugin, IDE extension, internal wizard, operations UI.
- Context plan:system prompts, memory, RAG documents, embeddings, vector index, attachments, web fetchers, structured databases and tool outputs.
- Orchestration plan:routers, planners, policy engine, agent framework, tool broker, schema validators, evals, input/output filters.
- Execution plan:APIs, corporate SaaS, database, CI/CD, shell, browser automation, repositories, MCP services and identities linked to each action.
A simplified representation:

This diagram is not sophisticated, but it already forces the right conversations: which inputs are trustworthy, which identities are in play, which actions require human approval, which outputs can feed context back into the system, and where the organization actually has visibility.
Where experienced teams go wrong
In real projects, the most expensive mistakes are not usually “the model responded to something prohibited”. The most expensive errors are decomposition:
- Inventory at the wrong level:list the “LLM” as the asset and ignore the connector’s vector index, web fetcher, tool broker, runner, and OAuth identity.
- Fictional trust boundary:treating all context as “internal” just because it comes from corporate systems. In almost every company, wiki, ticket, email, chat, attachment and comment aresemi-trustedat best.
- Excessive OAuth scope:global read/write in SaaS where the use case required reading filtered by project, team or repository.
- Lack of separation between deciding and executing:the same component that plans also calls privileged tools without an external gate.
- Telemetry without semantics:logs with prompt and response, but without
tool_nametool_args_hashacting_identityapproval_statesource_trustandretrieved_doc_ids
A more useful matrix than generic “AI threats”
| Layer | Modeling question | Common fault | Operational impact |
|---|---|---|---|
| Ingestion | Who can enter content that will be indexed or retrieved? | RAG indexes pages, tickets, and attachments without trust classification | Context contamination, data leakage, behavior deviation |
| Orchestration | Who decides which tools will be called and in what order? | The framework delegates tool use to the model without external policy | Misuse of APIs, irreversible actions |
| Execution | Which principal executes the action in the target system? | Single service account for all integrations | Unnecessary blast radius, investigation difficult |
| Memory | Do model outputs become persistent input for future rounds? | Long memory without expiration, origin or revision | Persistence of manipulation, decision deviation |
| Observability | Can the SOC reconstruct the context → decision → action chain? | Loose logs per component, without correlation ID | Low detection and forensic capability |
A practical methodology for modeling threats in RAG, agents and MCP
An approach that works better than trying to “invent an AI framework from scratch” is to combine four complementary views:
- DFD/STRIDE to identify flows, trust zones and classic failures.
- MITRE ATLAS to map AI domain-specific techniques, including prompt injection and pipeline manipulation.
- NIST AI RMF / GenAI Profile to align risk, governance, measurement and controls to the lifecycle.
- SAIF/agentic controlsto enforce coverage of data, infrastructure, model, and application — not just the model.
In practice, the workflow that yields the most in architectural reviews is this:
1) Model assets as capabilities, not just components
“We have a retriever” is not very useful. What matters is the ability it introduces: search arbitrary documents, query structured data, open URLs, invoke APIs, operate with legacy identity, create tickets, publish code, trigger automations. The attacker thinks in capabilities. The threat model should also think.
2) Classify every context by provenance and mutability
A practical taxonomy:
- T0 — Reliable and controlled:versioned internal policy, schemas, allowlists, approved catalogs.
- T1 — Internal but changeable:wiki, tickets, chats, internal emails, attachments, runbooks.
- T2 — Known external:vendor documentation, approved sites, monitored feeds.
- T3 — Unreliable:user-submitted content, arbitrary pages, third-party attachments, open web.
Without this taxonomy, the system treats everything as plain text and the defender loses the chance to apply different controls to different sources.
3) Separate human identity, agent identity and tool identity
This point seems bureaucratic until the day an agent opens, changes and approves something “on behalf of the team”. In mature environments, the trail needs to preserve at least:
- who requested;
- which policy allowed;
- which component decided;
- which principal performed;
- on which target system the action occurred.
If two of these five elements are lost, containment and attribution go bad.
4) Treat tools and MCP servers as execution boundaries, not as innocent “plugins”
This is a recurring error in teams that are already mature in AppSec, but have not yet internalized the agentic standard. The tool broker or MCP server is not just a connector; it is an impact multiplier. If it calls an API, reads a secret, accesses a shell, navigates a browser or writes to a repository, it is at the same level of criticality as a privileged microservice.
5) Model the entire chain, not just the initial event
Prompt injection, in itself, is not impactful. Impact is what happens next: exfiltration, inappropriate action, workflow change, secret disclosure, process sabotage, operational denial, memory corruption, context poisoning, persistence between sessions. A good threat model always explains the conversion step between manipulation and damage.
6) Produce reusable threat cards
A simple JSON format helps standardize review and prioritization:
{
"threat_id": "AI-TM-017",
"title": "Prompt injection indireta em documento RAG levando a tool misuse",
"entrypoint": "Wiki interna indexada",
"preconditions": [
"Documento mutável por usuários de baixa confiança",
"Retriever sem classificação de origem",
"Agente com tool create_ticket no Jira"
],
"attack_chain": [
"Inserção de instruções maliciosas no documento",
"Recuperação do documento em contexto relevante",
"Modelo interpreta texto como instrução prioritária",
"Agente aciona tool com parâmetros manipulados"
],
"business_impact": [
"Abertura indevida de chamados",
"Movimentação operacional não autorizada",
"Possível vazamento via campos do ticket"
],
"detections": [
"retrieved_doc_trust=t1 e tool_invocation=true",
"action_semantic_drift > baseline",
"approval_bypassed=false?"
],
"mitigations": [
"policy engine externo ao modelo",
"least privilege por tool",
"aprovação humana para ações state-changing",
"markup/proveniência visível no contexto"
]
}
Continue to Part 2
This first article focuses on architecture, trust boundaries, identity separation, provenance, and methodology. The operational continuation covers how these design mistakes are actually exploited in enterprise copilots, including indirect prompt injection, memory contamination, tool abuse, MCP amplification, and the telemetry needed to detect them. Read part 2 here: AI Attack Chains in Enterprise Copilots: Prompt Injection, Memory Poisoning, Tool Abuse, and Detection.
References
- NIST AI Risk Management Framework (AI RMF 1.0):https://www.nist.gov/itl/ai-risk-management-framework
- NIST AI 600-1,Artificial Intelligence Risk Management Framework: Generative AI Profilehttps://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.600-1.pdf
- NIST Glossary,Indirect Prompt Injectionhttps://csrc.nist.gov/glossary/term/indirect_prompt_injection
- MITRE ATLAS:https://atlas.mitre.org/
- OWASP,Prompt Injectionhttps://owasp.org/www-community/attacks/PromptInjection
- OWASP GenAI Security Project,Agentic AI Threats and Mitigationshttps://genai.owasp.org/resource/agentic-ai-threats-and-mitigations
- Microsoft Security Blog,3 takeaways from red teaming 100 generative AI productshttps://www.microsoft.com/en-us/security/blog/2025/01/13/3-takeaways-from-red-teaming-100-generative-ai-products
- Microsoft Learn,AI red teaming training serieshttps://learn.microsoft.com/en-us/security/ai-red-team/training
- Google Secure AI Framework (SAIF):https://www.saif.google/secure-ai-framework
- Google Cloud,Secure AI Framework (SAIF)https://cloud.google.com/use-cases/secure-ai-framework
- Anthropic,Introducing the Model Context Protocolhttps://www.anthropic.com/news/model-context-protocol
- OASIS Open / CoSAI, Model Context Protocol Securityhttps://www.oasis-open.org/2026/01/27/coalition-for-secure-ai-releases-extensive-taxonomy-for-model-context-protocol-security
- NSA,Model Context Protocol (MCP): Security Design Considerations for AI-Driven Automationhttps://www.nsa.gov/Portals/75/documents/Cybersecurity/CSI_MCP_SECURITY.pdf
- PortSwigger Web Security Academy,Web LLM attackshttps://portswigger.net/web-security/llm-attacks
💜 Enjoyed this content? Support the blog with USDT (TRC20):
TX7obcjHQbDUXb4mGqoASEu1QFTKT2CFGG
