MCP Security: How to Secure Agents, Tools, and Servers
MCP security is no longer a niche topic
For much of 2024, the LLM security conversation focused on prompt injection, data leakage, and guardrails. In 2025 and 2026, the center of gravity shifted. The real question became more operational: what happens when a model stops being a response interface and starts using tools, querying systems, triggering workflows, and crossing trust boundaries?
That is where MCP security became a serious topic.
The Model Context Protocol (MCP) was introduced to standardize how assistants and agents connect models to tools, resources, and data sources. The value proposition is obvious: fewer ad hoc integrations, better interoperability, and faster connections between AI systems and GitHub, Slack, internal APIs, IDEs, browsers, and enterprise platforms. The problem is equally obvious: the same layer that reduces integration friction also expands the attack surface.
Treating MCP as a product convenience is no longer enough. In a real environment, MCP introduces a new class of risk where:
- the model influences security-relevant decisions
- tool descriptions influence operational behavior
- weak authorization becomes excessive access
- MCP servers become pivot points between systems
- untrusted data can alter tool invocation paths
In other words: MCP is not just another connector. It is an orchestration layer with direct security implications.
What MCP means in practical terms
Anthropic defines MCP as an open standard for bidirectional connections between AI systems and data sources or tools. In operational terms, that usually means an architecture where:
- a host or agent application runs the user-facing experience
- an MCP client manages connections to MCP servers
- an MCP server exposes tools, resources, or prompts
- the model decides when and how to invoke those capabilities
The productivity gain is clear. The risk shift is clear too.
A traditional application can encapsulate integrations in rigid code paths. With MCP, part of the invocation logic depends on:
- tool descriptions
- schemas
- model-visible context
- client/server authorization policy
- emergent agent behavior
That changes the threat model. The problem is not only that the model can be wrong. The problem is that the model can be wrong inside an environment that is allowed to act.
Why “MCP security” is a relevant search term now
There are four concrete signals.
1. Accelerating protocol adoption
Since Anthropic introduced MCP, it has moved from a promising specification to a recurring integration layer across the agent ecosystem. Official documentation, SDKs, public demos, integrations, and community discussions all grew quickly.
2. Official security guidance became necessary early
The MCP project itself now publishes security best practices, covering issues such as:
- confused deputy
- token passthrough
- SSRF in OAuth discovery flows
- session hijacking
- local server compromise
- authorization URL validation failures
- escalation through
stdio
When a young technology already needs that level of security guidance, it usually indicates two things: real adoption and real risk.
3. Institutional attention reached the protocol
In 2026, the NSA published Model Context Protocol: Security Design Considerations. CoSAI / OASIS also published dedicated guidance on MCP threats and mitigations. Institutional attention of that kind rarely appears while a standard is still irrelevant.
4. Offensive research moved beyond theory
Independent researchers and vendors are already demonstrating:
- excessive privilege exposure
- ambiguous authorization models
- indirect prompt injection in tool-enabled flows
- weak trust chains between client, server, and tools
- code execution and context abuse scenarios
That makes MCP security a strong intersection of editorial relevance and organic search demand: new enough to attract search, technical enough to differentiate an offensive security blog.
The most common mistake: treating MCP as only a prompt injection problem
Prompt injection still matters, but it is only one component of the problem space.
The real risk in MCP is the combination of:
- operational capability
- dynamic tool discovery
- weak or poorly modeled authorization
- untrusted context
- implicit delegation across components
If a model reads hostile instructions in a document but cannot do anything with them, the effect is often limited to the response. If it can:
- open a browser
- query a database
- write files
- call internal APIs
- send messages
- execute code
then injection stops being only a semantic issue and becomes an entry point for unwanted action.
The 7 risk classes that matter most in MCP security
Below is a practical cut of the issues that matter most for technical teams.
1. Excessive privilege in tools
This is the most common problem — and probably the most dangerous.
Teams often connect agents to overly broad tools because that accelerates demos and proofs of concept. The result is an agent with visibility and action capability far beyond what the task requires.
Examples:
- broad repository access when read-only access would be enough
- SaaS tokens with global scope instead of per-workspace or per-action scope
- local filesystem write access where controlled read access would suffice
- terminal integration exposed for tasks that should have been handled by narrower APIs
Operational impact: any tool-selection error, prompt injection, or context abuse now has a larger blast radius.
2. Tool poisoning and untrusted descriptions
CoSAI highlights an important point: tool descriptions and annotations should be treated as untrusted unless they come from a validated, trusted source.
That matters because the model makes decisions from text. If the text describing a tool is ambiguous, manipulated, misleading, or malicious, the agent’s decision process is compromised from the beginning.
In practice, that can happen through:
- inconsistent tool metadata
- misleading schemas
- descriptions that encourage out-of-scope use
- MCP servers from weakly vetted sources
- supply-chain paths without validation
Operational impact: the agent may choose the wrong tool, send dangerous parameters, or trust a capability that should never have been authorized.
3. Indirect prompt injection
This remains the classic LLM issue, but in MCP the consequences are larger.
An agent may consume content from:
- web pages
- markdown files
- tickets
- internal documents
- code comments
- RAG stores
- chat messages
If that content contains hostile instructions, it can influence the next tool call. In an MCP environment, that may mean more than “the model answered incorrectly.” It can mean:
- unauthorized resource reads
- exfiltration through an already authorized tool
- writes to the wrong destination
- API calls with manipulated parameters
- unintended action chaining
4. Confused deputy in MCP proxies
The official MCP security documentation explicitly calls out the confused deputy problem in OAuth-based scenarios.
This becomes especially relevant when an MCP proxy:
- uses a static
client_idto talk to a third-party API - allows dynamic registration of downstream clients
- depends on consent cookies from the upstream provider
- does not implement client-specific consent correctly
In that design, a malicious client may try to reuse consent previously granted by the user to the legitimate proxy.
Operational impact: token issuance and authorization can be abused across trust boundaries, especially in SaaS integrations.
5. Compromise of local MCP servers
The official documentation also addresses the risk of local MCP server compromise. This matters because many adoption paths begin on a developer workstation or in lightly governed environments.
A compromised local MCP server can:
- manipulate tool outputs
- alter descriptions
- observe or modify context
- serve as a pivot for local actions
- amplify the effect of indirect injection
This risk is often underestimated because “local” is incorrectly treated as equivalent to “safe.” It is not.
6. Weak traceability between agent, server, and tool
CoSAI recommends that requests remain traceable across the full chain: originating user or agent, intermediate servers, and the tools or services that actually execute the action.
If you cannot reconstruct that chain, two problems appear immediately:
- you cannot investigate incidents properly
- you cannot enforce accountability for action
That turns manageable failures into operational gray zones.
7. Code execution and automation beyond necessity
The NSA explicitly warns about actor-controlled input reaching execution environments that are not properly constrained, creating conditions for severe failure classes, including ACE.
Put plainly: do not treat LLM-generated commands, code, or workflows as trustworthy by default.
In particular, do not combine:
- a semantically flexible model
- a tool with execution capability
- broad credentials
- weak sandboxing
That is exactly the combination that turns probabilistic error into a real incident.
The right threat model for MCP
The wrong way to model MCP is to ask: “does the server require authentication?”
The right way is to ask:
- Who can register, discover, or introduce tools?
- Which data sources influence the model’s decisions?
- Which actions have external side effects?
- What permissions does each action actually require?
- How can hostile instructions move across context, model, client, and tool?
- How do we prove who initiated each action?
- What happens if an MCP server lies, fails, or is compromised?
If those questions do not have technically clear answers, the system is not ready.
A practical hardening framework for MCP security
Below is a pragmatic baseline for teams building or auditing MCP environments.
1. Minimize capability before talking about guardrails
The first control is architecture, not prompting.
Do this:
- prefer narrowly scoped tools over generic ones
- apply least privilege per integration
- separate read paths from write paths
- segregate credentials by workflow
- reduce terminal or code execution access to the strict minimum
Useful review question:
If the agent selects the wrong tool once, what is the worst possible external effect?
2. Treat tool metadata as sensitive input
Do not allow descriptions, schemas, and annotations into the system without governance.
Do this:
- maintain an allowlist of trusted MCP servers
- validate and review critical metadata
- version tools explicitly
- enforce provenance policy
- use attestation where appropriate
3. Isolate untrusted context
Content from the web, tickets, email, RAG stores, documents, and code comments should be treated as potentially hostile.
Do this:
- label context by origin
- separate trusted from untrusted data in the pipeline
- prevent arbitrary content from directly steering privileged tool selection
- require additional validation steps before destructive or external actions
4. Redesign authorization for agents
Models and agents do not fit cleanly into legacy authorization models.
Even so, the baseline still applies:
- explicit identity
- minimum scopes
- verifiable consent
- short-lived tokens
- separation between user, service, and agent credentials
For OAuth-backed MCP proxies, carefully review scenarios involving:
- static client IDs
- dynamic registration
- consent-cookie reuse
- redirect URI validation
5. Strengthen end-to-end action observability
Chat logs are not enough.
You need an execution trail that answers:
- which prompt or context influenced the decision
- which tool was selected
- with which parameters
- under which identity or token
- on which MCP server
- with what external effect
Without that, you have automation without forensics.
6. Put real barriers in front of code and external actions
For any tool that:
- executes commands
- writes to critical systems
- calls APIs with external side effects
- moves sensitive data
consider these mandatory:
- appropriate sandboxing
- filesystem and network restrictions
- approval gates for selected action classes
- controlled egress policy
- timeouts and quotas
As CoSAI notes, containers should not be treated as a strong security boundary by themselves. Many demo architectures ignore that point.
7. Design for failure and compromise
Assume that one of these things will happen:
- a tool behaves differently than expected
- a description drives incorrect use
- external content attempts instruction injection
- an integration exposes too much privilege
The architecture must remain safe even in those scenarios.
What a Red Team should test in MCP environments
If you work on offensive assessments, this is a useful test cut.
Discovery and trust surface
- can an untrusted MCP server be introduced?
- is there a real allowlist or only convention?
- is tool metadata validated or merely consumed?
Agent decision surface
- does external content influence tool selection?
- does the model receive conflicting instructions from documents, comments, or pages?
- is trusted context separated from untrusted context?
Authorization surface
- are tokens overly broad?
- is there confusion between user, service, and agent identity?
- do consent prompts create fatigue that turns into “allow all” behavior?
Execution surface
- is there a path to terminal, filesystem, browser, or privileged API access?
- which external actions can be triggered without an extra validation step?
- is blast radius actually contained?
Supply-chain surface
- where do tools and servers come from?
- is there version, signature, origin, and review control?
- are third-party components inventoried?
Detection and response surface
- is the execution trail sufficient for investigation?
- can you prove which context led to which action?
- do alerts distinguish reads, writes, execution, and potential exfiltration?
Where many companies will fail in 2026
Most failures will not come from exotic cryptography or sophisticated zero-days. They will come from the usual pattern in a new wrapper:
- adopting too quickly
- trusting defaults that do not exist
- delegating governance to end users
- mixing convenience with privilege
- exposing powerful tools before identity, observability, and context separation are solved
MCP accelerates integration. It also accelerates security debt if the design starts without discipline.
Conclusion
MCP security matters because MCP moves LLMs from the role of interface to the role of operational intermediary between data, tools, and systems. That changes the nature of risk.
The central point is not merely that “LLMs can be fooled.” We already knew that.
The more important point is this: they can now be fooled inside architectures that read, decide, and act.
That is why the useful question for 2026 is not “are we going to use MCP?” In many contexts, the answer is already yes.
The useful question is:
what minimum security guarantees do you require before letting an agent touch real systems?
If your organization is connecting agents to tools, this is the moment to treat MCP as an AppSec, cloud security, IAM, detection, and architecture problem — not just as a productivity feature.
Quick MCP security checklist
- [ ] inventory every MCP server in use
- [ ] classify tools by read, write, execution, and potential exfiltration impact
- [ ] reduce privileges per tool and per workflow
- [ ] validate provenance and metadata for servers and tools
- [ ] separate trusted from untrusted context
- [ ] review OAuth and proxy flows for confused deputy conditions
- [ ] implement end-to-end execution tracing
- [ ] require additional approval for destructive or external actions
- [ ] apply real sandboxing to code and command execution
- [ ] exercise indirect prompt injection and tool misuse scenarios
References
- Anthropic. Introducing the Model Context Protocol. Nov. 25, 2024. https://www.anthropic.com/news/model-context-protocol
- Model Context Protocol. Security Best Practices. https://modelcontextprotocol.io/docs/tutorials/security/security_best_practices
- NSA. Model Context Protocol (MCP): Security Design Considerations. May 2026. https://www.nsa.gov/Portals/75/documents/Cybersecurity/CSI_MCP_SECURITY.pdf
- CoSAI / OASIS Open Project. Model Context Protocol (MCP) Security. Approved Jan. 8, 2026. https://www.coalitionforsecureai.org/wp-content/uploads/2026/03/model-context-protocol-security-1.pdf
- HiddenLayer Research. MCP: Model Context Pitfalls in an Agentic World. Apr. 10, 2025. https://www.hiddenlayer.com/research/mcp-model-context-pitfalls-in-an-agentic-world
Additional editorial options
More technical alternate title
MCP Security: Threat Modeling and Hardening for Tool-Enabled Agents
More SEO-oriented alternate title
MCP Security: Main Risks and How to Protect Model Context Protocol Agents
Featured image direction
- Family: map/system-centric
- Metaphor: control-plane map with broken trust routes
- Avoid: operator-at-workstation scenes, hoodies, generic neon code
- Palette: blue-gray + attack-path red + graphite
Possible future internal links
- indirect prompt injection in RAG
- autonomous agent security
- IAM for AI infrastructure
- threat modeling for LLM-enabled applications
💜 Enjoyed this content? Support the blog with USDT (TRC20):
TX7obcjHQbDUXb4mGqoASEu1QFTKT2CFGG
