What Is MCP Security? Risks, Attack Vectors, and Best Practices
Learn what MCP security means, the top Model Context Protocol attack vectors like prompt injection and token passthrough, and best practices to secure AI agents.

Sucharitha
AI agents aren't chatbots anymore. They hold credentials, call tools, and take real actions inside real systems, and Model Context Protocol is how most of them do it: Anthropic's December 2025 ecosystem update put MCP at more than 97 million monthly SDK downloads and over 10,000 active public servers, adopted across ChatGPT, Claude, Cursor, Gemini, and Microsoft Copilot. MCP security is the discipline of protecting that entire connection layer, since a single overlooked MCP server can hand an attacker the same access the agent itself was trusted with.

What Is MCP Security?
Model Context Protocol security, MCP security for short, covers every control protecting the connections between AI agents and the tools, data sources, and servers they reach through the protocol, spanning authentication, context handling, tool access, and runtime behavior. It's a subset of the broader discipline of AI agent security, specific to the protocol layer agents increasingly use to actually reach the tools they act through.
Why MCP Doesn't Enforce Security at the Protocol Level
MCP was designed for interoperability first, and it deliberately leaves most security decisions, authentication, authorization, input validation, to whoever implements a given client or server. That's a defensible design choice for a young, fast-moving standard, but it means MCP security risks are real by default rather than something a misconfiguration introduces on top of an otherwise-secure baseline.
Key Components: Context Handling, Tool Access Control, Identity and Session Management
Context handling governs what data and instructions actually reach an agent's reasoning; tool access control governs which actions an agent can take once connected; identity and session management governs who or what a given connection is actually authenticated as. Weakness in any one of these three undermines the other two, since an agent with strong tool access controls but no session validation is still exposed to an attacker who successfully impersonates a trusted session.

How MCP Works (Brief Architecture Primer)
Host, Client, Server Model
MCP defines three roles: a host application (the AI system a user interacts with), a client (the connector managing a specific server relationship), and a server (the component exposing tools, data, or resources). Each MCP server a host connects to is a new trust boundary, not an extension of an existing one.
Where Trust Boundaries Break Down
Trust boundaries break down at the handoff points: when a client accepts a server's tool descriptions without validation, when a host trusts a client's output without re-checking it, and when a server's own dependencies or third-party resources are treated as safe by default. Each handoff is a place where an attacker only needs to compromise one link, not the whole chain.
Why MCP Security Matters for Agentic AI
Expanding, Dynamic Attack Surface
Every new MCP server connection is a new piece of software with its own supply chain, its own authentication posture, and its own potential vulnerabilities, and MCP server security has to account for that surface expanding continuously as agents get wired into more tools, often without centralized review. The MCP attack vectors covered in this guide aren't a fixed, finite list either; new ones keep getting disclosed as researchers dig deeper into the protocol's real-world deployments.
Business Impact: Data Leakage, Unauthorized Actions, Compliance Risk
The business impact isn't abstract. A compromised MCP connection can leak customer data, trigger an unauthorized transaction, or process regulated data outside any approved workflow, each of which carries its own compliance exposure under frameworks like the NIST AI RMF or the OWASP Top 10 for LLM and Agentic Applications.
Core MCP Security Risks and Attack Vectors
Cataloging MCP vulnerabilities by category matters more than treating them as isolated incidents, since the same underlying patterns keep recurring across different vendors and different servers. One named example worth knowing specifically: CVE-2025-49596, a critical flaw in Anthropic's own MCP Inspector developer tool, stemmed from a lack of authentication between the Inspector's client and its local proxy, letting a malicious website trigger cross-site request forgery against an unauthenticated local port and achieve remote code execution.
Prompt Injection (Direct vs. Indirect)
Direct prompt injection targets an agent's instructions head-on. Indirect prompt injection, embedded in a document, tool response, or website an agent processes, is the more common and more dangerous variant for MCP-connected agents specifically, since the agent never distinguishes an instruction it was given from one buried in content it merely retrieved.
Confused Deputy and Tool Shadowing
A confused deputy attack happens when an MCP server, holding more authority than the end user, gets manipulated into acting on an attacker's behalf. Tool shadowing, demonstrated by Invariant Labs, is a related but distinct pattern: a malicious server's tool description silently overrides the behavior of a different, trusted server's tool, redirecting something as ordinary as an email-sending function without ever compromising the trusted server itself.
Token Passthrough
Token passthrough occurs when an MCP server forwards a client's authentication token to a downstream service instead of validating it directly. The MCP specification explicitly prohibits this, since a token issued for one server that gets replayed against another breaks audience binding and creates exactly the conditions a confused deputy attack needs.
Session Hijacking and Impersonation
Session hijacking exploits weak session binding, letting an attacker capture or reuse a valid session identifier to impersonate a legitimate user or agent, often surfacing in poorly secured local development setups exposed to a broader network than intended.
SSRF via OAuth Metadata Discovery
MCP's own security documentation flags this directly: during OAuth metadata discovery, a client fetches URLs from fields like resource_metadata, authorization_servers, and token_endpoint that a malicious server controls, letting that server point the client at internal IP addresses or cloud metadata endpoints instead. This server-side request forgery pattern can turn a routine authentication handshake into a path toward stealing cloud credentials.
Tool Poisoning and Rug-Pull Attacks
Tool poisoning embeds hidden instructions inside a tool's name, description, or schema, content a model reads while deciding how to use the tool even though a human reviewing its visible function would never see it. A related timing variant, dubbed line jumping by Trail of Bits, exploits the fact that tool descriptions load during the initial handshake, letting a hidden instruction execute before any legitimate tool is ever invoked. A rug-pull attack is the delayed version of tool poisoning: a tool behaves honestly when first approved, then gets silently modified afterward.
Supply Chain Compromise
In September 2025, Koi Security discovered postmark-mcp, an npm package that impersonated a legitimate email-sending MCP server. Its first fifteen published versions were exact, harmless copies of the real code; version 1.0.16 added a single line that silently BCC'd every outgoing email to an external address, exposing communications from an estimated 1,500 organizations before it was pulled. It remains one of the clearest illustrations of trust-then-poison as a supply chain pattern specific to MCP.
Why Traditional Security Tools Fail Here
Static Analysis vs. Runtime Behavior
Static analysis can catch a malicious pattern sitting in code, but it can't evaluate whether a tool's natural-language description contains an instruction aimed at manipulating a model. That's a semantic risk, not a syntax one, and it sits entirely outside what static scanning was built to catch.
Perimeter Controls vs. Agent-to-Tool Traffic
Traditional perimeter security assumes a request either belongs to a known application, or it doesn't. Agent-to-tool traffic doesn't fit that model cleanly, since a single agent session can legitimately call dozens of different tools across dozens of different trust boundaries in the course of one task; traffic a perimeter firewall has no context to evaluate meaningfully.

MCP Security Monitoring: What to Watch
Tool Usage Signals
Watch for tool calls outside an agent's established pattern, unexpected argument values, and any tool invocation that doesn't match the task an agent was actually given.
Context and Log Signals
Log tool descriptions themselves, not just tool invocations, since a line-jumping incident can only be reconstructed after the fact if the descriptor text that entered context was actually captured.
Session and Identity Signals
Watch for session reuse across unexpected origins, authentication attempts against unfamiliar authorization servers, and any pattern suggesting a token is being used somewhere other than where it was issued.
MCP Security Best Practices Checklist
Authentication and Authorization (OAuth 2.1, Token Validation)
Require OAuth 2.1 with PKCE for every MCP connection, and validate every token directly against its issuing authorization server rather than trusting a token's mere presence as proof of legitimacy. This is zero trust for AI agents applied concretely: no connection is assumed safe by default, regardless of where it originates. Routing every MCP connection through a centralized MCP gateway makes this kind of enforcement consistent across every application rather than relying on each team to implement it correctly on its own.
Least Privilege and Tool Scoping
Scope every MCP server connection to the minimum tools and data access a specific task requires, since a broadly permissioned connection turns a single compromised tool into access to everything else the agent can reach.
Prompt and Input Validation
Treat tool descriptions, retrieved documents, and third-party resources as untrusted input by default, validating them the same way an application would validate any other externally sourced data before it reaches a decision-making context.
Runtime Monitoring and Guardrails
Enforce policy at runtime, blocking an out-of-scope tool call or a suspicious session pattern as it's attempted, rather than relying solely on pre-deployment review to catch what a live agent might do. For an MCP connection's highest-risk actions, a human-in-the-loop checkpoint requiring explicit approval before execution is a stronger control than any automated filter alone.
Incident Response for Agentic Systems
Build a response plan specific to MCP compromise: identify every credential and downstream system a compromised connection touched, rotate all of it rather than the one credential presumed exposed, and review logs for exfiltration patterns rather than just the specific exploit disclosed.
Governance and Compliance Mapping (NIST AI RMF, ISO 42001)
Map MCP-specific findings to NIST AI RMF's Measure function and to ISO 42001's risk treatment requirements, giving security, legal, and compliance teams a shared vocabulary for MCP risk rather than three separate internal taxonomies.
Automated Red Teaming and Continuous Testing for MCP
Why Point-in-Time Audits Aren't Enough
A server that passed review last month can still be exposed to a technique disclosed today, and MCP's own vulnerability count keeps climbing month over month. Point-in-time audits only ever reflect the threat landscape as it existed on the day they ran.
Shift-Left Threat Modeling for MCP Connections
Threat modeling a new MCP connection before it goes live, mapping what it can reach and what could go wrong if it's compromised, catches structural risk earlier and cheaper than discovering the same gap after an incident.
How Akto Secures MCP Deployments
Discovery and Posture Management
Akto continuously discovers MCP servers, tools, and connections across an environment, building the inventory an AI-SPM program depends on and surfacing shadow AI connections nobody formally approved.
Continuous Red Teaming
Structured, ongoing adversarial testing targets tool poisoning, prompt injection, confused deputy scenarios, and token handling, mapped to the OWASP Top 10 for LLM and Agentic Applications so findings carry a shared, auditable vocabulary.
Real-Time Runtime Guardrails
Findings from that testing feed directly into runtime guardrails, enforcing least privilege and blocking out-of-scope tool calls in production rather than leaving enforcement as a policy recommendation nobody implements.
The Future of MCP Security
Where the Ecosystem Is Heading
MCP's own roadmap for 2026 centers on enterprise authentication, multi-agent coordination, and a curated registry with security ratings, closing gaps the current ecosystem exposes today. Expect the OWASP Top 10 for Agentic Applications and the OWASP MCP-specific taxonomy to keep expanding as more incidents get documented, and expect Gartner's continued attention to MCP as a named driver of the broader rise in agentic AI security incidents it has forecast through 2028.
Final Thoughts on MCP Security
MCP's growth from a niche protocol to infrastructure running through nearly every major AI platform happened faster than the security practices around it matured. Treating MCP security as a checklist item to satisfy once misses the point entirely: every new server connection is a new trust boundary, and the risks covered here, prompt injection, token passthrough, tool poisoning, supply chain compromise, are documented, recurring patterns rather than theoretical edge cases. Discovery, least privilege, runtime guardrails, and continuous red teaming aren't optional additions to an MCP deployment. They're what make the deployment defensible.
Frequently Asked Questions on MCP Security
How is MCP different from a traditional API in terms of security?
A traditional API has a single, well-defined trust boundary between client and server. MCP stacks multiple trust relationships, host to client, client to server, server to third-party resources, on top of an AI agent making autonomous decisions about which tools to call.
What are the biggest security risks in the Model Context Protocol?
Prompt injection, confused deputy attacks, token passthrough, session hijacking, SSRF via OAuth metadata discovery, tool poisoning and rug-pull attacks, and supply chain compromise through malicious or typosquatted MCP packages.
What is prompt injection in the context of MCP?
Direct prompt injection targets an agent's instructions explicitly. Indirect prompt injection hides malicious instructions in a document, tool response, or third-party resource an agent processes, which the agent can't reliably distinguish from a legitimate instruction.
What is a confused deputy attack in MCP?
It's when an MCP server, holding more authority than the end user driving it, gets manipulated into performing an action on behalf of an attacker, often enabled by token passthrough that lets a token intended for one server get used against another.
What is token passthrough, and why is it prohibited by the MCP spec?
Token passthrough is when a server forwards a client's authentication token to a downstream service instead of validating it directly. The MCP specification prohibits it because doing so breaks audience binding and enables confused deputy attacks.
Can MCP servers be hijacked or impersonated?
Yes, through session hijacking exploiting weak session binding, and through supply chain attacks like postmark-mcp, where a malicious package impersonated a legitimate MCP server and operated undetected for fifteen clean releases before turning malicious.
Why don't traditional security tools work for MCP-connected agents?
Static analysis can't evaluate whether a tool description contains a hidden instruction aimed at a model, and perimeter controls have no context for legitimate agent-to-tool traffic crossing many trust boundaries in a single task.
What should be included in an MCP security best practices checklist?
OAuth 2.1 authentication with direct token validation, least-privilege tool scoping, treating tool descriptions and retrieved content as untrusted input, runtime guardrails enforcing policy in real time, an MCP-specific incident response plan, and governance mapping to frameworks like NIST AI RMF and ISO 42001.
How do you monitor MCP servers and agents at runtime?
Track tool usage signals like unexpected argument values, log tool descriptions themselves rather than only invocations, and watch session and identity signals for token reuse across unfamiliar origins or authorization servers.
What is AI red teaming, and how does it apply to MCP?
AI red teaming is structured, adversarial testing simulating real attacks against an AI system. Applied to MCP, it means continuously testing tool poisoning, prompt injection, confused deputy scenarios, and token handling rather than relying on a single pre-launch review.
Does MCP support OAuth-based authentication?
Yes. MCP's current direction formalizes OAuth 2.1 with PKCE and resource indicators binding a token to the specific server it was issued for, closing off several of the token-related attack paths covered above.
How does MCP security relate to AI governance frameworks like NIST AI RMF or ISO 42001?
MCP-specific findings map into NIST AI RMF's Measure function as quantitative risk evidence, and into ISO 42001's risk treatment requirements as the technical substance a certification audit actually checks for.
What tools or platforms help secure MCP deployments?
Platforms combining continuous discovery, automated red teaming, and runtime guardrails, like Akto, cover the full lifecycle, while open-source frameworks like Garak and PyRIT support the testing layer specifically for teams building their own program.
What does the future of MCP security look like?
Expect formalized enterprise authentication, a curated registry with security ratings, and continued expansion of taxonomies like the OWASP Top 10 for Agentic Applications, as the ecosystem's security posture catches up to its adoption curve.
Important Links
Experience enterprise-grade Agentic Security solution

