AI assistants become much more useful when they can call tools. A tool may search a knowledge base, query a database, create a ticket or run a development workflow. The Model Context Protocol (MCP) is increasingly used to describe how models discover and use these tools. Its convenience also introduces a new security boundary: the model can influence software that has real permissions.
Why MCP security deserves attention
An MCP server may expose tool descriptions, input schemas, resources and actions. If authentication is weak, an attacker could discover sensitive capabilities. If authorization is broad, a low-privilege user may invoke an administrative function. If tool output is trusted blindly, a malicious result can inject instructions into the next model step.
Secure the server first
Run each MCP server with a clear owner and purpose. Keep the server separate from unrelated applications and restrict outbound network access. Require strong authentication and short-lived credentials. Validate every request on the server; never assume the model or client enforced the policy. Log identity, tool name, arguments, result status and approval decisions without recording secrets.
Tool design matters. Prefer small, single-purpose tools with strict schemas. Validate identifiers, lengths, formats and allowed values. Reject unknown fields. Separate read and write tools, and require explicit confirmation for actions involving payments, access changes, deletion, external messages or production infrastructure.
Protect against indirect prompt injection
MCP tools often return text from tickets, documents, web pages and repositories. That content is data, not authority. Mark untrusted results clearly and prevent them from changing system policy. Test with harmless documents that contain instructions such as “ignore previous rules” and confirm the agent does not call a privileged tool because of the text.
Identity and tenant isolation
Pass the authenticated user context to the authorization layer. Do not use one powerful service account for every user. Verify tenant, object ownership and action permission for every call. Test horizontal access by replacing an object ID with one belonging to another user. Also test whether tool discovery itself reveals names or descriptions that should remain private.
Monitoring and incident response
Alert on unusual tool sequences, repeated failed authorization, bulk reads, new servers and sudden changes in argument patterns. Keep the ability to revoke a server credential and disable one tool without taking down the entire assistant. Your incident plan should cover poisoned tool output, leaked secrets, compromised dependencies and malicious server updates.
Threat modelling an MCP deployment
Begin with a diagram rather than a scanner. Show the user, AI client, model, MCP server, identity provider, secrets store, downstream APIs and data sources. Mark which components are controlled by your organisation and which belong to vendors. For each boundary, ask what the receiver trusts and how that trust is verified.
Consider at least four attackers: an unauthenticated internet user, a legitimate low-privilege employee, a malicious content author and a compromised dependency. Their paths are different. The external attacker may probe exposed endpoints; the employee may abuse object identifiers; the content author may plant indirect instructions; and the dependency may alter tool behaviour during an update.
Write abuse cases in outcome language. “Read another customer’s tickets,” “send an email without confirmation,” and “retrieve a secret from an environment variable” are more useful than a generic goal such as “test prompt injection.” Outcome-based cases connect testing to business impact.
Authentication patterns
An MCP server should not accept identity merely because a client includes a username in a parameter. Use a trusted authentication mechanism and validate token issuer, audience, expiry and scope. For local servers, do not assume the local machine is automatically safe; other processes or browser content may attempt to reach listening ports.
Machine identities should be separate from human identities. A server calling a ticketing API should receive only the permissions needed for that workflow. Store secrets in a managed secret system and avoid placing them in prompts, configuration files or tool descriptions. Prefer short-lived credentials and establish a tested rotation process.
Authorization at three levels
First, authorize access to the MCP server. Second, authorize discovery and invocation of a particular tool. Third, authorize the resource and action inside that tool. Passing the first check must not imply permission for the other two.
For example, an employee may be allowed to use a customer-search tool but only for accounts assigned to their region. The downstream API should verify that object-level rule. If the MCP server merely forwards any account ID selected by the model, a simple identifier change can become a cross-tenant exposure.
Designing safer tools
Tool names and descriptions influence model selection, so make them precise. A tool called manage_account is ambiguous; separate read_account_summary from request_account_status_change. Narrow tools produce clearer logs and smaller permission scopes.
Use schemas with required fields, length limits, enumerated values and explicit formats. Reject additional properties unless they are necessary. Normalize and validate URLs, file paths and identifiers. For command or query tools, expose a constrained business operation instead of accepting arbitrary shell commands or database statements.
Return structured results where possible. Clearly separate status, data and error fields so that untrusted text is less likely to be interpreted as control information. Limit output size and remove secrets before results return to the model.
Human approval that adds value
Approval prompts should explain the exact action, target and impact. “Allow tool?” is not meaningful. “Send this message to 420 external recipients?” gives the reviewer a decision. Bind approval to the final arguments so the model cannot change them after consent.
Use approval for irreversible or high-impact actions, not every read. Too many prompts create fatigue and encourage automatic acceptance. Risk-based policy can consider action type, data sensitivity, destination and user role.
Deployment and supply-chain controls
Pin dependencies and verify the source of server packages. Review updates before production, especially when tool descriptions or requested permissions change. Run servers with minimal operating-system privileges, read-only filesystems where practical and restricted network egress.
Separate development, testing and production credentials. A developer testing a community MCP server should not accidentally expose production secrets from their environment. Maintain an allowlist of approved servers and a process for removing abandoned integrations.
A realistic testing engagement
Testing should begin with inventory and configuration review, then progress to authenticated abuse cases. Assess discovery, authentication, authorization, schema handling, prompt injection, output handling, secret exposure and logging. Use test accounts and reversible actions. Coordinate any production testing carefully because an agent can trigger downstream effects quickly.
The final report should include the tool call, identity, arguments, server decision, downstream result and business impact. A model screenshot alone does not prove whether a control succeeded or failed.
Testing checklist
- Inventory every MCP server, tool and resource.
- Review authentication, authorization and secret storage.
- Test malformed arguments and replayed requests.
- Test prompt injection through tool results.
- Check tenant isolation and object-level authorization.
- Review dependency provenance and update controls.
- Verify logs, alerts and credential revocation.
Repeat the checklist whenever a new tool, model, data source or permission is introduced. Small configuration changes can alter the effective trust boundary.
Detox can combine API, web and AI workflow testing through its cyber security services. Related context is available in the AI agent security testing article.
Questions for a design review
What can the server reach?
List every downstream API, file location, database and network destination available to the process. Compare that access with the declared business purpose and remove permissions that exist only for convenience.
Who makes the final decision?
Identify the component that authorizes each action. If the answer is the model or a natural-language prompt, move the decision into deterministic server-side policy.
Can results change behaviour?
Trace untrusted text from documents, web pages and API responses into subsequent model calls. Label it as data, constrain its influence and test indirect prompt injection.
What is reversible?
Separate harmless reads from actions that send, modify, delete or execute. High-impact operations need bounded arguments, meaningful confirmation and a reliable rollback or recovery path.
Can the event be reconstructed?
Ensure logs connect the human identity, agent, server, tool, arguments, authorization outcome and downstream response. Mask secrets while preserving evidence needed for investigation.
These questions are most useful when answered with evidence: configuration screenshots, access reports, log samples, recovery results and named owners. Record decisions and dates so the review becomes an improvement programme rather than a one-time discussion.
Example: a support automation
Consider an agent that reads support tickets and can issue account credits. Ticket text is untrusted, while the credit tool changes a financial record. The MCP server should expose ticket reading separately from credit approval, pass the employee identity to both services and limit credit amount by role. A malicious ticket that tells the agent to ignore policy must not alter those controls. Testing should verify the prompt response, the proposed tool call, the final server decision and the ledger result. This example shows why evaluating only the model’s words misses the most important security boundary.
Operational metrics
Track registered servers, tools without owners, credentials older than policy, high-impact actions lacking approval, denied authorization attempts and time to revoke a compromised integration. Measure remediation through repeated tests rather than a reduction in reported errors. A quiet server can still be unsafe if nobody is monitoring it.
From pilot to production
Pilot
Select one low-impact workflow, use synthetic data and expose only read tools. Capture every invocation and compare the agent’s proposed action with the server’s authorization decision. The responsible team should retain configuration evidence, test results, named exceptions and a completion date. Before moving to the next stage, confirm that the control works in practice and that operational teams know how to support it.
Controlled rollout
Introduce a small user group, production-like identities and explicit support ownership. Test credential rotation, server disablement and failure of downstream APIs before expanding access. The responsible team should retain configuration evidence, test results, named exceptions and a completion date. Before moving to the next stage, confirm that the control works in practice and that operational teams know how to support it.
Production gate
Require threat-model approval, resolved critical findings, documented monitoring and a rollback plan. Freeze tool schemas during the initial release window so behaviour remains understandable. The responsible team should retain configuration evidence, test results, named exceptions and a completion date. Before moving to the next stage, confirm that the control works in practice and that operational teams know how to support it.
Ongoing assurance
Review new servers and tools, repeat abuse cases after updates and investigate unusual action sequences. Retire integrations whose owner or business purpose disappears. The responsible team should retain configuration evidence, test results, named exceptions and a completion date. Before moving to the next stage, confirm that the control works in practice and that operational teams know how to support it.
A roadmap is useful because it creates order, not because every organisation must follow identical dates. Adjust sequencing for business impact, dependencies and available expertise, while keeping ownership and verification explicit.
FAQ
Related implementation guidance: review the cloud API security testing checklist and zero-trust controls for AI agents.
Is MCP itself insecure?
No protocol removes the need for secure implementation. Risk depends on server configuration, tool permissions, authentication and how untrusted content is handled.
Should every tool require human approval?
Not necessarily. Low-risk, reversible reads can be automated. High-impact or irreversible actions should have confirmation, policy checks or human approval.
What is the most important first control?
Use server-side authorization with least-privilege identities. A prompt rule cannot replace it.
Can an MCP server run safely on an employee laptop?
It can, but local deployment needs process isolation, restricted listening interfaces, controlled secrets and clear update management. Production workflows are usually easier to govern on managed infrastructure.
How should tool output be treated?
Treat it as untrusted data. Validate structured fields, limit size, remove secrets and ensure text from an external source cannot rewrite system policy.
What should happen when a tool fails?
Fail closed for sensitive actions. Return a clear error, avoid exposing stack traces or secrets, and do not let the model improvise a higher-privilege alternative.
Conclusion
MCP can make enterprise AI practical, but tool access must be treated like an API security boundary. Keep servers owned, authenticated, narrowly scoped and observable, then validate them with realistic security testing.
The safest MCP deployments are intentionally boring: capabilities are explicit, identities are narrow, approvals are understandable and every consequential action can be reconstructed from logs. That discipline lets teams gain the value of connected AI without turning convenience into uncontrolled authority.