Zero trust is often summarised as “never trust, always verify.” For AI agents, that principle applies to every action, not only the person who started the conversation. An agent may change context, retrieve new data or request a tool it did not need at the beginning of a task.
Give agents explicit identities
Every production agent should have an owner, purpose, environment and identity. Avoid shared credentials across unrelated workflows. Bind permissions to the user, agent, resource and action. Use short-lived tokens and make revocation simple.
Least privilege is a runtime decision
An agent may need read access to one system and write access to another. Expose only the tools required for the current task. Validate arguments on the server and require confirmation for irreversible actions. Record the decision that permitted each tool call.
Segment data and tools
Separate tenants, environments and sensitivity levels. Keep development agents away from production secrets. Treat retrieved content and tool output as untrusted. Prevent the agent from turning an untrusted instruction into a privileged command.
Continuous verification
Monitor unusual tool sequences, bulk reads, repeated denied actions and new integrations. Re-evaluate policy when the user, data source, task or risk changes. Run regression tests after model, prompt and connector updates.
Detox’s security testing services can assess agent identity, API authorization and runtime abuse cases. Link to the AI agent security testing guide for implementation examples.
Applying zero trust across an agent workflow
Zero trust for AI agents is not a new product category. It is an architectural discipline applied to identities, context, tools and data. Each step below asks the system to prove that an action is appropriate at the moment it occurs, even when the original user is legitimate.
1. Agent registration
During review, record owner, purpose, environment, model, tools and data sources. Keep evidence that unregistered agents cannot receive production credentials. This prevents a policy statement from being mistaken for a working control and gives the remediation owner a clear acceptance test.
2. Human identity
A practical test should preserve the authenticated user through every downstream request. The expected result is that shared service accounts do not erase accountability. Record exceptions with an owner and expiry date; undocumented exceptions tend to become permanent exposure.
3. Agent identity
Ask the responsible team to issue a distinct workload identity with short-lived credentials. Validate the answer in a representative environment so that one compromised workflow does not expose every agent. Where the control fails, capture business impact as well as the technical weakness.
4. Tool allowlisting
For this area, show the agent only capabilities required for the current task. Reviewers should be able to demonstrate that prompt manipulation cannot discover broad administrative functions. Re-test after material architecture or supplier changes because the effective boundary may have moved.
5. Object authorization
The control objective is straightforward: verify tenant, owner and resource permission server-side. Evidence should show that model-selected identifiers cannot cross data boundaries. If several teams share responsibility, name one person who coordinates the final decision and follow-up.
6. Action authorization
Do not rely only on documentation; apply policy to reads, writes, exports, messages and execution separately. A successful check confirms that access to a tool does not imply permission for every operation. Include negative tests, since secure behaviour is often revealed by how the system rejects an invalid or unauthorised request.
7. Context minimisation
During review, send only the data required for the immediate decision. Keep evidence that prompts and memory do not become unnecessary data warehouses. This prevents a policy statement from being mistaken for a working control and gives the remediation owner a clear acceptance test.
8. Retrieval isolation
A practical test should filter documents by identity and tenant before model access. The expected result is that citations and embeddings cannot leak another user’s records. Record exceptions with an owner and expiry date; undocumented exceptions tend to become permanent exposure.
9. Untrusted content
Ask the responsible team to label and constrain instructions found in files, websites and tool output. Validate the answer in a representative environment so that external text cannot override trusted policy. Where the control fails, capture business impact as well as the technical weakness.
10. Argument validation
For this area, use narrow schemas and business-rule checks on every tool call. Reviewers should be able to demonstrate that generated values cannot become arbitrary commands or queries. Re-test after material architecture or supplier changes because the effective boundary may have moved.
11. Approval binding
The control objective is straightforward: show the final action, target and impact before consent. Evidence should show that arguments cannot change after a reviewer approves. If several teams share responsibility, name one person who coordinates the final decision and follow-up.
12. Runtime monitoring
Do not rely only on documentation; detect unusual sequences, bulk reads and repeated denied actions. A successful check confirms that security teams can identify abuse while it is occurring. Include negative tests, since secure behaviour is often revealed by how the system rejects an invalid or unauthorised request.
13. Revocation
During review, disable one user, agent, tool or credential independently. Keep evidence that containment does not require shutting down the entire platform. This prevents a policy statement from being mistaken for a working control and gives the remediation owner a clear acceptance test.
14. Regression testing
A practical test should repeat abuse cases after model, prompt and connector changes. The expected result is that security behaviour remains measurable across releases. Record exceptions with an owner and expiry date; undocumented exceptions tend to become permanent exposure.
15. Governance review
Ask the responsible team to reassess permissions, owners and business value on a schedule. Validate the answer in a representative environment so that agents that are no longer justified lose access. Where the control fails, capture business impact as well as the technical weakness.
Taken together, these checks create a defensible baseline. Prioritise findings that enable unauthorised access, sensitive-data exposure, irreversible action or loss of recovery capability. Assign dates, validate fixes and keep the evidence with the system’s security record.
Example: an HR knowledge agent
An HR agent may answer policy questions and retrieve an employee’s own records. It should not expose another employee’s documents simply because a prompt includes their name. Retrieval must filter by authenticated identity before content reaches the model. A separate tool may submit a leave request, while salary changes remain unavailable. Logs should connect the employee, agent, retrieved documents and submitted action. This design limits both accidental mistakes and deliberate prompt manipulation.
Policy outside the prompt
System prompts can tell an agent how to behave, but they are probabilistic instructions. Access policy belongs in code and infrastructure that can make deterministic decisions. The server should reject an unauthorised object even if the model strongly requests it. Sensitive tools should validate final arguments and approval. This separation also improves audits because reviewers can inspect rules without interpreting natural-language conversation.
Handling changing context
An agent’s risk can change during one session. A harmless research task may become a request to export data or send an external message. Re-evaluate authorization when the tool, target, data sensitivity or destination changes. Do not assume approval at the start of a conversation covers every later action. Expire cached permissions and bind confirmation to the exact operation.
Measuring zero-trust coverage
Track agents with registered owners, distinct identities, scoped tools, server-side authorization and tested revocation. Count high-impact actions with meaningful approval and retrieval sources with tenant filtering. Review denied actions and unusual sequences for both attacks and design problems. Metrics should reveal where implicit trust remains, not merely how many policies exist.
A deployment sequence
Register
Document the agent, owner, users, data and intended actions. Deny production credentials until this record and risk classification exist. The responsible team should retain configuration evidence, test results, named exceptions and a completion date. Before moving to the next stage, confirm that the control works in practice and that operational teams know how to support it.
Constrain
Issue narrow identities, expose minimum tools, filter retrieval and implement deterministic authorization. Add approval where impact is high. The responsible team should retain configuration evidence, test results, named exceptions and a completion date. Before moving to the next stage, confirm that the control works in practice and that operational teams know how to support it.
Observe
Log context needed to reconstruct actions, alert on abnormal sequences and test revocation. Avoid recording secrets merely for troubleshooting convenience. The responsible team should retain configuration evidence, test results, named exceptions and a completion date. Before moving to the next stage, confirm that the control works in practice and that operational teams know how to support it.
Reassess
Repeat security tests after model, prompt, tool and data changes. Remove permissions and agents that no longer deliver justified value. The responsible team should retain configuration evidence, test results, named exceptions and a completion date. Before moving to the next stage, confirm that the control works in practice and that operational teams know how to support it.
A roadmap is useful because it creates order, not because every organisation must follow identical dates. Adjust sequencing for business impact, dependencies and available expertise, while keeping ownership and verification explicit.
Testing zero trust rather than trusting the diagram
Architecture diagrams show intended boundaries; testing reveals effective ones. Use two roles and two tenants, then attempt to retrieve, modify and export objects across those boundaries. Introduce untrusted instructions through documents and tool output. Change the requested action after approval and verify that consent is invalidated. Revoke the agent identity during an active session and confirm that subsequent calls fail. Finally, compare logs with the test timeline. If responders cannot connect the user, context, tool and result, continuous verification is incomplete even when access was denied.
Repeat these tests against direct APIs as well as the conversational interface. A refusal displayed by the agent does not prove that the backend rejected the action. Inspect tool calls, authorization records and business-system state. Include failure cases such as expired tokens, unavailable policy services and malformed tool output. Sensitive operations should fail closed, generate useful evidence and avoid revealing secrets in errors. The results become regression tests for future model and connector updates.
FAQ
Validate the architecture with the AI agent red teaming checklist and apply the specialised MCP server security guide to tool connections.
Is a firewall enough for agent security?
No. The most important decisions happen at identity, data and action boundaries inside the application.
Should agents be allowed to act autonomously?
Low-risk, reversible tasks can be automated. High-impact actions should use policy checks, confirmation or human approval.
How often should policies be reviewed?
Review when capabilities change and at least quarterly for production agents.
Conclusion
Zero trust makes AI agents safer by making every action attributable, limited and verifiable. Treat the agent as a software actor with changing context—not as a trusted employee.
Start by removing one implicit trust assumption. Replace a shared credential with an agent identity, move an access rule from the prompt into the API, or require approval for an external action. Test the change with a low-privilege user and untrusted document. Small, verifiable improvements create a stronger foundation than a large policy that applications cannot enforce. As agents gain new tools, repeat the same questions: who is acting, on whose behalf, against which resource, under what current evidence and with what recovery path?