Skip to content

Detox Technologies

Prompt Injection Testing Methodology for LLM Apps and AI Agents

Prompt injection testing evaluates whether untrusted instructions can change an AI system’s behaviour, expose sensitive information or trigger an unauthorised action. It is not simply a collection of jailbreak prompts. A useful assessment tests the complete application: the model, system instructions, retrieval pipeline, memory, tools, APIs, identities, approval gates and the business systems an AI agent can reach.

This methodology is designed for security teams, developers and product owners testing LLM applications, RAG systems, copilots and autonomous agents. It focuses on repeatable evidence and business impact rather than spectacular but low-value model outputs.

What is prompt injection?

NIST defines prompt injection as an attack that exploits the combination of untrusted input with a prompt created by a higher-trust party. In practice, the model receives instructions and data in the same context and may fail to preserve the intended trust boundary between them.

OWASP distinguishes two important forms:

  • Direct prompt injection: a user submits instructions intended to override or redirect the application’s expected behaviour.
  • Indirect prompt injection: malicious instructions are placed in external content—such as a webpage, document, email, ticket, code repository or tool response—that the AI system later reads.

The impact depends on the application’s authority. A text-only assistant may produce an incorrect answer. An agent with email, customer-data or cloud access could disclose information or perform an unauthorised action. Testing must therefore measure what the system can say, access and do.

Prompt injection vs jailbreaking

The terms overlap but should not be treated as identical. Jailbreaking generally aims to bypass model safety restrictions. Prompt injection is broader: it manipulates an LLM-integrated application so that untrusted input influences higher-trust instructions or workflows.

A test that makes a model produce prohibited text may demonstrate a safety-control weakness. A test that causes a support agent to retrieve another customer’s record demonstrates an application-security failure with a clearer confidentiality and authorisation impact. Both matter, but they require different remediation.

Before testing: define safe scope and success criteria

Prompt injection tests should run under written authorisation, preferably in a representative staging environment. Use synthetic records and test accounts whenever possible. If production testing is necessary, agree on rate limits, prohibited actions, emergency contacts and stop conditions.

Document the following before sending the first payload:

  • Models, versions, prompts and environments in scope.
  • User roles, tenants and agent identities.
  • RAG data sources, uploads, memory and external content.
  • Tools, plugins, MCP servers, APIs and downstream systems.
  • High-impact actions that require approval.
  • Sensitive data classes and test substitutes.
  • Expected logging, alerting and evidence retention.

Define success as an observable security outcome. “The model followed my instruction” is often too vague. Better criteria include: a restricted tool was invoked, another tenant’s synthetic record was returned, a human approval was bypassed, a test secret appeared in output, or an injected instruction persisted into a later session.

A practical prompt injection testing methodology

Step 1: map the AI system and its trust boundaries

Draw the path from user input to final action. A typical flow may be user → web interface → orchestration layer → model → retrieval service → tool gateway → business API. Mark where content changes trust level and where authorisation decisions occur.

Inventory every source the model can read: chat messages, uploaded files, vector-search results, web pages, emails, tickets, database records, tool output, image text and conversation memory. Indirect injection becomes possible wherever an attacker can influence one of these sources.

Step 2: establish a clean behavioural baseline

Run normal workflows before adversarial testing. Record the expected response, tool calls, retrieved documents, approval prompts, latency and logs. Without a baseline, testers may mistake normal model variability for a security finding.

Use deterministic settings where practical and repeat important cases. Store the model version, system configuration and timestamp so a failed case can be reproduced after remediation or a model update.

Step 3: test direct prompt injection

Begin with harmless requests that conflict with application rules. Vary the position, wording and format rather than relying on one famous payload. Useful families include instruction override, role reassignment, context extraction, multi-turn manipulation, payload splitting, multilingual variants and encoded or obfuscated instructions.

Observe more than the visible answer. Confirm whether hidden retrieval occurred, whether a tool call was prepared, whether sensitive context reached the model and whether the server accepted any resulting action. A refusal message is not proof of safety if the backend operation still occurred.

Step 4: test indirect prompt injection

Place benign test instructions inside each supported external source. Examples include a test document asking the agent to reveal a synthetic marker, a webpage requesting an unauthorised tool, or an email attempting to redirect a workflow. Do not use real secrets or destructive commands.

Test visible and non-obvious locations that the system genuinely processes: document metadata, quoted email chains, HTML attributes, OCR-readable image text, code comments and tool-response fields. The objective is not filter evasion for its own sake; it is to determine whether untrusted content is treated as authority.

Step 5: test RAG and retrieval boundaries

For RAG systems, evaluate document ingestion, retrieval and generation separately. Can an unauthorised user add content to a trusted corpus? Can one tenant retrieve another tenant’s chunks? Can a poisoned document dominate ranking? Can citations make malicious content appear trustworthy?

Use unique synthetic markers to trace which document reached the model. Test with different users, roles, tenants and query styles. Our RAG security testing checklist provides additional coverage for poisoning, isolation and retrieval controls.

Step 6: test agent tools and authorisation

Prompt injection becomes materially more dangerous when an agent can act. Test whether the model can select an unapproved tool, modify arguments, access another user’s object, reuse stale credentials or exceed the initiating user’s permissions.

Every action must be authorised by deterministic server-side policy. System prompts are not access-control mechanisms. Repeat tests through the underlying APIs because an apparently safe conversation can still sit on top of an insecure tool endpoint. Dedicated API penetration testing is valuable when tools expose sensitive business functions.

Step 7: test approval gates and high-impact actions

Identify operations such as sending external messages, approving refunds, changing permissions, executing code or deleting data. Test whether an injected instruction can skip, obscure or pre-populate the approval step. Verify that the approver sees the exact action, target and relevant parameters.

Approval should be bound to one specific operation and expire after use. Test replay, parameter changes after approval and chained actions where a low-risk first step creates a high-risk second step.

Step 8: test data leakage and prompt extraction

Seed the environment with synthetic secrets representing API keys, customer data and confidential instructions. Attempt to expose them through direct requests, indirect content, error messages, citations, encoded output and tool arguments.

Inspect more than chat output. Sensitive information may appear in traces, analytics, logs, model-provider requests, vector stores or third-party plugins. Review the surrounding LLM API implementation using an LLM API security checklist.

Step 9: test multi-turn, memory and persistence

Some injections only become effective after several turns. Test whether a malicious instruction persists after the original content leaves the context window, contaminates long-term memory or affects another user. Confirm that deleted conversations and documents no longer influence retrieval.

For agents, test whether injected state survives task hand-offs or moves between planner, executor and reviewer components. A secure system should maintain provenance and apply policy at every step, not only at the first prompt.

Step 10: test multimodal inputs

If the system processes images, audio, PDFs or video, repeat core cases through each modality. Instructions may be hidden in OCR-readable text, document layers or transcribed audio. Compare what the human reviewer sees with what the model receives.

Multimodal testing must remain tied to impact. A hidden phrase that changes a harmless summary is different from one that triggers a tool or exposes data.

Step 11: validate detection and response

Generate recognisable test events and verify that logs contain the user identity, agent identity, source content, retrieved documents, tool calls, policy decisions and outcome. Confirm that defenders can revoke tokens, disable the affected agent and identify impacted sessions.

Detection should focus on unauthorised outcomes and unusual behaviour, not only known prompt strings. Attack wording changes quickly, while privilege escalation, cross-tenant access and abnormal tool use remain observable.

How to rate prompt injection findings

Severity should reflect demonstrated business impact and the reliability of the attack. Consider:

  • Attacker access required and whether the injection can be placed remotely.
  • Data sensitivity and tenant boundaries affected.
  • Tools and permissions available to the compromised workflow.
  • Need for human interaction or approval.
  • Repeatability across users, sessions and model versions.
  • Quality of logging, containment and recovery.

A strange response with no security impact may be informational. Reliable cross-tenant data access, unauthorised code execution or high-impact tool use can be critical. Keep model-safety findings separate from application-security findings so the correct engineering owner can respond.

Controls to verify—not merely recommend

  • Least privilege: agents and tools receive only the permissions required for the task.
  • Server-side authorisation: every object and action is checked independently of the model.
  • Untrusted-content handling: retrieved text and tool output are treated as data, with clear provenance.
  • Strict tool schemas: arguments are allowlisted and validated against business rules.
  • Human approval: sensitive actions require specific, tamper-resistant confirmation.
  • Output validation: downstream systems never execute model output without deterministic checks.
  • Monitoring and revocation: teams can detect abuse and rapidly disable identities or integrations.

OWASP notes that prompt injection cannot be treated as a problem solved by one filter or stronger system-prompt wording. Defence in depth limits the impact when model-level controls fail.

Evidence and retesting

For each confirmed issue, record prerequisites, test identity, input source, exact payload, retrieved context, model and configuration, tool calls, authorisation result, final outcome and timestamp. Preserve the minimum evidence needed for engineering teams to reproduce the issue without storing unnecessary sensitive data.

After remediation, repeat the original case and nearby variants. Add the test to a regression suite that runs after changes to models, prompts, retrieval sources, tools and permissions. NIST’s recent AI-agent security work reinforces the need for continuing evaluation rather than a one-time test.

When to use an independent AI security assessment

Independent testing is especially useful when an AI system handles sensitive data, serves multiple tenants, uses external content, calls business tools or can make consequential decisions. Detox Technologies’ AI security testing and red teaming services combine adversarial prompt testing with application, API, identity, cloud and workflow assessment.

For agent-specific scenarios, use our AI agent red teaming checklist. The goal is not to prove that a model will never make a mistake. It is to verify that a manipulated model cannot turn that mistake into unauthorised access, data exposure or an uncontrolled business action.

Authoritative references

Discover more from Detox Technologies

Subscribe now to keep reading and get access to the full archive.

Continue reading

Verified by MonsterInsights