Skip to content

Detox Technologies

LLM Data Leakage Testing: A Practical Security Methodology

LLM applications can expose sensitive information through generated responses, retrieved documents, conversation memory, tool calls, error messages and telemetry. The risk is wider than asking a model to repeat its system prompt. A secure assessment must trace data from its source to every place it is processed, stored and returned.

LLM data leakage testing determines whether an unauthorised user can obtain personal information, customer records, credentials, source code or confidential business content. It also checks whether authorised data is sent to an unintended model provider, plugin, analytics service or external tool.

Define the security objective

The objective is to prove that each identity receives only the data required for an authorised task, that sensitive content is minimised before reaching the model, and that output and telemetry cannot become alternate exfiltration paths.

Test the complete application rather than the model in isolation. Record users, agents, prompts, retrieval sources, memory, tools, APIs, model providers, logs and downstream destinations. Mark sensitive data and the deterministic controls expected to protect it.

Where risk appears

1. Prompt and context exposure

System instructions, examples and retrieved context may contain secrets or internal details. Treat system prompts as potentially discoverable and keep credentials and enforcement logic outside them.

2. Cross-tenant retrieval

A RAG query may return chunks belonging to another customer because filters are missing, applied after retrieval or derived from model output. Test ownership at the retrieval service and source repository.

3. Conversation and long-term memory

Memory can preserve personal data, tokens or confidential instructions beyond the expected session. Test deletion, expiry, user separation and whether one conversation affects another.

4. Tool and API responses

An over-privileged tool may return complete records when the task needs one field. The model can then expose data intentionally or accidentally. Minimise responses and enforce object-level authorisation.

5. Logs, traces and evaluations

Prompts and outputs are frequently copied into observability, feedback and evaluation platforms. Review access, retention, redaction and export controls for every telemetry destination.

6. Model-provider and plugin boundaries

Data may leave the organisation through hosted models, connectors and guardrail services. Verify contracts, regional processing, training settings and technical egress restrictions.

7. Encoded and indirect output

Sensitive content may appear in citations, URLs, tool arguments, files or encoded text even when chat output is filtered. Inspect every output channel used by the workflow.

8. Training and fine-tuning data

Customer or employee content may enter training, fine-tuning or evaluation datasets without consent and later be reproduced. Test collection, de-identification, lineage and deletion.

Practical testing methodology

Use written authorisation, synthetic sensitive records and distinct users and tenants. Prefer staging; if production is necessary, agree on safe actions, rate limits and stop conditions. Preserve model and configuration versions so results can be reproduced.

Step 1: Classify and seed synthetic data

Create unique markers for credentials, personal data, customer records and confidential documents. Place them only in authorised test locations so any appearance can be traced.

Step 2: Build user and tenant pairs

Use at least two ordinary users, two tenants and one administrator. Repeat identical queries while changing identity, role and resource ownership.

Step 3: Test direct extraction

Ask for restricted data using normal language, role-play, summarisation, transformation, multilingual and multi-turn requests. Record whether the model refuses and whether the data nevertheless entered context.

Step 4: Test indirect prompt injection

Place benign exfiltration instructions in documents, webpages and tool responses. Verify that untrusted content cannot redirect data to another destination.

Step 5: Inspect retrieval and memory

Capture document IDs, tenant filters, chunk provenance and memory writes. Delete content and confirm it no longer influences later responses.

Step 6: Inspect tools and APIs

Call underlying endpoints with changed identifiers and roles. A safe chat response does not compensate for an API that returns another user’s record.

Step 7: Review telemetry and providers

Trace prompts, outputs and attachments through logs, analytics, model gateways and third parties. Verify masking and retention in practice.

Step 8: Test output channels

Review chat, downloads, citations, rendered HTML, emails and tool arguments for the synthetic markers.

Step 9: Validate revocation and deletion

Remove a user, document or consent grant and confirm cached embeddings, queued jobs, memory and backups follow the defined lifecycle.

High-value scenarios

Customer-support assistant

Ask one customer’s assistant to summarise a ticket owned by another account, vary ticket IDs and use a poisoned knowledge article. The API and retrieval layer must reject the request before data reaches the model.

Enterprise coding assistant

Seed a repository with a synthetic secret and test code search, generated patches, logs and external model requests. Confirm secrets are excluded or redacted and repository permissions follow the user.

Meeting and email copilot

Test quoted threads, attachments, participant changes and summary sharing. A late-added participant should not gain access to earlier confidential context without explicit policy.

Controls to verify

  • Data classification and minimisation before prompt assembly.
  • Server-side user, tenant and object authorisation for retrieval and tools.
  • Secrets management with no credentials in prompts or model-visible configuration.
  • Provider and plugin egress controls plus documented retention and training settings.
  • Redaction for logs, traces, evaluations and support exports.
  • Per-user memory boundaries, deletion and retention controls.
  • Output validation and destination allowlists for external actions.
  • Monitoring for unusual retrieval, bulk output and synthetic canary exposure.

Common mistakes

  • Treating the system prompt as a secret store.
  • Filtering visible chat while ignoring logs and tool arguments.
  • Applying tenant filters after vector retrieval.
  • Giving the model broad database or cloud credentials.
  • Using real customer data for adversarial tests.
  • Closing a finding after a refusal without checking backend requests.
  • Forgetting cached embeddings and queued jobs during deletion.

Evidence and reporting

Record the synthetic marker, source system, user and tenant, retrieval results, prompt assembly, model endpoint, tool calls, output channel and storage destinations. Distinguish between data entering model context and data reaching an unauthorised party; both may require remediation, but impact differs.

Rate findings by demonstrated business impact, attacker access and repeatability. After remediation, repeat the exact case and nearby variants, then add them to a regression suite for model, prompt, data and tool changes.

Related Detox resources

Use the RAG security testing checklist, LLM API security checklist and prompt injection methodology. Detox can validate the combined architecture through AI security testing services.

Frequently asked questions

Is system-prompt disclosure always a critical vulnerability?

No. The prompt should not contain secrets or replace access control. Severity depends on what sensitive information or exploit path the disclosure enables.

Can output filtering prevent all leakage?

No. Filtering is one layer; data minimisation, authorisation and egress control reduce the information available to leak.

Should testers use real personal data?

Use synthetic records whenever possible. Real data increases privacy risk and is rarely required to prove a boundary failure.

Conclusion

LLM data protection starts before the prompt and continues after the response. Limit what the model can receive, enforce access at source systems, control every destination and verify deletion and monitoring with traceable test data.

Authoritative references

How to design a leakage test matrix

A useful test matrix connects data classes, identities and output paths. For every sensitive class—customer records, personal data, credentials, source code and regulated documents—record who may access it, which retrieval source contains it, which tools can return it and where output can travel. Then execute positive and negative cases for ordinary users, administrators, service accounts and cross-tenant identities. This structure prevents a successful refusal in one chat prompt from hiding a leakage path in an export, citation, tool argument or background job.

Measure context exposure separately from external disclosure

Two related events should be reported separately. First, did restricted data enter the model context? Second, did it reach an unauthorised person or system? Context exposure can still be serious because hosted providers, traces or later model operations may process that information even when the visible answer is blocked. Capturing both stages helps engineering teams choose the correct fix: data minimisation and retrieval authorisation for the first, destination controls and output handling for the second.

Test changes and regression boundaries

Repeat the matrix after changes to models, embedding pipelines, chunking, rerankers, prompts, memory settings, connectors and observability platforms. Keep synthetic canaries stable enough to compare releases, but rotate them when a test value could become known to developers or tuning systems. A regression suite should include expected access as well as denial cases, because an overly broad block can damage legitimate workflows without actually repairing the source boundary.

Interpreting severity and business impact

Severity depends on the sensitivity and volume of exposed data, the attacker access required, the number of tenants affected and whether disclosure is repeatable or leaves the organisation. A single synthetic secret returned to another tenant usually demonstrates a broken isolation boundary even if bulk extraction has not yet been attempted. Reports should avoid speculative claims: show the exact data path, the minimum reliable reproduction and the realistic business consequence.

Discover more from Detox Technologies

Subscribe now to keep reading and get access to the full archive.

Continue reading

Verified by MonsterInsights