Skip to content
Talk to a Security Expert
AI / LLM SECURITY PENETRATION TESTOWASP TOP 10 FOR LLM APPLICATIONSSAMPLE REPORT

AI / LLM Security Penetration Testing Report — Sample

NimbusMind AI, Inc. — AI/LLM Application Security Assessment. A fictional AI/LLM application penetration test of NimbusMind Copilot, an AI-powered customer-support and knowledge-assistant platform operated by NimbusMind AI, Inc., showing how TMG Security structures scope, methodology, prompt-injection and jailbreak testing, agent/tool-use security, RAG and vector-store testing, findings, risk analysis, and remediation guidance for an AI/LLM assessment informed by the OWASP Top 10 for LLM Applications.

Illustrative Sample Fictional Organization Not an Actual Client Assessment

NimbusMind AI, Inc. and the NimbusMind Copilot AI platform are fictional. No real model, backend, agent tool, or client environment was built, deployed, accessed, or tested, and no real customer data was involved in the creation of this document. TMG Security is not affiliated with, endorsed by, or certified by OWASP or any foundation-model provider.

01 · OVERVIEW

AI LLM Security Penetration Testing — Overview

An AI/LLM application penetration test is an authorized, hands-on security assessment that combines unauthenticated black-box testing of public-facing chat surfaces, authenticated grey-box testing of the production chat and developer-support assistant variants, and assumed-breach testing from a low-privilege compromised account to identify exploitable weaknesses across the AI/LLM-specific attack surface. This illustrative sample shows how TMG Security structures that deliverable for a conversational AI product built on a third-party foundation-model API with a retrieval-augmented generation (RAG) knowledge base, an agentic function-calling tool layer, and multimodal image-input support.

This sample report demonstrates TMG Security's approach for a fictional organization, NimbusMind AI, Inc., and its fictional product, NimbusMind Copilot — an AI-powered customer-support and knowledge-assistant platform. The assessment was informed by, but not certified against, the OWASP Top 10 for LLM Applications and the OWASP GenAI Security Project's terminology, combined with hands-on testing specific to the target's agentic, RAG-backed architecture. This sample illustrates the depth of a typical TMG Security AI / LLM Security Penetration Testing engagement.

02 · WHAT THIS SAMPLE DEMONSTRATES

What This AI LLM Security Penetration Testing Sample Demonstrates

This sample AI / LLM Security Penetration Testing report is a fictional, illustrative deliverable prepared for an OWASP Top 10 for LLM Applications-informed assessment — not to describe any real organization's security posture.

  • How TMG defines scope, rules of engagement, and testing objectives for a conversational AI product and its agentic tool layer
  • How prompt injection and jailbreak resistance are tested across direct, indirect (RAG-sourced), and multimodal input vectors
  • How the RAG retrieval service, vector store, and knowledge-base content trust boundary are assessed
  • How the agent's exposed function-calling tool registry is tested for authorization, parameter-validation, and excessive-agency gaps
  • How sensitive data disclosure, PII handling, and cross-session context leakage are evaluated end to end
  • How model/API abuse, rate limiting, logging, and AI-specific detection controls are assessed as defense-in-depth
  • How individual findings are cross-referenced into realistic, non-destructive composite attack chains

None of these samples is a certification, an attestation, an actual client engagement, or proof of any compliance or security status. This report is labeled as illustrative throughout and carries a full disclaimer on this page and within the sample PDF.

03 · APPLICATION & PLATFORM PROFILE

Application & platform profile

The fictional NimbusMind Copilot platform is built on a third-party foundation-model API (fictional "ModelArc LLM Platform") and is architected as a layered conversational AI system: client surfaces (an authenticated web chat, an unauthenticated pre-signup marketing widget, a developer-support assistant variant, and an internal staff-facing ticket console) call into an API gateway and agent orchestration/planning layer, which coordinates a guardrail classifier, a RAG retrieval service backed by a shared vector store, and a 14-tool function-calling registry, plus a multimodal vision-extraction pipeline for image uploads. Downstream, these components read from and write to a multi-tenant customer/ticket data store, object storage for uploads and transcripts, a fine-tuning and feedback pipeline, and a SIEM/guardrail log store.

OrganizationNimbusMind AI, Inc. (Fictional)
Product AssessedNimbusMind Copilot — AI-powered customer-support & knowledge assistant
LLM PlatformThird-party foundation-model API (fictional "ModelArc LLM Platform")
Assessment TypeAI/LLM Application Penetration Test — Black-Box, Grey-Box & Assumed-Breach
Assessment Period02 September 2026 – 24 September 2026
Testing Domains / Test Cases18 AI/LLM testing domains · 137 fictional test cases executed
04 · TESTING SCOPE & RULES OF ENGAGEMENT

Testing scope & rules of engagement

In scope

  • The Chat Orchestration Service — primary authenticated chat API and system-prompt/guardrail pipeline
  • The Knowledge Retrieval (RAG) Service — vector store, embedding pipeline, and retrieval-ranking logic
  • The Agent Tool Registry — 14 function-calling tools exposed to the assistant
  • The Multimodal Ingestion Pipeline — image-upload endpoint, vision-extraction, and OCR-adjacent processing
  • The Feedback & Fine-Tuning Pipeline — thumbs-up/down and free-text feedback ingestion and model-promotion workflow
  • The unauthenticated Marketing-Site Chat Widget and the Developer-Support Assistant variant
  • The Internal Agent Console, identity & session layer, and guardrail logging/SIEM integration

Out of scope

  • The underlying foundation-model provider's own infrastructure and training pipeline (treated as a trusted third party)
  • Physical security and any real-world office or data-center facilities
  • Denial-of-service testing beyond bounded, non-destructive rate-limiting checks
  • Social engineering of real personnel (no real personnel exist in this fictional scenario)

Testing was grey-box (authenticated low-privilege test accounts provided) plus black-box testing of unauthenticated surfaces, over the fictional window of 02 September 2026 – 24 September 2026 (fictional business hours). Destructive testing was not permitted; all account-action and data-modification tests used dedicated fictional test accounts only, and model-interaction volume was bounded and rate-aware. A 30-day retest window is included in scope at no additional fictional cost.

05 · METHODOLOGY & FRAMEWORK CONTEXT

AI LLM Security Penetration Testing Methodology & Framework Context

TMG Security's fictional AI / LLM Security Penetration Testing methodology combines industry-discussed practice for generative-AI security assessment — informed by, but not certified against, the OWASP Top 10 for LLM Applications, the OWASP GenAI Security Project's terminology, and the MITRE ATLAS adversarial-ML knowledge base — with hands-on testing specific to the target's agentic, RAG-backed architecture. Testing proceeded through six phases: reconnaissance & attack-surface mapping; prompt-level & guardrail testing; authenticated application & tool-layer testing; assumed-breach / compromised-session testing; platform, supply-chain & governance review; and attack-chain construction & reporting.

TMG Security is not affiliated with, endorsed by, reviewed by, or certified by OWASP, MITRE, any foundation-model provider, or any other standards organization. References to industry frameworks describe methodology influences only and do not imply certification against them.

06 · PROMPT INJECTION & RAG CONTENT TRUST (D01–D02)

Prompt injection, jailbreak resistance & RAG content trust

Testing exercised the guardrail classifier and system-prompt adherence against a broad library of documented jailbreak techniques — layered persona/role-play, encoding and translation-based bypass, token-smuggling, and multilingual coverage gaps — across both single-turn and extended multi-turn conversations (CWE-1427 — Improper Neutralization of Input Used for LLM Prompting). TMG Security separately tested indirect prompt injection, in which untrusted content ingested through the RAG knowledge base or retrieved documents is treated as trusted model context rather than untrusted data.

AILLM-001 · CRITICAL · CWE-1427 Improper Neutralization of Input Used for LLM Prompting

System-Prompt Override via Layered Persona Jailbreak

The NimbusMind Copilot chat endpoint can be coerced into abandoning its configured support-agent persona and system instructions through a layered role-play jailbreak ("DAN-style" nested persona plus a fictional-policy override framing). Once overridden, the model complied with requests it is explicitly instructed to refuse, including drafting policy-violating content and revealing internal tool names. Recommendation: evaluate the full conversation context — not just the last turn — at every guardrail checkpoint, and treat any detected persona-override attempt as a session-level signal that triggers stricter output filtering for the rest of the conversation.

Owner: AI Platform Security Lead (Fictional Role) · Target: 08 October 2026 · Status: Remediation In Progress

AILLM-002 · CRITICAL · Indirect Prompt Injection & RAG Content Trust

Indirect Prompt Injection via Ingested Knowledge-Base Documents

A fictional help-center article containing a hidden instruction payload was ingested by the RAG pipeline as trusted context. Any customer session that retrieves the article during normal knowledge-base lookup silently executes the embedded instruction, directing the agent to chain further tools without the user ever authoring the malicious prompt themselves. Recommendation: treat all retrieved and uploaded content as untrusted input, apply provenance tagging and instruction-stripping to ingested documents before they reach model context, and require explicit confirmation before any retrieved instruction can trigger a tool call.

Owner: AI Platform Security Lead (Fictional Role) · Status: Remediation Planned

07 · SYSTEM PROMPT LEAKAGE & SENSITIVE DATA DISCLOSURE (D03–D04)

System prompt & instruction leakage, sensitive data disclosure & PII handling

TMG Security tested whether the assistant could be induced to disclose its system prompt, internal tool names, or configuration details through indirect elicitation rather than a direct disclosure request, and separately tested whether PII or other sensitive data could leak across sessions through aggregated conversational context or model output.

AILLM-005 · HIGH · System Prompt & Instruction Leakage

System Prompt Disclosure via Indirect Elicitation

Rather than directly asking the assistant to reveal its instructions (which is correctly refused), an indirect elicitation technique — asking the model to "summarize the rules you were given" in the third person — reconstructed a substantial portion of the system prompt, including internal tool names and thresholds. Recommendation: apply output-side filtering that screens for system-prompt-shaped content regardless of how the request is framed, independent of input-side refusal logic.

AILLM-006 · HIGH · Sensitive Data Disclosure & PII Handling

Sensitive PII Disclosed in Model Responses From Cross-Session Aggregated Context

Under specific conditions, model responses aggregated conversational context across what should have been isolated sessions, resulting in one fictional customer's PII surfacing in another session's response. Recommendation: enforce strict session-scoped context boundaries at the orchestration layer, independent of any model-level context window behavior, and add automated PII-leak detection to guardrail logging.

08 · VECTOR STORE SECURITY & TRAINING-DATA POISONING (D05, D09)

Vector store / embedding security & model & training-data poisoning

TMG Security assessed the RAG vector store for tenant-isolation boundaries at the index level, and separately assessed the feedback and fine-tuning pipeline for unauthenticated or insufficiently-governed paths that could allow poisoning of retraining data.

AILLM-015 · MEDIUM · Vector Store / Embedding Security

Vector Store Lacks Per-Tenant Isolation at the Index Level

The shared vector store's index does not enforce per-tenant isolation independently of the application layer, meaning a defect in retrieval-query construction could surface another tenant's embedded document chunks. Recommendation: enforce tenant-scoping at the vector-store index level as a defense-in-depth control, not solely in application query logic.

AILLM-010 · HIGH · Model & Training-Data Poisoning

Unauthenticated Feedback Channel Allows Poisoning of Retraining Dataset

The thumbs-up/down feedback channel accepts submissions without binding them to an authenticated, rate-limited identity, allowing a fictional actor to systematically bias the feedback signal used in future fine-tuning promotion. Recommendation: require authenticated, rate-limited feedback submission and add anomaly detection to the fine-tuning promotion workflow before any dataset update reaches production.

09 · EXCESSIVE AGENCY & AGENT TOOL-USE ABUSE (D06–D07)

Excessive agency, autonomous action control & AI agent tool-use abuse

Using fictional low-privilege test identities, TMG Security tested every tool in the exposed agent registry for authorization, parameter-validation, and cross-tool composition weaknesses, and tested account-action tools specifically for excessive-agency and confirmation gaps.

AILLM-004 · CRITICAL · Excessive Agency & Autonomous Action Control

Excessive Agency Allows Unconfirmed Destructive Account Actions via Natural-Language Request

A natural-language request phrased as an ordinary support ask can trigger an irreversible account action — such as a subscription cancellation — with no explicit confirmation step, no human-in-the-loop gate, and no distinguishable difference in agent behavior between a routine query and a destructive command. Recommendation: require an explicit, out-of-band confirmation step for any tool call with irreversible or financial impact, regardless of how confidently the model interprets user intent.

AILLM-007 · HIGH · AI Agent & Tool-Use Abuse

Agent Tool-Chaining Allows Privilege Escalation Beyond Any Single Tool's Scope

No individual tool in the 14-tool registry is over-privileged in isolation, but the agent's planning layer can chain multiple tools together to reach an effective privilege beyond what any single tool call would allow, including reading data one tool alone could not reach. Recommendation: treat the agent's multi-step planning layer itself as a security boundary, with composition-aware authorization checks rather than per-tool checks alone.

10 · INSECURE TOOL / API ACCESS & SSRF (D08, D11)

Insecure tool / function-calling API access, BOLA & SSRF via AI features

TMG Security tested the mobile-facing and agent-facing tool APIs directly for broken object-level authorization, and tested any tool capable of fetching a URL on the model's behalf for server-side request forgery risk.

AILLM-003 · CRITICAL · CWE-639 Authorization Bypass Through User-Controlled Key

Broken Object-Level Authorization in Ticket-Lookup Tool Enables Cross-Tenant Data Access

The agent's ticket-lookup tool accepts a ticket identifier supplied through natural-language conversation without independently verifying that the requesting session's tenant owns that ticket, enabling full cross-tenant ticket data disclosure — including PII and billing references — through ordinary conversational requests. Recommendation: enforce tenant-scoped authorization at the tool-implementation layer for every parameter that references another object, independent of how the parameter value was derived (typed input vs. model-generated).

AILLM-009 · HIGH · SSRF & Server-Side Data Exfiltration via AI Features

SSRF via Unrestricted URL-Fetch Tool Used for 'Summarize This Link' Feature

The "summarize this link" feature's URL-fetch tool performs no destination validation, allowing a fictional link-local or internal-network URL to be fetched and its response summarized back to the requester. Recommendation: apply an allowlist and destination-validation layer (blocking link-local and private-network ranges) to any tool capable of making an outbound request on the model's behalf.

11 · MODEL/API ABUSE & MEMORY ATTACKS (D12–D13)

Model / API abuse, rate limiting, resource exhaustion & memory / context-window attacks

TMG Security tested per-session and per-account rate limiting on model-interaction volume, and tested whether concurrent session load could cause cross-user context to bleed within the model's active context window.

AILLM-011 · HIGH · Model / API Abuse, Rate Limiting & Resource Exhaustion

Missing Per-Session Rate Limiting Enables Model-Abuse and Cost-Amplification

No per-session rate limit governs model-interaction volume on the authenticated chat endpoint, allowing a single fictional account to drive materially higher inference cost than any legitimate usage pattern would produce, with no automated cost-anomaly alert. Recommendation: apply per-session and per-account rate limiting and cost-anomaly detection tuned to the AI feature's expected usage envelope.

AILLM-012 · HIGH · Memory & Context-Window Attacks

Cross-User Context Bleed Under Concurrent Session Load

Under sustained concurrent session load, TMG Security observed conditions under which one fictional user's conversational context could bleed into another's active session. Recommendation: enforce hard per-session context isolation at the orchestration layer, with load and concurrency testing specifically targeting the context-assembly code path.

12 · MULTIMODAL INPUT & OUTPUT HANDLING (D14, D10)

Multimodal input risks & output handling / insecure AI-generated content

TMG Security tested the multimodal image-upload pipeline for hidden instruction text embedded in images, and tested whether AI-generated output is safely handled by downstream systems that render it, such as the internal ticket console.

AILLM-008 · HIGH · Multimodal Input Risks (Image / File Upload)

Prompt Injection via Text Embedded in Uploaded Image (Multimodal Bypass)

The vision-extraction pipeline passes text recovered from an uploaded image into the model's context as trusted input, allowing a screenshot-styled image containing hidden instruction text to bypass text-based guardrail filtering entirely. Recommendation: apply the same untrusted-content treatment (provenance tagging, instruction-stripping) to text extracted from images as is applied to retrieved documents.

AILLM-016 · MEDIUM · Output Handling & Insecure AI-Generated Content

Unsafe Rendering of AI-Generated Markdown Enables Stored Content Injection in Ticket Notes

AI-generated markdown summaries are rendered into the internal ticket console without output sanitization, allowing a crafted conversational input to result in stored content injection visible to staff reviewing the ticket. Recommendation: treat all AI-generated content as untrusted output requiring the same sanitization applied to any other user-supplied content before rendering.

13 · SUPPLY CHAIN & AI-SPECIFIC LOGGING (D15–D16)

Supply chain & model/plugin dependency security, logging & AI-specific detection

TMG Security reviewed AI-stack dependency versions and advisories, and assessed whether guardrail-refusal and injection-attempt logging would support timely detection of a sustained attack pattern rather than only individual events.

AILLM-013 · HIGH · Supply Chain & Model/Plugin Dependency Security

Unpinned Third-Party Prompt-Template Library With Known Advisory in Production

A third-party prompt-template library used by the orchestration layer is deployed unpinned in production and currently resolves to a version with a published security advisory. Recommendation: pin all AI-stack dependencies to specific, reviewed versions and integrate advisory scanning into the deployment pipeline for prompt-related libraries specifically, not only general application dependencies.

AILLM-014 · HIGH · Logging, Monitoring & AI-Specific Detection

No Alerting on Repeated Jailbreak or Injection Attempts From the Same Account

Individual guardrail refusals are logged, but no aggregated alerting exists for a single account generating dozens of jailbreak or injection variants in sequence — the exact pattern an iterative attacker would use. Recommendation: add attempt-pattern alerting on repeated guardrail refusals from a single identity within a bounded time window, escalating to session-level restriction.

14 · PRIVACY GOVERNANCE & AI-ENDPOINT AUTHENTICATION (D17–D18)

Privacy, data governance & retention, authentication & session for AI endpoints

TMG Security reviewed data-retention practice for chat transcripts against documented policy, and tested session and authentication controls specific to the AI chat surfaces, including token lifecycle behavior around password reset.

AILLM-022 · MEDIUM · Privacy, Data Governance & Retention

Chat Transcripts Retained Indefinitely With No Documented Retention Policy Enforcement

Chat transcripts containing customer conversational data are retained indefinitely with no automated enforcement of any documented retention window. Recommendation: define and technically enforce a retention schedule for conversational data consistent with the platform's privacy documentation.

AILLM-019 · MEDIUM · Authentication, Session & Authorization for AI Endpoints

Session Tokens for Chat Widget Do Not Expire on Password Reset

Existing chat-widget session tokens remain valid after a fictional user resets their password, rather than being invalidated as part of the reset flow. Recommendation: invalidate all active chat-session tokens for an account whenever a password reset or credential change occurs.

15 · FINDINGS SUMMARY

AI LLM Security Penetration Testing — Findings Summary

This fictional assessment identified 45 illustrative findings across NimbusMind Copilot's eighteen AI/LLM testing domains, drawn from 137 fictional test cases. Every finding, severity rating, and evidence reference below is fictional and specific to this AI / LLM Security Penetration Testing engagement.

4
Critical
10
High
15
Medium
9
Low
7
Informational
Full illustrative finding register — fictional sample data, Appendix C of the sample PDF.
IDTitleSeverity
AILLM-001System-Prompt Override via Layered Persona JailbreakCritical
AILLM-002Indirect Prompt Injection via Ingested Knowledge-Base DocumentsCritical
AILLM-003Broken Object-Level Authorization in Ticket-Lookup Tool Enables Cross-Tenant Data AccessCritical
AILLM-004Excessive Agency Allows Unconfirmed Destructive Account Actions via Natural-Language RequestCritical
AILLM-005System Prompt Disclosure via Indirect ElicitationHigh
AILLM-006Sensitive PII Disclosed in Model Responses From Cross-Session Aggregated ContextHigh
AILLM-007Agent Tool-Chaining Allows Privilege Escalation Beyond Any Single Tool's ScopeHigh
AILLM-008Prompt Injection via Text Embedded in Uploaded Image (Multimodal Bypass)High
AILLM-009SSRF via Unrestricted URL-Fetch Tool Used for 'Summarize This Link' FeatureHigh
AILLM-010Unauthenticated Feedback Channel Allows Poisoning of Retraining DatasetHigh
AILLM-011Missing Per-Session Rate Limiting Enables Model-Abuse and Cost-AmplificationHigh
AILLM-012Cross-User Context Bleed Under Concurrent Session LoadHigh
AILLM-013Unpinned Third-Party Prompt-Template Library With Known Advisory in ProductionHigh
AILLM-014No Alerting on Repeated Jailbreak or Injection Attempts From the Same AccountHigh
AILLM-015Vector Store Lacks Per-Tenant Isolation at the Index LevelMedium
AILLM-016Unsafe Rendering of AI-Generated Markdown Enables Stored Content Injection in Ticket NotesMedium
AILLM-017Language-Switching Bypass of Content-Policy FiltersMedium
AILLM-018RAG Source Documents Lack Freshness/Trust Scoring, Allowing Stale or Superseded Guidance to SurfaceMedium
AILLM-019Session Tokens for Chat Widget Do Not Expire on Password ResetMedium
AILLM-020Tool Parameter Schema Accepts Overly Broad Free-Text Where a Constrained Enum Should Be UsedMedium
AILLM-021Debug/Verbose Mode Response Header Discloses Model Version and Internal Pipeline Stage TimingsMedium
AILLM-022Chat Transcripts Retained Indefinitely With No Documented Retention Policy EnforcementMedium
AILLM-023Uploaded File Type Validation Relies on Client-Supplied MIME Type OnlyMedium
AILLM-024Agent Retries Failed Tool Calls With Escalating Privilege Rather Than Fixed ScopeMedium
AILLM-025No Cost/Usage Anomaly Detection for Individual Tenant AccountsMedium
AILLM-026Error Messages Reveal Internal Function/Tool Names on Malformed Tool-Call FailuresMedium
AILLM-027No Human Review Gate Before Fine-Tuning Dataset Promotion to ProductionMedium
AILLM-028AI-Generated Code Snippets in Developer-Support Responses Not Flagged as UnverifiedMedium
AILLM-029Incomplete Audit Trail for Tool Calls — Parameters Logged, Response Bodies TruncatedMedium
AILLM-030Assistant Discloses Its Own Model-Family Name When Directly AskedLow
AILLM-031No CAPTCHA or Bot-Mitigation on Unauthenticated Marketing-Site Chat WidgetLow
AILLM-032Verbose Client-Side Console Logging of Assembled Prompt Context in Non-Production BuildLow
AILLM-033Retrieved Document Snippets Shown to User Lack a Clear Visual Distinction From Model-Generated TextLow
AILLM-034Chat API Accepts Requests With Expired-But-Not-Yet-Purged Bearer Tokens for a Short Grace WindowLow
AILLM-035Uploaded Images Retain EXIF Metadata Through the Ingestion PipelineLow
AILLM-036Vector-Database Client Library Two Minor Versions Behind Latest, No Known AdvisoryLow
AILLM-037No User-Facing Log of Autonomous Actions Taken by the Assistant on the User's BehalfLow
AILLM-038Guardrail Refusal Reason Codes Are Inconsistent Across Services, Complicating Aggregated ReportingLow
AILLM-039Observation: Data-Processing Addendum Language Has Not Been Updated to Reference the AI Feature SetInformational
AILLM-040Observation: Refusal Messages Are Not Localized, Defaulting to English Regardless of Conversation LanguageInformational
AILLM-041Observation: Tool Registry Documentation Is Current and Matches Deployed Tool SetInformational
AILLM-042Observation: Guardrail-Refusal Events Are Consistently Timestamped and Immutable in StorageInformational
AILLM-043Observation: Financial Identifiers Are Consistently Masked in the Chat UI LayerInformational
AILLM-044Observation: MFA Is Enforced for All Internal Support-Console Administrator AccountsInformational
AILLM-045Observation: Embedding Model Version Is Pinned and Consistently Applied Across the CorpusInformational
16 · ATTACK CHAIN ANALYSIS

Attack chain analysis

Individual findings were cross-referenced to construct four realistic, non-destructive fictional attack chains, demonstrating how lower-friction weaknesses combine into materially higher business impact than any single finding suggests in isolation. All four chains below are fictional, illustrative, and non-destructive; no real system, model, or data was accessed, modified, or exploited in their construction or validation.

AC-01 — Public Content → Session Hijack → Data Exfiltration

Step 1: Attacker publishes a poisoned help-center article containing a hidden instruction payload (AILLM-002). Step 2: Any customer session that retrieves the article during normal RAG lookup silently executes the embedded instruction. Step 3: The embedded instruction directs the agent to chain the search_customer, get_ticket_details, and send_email tools (AILLM-007) to exfiltrate the victim's own session/ticket data to an attacker-controlled address. Step 4: Because tool-response logging is truncated (AILLM-029), TMG's fictional incident responders cannot fully reconstruct what data left the platform without the underlying fix.

Business impact: A single content-publishing capability, combined with under-governed agent tool-chaining and incomplete audit logging, escalates to session-level data exfiltration against any customer who merely asks a question that retrieves the poisoned article.

AC-02 — Jailbreak → Cross-Tenant BOLA → Full Ticket Disclosure

Step 1: Attacker establishes a persona-override jailbreak against the support chatbot (AILLM-001), suspending the assistant's normal refusal behavior for the remainder of the session. Step 2: Within the jailbroken session, the attacker requests ticket lookups using guessed or sequential ticket IDs belonging to other fictional tenants. Step 3: The get_ticket_details tool's missing tenant-scoping check (AILLM-003) returns full cross-tenant ticket content, including PII and billing references, with no additional authorization barrier.

Business impact: Combining a behavioral jailbreak with a tool-layer authorization gap allows an attacker to move from "convince the chatbot to misbehave" directly to "read any customer's support history" in a single session.

AC-03 — Crafted Image Upload → Multimodal Injection → Excessive-Agency Account Action

Step 1: Attacker uploads a screenshot-styled image containing hidden instruction text, exploiting the multimodal injection gap (AILLM-008). Step 2: The vision-extraction pipeline passes the hidden instruction into the model's context as trusted input. Step 3: The injected instruction phrases a subscription-cancellation request in a way the excessive-agency gap (AILLM-004) executes immediately with no confirmation step.

Business impact: This chain demonstrates that the multimodal injection surface is not merely a content-disclosure risk: combined with under-governed autonomous tools, it can trigger real, irreversible account and billing actions with no user awareness at the time of upload.

AC-04 — Undetected Iterative Jailbreak Attempts → System-Prompt Reconstruction → Targeted Tool Abuse

Step 1: Attacker iterates through dozens of jailbreak and elicitation variants against a single account with no aggregated alerting to flag the pattern (AILLM-014). Step 2: Using indirect elicitation techniques, the attacker gradually reconstructs the system prompt, learning internal tool names and thresholds (AILLM-005). Step 3: Armed with the exact tool name and parameter structure, the attacker crafts a malformed update_ticket_priority call that bypasses weak parameter validation (AILLM-020) to disrupt ticket routing.

Business impact: This chain illustrates how a purely detective control gap (no alerting on repeated attempts) removes the operational visibility that would otherwise have interrupted the reconnaissance-to-exploitation path well before the final tool-abuse step.

17 · BUSINESS IMPACT ANALYSIS

Business impact analysis

The findings in this assessment, individually and in combination, translate into concrete fictional business risk across five categories, supporting prioritized investment decisions by NimbusMind AI, Inc.'s fictional leadership.

Impact CategoryPriorityFictional Business Narrative
Data Breach / PrivacyCriticalCross-tenant and session-level data disclosure findings (AILLM-002, AILLM-003, AILLM-006) would, in a real deployment, constitute a reportable breach affecting every tenant, with attendant regulatory notification and contractual liability exposure.
Brand & Customer TrustHighA publicly jailbreakable customer-facing chatbot (AILLM-001, AILLM-016) risks screenshot-driven reputational damage and erodes customer confidence in the AI product.
Financial / Revenue AssuranceHighUnconfirmed destructive account actions (AILLM-004, AILLM-011, AILLM-025) directly threaten recurring revenue; uncontrolled inference cost and usage abuse threaten gross margin on AI features.
Regulatory / ComplianceMediumRetention-policy drift and incomplete AI-specific data-processing documentation (AILLM-022, AILLM-039) create audit and regulatory exposure, particularly for a platform processing customer conversational data at scale.
Operational / Detection ReadinessMediumGaps in attempt-pattern alerting and audit-log completeness (AILLM-014, AILLM-029) would materially slow incident detection and breach-scope determination.
18 · REMEDIATION ROADMAP

Remediation roadmap

The fictional sample remediation roadmap sequences all 45 findings across four phases, prioritizing the four Critical findings and the highest-risk High findings first. All dates and timelines shown are illustrative.

0–14 days

  • System-prompt override via layered persona jailbreak (AILLM-001)
  • Indirect prompt injection via ingested knowledge-base documents (AILLM-002)
  • Broken object-level authorization in ticket-lookup tool (AILLM-003)
  • Excessive agency allowing unconfirmed destructive account actions (AILLM-004)

15–30 days

  • System prompt disclosure via indirect elicitation (AILLM-005)
  • Sensitive PII disclosure from cross-session context (AILLM-006)
  • Agent tool-chaining privilege escalation (AILLM-007)
  • Multimodal prompt injection via uploaded image (AILLM-008)
  • SSRF via unrestricted URL-fetch tool (AILLM-009)
  • Unauthenticated feedback channel poisoning risk (AILLM-010)
  • Missing per-session rate limiting (AILLM-011)
  • Cross-user context bleed under concurrent load (AILLM-012)
  • Unpinned third-party prompt-template library (AILLM-013)
  • No alerting on repeated jailbreak attempts (AILLM-014)

31–90 days

  • All 15 Medium findings (AILLM-015 through AILLM-029), covering vector-store isolation, output sanitization, session-token lifecycle, and tool parameter validation
  • All 9 Low findings (AILLM-030 through AILLM-038), covering hygiene, logging consistency, and dependency freshness

Continuous / informational backlog

  • 7 informational observations (AILLM-039 through AILLM-045), including 5 confirmed positive controls tracked for continued monitoring
19 · POSITIVE SECURITY CONTROLS & TESTING LIMITATIONS

Positive security controls & testing limitations

Consistent with TMG Security's standard reporting practice, this section documents fictional controls confirmed operating effectively — an important counterweight to the findings sections, since an AI security report that only lists problems gives an incomplete picture of NimbusMind AI, Inc.'s actual fictional posture. Positive observations, recorded in full as Informational findings AILLM-041 through AILLM-045, include: tool registry documentation that is current and matches the deployed tool set; guardrail-refusal events that are consistently timestamped and immutable in storage; financial identifiers that are consistently masked in the chat UI layer; MFA enforced for all internal support-console administrator accounts; and an embedding model version that is pinned and consistently applied across the corpus.

As with any time-boxed security assessment, this fictional engagement is subject to inherent limitations. AI/LLM behavior is probabilistic — a technique that succeeded during fictional testing may not succeed on every attempt, and a technique that failed may succeed under different phrasing, timing, or model-version conditions not covered in this assessment window. Testing was time-boxed to the assessment period stated in Section 04 and does not guarantee the absence of additional exploitable weaknesses outside the tested scope, techniques, or time window. Findings reflect the fictional target's configuration and the underlying foundation model's behavior at the time of testing; subsequent model-provider updates could change guardrail behavior without any change on the client's part. This report does not constitute a certification of compliance with any regulatory framework, industry standard, or contractual security requirement.

20 · IMPORTANT DISCLAIMER
Illustrative Sample — Not an Actual Client AI/LLM Penetration Test

NimbusMind AI, Inc. and the NimbusMind Copilot platform are fictional. They do not correspond to any real business or product, and any resemblance to an actual organization is purely coincidental. All tenant/account identifiers, API keys, session tokens, and model endpoint URLs referenced throughout this sample (including api.nimbusmind-example.com and similar) are fictional, non-resolving placeholders used only for illustration.

No real AI model, real customer environment, real production system, or real third-party model provider was tested, accessed, or queried in the creation of this document. No real user, personal information, payment data, or customer records were accessed, viewed, or processed. No real credentials — including API keys, session tokens, or account passwords — were used or disclosed anywhere in this report. All prompts, model responses, system-prompt excerpts, tool-call traces, and API request/response examples shown in this report are synthetic and constructed for illustration only.

All vulnerabilities, findings, risk ratings, CVSS-style scores, dates, and remediation statuses described in this report are fictional and were constructed for demonstration purposes. All attack scenarios and proofs-of-concept described in this report are illustrative, synthetic, and non-destructive; none were executed against a real or live AI system, and no real data was exfiltrated, modified, or exposed. No real exploitation, jailbreaking, or prompt injection occurred at any point in the preparation of this document, and no actual customer or business impact occurred.

TMG Security is not affiliated with, endorsed by, reviewed by, or certified by OWASP, any foundation-model provider, or any other standards organization unless explicitly stated otherwise. References to OWASP GenAI/LLM terminology describe conceptual alignment only and do not imply certification. This report does not represent an actual TMG Security client engagement and must not be relied upon as evidence of the security posture of any real organization, AI product, or model deployment.

21 · FREQUENTLY ASKED QUESTIONS

Frequently asked questions

What is an AI / LLM security penetration testing report?+

An AI / LLM Security Penetration Testing report is the formal deliverable produced after an authorized, hands-on security assessment of a conversational AI product built on a foundation-model API. It documents scope and methodology, prompt-injection and jailbreak testing, RAG and vector-store security, agent tool-use authorization testing, detailed findings with severity ratings and evidence, risk and business-impact analysis, and remediation guidance — giving AI platform and security leadership an evidence-based roadmap for closing the gaps found.

What does an AI / LLM penetration test cover?+

An AI/LLM penetration test covers the full conversational AI attack surface: prompt injection and jailbreak resistance across direct, indirect (RAG-sourced), and multimodal input vectors; system-prompt and instruction leakage; sensitive data disclosure and PII handling; vector-store and embedding security; excessive agency and autonomous action control; agent tool-use and function-calling API authorization; model and training-data poisoning; output handling of AI-generated content; SSRF via AI-driven URL-fetch features; rate limiting and resource-exhaustion abuse; memory and context-window attacks; supply-chain and model/plugin dependency security; and AI-specific logging, monitoring, privacy, and authentication controls — 18 testing domains in this sample.

What is the OWASP Top 10 for LLM Applications and how does it relate to this report?+

The OWASP Top 10 for LLM Applications is an industry-discussed list of common security risks in large-language-model applications, covering areas such as prompt injection, insecure output handling, training-data poisoning, and excessive agency. This sample report's testing domains and terminology are informed by, but not certified against, that list and the related OWASP GenAI Security Project terminology; TMG Security is not affiliated with, endorsed by, or certified by OWASP.

What is indirect prompt injection?+

Indirect prompt injection occurs when untrusted content ingested through a retrieval-augmented generation (RAG) knowledge base, a retrieved document, or an uploaded image contains hidden instructions that the model treats as trusted context, rather than the attacker directly typing a malicious prompt themselves. This sample report's AILLM-002 finding illustrates a poisoned help-center article used exactly this way.

What is excessive agency in an AI agent context?+

Excessive agency describes an AI agent being granted more autonomous capability — such as the ability to take irreversible or financial actions — than is appropriately governed by confirmation steps, human review, or authorization checks. This sample's AILLM-004 finding shows a natural-language request triggering an unconfirmed destructive account action.

How is agent tool-use / function-calling security tested?+

TMG Security tests every tool exposed to the agent's function-calling registry individually for authorization, parameter-validation, and object-level access-control weaknesses, then tests the agent's multi-step planning layer for cross-tool composition risk — where chaining several individually-safe tools together reaches a privilege no single tool call would allow, as shown in this sample's AILLM-007 finding.

Why isn't every prompt-injection finding rated Critical?+

TMG Security's fictional rating approach weighs whether a technique generalizes across many prompts and sessions versus a narrow condition, whether impact is contained to model output (content-level) or reaches a real tool or data action (system-level), and whether existing detective controls would likely catch active exploitation. A content-level jailbreak with no reachable tool impact is generally rated lower than a comparable technique that reaches a tool with real data or account impact.

What is RAG content trust and why does it matter?+

RAG content trust refers to whether a system treats content retrieved from a knowledge base or vector store as untrusted input requiring the same scrutiny as direct user input, or incorrectly treats it as trusted model context. Systems that skip this distinction are exposed to indirect prompt injection, as documented in this sample's AILLM-002 and D02 testing domain.

Is this an actual client penetration testing report?+

No. This is a fictional sample/demonstration report and does not represent an actual client engagement. NimbusMind AI, Inc. and the NimbusMind Copilot platform are fictional, and no real client environment, model, or backend was assessed.

22 · VIEW FULL SAMPLE PDF

View the full sample PDF

The complete 109-page illustrative report, including the full methodology, AI/LLM architecture and attack-surface overview, all 45 detailed findings with evidence panels, the attack chain analysis, business impact analysis, remediation roadmap, and every appendix — test case register, endpoint/tool register, finding register, evidence register, AI/LLM security control matrix, attack path register, remediation priority matrix, and test accounts/assumed roles register.

TMG_Security_AI_LLM_Penetration_Testing_Sample_Report_FINAL.pdf
AI / LLM SECURITY PENETRATION TESTING REPORT · OWASP TOP 10 FOR LLM APPLICATIONS · NimbusMind AI, Inc. · 109 Pages · Illustrative Sample
23 · REQUEST A SIMILAR ASSESSMENT

Request a similar assessment

Looking to validate your own AI or LLM application's security directly? TMG Security's AI/LLM Security Testing team can scope a real AI / LLM Security Penetration Testing engagement built with the same structure, methodology, and reporting depth shown in this sample — prompt injection and jailbreak testing, RAG and vector-store security, agent tool-use authorization testing, and a full sample findings report.