AI / LLM Security Penetration Testing Report — Sample
NimbusMind AI, Inc. — AI/LLM Application Security Assessment. A fictional AI/LLM application penetration test of NimbusMind Copilot, an AI-powered customer-support and knowledge-assistant platform operated by NimbusMind AI, Inc., showing how TMG Security structures scope, methodology, prompt-injection and jailbreak testing, agent/tool-use security, RAG and vector-store testing, findings, risk analysis, and remediation guidance for an AI/LLM assessment informed by the OWASP Top 10 for LLM Applications.
NimbusMind AI, Inc. and the NimbusMind Copilot AI platform are fictional. No real model, backend, agent tool, or client environment was built, deployed, accessed, or tested, and no real customer data was involved in the creation of this document. TMG Security is not affiliated with, endorsed by, or certified by OWASP or any foundation-model provider.
AI LLM Security Penetration Testing — Overview
An AI/LLM application penetration test is an authorized, hands-on security assessment that combines unauthenticated black-box testing of public-facing chat surfaces, authenticated grey-box testing of the production chat and developer-support assistant variants, and assumed-breach testing from a low-privilege compromised account to identify exploitable weaknesses across the AI/LLM-specific attack surface. This illustrative sample shows how TMG Security structures that deliverable for a conversational AI product built on a third-party foundation-model API with a retrieval-augmented generation (RAG) knowledge base, an agentic function-calling tool layer, and multimodal image-input support.
This sample report demonstrates TMG Security's approach for a fictional organization, NimbusMind AI, Inc., and its fictional product, NimbusMind Copilot — an AI-powered customer-support and knowledge-assistant platform. The assessment was informed by, but not certified against, the OWASP Top 10 for LLM Applications and the OWASP GenAI Security Project's terminology, combined with hands-on testing specific to the target's agentic, RAG-backed architecture. This sample illustrates the depth of a typical TMG Security AI / LLM Security Penetration Testing engagement.
What This AI LLM Security Penetration Testing Sample Demonstrates
This sample AI / LLM Security Penetration Testing report is a fictional, illustrative deliverable prepared for an OWASP Top 10 for LLM Applications-informed assessment — not to describe any real organization's security posture.
- How TMG defines scope, rules of engagement, and testing objectives for a conversational AI product and its agentic tool layer
- How prompt injection and jailbreak resistance are tested across direct, indirect (RAG-sourced), and multimodal input vectors
- How the RAG retrieval service, vector store, and knowledge-base content trust boundary are assessed
- How the agent's exposed function-calling tool registry is tested for authorization, parameter-validation, and excessive-agency gaps
- How sensitive data disclosure, PII handling, and cross-session context leakage are evaluated end to end
- How model/API abuse, rate limiting, logging, and AI-specific detection controls are assessed as defense-in-depth
- How individual findings are cross-referenced into realistic, non-destructive composite attack chains
None of these samples is a certification, an attestation, an actual client engagement, or proof of any compliance or security status. This report is labeled as illustrative throughout and carries a full disclaimer on this page and within the sample PDF.
Application & platform profile
The fictional NimbusMind Copilot platform is built on a third-party foundation-model API (fictional "ModelArc LLM Platform") and is architected as a layered conversational AI system: client surfaces (an authenticated web chat, an unauthenticated pre-signup marketing widget, a developer-support assistant variant, and an internal staff-facing ticket console) call into an API gateway and agent orchestration/planning layer, which coordinates a guardrail classifier, a RAG retrieval service backed by a shared vector store, and a 14-tool function-calling registry, plus a multimodal vision-extraction pipeline for image uploads. Downstream, these components read from and write to a multi-tenant customer/ticket data store, object storage for uploads and transcripts, a fine-tuning and feedback pipeline, and a SIEM/guardrail log store.
Testing scope & rules of engagement
In scope
- The Chat Orchestration Service — primary authenticated chat API and system-prompt/guardrail pipeline
- The Knowledge Retrieval (RAG) Service — vector store, embedding pipeline, and retrieval-ranking logic
- The Agent Tool Registry — 14 function-calling tools exposed to the assistant
- The Multimodal Ingestion Pipeline — image-upload endpoint, vision-extraction, and OCR-adjacent processing
- The Feedback & Fine-Tuning Pipeline — thumbs-up/down and free-text feedback ingestion and model-promotion workflow
- The unauthenticated Marketing-Site Chat Widget and the Developer-Support Assistant variant
- The Internal Agent Console, identity & session layer, and guardrail logging/SIEM integration
Out of scope
- The underlying foundation-model provider's own infrastructure and training pipeline (treated as a trusted third party)
- Physical security and any real-world office or data-center facilities
- Denial-of-service testing beyond bounded, non-destructive rate-limiting checks
- Social engineering of real personnel (no real personnel exist in this fictional scenario)
Testing was grey-box (authenticated low-privilege test accounts provided) plus black-box testing of unauthenticated surfaces, over the fictional window of 02 September 2026 – 24 September 2026 (fictional business hours). Destructive testing was not permitted; all account-action and data-modification tests used dedicated fictional test accounts only, and model-interaction volume was bounded and rate-aware. A 30-day retest window is included in scope at no additional fictional cost.
AI LLM Security Penetration Testing Methodology & Framework Context
TMG Security's fictional AI / LLM Security Penetration Testing methodology combines industry-discussed practice for generative-AI security assessment — informed by, but not certified against, the OWASP Top 10 for LLM Applications, the OWASP GenAI Security Project's terminology, and the MITRE ATLAS adversarial-ML knowledge base — with hands-on testing specific to the target's agentic, RAG-backed architecture. Testing proceeded through six phases: reconnaissance & attack-surface mapping; prompt-level & guardrail testing; authenticated application & tool-layer testing; assumed-breach / compromised-session testing; platform, supply-chain & governance review; and attack-chain construction & reporting.
TMG Security is not affiliated with, endorsed by, reviewed by, or certified by OWASP, MITRE, any foundation-model provider, or any other standards organization. References to industry frameworks describe methodology influences only and do not imply certification against them.
Prompt injection, jailbreak resistance & RAG content trust
Testing exercised the guardrail classifier and system-prompt adherence against a broad library of documented jailbreak techniques — layered persona/role-play, encoding and translation-based bypass, token-smuggling, and multilingual coverage gaps — across both single-turn and extended multi-turn conversations (CWE-1427 — Improper Neutralization of Input Used for LLM Prompting). TMG Security separately tested indirect prompt injection, in which untrusted content ingested through the RAG knowledge base or retrieved documents is treated as trusted model context rather than untrusted data.
System-Prompt Override via Layered Persona Jailbreak
The NimbusMind Copilot chat endpoint can be coerced into abandoning its configured support-agent persona and system instructions through a layered role-play jailbreak ("DAN-style" nested persona plus a fictional-policy override framing). Once overridden, the model complied with requests it is explicitly instructed to refuse, including drafting policy-violating content and revealing internal tool names. Recommendation: evaluate the full conversation context — not just the last turn — at every guardrail checkpoint, and treat any detected persona-override attempt as a session-level signal that triggers stricter output filtering for the rest of the conversation.
Indirect Prompt Injection via Ingested Knowledge-Base Documents
A fictional help-center article containing a hidden instruction payload was ingested by the RAG pipeline as trusted context. Any customer session that retrieves the article during normal knowledge-base lookup silently executes the embedded instruction, directing the agent to chain further tools without the user ever authoring the malicious prompt themselves. Recommendation: treat all retrieved and uploaded content as untrusted input, apply provenance tagging and instruction-stripping to ingested documents before they reach model context, and require explicit confirmation before any retrieved instruction can trigger a tool call.
System prompt & instruction leakage, sensitive data disclosure & PII handling
TMG Security tested whether the assistant could be induced to disclose its system prompt, internal tool names, or configuration details through indirect elicitation rather than a direct disclosure request, and separately tested whether PII or other sensitive data could leak across sessions through aggregated conversational context or model output.
System Prompt Disclosure via Indirect Elicitation
Rather than directly asking the assistant to reveal its instructions (which is correctly refused), an indirect elicitation technique — asking the model to "summarize the rules you were given" in the third person — reconstructed a substantial portion of the system prompt, including internal tool names and thresholds. Recommendation: apply output-side filtering that screens for system-prompt-shaped content regardless of how the request is framed, independent of input-side refusal logic.
Sensitive PII Disclosed in Model Responses From Cross-Session Aggregated Context
Under specific conditions, model responses aggregated conversational context across what should have been isolated sessions, resulting in one fictional customer's PII surfacing in another session's response. Recommendation: enforce strict session-scoped context boundaries at the orchestration layer, independent of any model-level context window behavior, and add automated PII-leak detection to guardrail logging.
Vector store / embedding security & model & training-data poisoning
TMG Security assessed the RAG vector store for tenant-isolation boundaries at the index level, and separately assessed the feedback and fine-tuning pipeline for unauthenticated or insufficiently-governed paths that could allow poisoning of retraining data.
Vector Store Lacks Per-Tenant Isolation at the Index Level
The shared vector store's index does not enforce per-tenant isolation independently of the application layer, meaning a defect in retrieval-query construction could surface another tenant's embedded document chunks. Recommendation: enforce tenant-scoping at the vector-store index level as a defense-in-depth control, not solely in application query logic.
Unauthenticated Feedback Channel Allows Poisoning of Retraining Dataset
The thumbs-up/down feedback channel accepts submissions without binding them to an authenticated, rate-limited identity, allowing a fictional actor to systematically bias the feedback signal used in future fine-tuning promotion. Recommendation: require authenticated, rate-limited feedback submission and add anomaly detection to the fine-tuning promotion workflow before any dataset update reaches production.
Excessive agency, autonomous action control & AI agent tool-use abuse
Using fictional low-privilege test identities, TMG Security tested every tool in the exposed agent registry for authorization, parameter-validation, and cross-tool composition weaknesses, and tested account-action tools specifically for excessive-agency and confirmation gaps.
Excessive Agency Allows Unconfirmed Destructive Account Actions via Natural-Language Request
A natural-language request phrased as an ordinary support ask can trigger an irreversible account action — such as a subscription cancellation — with no explicit confirmation step, no human-in-the-loop gate, and no distinguishable difference in agent behavior between a routine query and a destructive command. Recommendation: require an explicit, out-of-band confirmation step for any tool call with irreversible or financial impact, regardless of how confidently the model interprets user intent.
Agent Tool-Chaining Allows Privilege Escalation Beyond Any Single Tool's Scope
No individual tool in the 14-tool registry is over-privileged in isolation, but the agent's planning layer can chain multiple tools together to reach an effective privilege beyond what any single tool call would allow, including reading data one tool alone could not reach. Recommendation: treat the agent's multi-step planning layer itself as a security boundary, with composition-aware authorization checks rather than per-tool checks alone.
Insecure tool / function-calling API access, BOLA & SSRF via AI features
TMG Security tested the mobile-facing and agent-facing tool APIs directly for broken object-level authorization, and tested any tool capable of fetching a URL on the model's behalf for server-side request forgery risk.
Broken Object-Level Authorization in Ticket-Lookup Tool Enables Cross-Tenant Data Access
The agent's ticket-lookup tool accepts a ticket identifier supplied through natural-language conversation without independently verifying that the requesting session's tenant owns that ticket, enabling full cross-tenant ticket data disclosure — including PII and billing references — through ordinary conversational requests. Recommendation: enforce tenant-scoped authorization at the tool-implementation layer for every parameter that references another object, independent of how the parameter value was derived (typed input vs. model-generated).
SSRF via Unrestricted URL-Fetch Tool Used for 'Summarize This Link' Feature
The "summarize this link" feature's URL-fetch tool performs no destination validation, allowing a fictional link-local or internal-network URL to be fetched and its response summarized back to the requester. Recommendation: apply an allowlist and destination-validation layer (blocking link-local and private-network ranges) to any tool capable of making an outbound request on the model's behalf.
Model / API abuse, rate limiting, resource exhaustion & memory / context-window attacks
TMG Security tested per-session and per-account rate limiting on model-interaction volume, and tested whether concurrent session load could cause cross-user context to bleed within the model's active context window.
Missing Per-Session Rate Limiting Enables Model-Abuse and Cost-Amplification
No per-session rate limit governs model-interaction volume on the authenticated chat endpoint, allowing a single fictional account to drive materially higher inference cost than any legitimate usage pattern would produce, with no automated cost-anomaly alert. Recommendation: apply per-session and per-account rate limiting and cost-anomaly detection tuned to the AI feature's expected usage envelope.
Cross-User Context Bleed Under Concurrent Session Load
Under sustained concurrent session load, TMG Security observed conditions under which one fictional user's conversational context could bleed into another's active session. Recommendation: enforce hard per-session context isolation at the orchestration layer, with load and concurrency testing specifically targeting the context-assembly code path.
Multimodal input risks & output handling / insecure AI-generated content
TMG Security tested the multimodal image-upload pipeline for hidden instruction text embedded in images, and tested whether AI-generated output is safely handled by downstream systems that render it, such as the internal ticket console.
Prompt Injection via Text Embedded in Uploaded Image (Multimodal Bypass)
The vision-extraction pipeline passes text recovered from an uploaded image into the model's context as trusted input, allowing a screenshot-styled image containing hidden instruction text to bypass text-based guardrail filtering entirely. Recommendation: apply the same untrusted-content treatment (provenance tagging, instruction-stripping) to text extracted from images as is applied to retrieved documents.
Unsafe Rendering of AI-Generated Markdown Enables Stored Content Injection in Ticket Notes
AI-generated markdown summaries are rendered into the internal ticket console without output sanitization, allowing a crafted conversational input to result in stored content injection visible to staff reviewing the ticket. Recommendation: treat all AI-generated content as untrusted output requiring the same sanitization applied to any other user-supplied content before rendering.
Supply chain & model/plugin dependency security, logging & AI-specific detection
TMG Security reviewed AI-stack dependency versions and advisories, and assessed whether guardrail-refusal and injection-attempt logging would support timely detection of a sustained attack pattern rather than only individual events.
Unpinned Third-Party Prompt-Template Library With Known Advisory in Production
A third-party prompt-template library used by the orchestration layer is deployed unpinned in production and currently resolves to a version with a published security advisory. Recommendation: pin all AI-stack dependencies to specific, reviewed versions and integrate advisory scanning into the deployment pipeline for prompt-related libraries specifically, not only general application dependencies.
No Alerting on Repeated Jailbreak or Injection Attempts From the Same Account
Individual guardrail refusals are logged, but no aggregated alerting exists for a single account generating dozens of jailbreak or injection variants in sequence — the exact pattern an iterative attacker would use. Recommendation: add attempt-pattern alerting on repeated guardrail refusals from a single identity within a bounded time window, escalating to session-level restriction.
Privacy, data governance & retention, authentication & session for AI endpoints
TMG Security reviewed data-retention practice for chat transcripts against documented policy, and tested session and authentication controls specific to the AI chat surfaces, including token lifecycle behavior around password reset.
Chat Transcripts Retained Indefinitely With No Documented Retention Policy Enforcement
Chat transcripts containing customer conversational data are retained indefinitely with no automated enforcement of any documented retention window. Recommendation: define and technically enforce a retention schedule for conversational data consistent with the platform's privacy documentation.
Session Tokens for Chat Widget Do Not Expire on Password Reset
Existing chat-widget session tokens remain valid after a fictional user resets their password, rather than being invalidated as part of the reset flow. Recommendation: invalidate all active chat-session tokens for an account whenever a password reset or credential change occurs.
AI LLM Security Penetration Testing — Findings Summary
This fictional assessment identified 45 illustrative findings across NimbusMind Copilot's eighteen AI/LLM testing domains, drawn from 137 fictional test cases. Every finding, severity rating, and evidence reference below is fictional and specific to this AI / LLM Security Penetration Testing engagement.
| ID | Title | Severity |
|---|---|---|
| AILLM-001 | System-Prompt Override via Layered Persona Jailbreak | Critical |
| AILLM-002 | Indirect Prompt Injection via Ingested Knowledge-Base Documents | Critical |
| AILLM-003 | Broken Object-Level Authorization in Ticket-Lookup Tool Enables Cross-Tenant Data Access | Critical |
| AILLM-004 | Excessive Agency Allows Unconfirmed Destructive Account Actions via Natural-Language Request | Critical |
| AILLM-005 | System Prompt Disclosure via Indirect Elicitation | High |
| AILLM-006 | Sensitive PII Disclosed in Model Responses From Cross-Session Aggregated Context | High |
| AILLM-007 | Agent Tool-Chaining Allows Privilege Escalation Beyond Any Single Tool's Scope | High |
| AILLM-008 | Prompt Injection via Text Embedded in Uploaded Image (Multimodal Bypass) | High |
| AILLM-009 | SSRF via Unrestricted URL-Fetch Tool Used for 'Summarize This Link' Feature | High |
| AILLM-010 | Unauthenticated Feedback Channel Allows Poisoning of Retraining Dataset | High |
| AILLM-011 | Missing Per-Session Rate Limiting Enables Model-Abuse and Cost-Amplification | High |
| AILLM-012 | Cross-User Context Bleed Under Concurrent Session Load | High |
| AILLM-013 | Unpinned Third-Party Prompt-Template Library With Known Advisory in Production | High |
| AILLM-014 | No Alerting on Repeated Jailbreak or Injection Attempts From the Same Account | High |
| AILLM-015 | Vector Store Lacks Per-Tenant Isolation at the Index Level | Medium |
| AILLM-016 | Unsafe Rendering of AI-Generated Markdown Enables Stored Content Injection in Ticket Notes | Medium |
| AILLM-017 | Language-Switching Bypass of Content-Policy Filters | Medium |
| AILLM-018 | RAG Source Documents Lack Freshness/Trust Scoring, Allowing Stale or Superseded Guidance to Surface | Medium |
| AILLM-019 | Session Tokens for Chat Widget Do Not Expire on Password Reset | Medium |
| AILLM-020 | Tool Parameter Schema Accepts Overly Broad Free-Text Where a Constrained Enum Should Be Used | Medium |
| AILLM-021 | Debug/Verbose Mode Response Header Discloses Model Version and Internal Pipeline Stage Timings | Medium |
| AILLM-022 | Chat Transcripts Retained Indefinitely With No Documented Retention Policy Enforcement | Medium |
| AILLM-023 | Uploaded File Type Validation Relies on Client-Supplied MIME Type Only | Medium |
| AILLM-024 | Agent Retries Failed Tool Calls With Escalating Privilege Rather Than Fixed Scope | Medium |
| AILLM-025 | No Cost/Usage Anomaly Detection for Individual Tenant Accounts | Medium |
| AILLM-026 | Error Messages Reveal Internal Function/Tool Names on Malformed Tool-Call Failures | Medium |
| AILLM-027 | No Human Review Gate Before Fine-Tuning Dataset Promotion to Production | Medium |
| AILLM-028 | AI-Generated Code Snippets in Developer-Support Responses Not Flagged as Unverified | Medium |
| AILLM-029 | Incomplete Audit Trail for Tool Calls — Parameters Logged, Response Bodies Truncated | Medium |
| AILLM-030 | Assistant Discloses Its Own Model-Family Name When Directly Asked | Low |
| AILLM-031 | No CAPTCHA or Bot-Mitigation on Unauthenticated Marketing-Site Chat Widget | Low |
| AILLM-032 | Verbose Client-Side Console Logging of Assembled Prompt Context in Non-Production Build | Low |
| AILLM-033 | Retrieved Document Snippets Shown to User Lack a Clear Visual Distinction From Model-Generated Text | Low |
| AILLM-034 | Chat API Accepts Requests With Expired-But-Not-Yet-Purged Bearer Tokens for a Short Grace Window | Low |
| AILLM-035 | Uploaded Images Retain EXIF Metadata Through the Ingestion Pipeline | Low |
| AILLM-036 | Vector-Database Client Library Two Minor Versions Behind Latest, No Known Advisory | Low |
| AILLM-037 | No User-Facing Log of Autonomous Actions Taken by the Assistant on the User's Behalf | Low |
| AILLM-038 | Guardrail Refusal Reason Codes Are Inconsistent Across Services, Complicating Aggregated Reporting | Low |
| AILLM-039 | Observation: Data-Processing Addendum Language Has Not Been Updated to Reference the AI Feature Set | Informational |
| AILLM-040 | Observation: Refusal Messages Are Not Localized, Defaulting to English Regardless of Conversation Language | Informational |
| AILLM-041 | Observation: Tool Registry Documentation Is Current and Matches Deployed Tool Set | Informational |
| AILLM-042 | Observation: Guardrail-Refusal Events Are Consistently Timestamped and Immutable in Storage | Informational |
| AILLM-043 | Observation: Financial Identifiers Are Consistently Masked in the Chat UI Layer | Informational |
| AILLM-044 | Observation: MFA Is Enforced for All Internal Support-Console Administrator Accounts | Informational |
| AILLM-045 | Observation: Embedding Model Version Is Pinned and Consistently Applied Across the Corpus | Informational |
Attack chain analysis
Individual findings were cross-referenced to construct four realistic, non-destructive fictional attack chains, demonstrating how lower-friction weaknesses combine into materially higher business impact than any single finding suggests in isolation. All four chains below are fictional, illustrative, and non-destructive; no real system, model, or data was accessed, modified, or exploited in their construction or validation.
AC-01 — Public Content → Session Hijack → Data Exfiltration
Step 1: Attacker publishes a poisoned help-center article containing a hidden instruction payload (AILLM-002). Step 2: Any customer session that retrieves the article during normal RAG lookup silently executes the embedded instruction. Step 3: The embedded instruction directs the agent to chain the search_customer, get_ticket_details, and send_email tools (AILLM-007) to exfiltrate the victim's own session/ticket data to an attacker-controlled address. Step 4: Because tool-response logging is truncated (AILLM-029), TMG's fictional incident responders cannot fully reconstruct what data left the platform without the underlying fix.
Business impact: A single content-publishing capability, combined with under-governed agent tool-chaining and incomplete audit logging, escalates to session-level data exfiltration against any customer who merely asks a question that retrieves the poisoned article.
AC-02 — Jailbreak → Cross-Tenant BOLA → Full Ticket Disclosure
Step 1: Attacker establishes a persona-override jailbreak against the support chatbot (AILLM-001), suspending the assistant's normal refusal behavior for the remainder of the session. Step 2: Within the jailbroken session, the attacker requests ticket lookups using guessed or sequential ticket IDs belonging to other fictional tenants. Step 3: The get_ticket_details tool's missing tenant-scoping check (AILLM-003) returns full cross-tenant ticket content, including PII and billing references, with no additional authorization barrier.
Business impact: Combining a behavioral jailbreak with a tool-layer authorization gap allows an attacker to move from "convince the chatbot to misbehave" directly to "read any customer's support history" in a single session.
AC-03 — Crafted Image Upload → Multimodal Injection → Excessive-Agency Account Action
Step 1: Attacker uploads a screenshot-styled image containing hidden instruction text, exploiting the multimodal injection gap (AILLM-008). Step 2: The vision-extraction pipeline passes the hidden instruction into the model's context as trusted input. Step 3: The injected instruction phrases a subscription-cancellation request in a way the excessive-agency gap (AILLM-004) executes immediately with no confirmation step.
Business impact: This chain demonstrates that the multimodal injection surface is not merely a content-disclosure risk: combined with under-governed autonomous tools, it can trigger real, irreversible account and billing actions with no user awareness at the time of upload.
AC-04 — Undetected Iterative Jailbreak Attempts → System-Prompt Reconstruction → Targeted Tool Abuse
Step 1: Attacker iterates through dozens of jailbreak and elicitation variants against a single account with no aggregated alerting to flag the pattern (AILLM-014). Step 2: Using indirect elicitation techniques, the attacker gradually reconstructs the system prompt, learning internal tool names and thresholds (AILLM-005). Step 3: Armed with the exact tool name and parameter structure, the attacker crafts a malformed update_ticket_priority call that bypasses weak parameter validation (AILLM-020) to disrupt ticket routing.
Business impact: This chain illustrates how a purely detective control gap (no alerting on repeated attempts) removes the operational visibility that would otherwise have interrupted the reconnaissance-to-exploitation path well before the final tool-abuse step.
Business impact analysis
The findings in this assessment, individually and in combination, translate into concrete fictional business risk across five categories, supporting prioritized investment decisions by NimbusMind AI, Inc.'s fictional leadership.
| Impact Category | Priority | Fictional Business Narrative |
|---|---|---|
| Data Breach / Privacy | Critical | Cross-tenant and session-level data disclosure findings (AILLM-002, AILLM-003, AILLM-006) would, in a real deployment, constitute a reportable breach affecting every tenant, with attendant regulatory notification and contractual liability exposure. |
| Brand & Customer Trust | High | A publicly jailbreakable customer-facing chatbot (AILLM-001, AILLM-016) risks screenshot-driven reputational damage and erodes customer confidence in the AI product. |
| Financial / Revenue Assurance | High | Unconfirmed destructive account actions (AILLM-004, AILLM-011, AILLM-025) directly threaten recurring revenue; uncontrolled inference cost and usage abuse threaten gross margin on AI features. |
| Regulatory / Compliance | Medium | Retention-policy drift and incomplete AI-specific data-processing documentation (AILLM-022, AILLM-039) create audit and regulatory exposure, particularly for a platform processing customer conversational data at scale. |
| Operational / Detection Readiness | Medium | Gaps in attempt-pattern alerting and audit-log completeness (AILLM-014, AILLM-029) would materially slow incident detection and breach-scope determination. |
Remediation roadmap
The fictional sample remediation roadmap sequences all 45 findings across four phases, prioritizing the four Critical findings and the highest-risk High findings first. All dates and timelines shown are illustrative.
0–14 days
- System-prompt override via layered persona jailbreak (AILLM-001)
- Indirect prompt injection via ingested knowledge-base documents (AILLM-002)
- Broken object-level authorization in ticket-lookup tool (AILLM-003)
- Excessive agency allowing unconfirmed destructive account actions (AILLM-004)
15–30 days
- System prompt disclosure via indirect elicitation (AILLM-005)
- Sensitive PII disclosure from cross-session context (AILLM-006)
- Agent tool-chaining privilege escalation (AILLM-007)
- Multimodal prompt injection via uploaded image (AILLM-008)
- SSRF via unrestricted URL-fetch tool (AILLM-009)
- Unauthenticated feedback channel poisoning risk (AILLM-010)
- Missing per-session rate limiting (AILLM-011)
- Cross-user context bleed under concurrent load (AILLM-012)
- Unpinned third-party prompt-template library (AILLM-013)
- No alerting on repeated jailbreak attempts (AILLM-014)
31–90 days
- All 15 Medium findings (AILLM-015 through AILLM-029), covering vector-store isolation, output sanitization, session-token lifecycle, and tool parameter validation
- All 9 Low findings (AILLM-030 through AILLM-038), covering hygiene, logging consistency, and dependency freshness
Continuous / informational backlog
- 7 informational observations (AILLM-039 through AILLM-045), including 5 confirmed positive controls tracked for continued monitoring
Positive security controls & testing limitations
Consistent with TMG Security's standard reporting practice, this section documents fictional controls confirmed operating effectively — an important counterweight to the findings sections, since an AI security report that only lists problems gives an incomplete picture of NimbusMind AI, Inc.'s actual fictional posture. Positive observations, recorded in full as Informational findings AILLM-041 through AILLM-045, include: tool registry documentation that is current and matches the deployed tool set; guardrail-refusal events that are consistently timestamped and immutable in storage; financial identifiers that are consistently masked in the chat UI layer; MFA enforced for all internal support-console administrator accounts; and an embedding model version that is pinned and consistently applied across the corpus.
As with any time-boxed security assessment, this fictional engagement is subject to inherent limitations. AI/LLM behavior is probabilistic — a technique that succeeded during fictional testing may not succeed on every attempt, and a technique that failed may succeed under different phrasing, timing, or model-version conditions not covered in this assessment window. Testing was time-boxed to the assessment period stated in Section 04 and does not guarantee the absence of additional exploitable weaknesses outside the tested scope, techniques, or time window. Findings reflect the fictional target's configuration and the underlying foundation model's behavior at the time of testing; subsequent model-provider updates could change guardrail behavior without any change on the client's part. This report does not constitute a certification of compliance with any regulatory framework, industry standard, or contractual security requirement.
NimbusMind AI, Inc. and the NimbusMind Copilot platform are fictional. They do not correspond to any real business or product, and any resemblance to an actual organization is purely coincidental. All tenant/account identifiers, API keys, session tokens, and model endpoint URLs referenced throughout this sample (including api.nimbusmind-example.com and similar) are fictional, non-resolving placeholders used only for illustration.
No real AI model, real customer environment, real production system, or real third-party model provider was tested, accessed, or queried in the creation of this document. No real user, personal information, payment data, or customer records were accessed, viewed, or processed. No real credentials — including API keys, session tokens, or account passwords — were used or disclosed anywhere in this report. All prompts, model responses, system-prompt excerpts, tool-call traces, and API request/response examples shown in this report are synthetic and constructed for illustration only.
All vulnerabilities, findings, risk ratings, CVSS-style scores, dates, and remediation statuses described in this report are fictional and were constructed for demonstration purposes. All attack scenarios and proofs-of-concept described in this report are illustrative, synthetic, and non-destructive; none were executed against a real or live AI system, and no real data was exfiltrated, modified, or exposed. No real exploitation, jailbreaking, or prompt injection occurred at any point in the preparation of this document, and no actual customer or business impact occurred.
TMG Security is not affiliated with, endorsed by, reviewed by, or certified by OWASP, any foundation-model provider, or any other standards organization unless explicitly stated otherwise. References to OWASP GenAI/LLM terminology describe conceptual alignment only and do not imply certification. This report does not represent an actual TMG Security client engagement and must not be relied upon as evidence of the security posture of any real organization, AI product, or model deployment.
Frequently asked questions
What is an AI / LLM security penetration testing report?+
An AI / LLM Security Penetration Testing report is the formal deliverable produced after an authorized, hands-on security assessment of a conversational AI product built on a foundation-model API. It documents scope and methodology, prompt-injection and jailbreak testing, RAG and vector-store security, agent tool-use authorization testing, detailed findings with severity ratings and evidence, risk and business-impact analysis, and remediation guidance — giving AI platform and security leadership an evidence-based roadmap for closing the gaps found.
What does an AI / LLM penetration test cover?+
An AI/LLM penetration test covers the full conversational AI attack surface: prompt injection and jailbreak resistance across direct, indirect (RAG-sourced), and multimodal input vectors; system-prompt and instruction leakage; sensitive data disclosure and PII handling; vector-store and embedding security; excessive agency and autonomous action control; agent tool-use and function-calling API authorization; model and training-data poisoning; output handling of AI-generated content; SSRF via AI-driven URL-fetch features; rate limiting and resource-exhaustion abuse; memory and context-window attacks; supply-chain and model/plugin dependency security; and AI-specific logging, monitoring, privacy, and authentication controls — 18 testing domains in this sample.
What is the OWASP Top 10 for LLM Applications and how does it relate to this report?+
The OWASP Top 10 for LLM Applications is an industry-discussed list of common security risks in large-language-model applications, covering areas such as prompt injection, insecure output handling, training-data poisoning, and excessive agency. This sample report's testing domains and terminology are informed by, but not certified against, that list and the related OWASP GenAI Security Project terminology; TMG Security is not affiliated with, endorsed by, or certified by OWASP.
What is indirect prompt injection?+
Indirect prompt injection occurs when untrusted content ingested through a retrieval-augmented generation (RAG) knowledge base, a retrieved document, or an uploaded image contains hidden instructions that the model treats as trusted context, rather than the attacker directly typing a malicious prompt themselves. This sample report's AILLM-002 finding illustrates a poisoned help-center article used exactly this way.
What is excessive agency in an AI agent context?+
Excessive agency describes an AI agent being granted more autonomous capability — such as the ability to take irreversible or financial actions — than is appropriately governed by confirmation steps, human review, or authorization checks. This sample's AILLM-004 finding shows a natural-language request triggering an unconfirmed destructive account action.
How is agent tool-use / function-calling security tested?+
TMG Security tests every tool exposed to the agent's function-calling registry individually for authorization, parameter-validation, and object-level access-control weaknesses, then tests the agent's multi-step planning layer for cross-tool composition risk — where chaining several individually-safe tools together reaches a privilege no single tool call would allow, as shown in this sample's AILLM-007 finding.
Why isn't every prompt-injection finding rated Critical?+
TMG Security's fictional rating approach weighs whether a technique generalizes across many prompts and sessions versus a narrow condition, whether impact is contained to model output (content-level) or reaches a real tool or data action (system-level), and whether existing detective controls would likely catch active exploitation. A content-level jailbreak with no reachable tool impact is generally rated lower than a comparable technique that reaches a tool with real data or account impact.
What is RAG content trust and why does it matter?+
RAG content trust refers to whether a system treats content retrieved from a knowledge base or vector store as untrusted input requiring the same scrutiny as direct user input, or incorrectly treats it as trusted model context. Systems that skip this distinction are exposed to indirect prompt injection, as documented in this sample's AILLM-002 and D02 testing domain.
Is this an actual client penetration testing report?+
No. This is a fictional sample/demonstration report and does not represent an actual client engagement. NimbusMind AI, Inc. and the NimbusMind Copilot platform are fictional, and no real client environment, model, or backend was assessed.
View the full sample PDF
The complete 109-page illustrative report, including the full methodology, AI/LLM architecture and attack-surface overview, all 45 detailed findings with evidence panels, the attack chain analysis, business impact analysis, remediation roadmap, and every appendix — test case register, endpoint/tool register, finding register, evidence register, AI/LLM security control matrix, attack path register, remediation priority matrix, and test accounts/assumed roles register.
Request a similar assessment
Looking to validate your own AI or LLM application's security directly? TMG Security's AI/LLM Security Testing team can scope a real AI / LLM Security Penetration Testing engagement built with the same structure, methodology, and reporting depth shown in this sample — prompt injection and jailbreak testing, RAG and vector-store security, agent tool-use authorization testing, and a full sample findings report.
