AI/LLM Application Security Checklist
A practical, pre-deployment checklist for AI and LLM applications — what to verify across inputs, retrieval, tools, output, and high-impact actions before a feature ships.
What this checklist is
Most reported LLM security incidents are not model failures — they are ordinary application security failures in a system that happens to contain a model. This checklist turns that finding into something you can run against a real system: fourteen checks, grouped into seven categories, covering the request path from input through retrieval, tools, output, and logging. It is built for pre-deployment review and security testing of LLM applications, including those with tool or function-calling.
About this checklist. This is TMG Security’s own condensed, practical reading of the checklist published in TMG’s full AI/LLM security analysis — not an official OWASP, NIST, or MITRE checklist, and not a certification. It does not cover multi-agent systems, autonomous-agent risk classification, or memory-poisoning in depth — see “Scope and TMG-ARC” below.
AI/LLM application security checklist
Categories and check wording are condensed directly from TMG’s full AI/LLM security analysis — the “why it matters” column is TMG’s own condensed reasoning drawn from that same analysis, not a quotation from any external standard.
| Category | Check | Why it matters |
|---|---|---|
| Input & Context Mapping | Map every input the model can read, including retrieved and machine-generated content | Every hop in the request path is a trust transition — you cannot secure what you have not mapped. |
| Input & Context Mapping | Label which of those sources are attacker-influenceable | Where anyone could have written the content and the model is permitted to call tools after reading it, that is an untrusted-input-to-privileged-action path. |
| Retrieval / RAG Access Control | Confirm retrieval is filtered by the requesting user’s entitlements at query time | A retrieval layer that cannot answer “which documents is this specific user allowed to match against?” before the query runs does not have document-level authorization. It has search. |
| System Instructions & Disclosure Hygiene | Confirm system instructions contain no credentials, internal endpoints, or business rules that would be damaging if disclosed | Anything placed in reachable context that a user is not entitled to see, the model will summarise and surface — a system prompt is reachable context too. |
| System Instructions & Disclosure Hygiene | Confirm error paths don’t surface stack traces or internal identifiers through the model’s reply | The model does not distinguish an internal error message from any other content it can read — if it can read it, it can surface it. |
| Tool / Function-Calling Authorization | Inventory every tool the model can invoke, in every conversation state | A model can only do what its tools let it do — an incomplete inventory is an incomplete threat model. |
| Tool / Function-Calling Authorization | Review each tool’s permissions independently of the model | A tool that accepts any identifier and returns a full record — rather than resolving the caller’s own record from session context — is a finding the model’s own behaviour will never reveal. |
| Tool / Function-Calling Authorization | Test authorization at the tool and API layer directly, bypassing the model | The model should not be the authorization layer. Authorization must be enforced in the tool layer, where it can actually be tested. |
| Tool / Function-Calling Authorization | Apply least privilege to the tool layer, then re-test with a low-privilege account | Tools are APIs; the same object-level authorization discipline applies, and only re-testing under a low-privilege account proves it holds. |
| Output Handling | Validate and encode model output before it is rendered or executed | Model output is attacker-influenceable content, not trusted markup — treat it like any other untrusted output at render time. |
| High-Impact Action Controls | Require explicit confirmation for irreversible or high-value actions | Confirmation only works when the user can evaluate the specific action and its parameters — a vague plan summary is not confirmation. |
| High-Impact Action Controls | Confirm tool responses return only the fields the model needs, not full objects | A tool that returns a full object when the model needed one field hands more than the request required to the model — and to anything reading its output. |
| High-Impact Action Controls | Confirm conversation memory cannot persist one user’s data into another user’s context | Memory that crosses user boundaries turns a single-session mistake into a standing disclosure. |
| Logging & Identity | Log tool invocations with the requesting identity, not the service identity | Logging the service identity instead of the requesting user’s makes incident reconstruction — and accountability — impossible. |
Rows 11–14 are TMG’s own reorganization of disclosure causes and defensive controls already set out in TMG’s full AI/LLM security analysis — not new findings, statistics, or external facts.
How to use this checklist
The model influences behaviour. The application determines consequence. Testing an LLM feature in isolation produces findings about the model; testing the system produces findings about the business. Run each row against the actual system, not only through the chat interface — the interesting findings are behind the tool boundary, and reaching them requires reading the integration code, not just talking to the model.
Practical testing guidance
Assessments that only exercise the chat surface tend to under-report. Exercise the tool and API layer directly, with a low-privilege account, and confirm authorization and data-handling behaviour independently of what the model says or refuses to do. Where a check depends on identity — logging, memory scoping, tool authorization — confirm which identity actually reaches the downstream system, not which identity the model claims to be acting as.
Scope and TMG-ARC
This checklist covers LLM applications generally, including those with tool or function-calling. It is not a multi-agent security framework, an autonomous-agent risk classification framework, a memory-poisoning framework, a formal certification checklist, or an official OWASP, NIST, or MITRE publication.
For deeper agentic and autonomous-agent risk classification — multi-agent communication, persistent memory, autonomous planning — see TMG’s Agentic AI Risk Classification Framework (TMG-ARC). This checklist and TMG-ARC are independent; neither depends on the other.
Limitations and attribution
- This is TMG Security’s own practical checklist, condensed from TMG’s published AI/LLM security analysis — it is not an official OWASP, NIST, or MITRE checklist, and TMG does not claim their endorsement, affiliation, or authorship of those frameworks.
- This checklist is not a certification. Completing it does not constitute a certified or industry-standard assessment.
- It does not cover authentication (session/login mechanics), model or API configuration, rate limiting or abuse controls, or third-party integration risk — these are not addressed in TMG’s underlying analysis and are not included here.
- Memory is covered only by the single row above (row 13). This checklist does not address memory poisoning, autonomous planning, or multi-agent communication — see “Scope and TMG-ARC.”
- This checklist contains no new statistics, CVE/CWE figures, or external data beyond what TMG’s underlying analysis already states.
References
At a glance

