The most dangerous AI security review is the one that asks only, “Is the model provider secure?”
That question matters. It is also a small part of the system. An AI workflow can use a well-secured model and still expose customer data through logs, retrieve documents a user should not see, execute an unsafe tool call, leak a credential in an error message, or repeat a payment after a timeout.
The security boundary is the complete path from input to action:
User or trigger -> integration -> orchestration -> retrieved data -> model -> validation -> tool -> downstream system -> logs -> human review
Every arrow is a trust boundary. Every component can fail independently.
This checklist is designed for AI agents, chatbots, retrieval-augmented generation, document workflows, and AI-assisted automations. It is not a certification and does not create legal compliance. Use it to structure a review, assign owners, collect evidence, and decide whether a system is ready for production.
Start with risk, not controls
Most teams copy a security checklist and mark boxes. That is incomplete because the same control has different importance in different systems. A policy chatbot that cites public documents has a smaller blast radius than an agent that can refund orders, send emails, or change bank details.
Classify the deployment across four dimensions:
Dimension Low risk Higher risk
Data Public or non-sensitive Personal, financial, health, confidential, credentials
Action Read, draft, summarise Send, approve, pay, delete, publish, change access
Reach Internal, low volume External, high volume, multi-tenant
Reversibility Easy to review and undo Irreversible, time-sensitive, legally consequential
Controls should increase with potential harm, not with how impressive the model sounds.
Phase 1: ownership and system inventory
1. Name a business owner and a technical owner
The business owner approves purpose, acceptable error, and escalation. The technical owner controls architecture, releases, and incidents. “The AI team” is not an accountable name.
2. Define the permitted purpose
Document what the system may do, for whom, and under which conditions. Purpose limits make later requests easier to reject. A support assistant should not quietly become an employee-screening tool because the underlying model can classify text.
3. Draw the data and action flow
Include every source, processor, storage location, model, tool, queue, log, and human interface. Record direction and data category. If the team cannot draw the system, it cannot credibly secure it.
4. Maintain an AI asset register
Track owner, environment, model and version, prompts, tools, data sources, vendors, deployment date, risk tier, and review date. Shadow agents are configuration drift with a friendly interface.
5. Create an action inventory
List each tool and action separately: search records, read an order, draft email, send email, update address, issue refund. “CRM access” is too broad for threat modelling.
Phase 2: data protection and privacy
6. Minimise data before the model call
Send only fields needed for the task. Redact or tokenise identifiers when possible. Do not attach an entire customer record to make prompt design easier.
7. Classify data at every hop
Labels should survive ingestion, retrieval, model processing, logs, and exports. A source classified confidential does not become public because its content was transformed into an embedding or summary.
8. Establish a lawful and documented basis
Privacy obligations depend on jurisdiction and use. Record purpose, data categories, retention, access, user notices, and required assessments. Obtain qualified advice for personal or regulated data.
9. Enforce retrieval permissions at query time
Filtering documents only during ingestion is not enough. User identity and entitlements must restrict retrieval for every request. Otherwise one shared vector index can become a cross-department data leak.
10. Isolate tenants
Multi-tenant systems need isolation in storage, retrieval, caching, logs, and administrative tools. Test with deliberately similar records across tenants; do not rely on a naming convention.
11. Define retention and deletion by component
The model provider, workflow platform, database, object store, monitoring tool, backups, and support exports may all retain data differently. “Thirty-day retention” is meaningless unless it names the system.
12. Verify deletion
Test that deletion propagates where required. A UI disappearing is not proof that logs, embeddings, backups, or vendor copies were handled correctly.
Phase 3: identity, credentials, and access
13. Use separate identities for people and workloads
Agents should use service identities, not an employee’s account. This improves revocation, attribution, rotation, and least privilege.
14. Grant the smallest tool scope
Separate read, draft, and write permissions. Limit resource types, records, regions, and transaction values where the downstream system supports it.
15. Propagate user authority
An assistant must not gain more access than the person requesting work. For delegated actions, capture who initiated, approved, and executed the operation.
16. Put secrets in a managed store
Never place API keys in prompts, workflow notes, source code, or chat history. Use a secrets manager with environment separation, rotation, and access logs.
17. Separate development, staging, and production
Use different credentials and data. A staging agent should be physically unable to send production email or modify production records.
18. Require strong administrator access
Use SSO, MFA, role-based access, session controls, and periodic access review for orchestration, model, cloud, and monitoring consoles.
Phase 4: prompts, retrieval, and model boundaries
19. Treat all external content as untrusted
User text, web pages, email, documents, tool output, and retrieved knowledge can contain instructions designed to override the workflow. Indirect prompt injection is a data problem as much as a prompt problem.
20. Keep enforcement outside the prompt
A system prompt can guide behaviour; it cannot securely enforce permissions, transaction limits, or privacy. OWASP explicitly warns against relying on system prompts for strict behaviour control. Use deterministic code and downstream access controls.
21. Separate instructions from data
Use structured fields and explicit boundaries. Label retrieved content as evidence, not instruction. This reduces ambiguity, though it does not eliminate injection.
22. Constrain model outputs
Use schemas, enumerations, length limits, and validators. Reject unexpected fields. Free text should never be concatenated directly into SQL, shell commands, HTML, or API parameters.
23. Pin and review model versions
Record the exact model identifier where possible. Run regression tests before changing versions, prompts, retrieval, or tools. A vendor upgrade is a software change.
24. Route by risk and capability
Use smaller or cheaper models for bounded low-risk tasks only after evaluation. Escalate ambiguous or high-impact cases to a stronger model or a person. Routing is a security and cost control.
Phase 5: tool and action safety
25. Allowlist tools and parameters
Expose only required functions. Validate destinations, resource identifiers, amounts, file types, and URLs. Do not give a general HTTP client to an agent when three specific APIs are sufficient.
26. Add human approval for consequential actions
External messages, payments, refunds, deletions, account changes, and publication may require approval based on impact. The reviewer needs evidence and a clear diff, not a vague “approve?” button.
27. Make operations idempotent
Retries must not duplicate emails, invoices, bookings, or payments. Use idempotency keys, transaction references, and reconciliation.
28. Limit rate, spend, and duration
Set per-user and system quotas, recursion limits, tool-call ceilings, timeouts, and budget alerts. OWASP describes unbounded consumption as a distinct LLM application risk.
29. Design for partial failure
If step four fails after step three succeeded, the system needs a known state and recovery path. Compensating actions should be explicit and tested.
30. Provide a kill switch
Operators need to disable a workflow, tool, model, tenant, or action type without waiting for a full deployment. Test the switch during a rehearsal.
Phase 6: output, interface, and human factors
31. Validate output before display or execution
Sanitise HTML, links, files, code, and structured payloads. Treat model output as untrusted even when the input is trusted.
32. Show provenance where decisions depend on evidence
Display source links, document titles, dates, and quoted evidence within copyright and access limits. A citation is not proof by itself; it makes verification possible.
33. Communicate system identity and limits
Users should know when they are interacting with AI and how to reach a person. Certain uses and jurisdictions impose specific transparency duties; obtain legal review.
34. Prevent approval fatigue
If every action requires approval, reviewers click through mechanically. Gate by risk, sample low-risk cases, and make the review screen fast and informative.
Phase 7: logging, monitoring, and detection
35. Log consequential events, not every secret
Record initiator, tool, parameters at a safe level, approval, result, model version, and correlation ID. Redact credentials and sensitive payloads. Logging everything can turn observability into a breach multiplier.
36. Monitor behaviour and cost
Alert on unusual tool frequency, new destinations, permission errors, repeated refusals, token spikes, retrieval anomalies, and high escalation. Cost anomalies can be an early security signal.
37. Make runs traceable end to end
Use correlation IDs across workflow, model, API, and downstream record. Without traceability, teams cannot reconstruct a partial failure.
38. Protect and limit log access
Logs need retention, encryption, role-based access, export controls, and deletion. Support personnel should not browse raw customer prompts by default.
Phase 8: vendors, dependencies, and supply chain
39. Review provider terms and controls
Confirm training use, retention, regions, subprocessors, encryption, incident notice, access controls, and deletion. Marketing language is not a data-processing agreement.
40. Inventory dependencies
Track models, SDKs, workflow nodes, containers, plugins, and data providers. Pin versions where appropriate, scan dependencies, and review community connectors before granting credentials.
41. Plan portability and exit
Export prompts, configurations, datasets, logs, and documentation. Know how credentials are revoked and data is deleted when a vendor or agency relationship ends.
42. Reassess material changes
New models, tools, data sources, regions, user groups, and actions can change the risk classification. Security approval applies to a defined system, not to the word “AI.”
Testing before launch
A checklist without tests creates confidence, not evidence. Build a test set containing normal cases and deliberate failures:
Direct and indirect prompt injection.
Requests outside the user’s permission.
Cross-tenant retrieval attempts.
Malformed and oversized input.
Hostile files and unsafe links.
Tool timeout, rate limit, and partial failure.
Duplicate trigger and replay.
Missing, conflicting, and outdated evidence.
Attempts to exceed transaction or spend limits.
Requests for a person, deletion, or data export.
Record expected behaviour. “The agent should handle it safely” is not testable. Specify refuse, ask for clarification, route to review, use a fallback, or stop.
Composite incident scenario
An accounts-payable workflow reads emailed invoices and proposes ERP entries. A supplier PDF contains hidden text instructing the model to replace bank details and mark the invoice urgent. The model follows the document instruction.
One prompt-level defence might fail. Layered controls prevent harm: retrieved content is treated as data; bank-detail changes are prohibited in this workflow; supplier identity is checked against the ERP; output fields are validated; any mismatch enters a review queue; the service account cannot edit supplier banking; and the event is logged.
That is defence in depth. The model does not need perfect judgement when the architecture refuses an unsafe action.
Incident response for AI workflows
Extend the existing incident process rather than inventing a separate security organisation. The playbook should cover:
1. Detect: alert, user report, audit anomaly, or evaluation failure.
2. Contain: disable the affected tool, credential, workflow, model, or tenant.
3. Preserve: retain relevant traces without spreading sensitive data.
4. Assess: identify data, actions, users, and downstream systems affected.
5. Recover: rotate credentials, revert configuration, reconcile actions, and restore safely.
6. Notify: follow contractual, legal, regulatory, and user obligations.
7. Learn: add regression tests and update controls, documentation, and ownership.
Rehearse at least one scenario before launch. A kill switch that nobody can find is not a control.
Governance and regulatory context in 2026
NIST AI RMF organises work around Govern, Map, Measure, and Manage; its Generative AI Profile addresses risks specific to generative systems. ISO/IEC 42001 specifies an AI management system for establishing and continually improving governance. OWASP GenAI provides application-level threat guidance.
For organisations operating in or serving the EU, the AI Act applies in stages. The European Commission confirmed that provisions entered application over time after the Act entered into force on 1 August 2024. Do not rely on a blog summary to determine obligations; classify the use case with qualified counsel and current official guidance.
Frameworks overlap but are not interchangeable. A security control can support governance. It does not automatically satisfy privacy, sector, employment, consumer, or AI-specific law.
Go-live decision
Do not launch until:
Owners and on-call contacts are named.
The data and action map matches the deployed system.
Non-negotiable permissions are technically enforced.
Test evidence meets approved thresholds.
Monitoring, alerts, runbooks, backup, rollback, and kill switches are tested.
Users and reviewers are trained.
Known residual risks are accepted by an authorised owner.
The next review date and change triggers are scheduled.
Frequently asked questions
What is the biggest security risk in AI automation?
There is no universal single risk. Prompt injection is important, but over-privileged tools, sensitive logs, broken retrieval permissions, unsafe output handling, and missing incident controls can be equally serious. Risk depends on data, actions, reach, and reversibility.
Can prompt engineering prevent prompt injection?
No. Prompt design can reduce some failures but cannot enforce security. Use least-privilege tools, deterministic validation, isolation, approval gates, monitoring, and safe failure paths.
Should an AI agent have write access?
Only when write access creates justified value and the action is tightly scoped, validated, observable, reversible where possible, and approval-gated according to risk. Start read-only or draft-only.
Is self-hosting automatically more secure?
No. It provides control but transfers patching, configuration, backups, monitoring, and incident responsibility to your organisation. Managed services can be safer for teams without mature operations; requirements determine the choice.
What should be logged?
Log enough to attribute and reconstruct consequential actions: initiator, approval, tool, safe parameters, result, timing, model/configuration version, and correlation ID. Redact secrets and minimise sensitive content.
How often should security be reviewed?
Review on a fixed cadence and after material change: new model, tool, data source, user group, action, vendor, region, or incident. High-risk systems need more frequent monitoring and testing.
Does this checklist make an AI system compliant?
No. Compliance depends on jurisdiction, sector, role, data, and use. This checklist supports a structured security review; obtain qualified legal, privacy, security, and domain advice.
Conclusion and next steps
Secure AI automation is ordinary security engineering plus several new failure modes. The practical response is not fear or a longer system prompt. It is clear ownership, minimised data, least privilege, untrusted-input handling, constrained tools, validated outputs, traceability, testing, and rehearsed recovery.
Take one live workflow and complete the inventory first. Then test the highest-impact action. Wizora Studio can help map that data and action path through its AI automation services, but the useful deliverable should be a control plan your own team can inspect and operate.
Evidence pack for the security review
A control is stronger when the team can show evidence. Assemble a compact release pack rather than relying on meeting notes:
Current architecture, data-flow, and action-flow diagrams.
Asset register with owners, versions, vendors, data classes, and review dates.
Role and service-account permission exports.
Secrets inventory showing storage and rotation ownership, never secret values.
Model, prompt, tool, and retrieval configuration versions.
Test plan, test data description, results, unresolved failures, and accepted residual risk.
Screenshots or exports of alerts, budgets, rate limits, and kill switches.
Runbooks for common failures and one completed incident rehearsal.
Vendor terms, data-processing documents, subprocessors, and deletion procedure.
Launch approval, scope, expiry or review date, and rollback criteria.
This pack improves incident response and future change review. It also prevents knowledge from disappearing when the original builder leaves.
Minimum viable controls for a low-risk pilot
Small teams sometimes react to enterprise checklists by doing nothing. A contained pilot still needs a floor: named owners; non-production or minimised data; separate credentials; read-only or draft-only tools; a strict spend and run limit; representative testing; redacted logs; an easy stop mechanism; and human review of every outcome.
That is not sufficient for broad production use. It is enough to ensure the pilot tests the workflow without creating uncontrolled exposure. Expansion should add controls before adding data, users, tools, or autonomy.
Security review questions for executives
Executives do not need to inspect prompt syntax, but they should ask whether the organisation can answer:
1. What is the worst action this system can take with its current permissions?
2. How quickly can that action be detected and stopped?
3. Which sensitive data reaches external processors or logs?
4. Who has authority to accept residual risk?
5. What would make us suspend the system?
These questions translate technical controls into business exposure. If the answer is “the vendor handles it,” responsibility has been outsourced without being understood.
For higher-risk deployments, commission an independent review before launch and after major changes. Independence does not guarantee safety, but it challenges shared assumptions between the sponsor and builder. Scope the review to the real system: identity, retrieval, tools, logs, deployment, and operating procedures, not merely a model endpoint.
Track remediation to closure, assign dates and owners, and retest controls rather than accepting screenshots as proof.
Security also needs usability. Operators bypass controls that make legitimate work impossible, and reviewers approve blindly when the interface hides context. Test controls with the people who will use them under realistic time pressure. Measure how long approval takes, whether evidence is understandable, and whether escalation reaches an available owner. A technically correct gate that creates an unmanageable backlog will be weakened later. Good control design makes the safe path the easiest path, while preserving a clear record of exceptional access and emergency changes.
Sources
NIST, AI RMF Generative AI Profile.
OWASP GenAI, Top 10 for LLM Applications.
ISO, ISO/IEC 42001:2023.
European Commission, Safer and more transparent AI.
ICO, Guidance on AI and data protection.
Review date: 11 August 2026. This article is technical and operational guidance, not legal advice.
AI Automation Security Checklist: 42 Controls Before Production: a practical decision framework
AI Automation Security Checklist: 42 Controls Before Production should be evaluated against the real problem, the intended audience, the systems involved, and the level of human review required. The right approach is the one that makes the workflow more useful and more inspectable, not the one that simply adds another tool or trend to the stack.
Key topics to cover
- AI Automation Security Checklist: 42 Controls Before Production guide
- AI Automation Security Checklist: 42 Controls Before Production best practices
- AI Automation Security Checklist implementation
- AI Automation Security Checklist examples
Use these topics as supporting language only when they answer a real question in the article. Explain the implementation choices in plain language, distinguish a reliable workflow from a prototype, and qualify claims that depend on the project scope, data quality, vendor limits, or operating model.
Questions readers should ask
- how AI Automation Security Checklist: 42 Controls Before Production works
- how to use AI Automation Security Checklist: 42 Controls Before Production
- what are the 42 controls in the AI automation security checklist
Limits, evidence, and next steps
Results depend on the workflow, inputs, integrations, security requirements, and review process. Do not treat this guide as a guarantee of cost, speed, rankings, compliance, or business outcomes. Document the assumptions, define what will be measured, and keep a clear stopping or escalation condition.
If you want to map the topic to a real project, review the relevant Wizora service or contact Wizora Studio with the current process, constraints, and desired outcome.



