AI Agent Governance Checklist
50 requirements before an AI agent touches production. Every item your legal, compliance, and security teams need to sign off on.
All four phases, all 50 items, free. We run this checklist on our own production agents. Use it, share it, adapt it.
Phase 1: Governance Design
Before writing a single line of agent code
- 1SOUL.md drafted - all six sections present and non-empty
- 2Principal hierarchy defined - who can instruct the agent, ranked by authority
- 3Tier 1 / Tier 2 / Tier 3 authority breakdown documented with specific examples
- 4Scope boundaries stated in specific, falsifiable terms - not vague role descriptions
- 5Out-of-scope categories explicitly listed - not implied by omission
- 6Escalation triggers defined for: ambiguity, irreversibility, cost threshold, scope boundary, uncertainty
- 7Escalation recipient designated by name or role - testable
- 8Escalation channel confirmed working - test message sent
- 9Escalation timeout behavior defined - what the agent does if no response in X minutes
- 10Memory lifespan policies documented per data category
- 11Session-end clearing rules specified - what must not persist
- 12Memory access controls defined - who can read the agent's persistent memory
- 13Agent owner designated by name - the person accountable for behavior
- 14Steward roles assigned: Clarity, Execution, Narrative, Access, Integrity
- 15Audit trail mechanism defined and implemented
Phase 2: Staging Validation
Prove the governance works before production sees it
- 16Agent tested against explicit out-of-scope request types - correct refusal behavior confirmed
- 17Escalation pathway tested end-to-end: trigger, notification, receipt confirmed
- 18Escalation timeout tested - what happens when the recipient doesn't respond
- 19Memory clearing tested - session end, then confirm nothing persists that should not
- 20Kill switch tested - can be activated in under 5 minutes without agent cooperation
- 21Credential revocation tested - agent loses capability immediately upon revocation
- 22Audit trail reviewed - verify it captures the information an investigation would need
- 23Cost baseline established - know what normal spend looks like for this agent
- 24Volume baseline established - normal actions-per-hour rate documented
- 25Scope monitoring implemented - requests categorized against the scope taxonomy
- 26Alert thresholds set for cost anomaly, volume anomaly, and scope violation rate
- 27Tool invocation logging confirmed working
- 28Principal hierarchy tested with conflicting instructions - correct priority behavior confirmed
- 29Boundary behavior tested - agent handles edge-of-scope requests correctly
- 30External input validation tested - agent handles adversarial and injected content correctly
Phase 3: Production Launch
The paper trail that makes the system auditable
- 31SOUL.md version-controlled and tagged at launch version
- 32System prompt version-controlled and tagged at launch version
- 33Configuration version-controlled - any setting that affects agent authority or scope
- 34Change management process defined - who approves changes to authority, scope, or configuration
- 35Incident response playbook finalized and reviewed by the Integrity steward
- 36Communication templates pre-written for incident scenarios
- 37Legal review completed, if the agent takes actions with external legal implications
- 38Security review completed - Access steward confirmed minimum necessary permissions
- 39Compliance review completed (GDPR, CCPA, industry-specific) where applicable
- 40Rollout plan defined - limited initial deployment before full rollout
Phase 4: Ongoing Operations
Governance is a practice, not a launch artifact
- 41Monitoring dashboards configured and reviewed by the Execution steward
- 42Alert routing configured - who receives which alerts, at what hours
- 43Human review sampling established - 5% of outputs reviewed monthly
- 44Quarterly governance review scheduled: SOUL.md, permissions, scope
- 45Post-incident review template ready before the first incident
- 46Agent behavior drift monitoring - a mechanism to detect gradual scope expansion
- 47Dependency monitoring - alert if tools or APIs the agent depends on change behavior
- 48Model update protocol - procedure for when the underlying LLM version changes
- 49Decommission plan documented - how the agent will be retired when no longer needed
- 50Governance documentation accessible to legal, compliance, and security teams
Why Phase 1 matters most: These first 15 items are the governance foundation. If any are missing at launch, Phases 2-4 cannot compensate. A SOUL.md written after incidents have already occurred is documentation, not governance.
Need this implemented, not just listed?
We design and operate agent governance for organizations that run AI in production. If your team needs the system behind the checklist, talk to us.
Start the conversationOne email. One idea.
Every other Thursday.
Field notes on AI × organizational design. No promotion. No filler. Unsubscribe with a single click whenever it stops earning its place in your inbox.