Appearance
16. Production Deployment Checklist
Use this checklist before considering an agent production-ready.
Task and Autonomy
- [ ] The task boundary is explicitly defined.
- [ ] Success criteria are measurable.
- [ ] Out-of-scope actions are listed.
- [ ] Autonomy level is chosen intentionally.
- [ ] Human approval points are defined for risky actions.
Context and Model
- [ ] System instruction defines role, constraints, and output protocol.
- [ ] Task description includes goal and success criteria.
- [ ] Context budget is defined before execution.
- [ ] Summarization or truncation strategy exists.
- [ ] Critical constraints are protected from loss.
- [ ] Model choice is justified for task, latency, cost, and reliability.
Structured Output
- [ ] Output protocol is explicit.
- [ ] Final-answer intent is required.
- [ ] Invalid output is rejected, not executed.
- [ ] Retry limit for malformed output is defined.
- [ ] Raw model output is logged separately from parsed intent.
Tools and Execution
- [ ] Tool set is minimal and task-relevant.
- [ ] Every tool has a clear description and argument contract.
- [ ] Tool allowlist is enforced.
- [ ] Argument validation is enforced.
- [ ] Permissions follow least privilege.
- [ ] Mutation tools are separated from read-only tools.
- [ ] Irreversible actions require approval or strong policy.
- [ ] Execution environment is sandboxed or otherwise bounded.
- [ ] Resource limits are enforced: time, memory, network, cost.
Loop Control
- [ ] Maximum iteration cap is enforced.
- [ ] Token budget is enforced.
- [ ] Cost budget is enforced.
- [ ] Timeout is enforced.
- [ ] Repeated-action detection exists.
- [ ] Stop reason is recorded for every run.
- [ ] Retry policy distinguishes transient errors from policy violations.
Memory and State
- [ ] Current state is explicit.
- [ ] Large artifacts are stored outside context and referenced by keys.
- [ ] Retrieved information is scoped and relevant.
- [ ] Stale state is superseded or expired.
- [ ] Sensitive state is access-controlled.
Safety and Observability
- [ ] Guardrails fail closed.
- [ ] Guardrail decisions are logged.
- [ ] Final answers are verified where possible.
- [ ] Every run has a stable run ID.
- [ ] Every step is traceable.
- [ ] Token, latency, and cost metrics are recorded.
- [ ] Failure paths are tested, not only happy paths.