Skip to content

16. Production Deployment Checklist ​

Use this checklist before considering an agent production-ready.

Task and Autonomy ​

  • [ ] The task boundary is explicitly defined.
  • [ ] Success criteria are measurable.
  • [ ] Out-of-scope actions are listed.
  • [ ] Autonomy level is chosen intentionally.
  • [ ] Human approval points are defined for risky actions.

Context and Model ​

  • [ ] System instruction defines role, constraints, and output protocol.
  • [ ] Task description includes goal and success criteria.
  • [ ] Context budget is defined before execution.
  • [ ] Summarization or truncation strategy exists.
  • [ ] Critical constraints are protected from loss.
  • [ ] Model choice is justified for task, latency, cost, and reliability.

Structured Output ​

  • [ ] Output protocol is explicit.
  • [ ] Final-answer intent is required.
  • [ ] Invalid output is rejected, not executed.
  • [ ] Retry limit for malformed output is defined.
  • [ ] Raw model output is logged separately from parsed intent.

Tools and Execution ​

  • [ ] Tool set is minimal and task-relevant.
  • [ ] Every tool has a clear description and argument contract.
  • [ ] Tool allowlist is enforced.
  • [ ] Argument validation is enforced.
  • [ ] Permissions follow least privilege.
  • [ ] Mutation tools are separated from read-only tools.
  • [ ] Irreversible actions require approval or strong policy.
  • [ ] Execution environment is sandboxed or otherwise bounded.
  • [ ] Resource limits are enforced: time, memory, network, cost.

Loop Control ​

  • [ ] Maximum iteration cap is enforced.
  • [ ] Token budget is enforced.
  • [ ] Cost budget is enforced.
  • [ ] Timeout is enforced.
  • [ ] Repeated-action detection exists.
  • [ ] Stop reason is recorded for every run.
  • [ ] Retry policy distinguishes transient errors from policy violations.

Memory and State ​

  • [ ] Current state is explicit.
  • [ ] Large artifacts are stored outside context and referenced by keys.
  • [ ] Retrieved information is scoped and relevant.
  • [ ] Stale state is superseded or expired.
  • [ ] Sensitive state is access-controlled.

Safety and Observability ​

  • [ ] Guardrails fail closed.
  • [ ] Guardrail decisions are logged.
  • [ ] Final answers are verified where possible.
  • [ ] Every run has a stable run ID.
  • [ ] Every step is traceable.
  • [ ] Token, latency, and cost metrics are recorded.
  • [ ] Failure paths are tested, not only happy paths.