Appearance
Appendix C: Failure Quick Map
| Symptom | First layer to inspect | What to check |
|---|---|---|
| Agent ignores system rule | Context | Is the rule present, early, and not truncated? |
| Agent invents a tool | Structured output / tools | Is tool allowlist validation enforced? |
| Tool arguments are wrong | Tool contract / validation | Are schemas and argument policies clear? |
| Agent loops | Control loop | Are iteration caps and repetition detection active? |
| Context overflow | Memory / context | Is summarization or truncation triggered early? |
| Unsafe action happens | Guardrails / permissions | Did a deterministic boundary fail? |
| Tool result is useless | Tool design | Does the observation give enough actionable state? |
| Final answer is unsupported | Final-answer verification | Is completion evidence required? |
| Cannot debug run | Observability | Are run IDs, spans, raw outputs, and stop reasons recorded? |
| Cost is too high | Loop / context / tools | Are budgets, context size, and tool calls controlled? |