Appearance
12. Layer 7: Observability & Evaluation
Problem Solved
Agent loops take many steps and behave non-deterministically. The final answer alone does not explain why the agent succeeded or failed. Observability rebuilds the execution path so you can see what happened at each step.
Service Provided to the Layer Above
Observability provides the evidence needed to debug, evaluate, cost, and improve agent systems.
Trace Model
An agent run should be represented as a tree of spans.
text
Run
|
+-- Step 1
| +-- Context assembly
| +-- Model call
| +-- Output parse
| +-- Guardrail check
| +-- Tool call
| +-- Observation capture
|
+-- Step 2
| +-- Context assembly
| +-- Model call
| +-- Output parse
| +-- Final answer check
|
+-- Stop reasonEach span records what happened at that stage. The span and attribute conventions below follow the OpenTelemetry GenAI semantic conventions (25), which standardize how LLM applications are traced.
A full trace is often the only way to explain a wrong decision. In one case the task was to generate a workflow that charges an electric car when the electricity price is low, but always keeps the battery above a 100-mile minimum range. The model's final output looked correct: a clean workflow with a sensible condition. The reasoning trace showed what actually happened. The model had reversed the range condition, so the car would start charging only when the range was already above 100 miles, the moment it does not need charging, instead of charging when the range drops below 100 miles. Reading the final answer alone would have hidden the flaw. The reasoning trace exposed it (47).
Useful Span Attributes
| Attribute | Purpose |
|---|---|
| Run ID | Correlates all events from one task |
| Step number | Orders loop iterations |
| Model name | Identifies which engine was used |
| Input token count | Measures context size |
| Output token count | Measures generation size |
| Latency | Measures performance |
| Cost | Measures economic consumption |
| Parsed intent | Shows what the harness understood |
| Validation result | Shows whether the intent was accepted |
| Tool name | Identifies executed capability |
| Tool status | Shows success, failure, timeout, or denial |
| Stop reason | Explains why the run ended |
The same span model applies to multi-step agent runs, where the trace must record each loop iteration and the model and tool calls inside it (26).
Metrics
| Metric category | Examples |
|---|---|
| Correctness | Task success rate, answer accuracy, tool selection accuracy |
| Efficiency | Steps to completion, tokens per task, cost per task |
| Reliability | Crash rate, retry rate, timeout rate, validation failure rate |
| Safety | Guardrail violations, denied actions, policy breaches |
| User impact | Latency, interruption rate, human escalation rate |
Debugging by Layer
When an agent fails, assign the failure to a layer.
| Symptom | Likely layer |
|---|---|
| Model ignores instructions | Context or model capability |
| Invalid tool arguments | Structured output or tool contract |
| Tool called but no useful result | Tool design or observation design |
| Agent repeats same action | Control loop or missing state change |
| Context grows too large | Memory/context budgeting |
| Unsafe action attempted | Guardrails or permissions |
| Cannot explain failure | Observability |
| Cost is too high | Loop control, context budget, tool usage |
Evaluation
Evaluation should happen at multiple levels.
| Level | Question |
|---|---|
| Component evaluation | Does each tool return correct, useful observations? |
| Output evaluation | Does the model produce valid intents reliably? |
| Task evaluation | Does the agent complete the end-to-end goal? |
| Safety evaluation | Does the system refuse or block unsafe actions? |
| Regression evaluation | Did a prompt, model, or tool change break existing behavior? |
| Judge evaluation | Can an LLM grader reliably score outputs (27)? |
Production telemetry and offline evaluation are complementary. Telemetry shows what happens in the field. Evaluation suites show whether changes improve or degrade behavior.
Benchmarks and Testing Frameworks
Frontier labs use a small set of reference benchmarks in their publications. Agent-specific benchmarks measure end-to-end agent behavior; general benchmarks measure the underlying model capabilities.
Agent Benchmarks (used in frontier lab papers)
| Benchmark | Lab | What It Measures | Key Paper |
|---|---|---|---|
| SWE-bench | Princeton NLP / OpenAI / DeepMind | Resolving real GitHub issues; used by OpenAI (o1), DeepMind (Gemini), Anthropic (Claude) | (28) |
| GAIA | Meta AI, EPFL | Real-world web search and information synthesis; used by Anthropic (Claude), Meta (Llama) | (29) |
| AgentBench | Tsinghua / Zhipu AI | LLM-as-agent across OS, browser, DB, KG, shopping, browsing, games | (30) |
| WebArena | UC San Diego | Realistic web browsing task automation; used by OpenAI, Anthropic | (31) |
| OSWorld | OSWorld Team | Desktop OS automation (Windows tasks); used by DeepMind, OpenAI | (32) |
Model Capability Benchmarks (used in frontier lab papers)
| Benchmark | Lab | What It Measures | Key Paper |
|---|---|---|---|
| MMLU | UC Santa Barbara | General knowledge across 57 subjects; used by OpenAI, Anthropic, Google, Meta | (33) |
| GSM8K | OpenAI | Grade-school math word problems; used by OpenAI (o1), Anthropic (Claude), Google (Gemini) | (34) |
| HumanEval | OpenAI | Python code generation; used by OpenAI, Anthropic, Google, Meta | (35) |
| GPQA | David Rein / Google | PhD-level science/math questions; used by OpenAI (o1), DeepMind | (36) |
| IFEval | Google DeepMind | Following instructions; used by Google (Gemini), Anthropic (Claude) | (37) |
| BBH | Google DeepMind | 23 challenging reasoning tasks; used by Google, OpenAI | (38) |
Evaluation Frameworks
| Framework | Purpose |
|---|---|
| LangSmith | Agent tracing, monitoring, evaluation, and deployment; SDKs for Python, TypeScript, Go, Java; OpenTelemetry support |
| Testcontainers | Containerized test dependencies for integration testing agent tool interactions |
| OpenTelemetry | Vendor-neutral observability standard; GenAI semantic conventions (25); integrates with LangSmith, LangFuse, and other platforms |
| pytest / unittest | Standard unit and integration testing frameworks; combine with agent harness for regression testing |
When building an evaluation pipeline, start with component-level tests (tools return correct observations), then task-level tests (agent completes end-to-end goals), and finally safety tests (system refuses or blocks unsafe actions). Use benchmarks as periodic checkpoints, not daily metrics.
Design Rules
- Assign every run a stable identifier.
- Record every loop step.
- Store raw model output separately from parsed intent.
- Record guardrail decisions, not only successful executions.
- Track token and cost metrics per step and per run.
- Make stop reasons explicit.
- Build replayability: from traces, a human should be able to reconstruct the decision path.
- Evaluate failure paths, not only happy paths.
Failure Modes
| Failure | Symptom | Fix |
|---|---|---|
| Final-log-only debugging | Only the final answer is visible | Step-level tracing |
| Missing raw output | Cannot see what model actually said | Preserve raw emissions |
| Uncorrelated logs | Tool logs and model logs cannot be joined | Run IDs and step IDs |
| No cost tracking | Surprises in spend | Token and cost accounting |
| Happy-path evaluation | Breaks appear only in production | Adversarial and regression tests |
| Unexplained stop | Agent stops but reason is unknown | Explicit stop-reason field |