Skip to content

12. Layer 7: Observability & Evaluation ​

Problem Solved ​

Agent loops take many steps and behave non-deterministically. The final answer alone does not explain why the agent succeeded or failed. Observability rebuilds the execution path so you can see what happened at each step.

Service Provided to the Layer Above ​

Observability provides the evidence needed to debug, evaluate, cost, and improve agent systems.

Trace Model ​

An agent run should be represented as a tree of spans.

text
Run
 |
 +-- Step 1
 |    +-- Context assembly
 |    +-- Model call
 |    +-- Output parse
 |    +-- Guardrail check
 |    +-- Tool call
 |    +-- Observation capture
 |
 +-- Step 2
 |    +-- Context assembly
 |    +-- Model call
 |    +-- Output parse
 |    +-- Final answer check
 |
 +-- Stop reason

Each span records what happened at that stage. The span and attribute conventions below follow the OpenTelemetry GenAI semantic conventions (25), which standardize how LLM applications are traced.

A full trace is often the only way to explain a wrong decision. In one case the task was to generate a workflow that charges an electric car when the electricity price is low, but always keeps the battery above a 100-mile minimum range. The model's final output looked correct: a clean workflow with a sensible condition. The reasoning trace showed what actually happened. The model had reversed the range condition, so the car would start charging only when the range was already above 100 miles, the moment it does not need charging, instead of charging when the range drops below 100 miles. Reading the final answer alone would have hidden the flaw. The reasoning trace exposed it (47).

Useful Span Attributes ​

AttributePurpose
Run IDCorrelates all events from one task
Step numberOrders loop iterations
Model nameIdentifies which engine was used
Input token countMeasures context size
Output token countMeasures generation size
LatencyMeasures performance
CostMeasures economic consumption
Parsed intentShows what the harness understood
Validation resultShows whether the intent was accepted
Tool nameIdentifies executed capability
Tool statusShows success, failure, timeout, or denial
Stop reasonExplains why the run ended

The same span model applies to multi-step agent runs, where the trace must record each loop iteration and the model and tool calls inside it (26).

Metrics ​

Metric categoryExamples
CorrectnessTask success rate, answer accuracy, tool selection accuracy
EfficiencySteps to completion, tokens per task, cost per task
ReliabilityCrash rate, retry rate, timeout rate, validation failure rate
SafetyGuardrail violations, denied actions, policy breaches
User impactLatency, interruption rate, human escalation rate

Debugging by Layer ​

When an agent fails, assign the failure to a layer.

SymptomLikely layer
Model ignores instructionsContext or model capability
Invalid tool argumentsStructured output or tool contract
Tool called but no useful resultTool design or observation design
Agent repeats same actionControl loop or missing state change
Context grows too largeMemory/context budgeting
Unsafe action attemptedGuardrails or permissions
Cannot explain failureObservability
Cost is too highLoop control, context budget, tool usage

Evaluation ​

Evaluation should happen at multiple levels.

LevelQuestion
Component evaluationDoes each tool return correct, useful observations?
Output evaluationDoes the model produce valid intents reliably?
Task evaluationDoes the agent complete the end-to-end goal?
Safety evaluationDoes the system refuse or block unsafe actions?
Regression evaluationDid a prompt, model, or tool change break existing behavior?

| Judge evaluation | Can an LLM grader reliably score outputs (27)? |

Production telemetry and offline evaluation are complementary. Telemetry shows what happens in the field. Evaluation suites show whether changes improve or degrade behavior.

Benchmarks and Testing Frameworks ​

Frontier labs use a small set of reference benchmarks in their publications. Agent-specific benchmarks measure end-to-end agent behavior; general benchmarks measure the underlying model capabilities.

Agent Benchmarks (used in frontier lab papers) ​

BenchmarkLabWhat It MeasuresKey Paper
SWE-benchPrinceton NLP / OpenAI / DeepMindResolving real GitHub issues; used by OpenAI (o1), DeepMind (Gemini), Anthropic (Claude)(28)
GAIAMeta AI, EPFLReal-world web search and information synthesis; used by Anthropic (Claude), Meta (Llama)(29)
AgentBenchTsinghua / Zhipu AILLM-as-agent across OS, browser, DB, KG, shopping, browsing, games(30)
WebArenaUC San DiegoRealistic web browsing task automation; used by OpenAI, Anthropic(31)
OSWorldOSWorld TeamDesktop OS automation (Windows tasks); used by DeepMind, OpenAI(32)

Model Capability Benchmarks (used in frontier lab papers) ​

BenchmarkLabWhat It MeasuresKey Paper
MMLUUC Santa BarbaraGeneral knowledge across 57 subjects; used by OpenAI, Anthropic, Google, Meta(33)
GSM8KOpenAIGrade-school math word problems; used by OpenAI (o1), Anthropic (Claude), Google (Gemini)(34)
HumanEvalOpenAIPython code generation; used by OpenAI, Anthropic, Google, Meta(35)
GPQADavid Rein / GooglePhD-level science/math questions; used by OpenAI (o1), DeepMind(36)
IFEvalGoogle DeepMindFollowing instructions; used by Google (Gemini), Anthropic (Claude)(37)
BBHGoogle DeepMind23 challenging reasoning tasks; used by Google, OpenAI(38)

Evaluation Frameworks ​

FrameworkPurpose
LangSmithAgent tracing, monitoring, evaluation, and deployment; SDKs for Python, TypeScript, Go, Java; OpenTelemetry support
TestcontainersContainerized test dependencies for integration testing agent tool interactions
OpenTelemetryVendor-neutral observability standard; GenAI semantic conventions (25); integrates with LangSmith, LangFuse, and other platforms
pytest / unittestStandard unit and integration testing frameworks; combine with agent harness for regression testing

When building an evaluation pipeline, start with component-level tests (tools return correct observations), then task-level tests (agent completes end-to-end goals), and finally safety tests (system refuses or blocks unsafe actions). Use benchmarks as periodic checkpoints, not daily metrics.

Design Rules ​

  1. Assign every run a stable identifier.
  2. Record every loop step.
  3. Store raw model output separately from parsed intent.
  4. Record guardrail decisions, not only successful executions.
  5. Track token and cost metrics per step and per run.
  6. Make stop reasons explicit.
  7. Build replayability: from traces, a human should be able to reconstruct the decision path.
  8. Evaluate failure paths, not only happy paths.

Failure Modes ​

FailureSymptomFix
Final-log-only debuggingOnly the final answer is visibleStep-level tracing
Missing raw outputCannot see what model actually saidPreserve raw emissions
Uncorrelated logsTool logs and model logs cannot be joinedRun IDs and step IDs
No cost trackingSurprises in spendToken and cost accounting
Happy-path evaluationBreaks appear only in productionAdversarial and regression tests
Unexplained stopAgent stops but reason is unknownExplicit stop-reason field