Skip to content

9. Layer 4: Control Loop ​

Problem Solved ​

A single model call is not enough for multi-step tasks. Layer 4 provides the deterministic loop that connects model decisions to tool execution and state updates.

This layer is the heart of the agent harness.

Service Provided to the Layer Above ​

The control loop turns a stateless model into a goal-directed process. It decides:

  • When to call the model
  • What context to provide
  • Whether an action may execute
  • How to record observations
  • When to retry
  • When to stop

Harness Responsibilities ​

ResponsibilityDescription
Context assemblyBuilds the model input for each step
Model invocationCalls the LLM engine
Output parsingExtracts the model’s intended action
ValidationChecks structure, permissions, and policy
Tool dispatchExecutes approved actions
Observation captureRecords tool results
State updateUpdates memory, counters, budgets, and artifacts
Termination controlEnforces stop conditions
Error handlingManages retries, fallbacks, and aborts
Telemetry emissionProduces trace events for observability

Core Loop ​

text
+-------------------------------------------------------------+
|                      AGENT LOOP                             |
|                                                             |
|  1. Build context                                           |
|  2. Call model                                              |
|  3. Parse output                                            |
|  4. Validate intent                                         |
|  5. Execute action, if approved                             |
|  6. Capture observation                                     |
|  7. Update state and context                                |
|  8. Check stop condition                                    |
|  9. Continue or stop                                        |
+-------------------------------------------------------------+

The control flow of the loop is deterministic: the same steps (build context, call model, parse, validate, execute, observe, update, check stop) always run in the same order. What is probabilistic is the model's output at each step, so the harness must handle varying intents, tool requests, and stop decisions smoothly.

ReAct loop ​

The ReAct pattern alternates reasoning and acting, as described in "ReAct: Synergizing Reasoning and Acting in Language Models" (19).

text
+---------+       +---------+       +-------------+
| Thought | ----> | Action  | ----> | Observation |
+---------+       +---------+       +-------------+
     ^                                     |
     |                                     |
     +----------------- next step <--------+

Conceptually:

StepMeaning
ThoughtThe model reasons about the current state and next move
ActionThe model requests a tool or final answer
ObservationThe environment returns a result
Next thoughtThe model reasons again using the new observation

The visible “thought” is not the essential part. The essential part is the cycle:

Decide, act, observe, update.

Some systems make reasoning explicit. Others keep it implicit. Architecturally, what matters is that each action is followed by feedback that changes the next decision.

Termination Conditions ​

A production agent must have explicit stop conditions.

Stop conditionPurpose
Final answer emittedTask completed
Maximum iterations reachedPrevents infinite loops
Token budget exhaustedPrevents context overflow
Cost budget exhaustedPrevents runaway spend
Repeated error detectedPrevents useless retry cycles
Timeout reachedPrevents stale or hung execution
Guardrail violationStops unsafe behavior
User cancellationHuman aborts the run

An iteration cap is essential because a probabilistic model can fall into a repeating cycle of similar actions with no natural stopping point (7). Without it, a single misbehaving agent can burn through unlimited tokens, cost, and time.

A small local model showed this directly. A read of a credentials file was blocked by the harness, and the model kept requesting the same denied read instead of finding another way. The guardrail held every time, so nothing leaked, but the run never recovered and the task failed. Repetition detection plus an explicit stop reason turn that silent spin into a bounded, visible failure, and a structured error that suggests an alternative (ask the user for the value) lets a small model recover instead of looping.

Retry Policy ​

Retries must be classified.

Error typeRetry strategy
Malformed model outputReturn structured validation error to model, limited retries
Transient tool failureBackoff and retry if idempotent
Permission denialDo not retry blindly; record and stop or ask for help
Repeated same actionBreak loop or escalate to human
Context overflowSummarize or compress before retry
Fatal policy violationStop immediately

Loop Invariants ​

A well-designed loop always keeps the following true:

  1. No tool executes without validation.
  2. Every executed action produces an observation.
  3. Every observation is recorded in the trace.
  4. Context is updated deterministically after each step.
  5. Budgets only decrease.
  6. Stop conditions are checked every iteration.
  7. The raw model output is preserved for audit.
  8. The parsed intent is separate from the raw output.

Failure Modes ​

FailureSymptomFix
Infinite loopSame step repeatsIteration cap, repetition detection
OscillationAgent alternates between two actionsState-change detection, planner reset
Context exhaustionContext grows until call failsBudgeting, summarization, truncation
Runaway costToo many model or tool callsCost caps, step caps, tool budgets
Silent failureTool fails but agent claims successObservation validation, final-answer checks
Error maskingRetry hides a real policy problemClassify errors, log stop reason
Premature stopAgent answers before task is completeFinal-answer verification, required evidence