Skip to content

6. Layer 1: Context & Messages ​

Problem Solved ​

The model engine has no memory. Before every call, the system must assemble the exact information the model should see. Layer 1 manages that assembly.

Context is the model’s working memory for one inference call. It is finite, ordered, and expensive.

Service Provided to the Layer Above ​

Layer 1 provides a bounded, ordered set of context that the model can use to make its next decision.

Core Components ​

ComponentPurpose
System instructionDefines role, constraints, output protocol, and safety boundaries
User taskStates the goal, inputs, and success criteria
Assistant historyRecords prior model decisions and reasoning
Tool observationsRecords results produced by tool execution
Tool declarationsDescribes available tools and their argument contracts
Token budgetTracks how much context can be used before the model call

A simplified context layout:

text
+-----------------------------+
| Context Window              |
|                             |
| [ System ]                 |
|  protected rules            |
|                             |
| [ Task ]                   |
|  goal & objectives          |
|                             |
| [ Tools ]                  |
|  contracts & capabilities   |
|                             |
| [ History ]                |
|  prior steps                |
|                             |
| [ Observation ]            |
|  latest env feedback        |
+-----------------------------+

The exact API format varies, but the architectural roles are stable.

Context Ordering ​

Models do not pay equal attention to every position in a long context. Information near the start and near the end is usually more noticeable than information buried in the middle (10).

Context regionRecommended content
BeginningSystem rules, safety constraints, output protocol
MiddleSummarized history, retrieved background, large reference material
EndLatest task state, most recent observation, immediate decision request

This does not mean middle content is useless. It means critical constraints should not depend on being remembered from the middle of a very long context.

Token Budgeting ​

Every agent needs a context budget.

The cost of a large context is not just money; it is latency. Running a model locally, on a laptop or a small GPU, makes it a delay you can feel: one question can take minutes to answer, because every token in the context is re-processed on each call. Under that constraint it is common to compact the session, that is, summarize the conversation, often instead of waiting. This was routine in 2023, when context windows were 4k to 8k tokens and you combined chunking, one-sentence summaries, and retrieval to fit the window. It was the 1980s of context: programming with a few kilobytes of RAM.

text
Total context budget
    - system instruction
    - task description
    - tool declarations
    - conversation / step history
    - retrieved memory
    - latest observations
    = reserve space for model output

If the budget is exceeded, the harness must choose a context operation:

OperationUse whenExample
TruncateOld details are no longer neededA conversation has 50 back-and-forth turns about debugging, but only the last 5 matter for the current decision.
SummarizeHistory is long but trends and decisions matter (11)After 20 steps of data analysis, replace the full log with: "Analyzed 3 datasets; trend: declining accuracy after step 12."
RetrieveOnly relevant external facts should be injectedThe model needs a specific API spec from a 100-page document (retrieve only the relevant section instead of the full text).
Compress observationsTool output is large but only a summary is neededThe model asked to read a text file larger than the context window; keep only the first 2000 and last 2000 characters.
Drop artifactsLarge objects should be referenced by state keys, not inlinedA generated CSV file (5MB) is stored externally; only the file path is kept in context.

The Retrieve operation is the context-side counterpart of retrieval-augmented generation (RAG), which injects externally retrieved documents into the prompt to ground the model's output (12).

Design Rules ​

  1. Protect system instructions from accidental truncation.
  2. Keep the latest observation close to the decision point.
  3. State the goal and success criteria explicitly.
  4. Budget context before calling the model, not after failure.
  5. Prefer summaries plus references over dumping all history.
  6. Make contradictions visible instead of silently accumulating conflicting statements.

Failure Modes ​

FailureSymptomFix
Context overflowAPI rejection or dropped informationBudgeting, summarization, truncation
Lost in the middleModel ignores constraints or evidenceReorder context, repeat critical constraints, retrieve key facts
Stale stateModel acts on outdated informationRefresh observations, timestamp state, summarize latest state
Observation bloatTool output consumes most of the contextCompress, paginate, or store large results externally
Contradictory historyModel receives conflicting instructionsResolve conflicts explicitly, mark superseded state
Missing task frameModel optimizes for plausible text rather than the goalAdd explicit success criteria and stop conditions