Appearance
5. Layer 0: Model Engine
Problem Solved
The model engine provides reasoning over natural language, text generation, and pattern matching. It turns a context into likely continuations.
Service Provided to the Layer Above
Given a context, the model produces text. That text may contain:
- Direct answers
- Reasoning traces (4)
- Requests for tool use
- Structured data
- Code-like plans
- Clarifying questions
The layer above must interpret that text. The model does not interpret its own output into safe system behavior.
Core Properties
| Property | Architectural consequence | Explanation |
|---|---|---|
| Stateless | The harness must supply all relevant state on every call | The model holds no memory between calls. Each call is independent, like a function that forgets everything after it returns. The harness must rebuild and resend the full context every time. |
| Probabilistic | The same context may produce different outputs | The model samples from a probability distribution, so identical inputs can yield different outputs. Agents must be built to handle this variation. |
| Bounded context | The harness must manage token budgets | The model only sees a fixed window of text. Anything outside this window is invisible, so the harness must decide what to keep, compress, or discard. |
| Text-native | All information must be serialized into context | The model only understands text. Images, code, structured data, and signals must all be encoded as text before the model can process them. Note: Vision models can process images natively. Plain LLMs cannot (they rely on image descriptions provided). |
| Non-executing | The model proposes actions; the harness executes them | The model generates text suggestions but cannot run code, call APIs, or modify files. The harness interprets and executes those suggestions safely. |
| Latency/cost bearing | Every loop iteration consumes time and money | Each call takes seconds and costs money. Agents must minimize round-trips and avoid unnecessary calls to stay practical. |
What the Model Does Not Do
The model engine does not, by itself:
- Remember previous API calls
- Maintain variables between calls
- Execute code or query databases
- Enforce permissions
- Validate its own output
- Guarantee termination
- Know when it is wrong unless the harness checks
The key distinction in this course is between the model and its harness: the model reasons, the harness controls.
Model Failure Modes
| Failure | Description | Architectural response |
|---|---|---|
| Hallucination | Model invents facts, tools, or results (5) | Verification, retrieval, deterministic checks |
| Format drift | Model stops following the required output structure | Strict output contracts, parsers, retries |
| Context neglect | Model ignores important information in long context | Context budgeting, salience ordering, summarization |
| Weak planning | Model chooses poor next steps (6) | Better task framing, planning prompts, step limits |
| Looping | Model repeats similar actions (7) | Iteration caps, repetition detection, state change checks |
| Overconfidence | Model claims success without evidence (8) | Final-answer verification, observation requirements |
Design Rule
The model is a reasoning component, not a full runtime. It does not manage state, enforce permissions, or guarantee termination. Those duties belong to the harness that wraps and directs it.
What Is a Token?
Text is split into tokens (chunks of characters, typically 1–4 characters each) through a process called tokenization (9). The model operates on tokens, not raw characters or words. A single token may represent a word, part of a word, or a punctuation mark.