Skip to content

5. Layer 0: Model Engine ​

Problem Solved ​

The model engine provides reasoning over natural language, text generation, and pattern matching. It turns a context into likely continuations.

Service Provided to the Layer Above ​

Given a context, the model produces text. That text may contain:

  • Direct answers
  • Reasoning traces (4)
  • Requests for tool use
  • Structured data
  • Code-like plans
  • Clarifying questions

The layer above must interpret that text. The model does not interpret its own output into safe system behavior.

Core Properties ​

PropertyArchitectural consequenceExplanation
StatelessThe harness must supply all relevant state on every callThe model holds no memory between calls. Each call is independent, like a function that forgets everything after it returns. The harness must rebuild and resend the full context every time.
ProbabilisticThe same context may produce different outputsThe model samples from a probability distribution, so identical inputs can yield different outputs. Agents must be built to handle this variation.
Bounded contextThe harness must manage token budgetsThe model only sees a fixed window of text. Anything outside this window is invisible, so the harness must decide what to keep, compress, or discard.
Text-nativeAll information must be serialized into contextThe model only understands text. Images, code, structured data, and signals must all be encoded as text before the model can process them. Note: Vision models can process images natively. Plain LLMs cannot (they rely on image descriptions provided).
Non-executingThe model proposes actions; the harness executes themThe model generates text suggestions but cannot run code, call APIs, or modify files. The harness interprets and executes those suggestions safely.
Latency/cost bearingEvery loop iteration consumes time and moneyEach call takes seconds and costs money. Agents must minimize round-trips and avoid unnecessary calls to stay practical.

What the Model Does Not Do ​

The model engine does not, by itself:

  • Remember previous API calls
  • Maintain variables between calls
  • Execute code or query databases
  • Enforce permissions
  • Validate its own output
  • Guarantee termination
  • Know when it is wrong unless the harness checks

The key distinction in this course is between the model and its harness: the model reasons, the harness controls.

Model Failure Modes ​

FailureDescriptionArchitectural response
HallucinationModel invents facts, tools, or results (5)Verification, retrieval, deterministic checks
Format driftModel stops following the required output structureStrict output contracts, parsers, retries
Context neglectModel ignores important information in long contextContext budgeting, salience ordering, summarization
Weak planningModel chooses poor next steps (6)Better task framing, planning prompts, step limits
LoopingModel repeats similar actions (7)Iteration caps, repetition detection, state change checks
OverconfidenceModel claims success without evidence (8)Final-answer verification, observation requirements

Design Rule ​

The model is a reasoning component, not a full runtime. It does not manage state, enforce permissions, or guarantee termination. Those duties belong to the harness that wraps and directs it.

What Is a Token? ​

Text is split into tokens (chunks of characters, typically 1–4 characters each) through a process called tokenization (9). The model operates on tokens, not raw characters or words. A single token may represent a word, part of a word, or a punctuation mark.