Appearance
14. How Performance Is Determined
Agent performance is determined by two major factors:
- The capability of the LLM engine.
- The quality of the harness around it.
Neither one alone is enough.
Model Influence
The model influences:
| Dimension | Model contribution |
|---|---|
| Reasoning quality | Ability to break tasks into steps and choose the next one |
| Instruction following | Following system rules and the output protocol |
| Tool selection | Choosing the right tool for the current state |
| Argument quality | Producing correct and complete tool parameters |
| Context use | Attending to relevant facts and ignoring noise |
| Efficiency | Reaching the answer in few useful steps |
| Calibration | Knowing when to ask for clarification or abort |
A stronger model can make certain failures less frequent, but it does not remove the need for architectural controls. In practice this gap is wide. Small models tend to insist on a sub-objective even when the action keeps failing, retrying the same blocked or failed call. Larger models, the strong ones released since mid 2025, follow instructions far more reliably, but they still occasionally issue a bad command. A stronger model lowers the frequency of these failures; it does not remove the controls that catch the rest.
Harness Influence
The harness influences:
| Dimension | Harness contribution |
|---|---|
| Reliability | Deterministic parsing, validation, retries, and stop conditions |
| Safety | Permissions, sandboxing, guardrails, approval gates |
| Cost control | Token budgets, iteration caps, tool limits |
| Latency | Context size, tool performance, retry policy |
| Memory | State management, retrieval, summarization |
| Diagnosability | Tracing, logging, evaluation, stop reasons |
| Termination | Guarantees that the loop ends |
A weak harness can make a strong model appear unreliable. A strong harness cannot make a weak model understand what it does not understand, but it can contain errors, expose them, and prevent unsafe side effects.
Interaction Between Model and Harness
The following matrix shows how model quality and harness quality combine to determine system behavior.
| Model Quality ↓ / Harness Quality → | Weak Harness | Strong Harness |
|---|---|---|
| Strong Model | Capable but unsafe, costly, or hard to debug | Reliable agent system |
| Weak Model | Unpredictable and fragile system | Limited reasoning but bounded, observable, and recoverable |
A strong harness cannot make a weak model understand concepts it has not learned, but it can contain errors, expose them through observability, and prevent unsafe side effects (20)(21). Conversely, a strong model with a weak harness will produce unreliable, unsafe, and costly behavior regardless of its reasoning capability.
Common Misconceptions
| Misconception | Correction |
|---|---|
| “The agent is the model.” | The agent is the model plus harness, tools, state, and controls |
| “Better prompting fixes unsafe behavior.” | Prompts influence behavior; guardrails enforce it |
| “Longer context means better memory.” | Context is a finite working buffer, not a memory system |
| “More tools make the agent smarter.” | Too many tools increase selection error and attack surface |
| “If the final answer looks good, the run was good.” | The path must be auditable and the answer verified |
| “Autonomy is always better.” | More autonomy requires more control, not less |