Skip to content

14. How Performance Is Determined ​

Agent performance is determined by two major factors:

  1. The capability of the LLM engine.
  2. The quality of the harness around it.

Neither one alone is enough.

Model Influence ​

The model influences:

DimensionModel contribution
Reasoning qualityAbility to break tasks into steps and choose the next one
Instruction followingFollowing system rules and the output protocol
Tool selectionChoosing the right tool for the current state
Argument qualityProducing correct and complete tool parameters
Context useAttending to relevant facts and ignoring noise
EfficiencyReaching the answer in few useful steps
CalibrationKnowing when to ask for clarification or abort

A stronger model can make certain failures less frequent, but it does not remove the need for architectural controls. In practice this gap is wide. Small models tend to insist on a sub-objective even when the action keeps failing, retrying the same blocked or failed call. Larger models, the strong ones released since mid 2025, follow instructions far more reliably, but they still occasionally issue a bad command. A stronger model lowers the frequency of these failures; it does not remove the controls that catch the rest.

Harness Influence ​

The harness influences:

DimensionHarness contribution
ReliabilityDeterministic parsing, validation, retries, and stop conditions
SafetyPermissions, sandboxing, guardrails, approval gates
Cost controlToken budgets, iteration caps, tool limits
LatencyContext size, tool performance, retry policy
MemoryState management, retrieval, summarization
DiagnosabilityTracing, logging, evaluation, stop reasons
TerminationGuarantees that the loop ends

A weak harness can make a strong model appear unreliable. A strong harness cannot make a weak model understand what it does not understand, but it can contain errors, expose them, and prevent unsafe side effects.

Interaction Between Model and Harness ​

The following matrix shows how model quality and harness quality combine to determine system behavior.

Model Quality ↓ / Harness Quality →Weak HarnessStrong Harness
Strong ModelCapable but unsafe, costly, or hard to debugReliable agent system
Weak ModelUnpredictable and fragile systemLimited reasoning but bounded, observable, and recoverable

A strong harness cannot make a weak model understand concepts it has not learned, but it can contain errors, expose them through observability, and prevent unsafe side effects (20)(21). Conversely, a strong model with a weak harness will produce unreliable, unsafe, and costly behavior regardless of its reasoning capability.

Common Misconceptions ​

MisconceptionCorrection
“The agent is the model.”The agent is the model plus harness, tools, state, and controls
“Better prompting fixes unsafe behavior.”Prompts influence behavior; guardrails enforce it
“Longer context means better memory.”Context is a finite working buffer, not a memory system
“More tools make the agent smarter.”Too many tools increase selection error and attack surface
“If the final answer looks good, the run was good.”The path must be auditable and the answer verified
“Autonomy is always better.”More autonomy requires more control, not less