Appearance
8. Layer 3: Tools & Actions
Problem Solved
Tools give the agent access to information and actions outside the model’s parameters.
Without tools, the system can only generate text. With tools, the system can search, compute, read, write, call APIs, run simulations, or change external state.
Service Provided to the Layer Above
Layer 3 gives the control loop controlled capabilities and returns observations that can be added to context.
Tool Contract
Every tool should have an explicit contract.
| Contract element | Purpose |
|---|---|
| Name | Stable identifier used by the model |
| Description | Tells the model when the tool is useful |
| Parameters | Defines required and optional inputs |
| Permissions | States what resources the tool may touch |
| Side-effect class | Indicates whether the tool reads, computes, mutates, or communicates |
| Output format | Defines what observation the harness will return |
| Failure behavior | Defines timeouts, retries, and error observations |
A tool is more than a callable function, because it sits at a policy boundary where the harness enforces permissions, validates arguments, checks risk class, and decides whether the requested action is safe to execute (16)(17).
Tool Risk Classes
| Risk class | Examples | Required controls | Unintended consequence |
|---|---|---|---|
| Read-only | Search, fetch record, list files | Input validation, rate limits | Extra cost and latency from too many queries (e.g., several searches before finding the right data). |
| Computation | Calculate, transform data, run pure simulation | Resource limits, timeout | Resource exhaustion or infinite loops from unbounded computation. |
| Mutation | Update database, edit file, create object | Sandbox, audit, rollback where possible | Data corruption or inconsistent state from partial or overlapping writes. |
| External communication | Send message, post request, publish event | Approval gates, allowlists, idempotency | Spam, wrong recipients, or duplicate messages if idempotency is missing. |
| Irreversible action | Delete, pay, deploy, terminate | Human approval or strong deterministic policy | Permanent data loss, financial loss, or outage from targeting the wrong object. |
Execution Environment
Tool execution should occur in an environment whose permissions are known.
| Environment property | Why it matters |
|---|---|
| Isolation | Prevents one agent task from damaging the host or other tasks |
| Permission scope | Limits file, network, database, and API access |
| Resource limits | Prevents runaway CPU, memory, time, or cost |
| Auditability | Records what was executed and what it touched |
| Rollback strategy | Supports recovery when a mutation is wrong |
The exact mechanism may be a local sandbox, container, remote executor, or managed service (18). The architectural requirement is bounded side effect.
Observations
An observation is the result of an action, returned to the agent’s context.
Good observations are:
- Structured
- Concise
- Actionable
- Explicit about success or failure
- Safe to include in context
- Enough to make the next decision
Bad observations are:
- Huge raw dumps with no summary
- Empty success messages with no useful state
- Errors that do not explain what constraint was violated
- Secrets or credentials leaked into context
- Partial results presented as complete
Observation design is part of agent design. A tool that returns useless observations forces the model to guess.
Design Rules
- Prefer few precise tools over many vague tools.
- Write tool descriptions for the model, not just for humans.
- Make read-only tools the default during exploration.
- Separate mutation tools from read tools.
- Return errors as observations, not only as logs.
- Store large artifacts outside context and reference them by state keys.
- Make irreversible actions require stronger approval.
- Enforce tool permissions in the harness, not in the prompt.
Failure Modes
| Failure | Symptom | Fix |
|---|---|---|
| Tool hallucination | Model requests a tool that does not exist | Tool allowlist and validation |
| Wrong arguments | Valid tool called with bad values | Argument schemas and policy checks |
| Observation bloat | Context fills with raw tool output | Summarization, pagination, state keys |
| Unsafe side effect | Tool mutates or deletes unexpected state | Sandboxing, permission scopes, approval gates |
| Timeout | Tool takes too long | Execution limits and cancellation |
| Partial mutation | Some changes succeed, others fail | Transactions, idempotency, rollback |
| Ambiguous failure | Model cannot recover from error | Structured error observations with next-step hints |