Appearance
11. Layer 6: Guardrails & Safety
Problem Solved
Guardrails put deterministic limits on a probabilistic system that can cause real side effects.
Prompting alone is not a safety architecture. A model may be asked to behave, but the harness must still enforce behavior.
Prompting vs. Guardrails
| Mechanism | Nature | Example |
|---|---|---|
| Prompting | Probabilistic request | “Only use allowed tools.” |
| Guardrail | Deterministic enforcement | Tool allowlist rejects unknown tools |
| Prompting | Probabilistic request | “Do not delete files.” |
| Guardrail | Deterministic enforcement | Filesystem permission denies deletion outside workspace |
| Prompting | Probabilistic request | “Output valid JSON.” |
| Guardrail | Deterministic enforcement | Schema validator rejects invalid output |
Reminding the model through a prompt is probabilistic, and the model can ignore it under a complex context or reasoning drift. Adversarial inputs show this failure at scale (24). The harness must enforce deterministic boundaries (allowlists, schemas, permission scopes), because only those guarantee that safety constraints are actually respected.
Guardrail Layers
Guardrails operate at multiple points along the agent's execution path. Each layer catches different classes of errors and violations.
| Layer | Guardrail examples |
|---|---|
| Input | Task validation, prohibited content filters, allowed resource scopes |
| Output | Schema validation, secret detection, policy checks, format enforcement |
| Execution | Sandboxing, file/network allowlists, non-root privileges, resource limits |
| Tool use | Tool allowlists, argument constraints, permission checks, rate limits |
| Final answer | Verification hooks, evidence requirements, answer format checks |
| Loop control | Iteration caps, cost caps, timeouts, repetition detection |
Fail Closed
A guardrail should fail closed.
If validation is uncertain, the safe behavior is to stop, ask for clarification, or escalate, not to guess and execute.
text
+----------------+ +----------------+ +----------------+
| Model output | --> | Guardrail check| --> | Uncertain? |
+----------------+ +----------------+ +----------------+
| yes | no
v v
+-----------+ +-----------+
| Stop / | | Execute / |
| escalate | | continue |
+-----------+ +-----------+Circuit Breakers
Circuit breakers stop the agent when continuing is more dangerous than stopping.
| Trigger | Action |
|---|---|
| Maximum steps exceeded | Stop and report partial state |
| Cost threshold exceeded | Pause or terminate |
| Repeated validation failure | Stop or request human help |
| Repeated tool error | Back off, then abort |
| Unsafe intent detected | Block action and log violation |
| Timeout exceeded | Cancel execution and preserve state |
Case Pattern: Missing Execution Boundary
Two destructive failures share the same root cause: the harness gave the agent more permission than the task needed.
One agent was asked to reorganize files. Mid-task it deleted an entire folder, with no stated reason. Another was modifying a database schema; when a migration would not run, it decided that dropping the database and starting from scratch was a reasonable fix.
In both cases the model's reasoning sounded plausible, and the action had a real, destructive effect. The fix was not better prompting. It was deterministic permission scoping outside the model:
- A read-only database role, so the agent could inspect but not change the data. Every mutation came back as a suggestion the operator ran by hand.
- A command-scanning allowlist that checks each proposed command for mutation before it runs. This is the pattern in tools such as Tirith, which Hermes Agent uses to block dangerous commands.
Even so, a guardrail that blocks one channel is not enough. After the deletion checks were in place, the same class of agent still destroyed data by overwriting files. Every path that can change state needs a boundary, not just the obvious one.
Architectural lesson:
Safety boundaries must exist outside the model, where they cannot be negotiated away by prompt pressure or reasoning drift. Cover every channel that can change state, and test each one.
Design Rules
- Enforce safety in the harness, not only in the system prompt.
- Use least privilege for every tool.
- Separate read permissions from write permissions.
- Require approval for irreversible actions.
- Validate before execution.
- Log every guardrail decision.
- Test guardrails with adversarial and malformed model outputs.
- Prefer reversible tools when possible.
Failure Modes
| Failure | Symptom | Fix |
|---|---|---|
| Prompt-only safety | Model occasionally violates instructions | Deterministic guardrails |
| Overbroad tool permission | One tool can do too much | Split tools, scope permissions |
| Guardrail bypass | Tool output or retrieved text influences unsafe action (23) | Sanitize data, separate instructions from data |
| Silent guardrail failure | Validation error ignored | Fail closed, alert on validation failure |
| No violation audit | Cannot reconstruct why agent stopped | Log guardrail decisions and stop reasons |
| Trusting final answer | Agent claims success without proof | Final-answer verification |