Skip to content

11. Layer 6: Guardrails & Safety ​

Problem Solved ​

Guardrails put deterministic limits on a probabilistic system that can cause real side effects.

Prompting alone is not a safety architecture. A model may be asked to behave, but the harness must still enforce behavior.

Prompting vs. Guardrails ​

MechanismNatureExample
PromptingProbabilistic request“Only use allowed tools.”
GuardrailDeterministic enforcementTool allowlist rejects unknown tools
PromptingProbabilistic request“Do not delete files.”
GuardrailDeterministic enforcementFilesystem permission denies deletion outside workspace
PromptingProbabilistic request“Output valid JSON.”
GuardrailDeterministic enforcementSchema validator rejects invalid output

Reminding the model through a prompt is probabilistic, and the model can ignore it under a complex context or reasoning drift. Adversarial inputs show this failure at scale (24). The harness must enforce deterministic boundaries (allowlists, schemas, permission scopes), because only those guarantee that safety constraints are actually respected.

Guardrail Layers ​

Guardrails operate at multiple points along the agent's execution path. Each layer catches different classes of errors and violations.

LayerGuardrail examples
InputTask validation, prohibited content filters, allowed resource scopes
OutputSchema validation, secret detection, policy checks, format enforcement
ExecutionSandboxing, file/network allowlists, non-root privileges, resource limits
Tool useTool allowlists, argument constraints, permission checks, rate limits
Final answerVerification hooks, evidence requirements, answer format checks
Loop controlIteration caps, cost caps, timeouts, repetition detection

Fail Closed ​

A guardrail should fail closed.

If validation is uncertain, the safe behavior is to stop, ask for clarification, or escalate, not to guess and execute.

text
+----------------+     +----------------+     +----------------+
| Model output   | --> | Guardrail check| --> | Uncertain?     |
+----------------+     +----------------+     +----------------+
                                                | yes        | no
                                                v            v
                                          +-----------+  +-----------+
                                          | Stop /    |  | Execute / |
                                          | escalate  |  | continue  |
                                          +-----------+  +-----------+

Circuit Breakers ​

Circuit breakers stop the agent when continuing is more dangerous than stopping.

TriggerAction
Maximum steps exceededStop and report partial state
Cost threshold exceededPause or terminate
Repeated validation failureStop or request human help
Repeated tool errorBack off, then abort
Unsafe intent detectedBlock action and log violation
Timeout exceededCancel execution and preserve state

Case Pattern: Missing Execution Boundary ​

Two destructive failures share the same root cause: the harness gave the agent more permission than the task needed.

One agent was asked to reorganize files. Mid-task it deleted an entire folder, with no stated reason. Another was modifying a database schema; when a migration would not run, it decided that dropping the database and starting from scratch was a reasonable fix.

In both cases the model's reasoning sounded plausible, and the action had a real, destructive effect. The fix was not better prompting. It was deterministic permission scoping outside the model:

  • A read-only database role, so the agent could inspect but not change the data. Every mutation came back as a suggestion the operator ran by hand.
  • A command-scanning allowlist that checks each proposed command for mutation before it runs. This is the pattern in tools such as Tirith, which Hermes Agent uses to block dangerous commands.

Even so, a guardrail that blocks one channel is not enough. After the deletion checks were in place, the same class of agent still destroyed data by overwriting files. Every path that can change state needs a boundary, not just the obvious one.

Architectural lesson:

Safety boundaries must exist outside the model, where they cannot be negotiated away by prompt pressure or reasoning drift. Cover every channel that can change state, and test each one.

Design Rules ​

  1. Enforce safety in the harness, not only in the system prompt.
  2. Use least privilege for every tool.
  3. Separate read permissions from write permissions.
  4. Require approval for irreversible actions.
  5. Validate before execution.
  6. Log every guardrail decision.
  7. Test guardrails with adversarial and malformed model outputs.
  8. Prefer reversible tools when possible.

Failure Modes ​

FailureSymptomFix
Prompt-only safetyModel occasionally violates instructionsDeterministic guardrails
Overbroad tool permissionOne tool can do too muchSplit tools, scope permissions
Guardrail bypassTool output or retrieved text influences unsafe action (23)Sanitize data, separate instructions from data
Silent guardrail failureValidation error ignoredFail closed, alert on validation failure
No violation auditCannot reconstruct why agent stoppedLog guardrail decisions and stop reasons
Trusting final answerAgent claims success without proofFinal-answer verification