AGENT 07 / 07
Reviewer
The vigilant security officer. It guards sensitive surfaces, enforces human approval gates on high-impact actions, detects prompt injection, and maintains immutable audit trails.
reviewer · autonomous
01What it does
Ironclad safety & governance
Human approval gates
Enforces mandatory operator approval before executing side-effect tools (save_contacts, git_push, bank transfers).
Hardware secret scanning
Scans all inputs and outputs for API keys, passwords, and private tokens before persistence.
Prompt injection defense
Detects and neutralizes jailbreak patterns in untrusted web content before it reaches LLM models.
SSRF & redirect guard
Blocks access to private RFC1918 subnets, metadata endpoints, and DNS rebinding attacks.
Verification receipts
Writes typed, immutable JSONL verification logs for complete auditability.
Regulatory compliance
Evaluates actions against GDPR, CCPA, and enterprise safety policies.
02How it works
From intercepted action to verified approval
Action interception
Interceps proposed worker actions and evaluates them against active policy rules.
Security scan
Runs regex secret scanners and prompt-injection detectors on all payloads.
Gate evaluation
If a policy boundary is crossed, pauses execution and generates a human approval request.
Receipt logging
Records a typed verification receipt to the append-only audit ledger.
Handoff or deny
Grants execution clearance upon approval or triggers a safe fail-closed denial.
03Under the hood
The runtime underneath
| policy engine | Fail-closed rule evaluation with configurable risk tiers |
| vault integration | AES-256-GCM hardware key isolation via Rust ring crate |
| receipt ledger | Append-only JSONL receipt ledger keyed by (kind, value) |
| SSRF filter | Per-hop redirect validation blocking loopback and cloud metadata IPs |
04Tools