AGENT 07 / 07

Reviewer

The vigilant security officer. It guards sensitive surfaces, enforces human approval gates on high-impact actions, detects prompt injection, and maintains immutable audit trails.

reviewerfathom worker

reviewer · autonomous

01What it does

Ironclad safety & governance

01

Human approval gates

Enforces mandatory operator approval before executing side-effect tools (save_contacts, git_push, bank transfers).

02

Hardware secret scanning

Scans all inputs and outputs for API keys, passwords, and private tokens before persistence.

03

Prompt injection defense

Detects and neutralizes jailbreak patterns in untrusted web content before it reaches LLM models.

04

SSRF & redirect guard

Blocks access to private RFC1918 subnets, metadata endpoints, and DNS rebinding attacks.

05

Verification receipts

Writes typed, immutable JSONL verification logs for complete auditability.

06

Regulatory compliance

Evaluates actions against GDPR, CCPA, and enterprise safety policies.

02How it works

From intercepted action to verified approval

1

Action interception

Interceps proposed worker actions and evaluates them against active policy rules.

2

Security scan

Runs regex secret scanners and prompt-injection detectors on all payloads.

3

Gate evaluation

If a policy boundary is crossed, pauses execution and generates a human approval request.

4

Receipt logging

Records a typed verification receipt to the append-only audit ledger.

5

Handoff or deny

Grants execution clearance upon approval or triggers a safe fail-closed denial.

03Under the hood

The runtime underneath

policy engineFail-closed rule evaluation with configurable risk tiers
vault integrationAES-256-GCM hardware key isolation via Rust ring crate
receipt ledgerAppend-only JSONL receipt ledger keyed by (kind, value)
SSRF filterPer-hop redirect validation blocking loopback and cloud metadata IPs

04Tools

Governance toolset

questionsave_contactsgit_pushmemory_absorb

Enterprise AI autonomy with zero compliance risk.