Architecture

Architecture

Fathom is a modular Rust runtime for universal autonomous workers, with a typed event protocol, asynchronous agent loop on tokio, TUI, HTTP, sessions, plugins, MCP, and LSP integration. This page covers the crate map, lifecycle, subagent tree, context management and transport boundaries.

Crate Map

The workspace layout. core depends on nothing; agent ties llm, tools, memory and persistence together; server and tui are the two outer shells.

workspace layout
fathom/
├── crates/
│   ├── core/          # fundamental types & domain logic
│   ├── llm/           # LLM provider abstraction
│   ├── agent/         # agent runtime, coordination, control plane
│   ├── tools/         # 48 base tools plus up to 5 CDP browser tools and up to 6 computer tools + MCP/LSP tools
│   ├── memory/        # long-term semantic memory + entity graph
│   ├── mcp/           # Model Context Protocol (client and server)
│   ├── persistence/   # storage (SQLite, PostgreSQL, connection pool, jobs, vault)
│   ├── server/        # Axum HTTP/SSE API, AG-UI bridge, computer relay
│   ├── tui/           # ratatui terminal interface
│   ├── lsp/           # Language Server Protocol integration
│   ├── governance/    # policy engine & audit trail
│   ├── supervisor/    # Docker container supervisor
│   └── desktop/       # GPUI native desktop integration
├── Bot/               # Fathom Bot — local-first React 19 + Electron chat client & Node harness
└── src/main.rs        # CLI entry point
dependency graph
core  ←──  llm  ←──  agent  ←──  server
↑          ↑          ↑           ↑
└── tools ─┴── mcp ───┴── tui ────┘
     ↑          └── persistence ──┘
     └── memory ──┘  (depends on core + llm)
crates/core

The foundation — depends on nothing. IDs (UUID v7), VirtualUri VFS resolver (skill://, rule://, memory://, artifact://, local://, xd://), messages & events, findings, AppConfig, export (PDF/HTML/JSON/DOCX), notifications, CRM sync, exact token counting.

crates/llm

Provider abstraction. LlmProvider trait with complete() and stream(), native Anthropic Messages API with prompt caching & extended thinking, CascadeProvider failover, retry with exponential backoff.

crates/agent

The runtime itself: agent loop, coordinator, compaction, prompt builder, tool executor with parallel safety, result budgets, control plane, IPC, process manager, doom-loop detector, hooks, session resume.

crates/tools

60 base tools plus up to 5 CDP browser tools, up to 6 computer tools, and optional LSP/MCP tools: high-precision hashline patching (edit), AST repo-mapping, git worktrees, compiler checks, DAP debugging, web, browser stealth, memory and coordination.

crates/memory

Long-term semantic memory + entity graph (mem0 / Memora model): SIMD-accelerated HNSW vector search, RDF Triples & GraphRAG traversal, facts, FTS5, embeddings, 5-outcome absorb pipeline, hybrid search, distillation, secret detection, and 2-way Obsidian vault sync.

crates/mcp

Model Context Protocol client and server with bounded connection lifecycle, reconnect, circuit breaker, refresh, and stdio/HTTP transports.

crates/persistence

SQLite in WAL mode with a 4-connection pool + PostgreSQL: sessions, agents, messages, findings, sub-tasks, contact DB, credentials vault, and durable jobs with self-healing retry.

crates/server

Axum HTTP API: REST endpoints for sessions, agents, jobs, memory, credentials, coworkers, schedules, governance, inbound webhooks (/webhooks/inbound), multiplexed WebSocket (/ws), OpenAPI 3.1 (/openapi.json); control-plane answers/approvals; SSE streaming; API-key auth, rate limiting, Prometheus metrics.

crates/tui

Ratatui interface: multi-agent tree view, PTY terminal streams, live streaming buffer, jobs and memory panels, operator control — answer questions, grant approvals, watch the fleet.

crates/lsp

Language Server Protocol client and adapter tool for IDE code intelligence and cross-file refactoring.

crates/governance

Fail-closed policy engine: declarative allow/deny rules, BudgetPolicy ($USD / token caps), ActionContext validation, audit trail, and credentials vault integration.

crates/supervisor

Docker container sandboxing & HostSandbox (Linux bwrap / macOS sandbox-exec) for per-agent isolated computer and process environments.

Agent Lifecycle

Every agent — top-level or spawned — runs the same eight-step loop. The loop is where caching, safety gates and budget discipline live.

01
Build the system prompt

PromptBuilder stacks three cache tiers — stable (identity, role, task), context (cwd, platform, model, date), volatile (tools, skills, memory). Depth-0 agents additionally receive the memory digest as a frozen snapshot, keeping the prefix cache warm.

02
Check the doom loop

DoomLoopDetector inspects recent history before another token is spent: three identical tool calls in a row stops the loop.

03
Estimate tokens → compact

Real BPE counting with tiktoken cl100k_base (CJK-aware heuristic fallback). When estimated tokens reach context_window × compact_threshold (default 0.50), compaction runs.

04
Call the LLM — streaming

Text deltas hit the UI live as events; tool calls are assembled from stream fragments. On stream failure the runtime falls back to a non-streaming complete() automatically.

05
Pass the gates

Role-based deny → approval flow for approval_tools (y/n in the TUI, POST /sessions/:id/approve via API) → PreToolUse subprocess hooks that can allow, deny or enrich.

06
Execute the batch

Read-only calls run in parallel, writes serialize. question goes to the operator, spawn_agent starts children, discovered contacts are autosaved and absorbed into memory, results are truncated and appended.

07
Loop

Repeat until the model stops issuing tool calls or max_iterations is reached.

08
Return

The final answer returns to the caller. At coordinator level this is where findings get synthesized into the report.

Coordinator Flow & Goal Mode

A research run is the same loop one level up: the Coordinator plans, fans out, collects under budgets, lets an LLM judge check the result against the goal, then synthesizes and ships.

01
Plan

The LLM decomposes the query into sub-tasks; sub-tasks are persisted to the database before anything spawns.

02
Fan-out

Sub-agents start as a JoinSet of tokio tasks — or as OS processes via ProcessManager when use_multiprocess = true.

03
Collect

Results come back as budget-capped summaries, keeping the coordinator context small enough to reason over the fleet.

04
Reflection

For lead generation, count-based: if the contact quota is missed, gap rounds run until the target is met.

05
Goal Mode — LLM judge

A judge compares the result against the original goal; concrete gaps become new gap-filling sub-tasks, up to replan_rounds rounds.

06
Synthesize

The LLM merges findings into a single report: index.md, summary.md and the findings/ directory.

07
Export + Notify

Results are absorbed into long-term memory, exported to PDF/HTML/JSON/DOCX, and announced via webhook, email or Telegram.

Multi-Process

By default all agents run in one process as tokio tasks. The hidden fathom worker … mode is reserved for coordinator-managed process isolation and is not a user-facing command.

unix socket IPC
Coordinator (process 1)
  │  Unix socket IPC (JSON lines)
  ├──spawn──► Worker (process 2)  ── agent_id_1
  ├──spawn──► Worker (process 3)  ── agent_id_2
  └──spawn──► Worker (process 4)  ── agent_id_3

events: progress · tool-call · completion
  • kill_on_drop — worker handles are spawned with kill_on_drop(true): a dropped handle terminates the child, so no orphaned processes survive a cancelled run.
  • 30 s startup timeout — a worker must open its socket within 30 seconds or it is reaped and reported.
  • SQLite WAL mode — concurrent access from every process against one database, pool of 4 connections.
  • Isolation — one agent crashing never takes down the fleet; resource limits apply per process.

Context Management

Five stages keep the context window honest. Compaction triggers when estimated_tokens ≥ context_window × compact_threshold (default 0.50). Long-term memory complements all of it: relevant facts arrive as a prompt digest instead of living in session context.

BPE token counting

Exact counting with tiktoken cl100k_base; CJK-aware heuristic fallback. You cannot manage what you estimate wrong.

Tool result truncation

50 KB / 2000 lines per tool output, 200 KB per-turn budget. Overflow is persisted to disk with a pointer, so nothing is lost.

Micro-compaction

Deduplicate tool results by hash and prune old outputs — fast, deterministic, no LLM call.

Full compaction

Hermes-style: an LLM summarizes the middle of the conversation; head + summary + tail remain in context.

Anti-thrashing cooldown

After an ineffective compression pass, the runtime backs off instead of compacting again and again.

Error Handling

Failure modes are designed for, not patched onto the happy path:

Failure modeDefense
Transient LLM failureRetry with exponential backoff (3 attempts), Retry-After respected
Looping agentDoom-loop detection: three identical tool calls → stop
Failed tool callCascading cancellation: a shell failure cancels sibling calls; cancellation tokens propagate through the whole agent tree
Destructive commandGuard blocks rm -rf /, mkfs, fork bombs before execution
Tool errorGraceful degradation: errors return to the model as tool results so it can adapt

Task Tree Ledger

A shared, durable, append-only JSONL journal for the whole task tree rooted at one session. The coordinator and every sub-agent read from and write to the same file. It serves two roles: (1) coordination records that keep a shared picture of swarm decisions, and (2) typed child-to-parent beacons (“letters home”) that carry high-signal, attention-worthy signals.

Record KindTypeDescription
contractCoordinationAn agreement between agents about task division or interfaces
decisionCoordinationA recorded decision point with rationale
factCoordinationA discovered fact shared across the tree
noteCoordinationGeneral annotation
milestoneBeaconProgress marker — “I finished stage 2”
partial_findingBeaconIntermediate result worth sharing early
blockerBeacon ⚡Wakes parent immediately — agent is stuck
questionBeacon ⚡Wakes parent immediately — needs input
delegation_constraintBeacon ⚡Directive: halt_fanout, cap_children, require_lane, block_surface
interface_contractBeaconDefines interface boundaries between sub-agents.

Ledger path: ~/.fathom/ledger/task_tree/<session_id>.jsonl. Text capped at 4,000 chars per row. Total file capped at 2 MiB. ⚡ = wakes parent immediately (attention kinds).

Self-Improvement Backlog

A durable Markdown ledger where agents record candidate improvements. Re-triggering the same issue bumps a recurrence count instead of creating a duplicate. An optional LLM pass performs semantic dedup so reformulated ideas fold into existing entries.

FieldDescription
fingerprintSHA-256 dedup key (first 24 hex chars of normalized summary|category|source)
statusopen or done
priorityhigh, med, or low
kindbug, improvement, or capability_idea
countRecurrence count (starts at 1, bumped on re-match)
summaryHuman-readable description
proposed_next_stepSuggested action

File path: ~/.fathom/ledger/improvement-backlog.md. Grooming cap: 30 items. Semantic dedup candidate cap: 20 items. Digest format: - [high] (x3 | verify_email) Fix MX fallback.

Verification Receipt Ledger

An append-only JSONL ledger of typed verification receipts. Each receipt is keyed by (kind, value) so a PASS on one kind can never silence a FAIL on a different kind. The latest receipt per typed key wins.

Receipt KindVerdictsDescription
email_syntaxPass/FailRFC 5322 syntax validation
email_domain_mxPass/Fail/InconclusiveDNS-over-HTTPS MX/A record lookup
email_smtpPass/Fail/InconclusiveSMTP port 25 RCPT TO probe
email_disposablePass/FailDisposable email detection (50+ providers)
phone_normalizePass/FailE.164 normalization via libphonenumber
social_profilePass/Fail/InconclusiveHTTP check of social profile URL
person_namePass/FailPerson name verification

Verdict: PossiblyLaundered flags shell-piped checks (e.g. | tail) where exit codes may be masked. File path: ~/.fathom/ledger/verify_receipts.jsonl. The autosave module gates Verification::Verified on SMTP PASS, Partial on syntax+MX PASS.

Data Flow Diagram

This diagram shows how data moves through the system during a typical research run, from the initial user query to the final synthesized output.

data flow
User Query
  │
  ▼
┌─────────────┐
│ Coordinator  │  Plans sub-tasks, assigns roles, selects models
└──────┬──────┘
     │  Fan-out (JoinSet of tokio tasks or OS processes)
     ▼
┌──────────┐ ┌──────────┐ ┌──────────┐ ┌──────────┐
│ Researcher│ │ Researcher│ │ Analyst  │ │ Verifier │
└────┬─────┘ └────┬─────┘ └────┬─────┘ └────┬─────┘
   │            │            │            │
   ▼            ▼            ▼            ▼
┌─────────────────────────────────────────────────┐
│                  Tool Layer                      │
│  web_search · web_fetch · exec · memory_search  │
│  extract_contacts · verify_email · code_symbols │
└────────────────────────┬────────────────────────┘
                       │  Raw results (truncated)
                       ▼
┌─────────────────────────────────────────────────┐
│              Findings Budget                     │
│  Each sub-agent summarizes within its budget     │
│  (context_window × 0.1 per agent)                │
└────────────────────────┬────────────────────────┘
                       │  Budget-capped summaries
                       ▼
┌─────────────┐
│ Coordinator  │  Collects, runs Goal Mode judge
└──────┬──────┘
     │  Gap-filling rounds (replan_rounds)
     ▼
┌─────────────┐
│  Synthesis   │  LLM merges findings into report
└──────┬──────┘
     │
     ▼
┌──────────────────────────────────────────────────┐
│                   Output                          │
│  index.md · summary.md · findings/ · sources.md  │
│  + PDF / HTML / JSON / DOCX export               │
│  + Memory absorb (long-term)                     │
│  + Notifications (webhook / email / Telegram)     │
└──────────────────────────────────────────────────┘

Security Architecture

Fathom has 12 security layers that run at different points in the agent lifecycle. These are not optional features — they are always active and cannot be disabled.

All 12 security layers

#LayerWhen it runsDescription
1SSRF guardBefore HTTP requestsBlocks requests to internal/private network addresses
2Prompt injection detectionOn all tool inputs12 regex + heuristic patterns catch prompt injection attempts
3Secret scanningOn all tool outputs8 secret types detected and redacted before they enter context
4Destructive command blockingBefore exec8 shell patterns blocked (rm -rf /, mkfs, fork bombs, etc.)
5Doom-loop detectionBefore each LLM call3 identical tool calls in a row stops the loop
6Role-based deny listBefore tool executionProfile-level tool denials (e.g. analyst can’t save_contacts)
7Approval flowBefore side effectsHuman-in-the-loop for approval_tools via TUI or API
8PreToolUse hooksBefore tool executionExternal subprocess hooks can allow, deny or enrich tool calls
9Tool output truncationAfter tool execution50 KB / 2000 lines per output; overflow persisted to disk
10Protected surfacesOn file operationsconfig.toml, memory.db, .research.db cannot be modified by agents
11API key authenticationOn HTTP requestsNon-loopback binds require FATHOM_API_KEYS
12Rate limitingOn HTTP requestsPer-key rate limits on the server API

SSRF guard

The SSRF guard intercepts every HTTP request before it leaves the process. It resolves the target hostname and checks the resulting IP against a blocklist: 10.0.0.0/8, 172.16.0.0/12, 192.168.0.0/16, 127.0.0.0/8, 169.254.0.0/16, ::1, fc00::/7. DNS rebinding is mitigated by resolving before connecting and verifying the resolved address before the TCP handshake completes.

Prompt injection defense

Twelve patterns detect prompt injection attempts in tool outputs and user inputs:

detected patterns
1.  "ignore previous instructions"
2.  "you are now a ..."
3.  "system: " or "[system]" prefix injection
4.  Base64-encoded instruction blocks
5.  "disregard all prior"
6.  Invisible Unicode direction overrides (BiDi)
7.  Markdown link injection with javascript: URIs
8.  Nested tool_call JSON in tool outputs
9.  "IMPORTANT:" overrides in scraped page content
10. Repeated newlines followed by instruction blocks
11. XML/CDATA wrapping to break output parsing
12. Unicode homoglyph substitution in keywords

When a pattern is detected, the output is sanitized and a warning is logged. The agent receives a cleaned version of the content.

Secret scanning

Every tool output passes through a secret scanner before entering the context window. Eight types are detected and redacted with a [REDACTED:type] marker:

Secret typePattern
AWS Access KeyAKIA…
AWS Secret Key40-char base64 after aws_secret_access_key
GitHub Tokengh[pousr]_…
Slack Tokenxox[bpoas]-…
Private Key—–BEGIN (RSA|EC|OPENSSH) PRIVATE KEY—–
JWTeyJ… (3-part base64 with dots)
Generic API Keyapi_key=, apikey:, Authorization: Bearer
Password in URL://user:password@host

Destructive command blocking

The exec tool has a pre-execution guard that blocks 8 categories of destructive commands:

blocked patterns
1.  rm -rf /          — recursive root deletion
2.  mkfs / dd if=     — filesystem destruction
3.  :(){ :|:& };:    — fork bombs
4.  chmod -R 777 /    — permission escalation
5.  wget | sh / curl | sh — remote code execution
6.  > /dev/sda        — raw disk writes
7.  shutdown / reboot / halt — system control
8.  iptables -F       — firewall flush

Protected surfaces

These files and directories are read-only to agents — they can be read but never written or deleted by any tool call:

protected paths
~/.fathom/config.toml      # main configuration
~/.fathom/memory.db        # long-term memory database
~/.fathom/profiles/        # profile definitions
/.research.db             # session database
/index.md                 # generated index

Context Window Management

The context window is the most constrained resource in any agent system. Fathom manages it through five stages, each progressively more aggressive. The goal: keep the agent productive without losing critical information.

The five stages in detail

01
BPE token counting

Exact counting with tiktoken cl100k_base; CJK-aware heuristic fallback for non-Latin scripts. The counter runs before every LLM call. You cannot manage what you estimate wrong — this is why approximate char/4 counting is not used.

02
Tool result truncation

Each tool output is capped at tool_output_max_bytes (default 50 KB) or 2000 lines. The per-turn budget is turn_budget_bytes (default 200 KB). Overflow is persisted to disk with a pointer, so the agent can reference it later without bloating the context.

03
Micro-compaction

Deduplicate tool results by content hash and prune older outputs that have been superseded. This is fast and deterministic — no LLM call required. Runs automatically when context fills up.

04
Full compaction (Hermes-style)

An LLM summarizes the middle of the conversation. The context becomes: head (system prompt + first messages) + summary (compressed middle) + tail (recent messages). This is the only stage that costs tokens.

05
Anti-thrashing cooldown

After a full compaction pass that fails to free enough space, the runtime enters a cooldown period instead of compacting again immediately. This prevents the pathological case of spending all tokens on compaction without making progress.

Configuration

SettingDefaultDescription
context_windowmodel defaultMaximum tokens for the model (128K or 256K)
compact_threshold0.50Trigger compaction at this fraction of context_window
tool_output_max_bytes50 KBPer-tool output cap
turn_budget_bytes200 KBTotal per-turn budget for all tool outputs

Window profiles

Two profiles are available depending on the model’s context window size:

Profilecontext_windowcompact_thresholdBest for
low128K tokens0.50 (64K)DeepSeek chat-compatible fast models and other OpenAI-compatible endpoints
max256K tokens0.50 (128K)DeepSeek chat and other OpenAI-compatible strong models

Long-term memory complements all of it: relevant facts arrive as a prompt digest instead of living in session context. This means the agent can “remember” thousands of facts without consuming context window space.

Multi-Process Architecture

By default all agents run as tokio tasks within a single process. With use_multiprocess = true in config.toml, each sub-agent becomes its own OS process with full isolation.

How use_multiprocess works

01
Coordinator spawns workers

When process isolation is enabled, the coordinator may launch the hidden fathom worker mode as a child process; users should start work with fathom run or fathom tui.

02
Unix socket connection

Each worker connects back to the coordinator over a Unix domain socket. Messages are newline-delimited JSON (JSONL). The protocol supports: progress events, tool-call events, completion events, and error events.

03
Independent agent loop

Each worker runs its own complete agent loop: prompt building, LLM calls, tool execution, context management. It shares nothing with other workers except the SQLite database (WAL mode, 4-connection pool).

04
Result collection

When a worker finishes, it sends a completion event over the socket. The coordinator collects the result and feeds it into the synthesis step.

Unix socket IPC protocol

IPC message format
→ Coordinator → Worker:   {"type":"start","task":"...","tools":[...]}
← Worker → Coordinator:   {"type":"progress","agent_id":"...","msg":"searching..."}
← Worker → Coordinator:   {"type":"tool_call","tool":"web_search","args":{...},"ms":342}
← Worker → Coordinator:   {"type":"completion","agent_id":"...","result":"..."}
← Worker → Coordinator:   {"type":"error","agent_id":"...","error":"timeout"}

kill_on_drop safety

Worker handles are wrapped in tokio::process::Command::kill_on_drop(true). This means:

  • If the coordinator process crashes, all workers are killed automatically by the OS.
  • If a worker handle is dropped (cancelled task), the child process is terminated immediately.
  • No orphaned worker processes can survive a cancelled run or a panic.
  • A 30-second startup timeout ensures dead workers are reaped quickly.

When to use multi-process vs single-process

AspectSingle-process (default)Multi-process
OverheadLow — tokio tasks are lightweightHigher — OS processes with IPC
IsolationShared memory, one panic kills allFull process isolation
Resource limitsPer-task (soft)Per-process (OS-enforced)
DebuggingHarder to isolate issuesEasier — each worker has its own stderr
Best forSmall swarms (≤4 agents), developmentLarge swarms (6+ agents), production