Архитектура
Fathom is a modular Rust runtime for universal autonomous workers, with a typed event protocol, asynchronous agent loop on tokio, TUI, HTTP, sessions, plugins, MCP, and LSP integration. This page covers the crate map, lifecycle, subagent tree, context management and transport boundaries.
Карта крейтов
The workspace layout. core depends on nothing; agent ties llm, tools, memory and persistence together; server and tui are the two outer shells.
fathom/
├── crates/
│ ├── core/ # fundamental types & domain logic
│ ├── llm/ # LLM provider abstraction
│ ├── agent/ # agent runtime, coordination, control plane
│ ├── tools/ # 48 base tools plus up to 5 CDP browser tools and up to 6 computer tools + MCP/LSP tools
│ ├── memory/ # long-term semantic memory + entity graph
│ ├── mcp/ # Model Context Protocol (client and server)
│ ├── persistence/ # storage (SQLite, PostgreSQL, connection pool, jobs, vault)
│ ├── server/ # Axum HTTP/SSE API, AG-UI bridge, computer relay
│ ├── tui/ # ratatui terminal interface
│ ├── lsp/ # Language Server Protocol integration
│ ├── governance/ # policy engine & audit trail
│ ├── supervisor/ # Docker container supervisor
│ └── desktop/ # GPUI native desktop integration
├── Bot/ # Fathom Bot — local-first React 19 + Electron chat client & Node harness
└── src/main.rs # CLI entry pointcore ←── llm ←── agent ←── server
↑ ↑ ↑ ↑
└── tools ─┴── mcp ───┴── tui ────┘
↑ └── persistence ──┘
└── memory ──┘ (depends on core + llm)The foundation — depends on nothing. IDs (UUID v7), VirtualUri VFS resolver (skill://, rule://, memory://, artifact://, local://, xd://), messages & events, findings, AppConfig, export (PDF/HTML/JSON/DOCX), notifications, CRM sync, exact token counting.
Provider abstraction. LlmProvider trait with complete() and stream(), native Anthropic Messages API with prompt caching & extended thinking, CascadeProvider failover, retry with exponential backoff.
The runtime itself: agent loop, coordinator, compaction, prompt builder, tool executor with parallel safety, result budgets, control plane, IPC, process manager, doom-loop detector, hooks, session resume.
60 base tools plus up to 5 CDP browser tools, up to 6 computer tools, and optional LSP/MCP tools: high-precision hashline patching (edit), AST repo-mapping, git worktrees, compiler checks, DAP debugging, web, browser stealth, memory and coordination.
Long-term semantic memory + entity graph (mem0 / Memora model): SIMD-accelerated HNSW vector search, RDF Triples & GraphRAG traversal, facts, FTS5, embeddings, 5-outcome absorb pipeline, hybrid search, distillation, secret detection, and 2-way Obsidian vault sync.
Model Context Protocol client and server with bounded connection lifecycle, reconnect, circuit breaker, refresh, and stdio/HTTP transports.
SQLite in WAL mode with a 4-connection pool + PostgreSQL: sessions, agents, messages, findings, sub-tasks, contact DB, credentials vault, and durable jobs with self-healing retry.
Axum HTTP API: REST endpoints for sessions, agents, jobs, memory, credentials, coworkers, schedules, governance, inbound webhooks (/webhooks/inbound), multiplexed WebSocket (/ws), OpenAPI 3.1 (/openapi.json); control-plane answers/approvals; SSE streaming; API-key auth, rate limiting, Prometheus metrics.
Ratatui interface: multi-agent tree view, PTY terminal streams, live streaming buffer, jobs and memory panels, operator control — answer questions, grant approvals, watch the fleet.
Language Server Protocol client and adapter tool for IDE code intelligence and cross-file refactoring.
Fail-closed policy engine: declarative allow/deny rules, BudgetPolicy ($USD / token caps), ActionContext validation, audit trail, and credentials vault integration.
Docker container sandboxing & HostSandbox (Linux bwrap / macOS sandbox-exec) for per-agent isolated computer and process environments.
Жизненный цикл агента
Every agent — top-level or spawned — runs the same eight-step loop. The loop is where caching, safety gates and budget discipline live.
PromptBuilder stacks three cache tiers — stable (identity, role, task), context (cwd, platform, model, date), volatile (tools, skills, memory). Depth-0 agents additionally receive the memory digest as a frozen snapshot, keeping the prefix cache warm.
DoomLoopDetector inspects recent history before another token is spent: three identical tool calls in a row stops the loop.
Real BPE counting with tiktoken cl100k_base (CJK-aware heuristic fallback). When estimated tokens reach context_window × compact_threshold (default 0.50), compaction runs.
Text deltas hit the UI live as events; tool calls are assembled from stream fragments. On stream failure the runtime falls back to a non-streaming complete() automatically.
Role-based deny → approval flow for approval_tools (y/n in the TUI, POST /sessions/:id/approve via API) → PreToolUse subprocess hooks that can allow, deny or enrich.
Read-only calls run in parallel, writes serialize. question goes to the operator, spawn_agent starts children, discovered contacts are autosaved and absorbed into memory, results are truncated and appended.
Repeat until the model stops issuing tool calls or max_iterations is reached.
The final answer returns to the caller. At coordinator level this is where findings get synthesized into the report.
Поток координатора и Goal Mode
A research run is the same loop one level up: the Coordinator plans, fans out, collects under budgets, lets an LLM judge check the result against the goal, then synthesizes and ships.
The LLM decomposes the query into sub-tasks; sub-tasks are persisted to the database before anything spawns.
Sub-agents start as a JoinSet of tokio tasks — or as OS processes via ProcessManager when use_multiprocess = true.
Results come back as budget-capped summaries, keeping the coordinator context small enough to reason over the fleet.
For lead generation, count-based: if the contact quota is missed, gap rounds run until the target is met.
A judge compares the result against the original goal; concrete gaps become new gap-filling sub-tasks, up to replan_rounds rounds.
The LLM merges findings into a single report: index.md, summary.md and the findings/ directory.
Results are absorbed into long-term memory, exported to PDF/HTML/JSON/DOCX, and announced via webhook, email or Telegram.
Мультипроцессность
By default all agents run in one process as tokio tasks. The hidden fathom worker … mode is reserved for coordinator-managed process isolation and is not a user-facing command.
Coordinator (process 1)
│ Unix socket IPC (JSON lines)
├──spawn──► Worker (process 2) ── agent_id_1
├──spawn──► Worker (process 3) ── agent_id_2
└──spawn──► Worker (process 4) ── agent_id_3
events: progress · tool-call · completion- kill_on_drop — worker handles are spawned with
kill_on_drop(true): a dropped handle terminates the child, so no orphaned processes survive a cancelled run. - 30 s startup timeout — a worker must open its socket within 30 seconds or it is reaped and reported.
- SQLite WAL mode — concurrent access from every process against one database, pool of 4 connections.
- Isolation — one agent crashing never takes down the fleet; resource limits apply per process.
Управление контекстом
Five stages keep the context window honest. Compaction triggers when estimated_tokens ≥ context_window × compact_threshold (default 0.50). Long-term memory complements all of it: relevant facts arrive as a prompt digest instead of living in session context.
BPE token counting
Exact counting with tiktoken cl100k_base; CJK-aware heuristic fallback. You cannot manage what you estimate wrong.
Tool result truncation
50 KB / 2000 lines per tool output, 200 KB per-turn budget. Overflow is persisted to disk with a pointer, so nothing is lost.
Micro-compaction
Deduplicate tool results by hash and prune old outputs — fast, deterministic, no LLM call.
Full compaction
Hermes-style: an LLM summarizes the middle of the conversation; head + summary + tail remain in context.
Anti-thrashing cooldown
After an ineffective compression pass, the runtime backs off instead of compacting again and again.
Обработка ошибок
Failure modes are designed for, not patched onto the happy path:
| Failure mode | Defense |
|---|---|
| Transient LLM failure | Retry with exponential backoff (3 attempts), Retry-After respected |
| Looping agent | Doom-loop detection: three identical tool calls → stop |
| Failed tool call | Cascading cancellation: a shell failure cancels sibling calls; cancellation tokens propagate through the whole agent tree |
| Destructive command | Guard blocks rm -rf /, mkfs, fork bombs before execution |
| Tool error | Graceful degradation: errors return to the model as tool results so it can adapt |
Журнал дерева задач
Общий долговременный append-only JSONL-журнал для всего дерева задач сессии. Координатор и субагенты синхронизируются через единый журнал: (1) координационные записи для единой картины решений роя, (2) типизированные маяки от дочерних агентов родительскому ("письма домой"), передающие ключевые сигналы, требующие приоритетного внимания.
| Тип записи | Тип | Описание |
|---|---|---|
contract | Координация | Соглашение между агентами о разделении задач или интерфейсах |
decision | Координация | Зафиксированная точка решения с обоснованием |
fact | Координация | Обнаруженный факт, расшаренный по дереву |
note | Координация | Общая аннотация |
milestone | Маячок | Маркер прогресса — "Я завершил этап 2" |
partial_finding | Маячок | Промежуточный результат, стоит поделиться заранее |
blocker | Beacon ⚡ | Будит родителя немедленно — агент застрял |
question | Beacon ⚡ | Будит родителя немедленно — нужен ввод |
delegation_constraint | Beacon ⚡ | Директива: halt_fanout, cap_children, require_lane, block_surface |
interface_contract | Beacon | Defines interface boundaries between sub-agents. |
Путь журнала: ~/.fathom/ledger/task_tree/<session_id>.jsonl. Текст ограничен 4 000 символов на строку. Файл ограничен 2 МиБ. ⚡ = будит родителя немедленно (внимание-типы).
Бэклог самоулучшения
Durable Markdown-журнал, где агенты записывают кандидатов на улучшение. Повторное срабатывание увеличивает счётчик рецидивов вместо создания дубликата. Опциональный LLM-проход выполняет семантическую дедупликацию.
| Поле | Описание |
|---|---|
fingerprint | SHA-256 ключ дедупликации (первые 24 hex-символа нормализованного summary|category|source) |
status | open или done |
priority | high, med или low |
kind | bug, improvement или capability_idea |
count | Счётчик рецидивов (начинается с 1, увеличивается при совпадении) |
summary | Человекочитаемое описание |
proposed_next_step | Предлагаемое действие |
Путь: ~/.fathom/ledger/improvement-backlog.md. Лимит: 30 записей. Лимит кандидатов для семантического дедупа: 20. Формат дайджеста: - [high] (x3 | verify_email) Fix MX fallback.
Журнал чеков верификации
Append-only JSONL-журнал типизированных чеков верификации. Каждый чек ключуется по (тип, значение) — PASS по одному типу никогда не заглушит FAIL по другому. Последний чек по ключу побеждает.
| Тип чека | Вердикты | Описание |
|---|---|---|
email_syntax | Pass/Fail | Валидация синтаксиса RFC 5322 |
email_domain_mx | Pass/Fail/Inconclusive | DNS-over-HTTPS MX/A запись |
email_smtp | Pass/Fail/Inconclusive | SMTP port 25 RCPT TO зонд |
email_disposable | Pass/Fail | Детекция одноразовых email (55+ провайдеров) |
phone_normalize | Pass/Fail | Нормализация E.164 через libphonenumber |
social_profile | Pass/Fail/Inconclusive | HTTP-проверка URL социального профиля |
person_name | Pass/Fail | Верификация имени человека |
Вердикт: PossiblyLaundered помечает shell-пайп проверки (напр. | tail), где коды возврата могут быть замаскированы. Путь: ~/.fathom/ledger/verify_receipts.jsonl. Автосейв ставит Verification::Verified при SMTP PASS, Partial при syntax+MX PASS.
Data Flow Diagram
This diagram shows how data moves through the system during a typical research run, from the initial user query to the final synthesized output.
User Query
│
▼
┌─────────────┐
│ Coordinator │ Plans sub-tasks, assigns roles, selects models
└──────┬──────┘
│ Fan-out (JoinSet of tokio tasks or OS processes)
▼
┌──────────┐ ┌──────────┐ ┌──────────┐ ┌──────────┐
│ Researcher│ │ Researcher│ │ Analyst │ │ Verifier │
└────┬─────┘ └────┬─────┘ └────┬─────┘ └────┬─────┘
│ │ │ │
▼ ▼ ▼ ▼
┌─────────────────────────────────────────────────┐
│ Tool Layer │
│ web_search · web_fetch · exec · memory_search │
│ extract_contacts · verify_email · code_symbols │
└────────────────────────┬────────────────────────┘
│ Raw results (truncated)
▼
┌─────────────────────────────────────────────────┐
│ Findings Budget │
│ Each sub-agent summarizes within its budget │
│ (context_window × 0.1 per agent) │
└────────────────────────┬────────────────────────┘
│ Budget-capped summaries
▼
┌─────────────┐
│ Coordinator │ Collects, runs Goal Mode judge
└──────┬──────┘
│ Gap-filling rounds (replan_rounds)
▼
┌─────────────┐
│ Synthesis │ LLM merges findings into report
└──────┬──────┘
│
▼
┌──────────────────────────────────────────────────┐
│ Output │
│ index.md · summary.md · findings/ · sources.md │
│ + PDF / HTML / JSON / DOCX export │
│ + Memory absorb (long-term) │
│ + Notifications (webhook / email / Telegram) │
└──────────────────────────────────────────────────┘Security Architecture
Fathom has 12 security layers that run at different points in the agent lifecycle. These are not optional features — they are always active and cannot be disabled.
All 12 security layers
| # | Layer | When it runs | Description |
|---|---|---|---|
| 1 | SSRF guard | Before HTTP requests | Blocks requests to internal/private network addresses |
| 2 | Prompt injection detection | On all tool inputs | 12 regex + heuristic patterns catch prompt injection attempts |
| 3 | Secret scanning | On all tool outputs | 8 secret types detected and redacted before they enter context |
| 4 | Destructive command blocking | Before exec | 8 shell patterns blocked (rm -rf /, mkfs, fork bombs, etc.) |
| 5 | Doom-loop detection | Before each LLM call | 3 identical tool calls in a row stops the loop |
| 6 | Role-based deny list | Before tool execution | Profile-level tool denials (e.g. analyst can’t save_contacts) |
| 7 | Approval flow | Before side effects | Human-in-the-loop for approval_tools via TUI or API |
| 8 | PreToolUse hooks | Before tool execution | External subprocess hooks can allow, deny or enrich tool calls |
| 9 | Tool output truncation | After tool execution | 50 KB / 2000 lines per output; overflow persisted to disk |
| 10 | Protected surfaces | On file operations | config.toml, memory.db, .research.db cannot be modified by agents |
| 11 | API key authentication | On HTTP requests | Non-loopback binds require FATHOM_API_KEYS |
| 12 | Rate limiting | On HTTP requests | Per-key rate limits on the server API |
SSRF guard
The SSRF guard intercepts every HTTP request before it leaves the process. It resolves the target hostname
and checks the resulting IP against a blocklist: 10.0.0.0/8, 172.16.0.0/12,
192.168.0.0/16, 127.0.0.0/8, 169.254.0.0/16, ::1,
fc00::/7. DNS rebinding is mitigated by resolving before connecting and verifying the
resolved address before the TCP handshake completes.
Prompt injection defense
Twelve patterns detect prompt injection attempts in tool outputs and user inputs:
1. "ignore previous instructions"
2. "you are now a ..."
3. "system: " or "[system]" prefix injection
4. Base64-encoded instruction blocks
5. "disregard all prior"
6. Invisible Unicode direction overrides (BiDi)
7. Markdown link injection with javascript: URIs
8. Nested tool_call JSON in tool outputs
9. "IMPORTANT:" overrides in scraped page content
10. Repeated newlines followed by instruction blocks
11. XML/CDATA wrapping to break output parsing
12. Unicode homoglyph substitution in keywordsWhen a pattern is detected, the output is sanitized and a warning is logged. The agent receives a cleaned version of the content.
Secret scanning
Every tool output passes through a secret scanner before entering the context window. Eight types
are detected and redacted with a [REDACTED:type] marker:
| Secret type | Pattern |
|---|---|
| AWS Access Key | AKIA… |
| AWS Secret Key | 40-char base64 after aws_secret_access_key |
| GitHub Token | gh[pousr]_… |
| Slack Token | xox[bpoas]-… |
| Private Key | —–BEGIN (RSA|EC|OPENSSH) PRIVATE KEY—– |
| JWT | eyJ… (3-part base64 with dots) |
| Generic API Key | api_key=, apikey:, Authorization: Bearer |
| Password in URL | ://user:password@host |
Destructive command blocking
The exec tool has a pre-execution guard that blocks 8 categories of destructive commands:
1. rm -rf / — recursive root deletion
2. mkfs / dd if= — filesystem destruction
3. :(){ :|:& };: — fork bombs
4. chmod -R 777 / — permission escalation
5. wget | sh / curl | sh — remote code execution
6. > /dev/sda — raw disk writes
7. shutdown / reboot / halt — system control
8. iptables -F — firewall flushProtected surfaces
These files and directories are read-only to agents — they can be read but never written or deleted by any tool call:
~/.fathom/config.toml # main configuration
~/.fathom/memory.db # long-term memory database
~/.fathom/profiles/ # profile definitions
/.research.db # session database
/index.md # generated index Context Window Management
The context window is the most constrained resource in any agent system. Fathom manages it through five stages, each progressively more aggressive. The goal: keep the agent productive without losing critical information.
The five stages in detail
Exact counting with tiktoken cl100k_base; CJK-aware heuristic fallback for non-Latin scripts. The counter runs before every LLM call. You cannot manage what you estimate wrong — this is why approximate char/4 counting is not used.
Each tool output is capped at tool_output_max_bytes (default 50 KB) or 2000 lines. The per-turn budget is turn_budget_bytes (default 200 KB). Overflow is persisted to disk with a pointer, so the agent can reference it later without bloating the context.
Deduplicate tool results by content hash and prune older outputs that have been superseded. This is fast and deterministic — no LLM call required. Runs automatically when context fills up.
An LLM summarizes the middle of the conversation. The context becomes: head (system prompt + first messages) + summary (compressed middle) + tail (recent messages). This is the only stage that costs tokens.
After a full compaction pass that fails to free enough space, the runtime enters a cooldown period instead of compacting again immediately. This prevents the pathological case of spending all tokens on compaction without making progress.
Configuration
| Setting | Default | Description |
|---|---|---|
context_window | model default | Maximum tokens for the model (128K or 256K) |
compact_threshold | 0.50 | Trigger compaction at this fraction of context_window |
tool_output_max_bytes | 50 KB | Per-tool output cap |
turn_budget_bytes | 200 KB | Total per-turn budget for all tool outputs |
Window profiles
Two profiles are available depending on the model’s context window size:
| Profile | context_window | compact_threshold | Best for |
|---|---|---|---|
low | 128K tokens | 0.50 (64K) | DeepSeek chat-compatible fast models and other OpenAI-compatible endpoints |
max | 256K tokens | 0.50 (128K) | DeepSeek chat and other OpenAI-compatible strong models |
Long-term memory complements all of it: relevant facts arrive as a prompt digest instead of living in session context. This means the agent can “remember” thousands of facts without consuming context window space.
Multi-Process Architecture
By default all agents run as tokio tasks within a single process. With use_multiprocess = true
in config.toml, each sub-agent becomes its own OS process with full isolation.
How use_multiprocess works
When process isolation is enabled, the coordinator may launch the hidden fathom worker mode as a child process; users should start work with fathom run or fathom tui.
Each worker connects back to the coordinator over a Unix domain socket. Messages are newline-delimited JSON (JSONL). The protocol supports: progress events, tool-call events, completion events, and error events.
Each worker runs its own complete agent loop: prompt building, LLM calls, tool execution, context management. It shares nothing with other workers except the SQLite database (WAL mode, 4-connection pool).
When a worker finishes, it sends a completion event over the socket. The coordinator collects the result and feeds it into the synthesis step.
Unix socket IPC protocol
→ Coordinator → Worker: {"type":"start","task":"...","tools":[...]}
← Worker → Coordinator: {"type":"progress","agent_id":"...","msg":"searching..."}
← Worker → Coordinator: {"type":"tool_call","tool":"web_search","args":{...},"ms":342}
← Worker → Coordinator: {"type":"completion","agent_id":"...","result":"..."}
← Worker → Coordinator: {"type":"error","agent_id":"...","error":"timeout"}kill_on_drop safety
Worker handles are wrapped in tokio::process::Command::kill_on_drop(true). This means:
- If the coordinator process crashes, all workers are killed automatically by the OS.
- If a worker handle is dropped (cancelled task), the child process is terminated immediately.
- No orphaned worker processes can survive a cancelled run or a panic.
- A 30-second startup timeout ensures dead workers are reaped quickly.
When to use multi-process vs single-process
| Aspect | Single-process (default) | Multi-process |
|---|---|---|
| Overhead | Low — tokio tasks are lightweight | Higher — OS processes with IPC |
| Isolation | Shared memory, one panic kills all | Full process isolation |
| Resource limits | Per-task (soft) | Per-process (OS-enforced) |
| Debugging | Harder to isolate issues | Easier — each worker has its own stderr |
| Best for | Small swarms (≤4 agents), development | Large swarms (6+ agents), production |