Architecture

The engine,
crate by crate.

13 workspace crates, hierarchical coordination, durable runtime services — one self-hosted binary.

12
workspace crates
51+11
built-in + browser & computer tools
7
search backends
20
parallel agents max

01System overview

13 crates, one runtime

Each layer is purpose-built. Together they form a self-hosted runtime for planning, delegation, tools, memory, governance, and operator interfaces.

coreconfig · bootstrap · errors
llmprovider abstraction · streaming
agentloop · fan-out · goal mode
tools63 built-in + browser + computer (up to 75)
mcpclient + server
persistenceSQLite WAL
serverHTTP · SSE
tuiterminal UI
memorysemantic store
Tokio JoinSet Swarm Coordinator DAG

Fig 2.1 — Swarm Coordinator: Tokio JoinSet DAG execution across 4 parallel CPU worker pods with fair-share token budgeting and disk-spill memory scaling.

02Deep dive

The runtime pillars

01

Multi-Agent Orchestration

A coordinator agent decomposes complex research tasks into parallel sub-tasks. Specialized researcher agents run concurrently, each exploring a different data source. A writer agent synthesizes findings into a unified, structured report.

Coordinator → Researcher × N → Writer
02

Context Management

Two-phase compaction keeps context windows clean: micro-deduplication removes redundant tool outputs, then LLM-powered summarization compresses long histories. An anti-thrashing cooldown prevents rapid token oscillation between turns.

Micro dedup → LLM summarize → Cooldown
03

Safety & Security

Built-in SSRF guards block internal network access. Prompt injection detection neutralizes adversarial inputs. Destructive command blocking prevents accidental data loss. Doom loop detection halts runaway agents.

SSRF · Injection · Destructive · Doom loop
04

Persistent State

SQLite with WAL mode provides crash-safe persistence. All contacts, messages, and research artifacts survive restarts. TOCTOU-safe contact deduplication ensures no duplicates even under concurrent writes.

SQLite WAL · Crash recovery · TOCTOU-safe
05

Extensibility

MCP (Model Context Protocol) lets you plug in external tools and data sources. Lifecycle hooks let you intercept agent actions. Custom search backends swap in your proprietary knowledge bases.

MCP · Hooks · Custom backends
06

Observability

Prometheus-compatible metrics expose latency, token usage, and success rates. Structured JSON logging captures every agent action. SSE event streaming gives real-time visibility into running tasks.

Prometheus · Structured logs · SSE
07

Isolated Worker Processes

Every agent runs as its own OS process, not a thread. Workers talk over Unix sockets, are kill-on-drop safe, and a crash-resume path re-attaches to unfinished sub-tasks after a stall — so a failed run recovers instead of restarting from zero.

OS processes · Unix IPC · crash resume
08

Tool Batching

The executor automatically partitions tool calls into parallel-safe and sequential groups. File tools detect path overlaps to prevent write conflicts. CPU-heavy tools run on dedicated threads while I/O tools fan out across the async runtime.

Auto-partition · Path overlap · Async I/O
09

Fail-Closed Capability

Context window resolution tracks evidence quality: Confirmed (verified from API), Asserted (from config), Unknown. Budget allocation fails closed — if the model's true window cannot be proven, a conservative default is used instead of risking overflow.

Confirmed · Asserted · Unknown · Fail-closed
10

Durable Improvement Backlog

Workers can record improvement candidates in a durable ledger. Re-triggering an issue raises its recurrence count; an optional LLM pass can semantically deduplicate reformulated ideas.

Durable ledger · Recurrence count · Optional dedup
11

Reflection & Pattern Register

After every run, the coordinator reflects on what worked and what didn't. Patterns are recorded in a durable register and injected into future sessions. Lead-gen runs use count-based reflection to check quota progress and plan gap-filling rounds.

Post-run reflection · Gap-filling · Durable register
12

Goal Mode

An LLM judge evaluates partial results against the original goal, decides whether more data is needed, and launches additional rounds (replan_rounds, default 1). Agents do not stop until the goal is satisfied or the budget is exhausted.

LLM judge · replan_rounds · Budget-aware
13

Prompt Caching

Three prompt tiers — stable (system + tools), context (session history), volatile (latest turn) — are assembled so that stable prefixes reuse cached KV blocks across turns, cutting TTFT and cost on long sessions.

Stable · Context · Volatile · KV reuse
14

Coordination Runtime

The runtime connects agents through IrcBus, AgentRegistry and hub messaging. Operators can steer a run, park and revive agents, submit structured batch work, hand off sessions, and manage in-process async jobs or daemons.

Hub · Steering · Park/revive · Jobs · Daemons

03Built on

Battle-tested Rust stack

tokioaxumratatuirusqlite (WAL)tokio-postgresscraperquick-xmltiktoken-rslopdfphonenumberlettrereqwest-eventsourceschemarsuuid v7

Rust workspace · edition 2021 · release builds with LTO + strip · Elastic License 2.0

Ready to go deeper?

Explore the full technical documentation — architecture decisions, API reference, and deployment guides.