Architecture
The engine,
crate by crate.
13 workspace crates, hierarchical coordination, durable runtime services — one self-hosted binary.
01System overview
13 crates, one runtime
Each layer is purpose-built. Together they form a self-hosted runtime for planning, delegation, tools, memory, governance, and operator interfaces.

Fig 2.1 — Swarm Coordinator: Tokio JoinSet DAG execution across 4 parallel CPU worker pods with fair-share token budgeting and disk-spill memory scaling.
02Deep dive
The runtime pillars
Multi-Agent Orchestration
A coordinator agent decomposes complex research tasks into parallel sub-tasks. Specialized researcher agents run concurrently, each exploring a different data source. A writer agent synthesizes findings into a unified, structured report.
Coordinator → Researcher × N → WriterContext Management
Two-phase compaction keeps context windows clean: micro-deduplication removes redundant tool outputs, then LLM-powered summarization compresses long histories. An anti-thrashing cooldown prevents rapid token oscillation between turns.
Micro dedup → LLM summarize → CooldownSafety & Security
Built-in SSRF guards block internal network access. Prompt injection detection neutralizes adversarial inputs. Destructive command blocking prevents accidental data loss. Doom loop detection halts runaway agents.
SSRF · Injection · Destructive · Doom loopPersistent State
SQLite with WAL mode provides crash-safe persistence. All contacts, messages, and research artifacts survive restarts. TOCTOU-safe contact deduplication ensures no duplicates even under concurrent writes.
SQLite WAL · Crash recovery · TOCTOU-safeExtensibility
MCP (Model Context Protocol) lets you plug in external tools and data sources. Lifecycle hooks let you intercept agent actions. Custom search backends swap in your proprietary knowledge bases.
MCP · Hooks · Custom backendsObservability
Prometheus-compatible metrics expose latency, token usage, and success rates. Structured JSON logging captures every agent action. SSE event streaming gives real-time visibility into running tasks.
Prometheus · Structured logs · SSEIsolated Worker Processes
Every agent runs as its own OS process, not a thread. Workers talk over Unix sockets, are kill-on-drop safe, and a crash-resume path re-attaches to unfinished sub-tasks after a stall — so a failed run recovers instead of restarting from zero.
OS processes · Unix IPC · crash resumeTool Batching
The executor automatically partitions tool calls into parallel-safe and sequential groups. File tools detect path overlaps to prevent write conflicts. CPU-heavy tools run on dedicated threads while I/O tools fan out across the async runtime.
Auto-partition · Path overlap · Async I/OFail-Closed Capability
Context window resolution tracks evidence quality: Confirmed (verified from API), Asserted (from config), Unknown. Budget allocation fails closed — if the model's true window cannot be proven, a conservative default is used instead of risking overflow.
Confirmed · Asserted · Unknown · Fail-closedDurable Improvement Backlog
Workers can record improvement candidates in a durable ledger. Re-triggering an issue raises its recurrence count; an optional LLM pass can semantically deduplicate reformulated ideas.
Durable ledger · Recurrence count · Optional dedupReflection & Pattern Register
After every run, the coordinator reflects on what worked and what didn't. Patterns are recorded in a durable register and injected into future sessions. Lead-gen runs use count-based reflection to check quota progress and plan gap-filling rounds.
Post-run reflection · Gap-filling · Durable registerGoal Mode
An LLM judge evaluates partial results against the original goal, decides whether more data is needed, and launches additional rounds (replan_rounds, default 1). Agents do not stop until the goal is satisfied or the budget is exhausted.
LLM judge · replan_rounds · Budget-awarePrompt Caching
Three prompt tiers — stable (system + tools), context (session history), volatile (latest turn) — are assembled so that stable prefixes reuse cached KV blocks across turns, cutting TTFT and cost on long sessions.
Stable · Context · Volatile · KV reuseCoordination Runtime
The runtime connects agents through IrcBus, AgentRegistry and hub messaging. Operators can steer a run, park and revive agents, submit structured batch work, hand off sessions, and manage in-process async jobs or daemons.
Hub · Steering · Park/revive · Jobs · Daemons03Built on
Battle-tested Rust stack
Rust workspace · edition 2021 · release builds with LTO + strip · Elastic License 2.0
Ready to go deeper?
Explore the full technical documentation — architecture decisions, API reference, and deployment guides.