AGENT 02 / 07
Searching
The discovery layer of the fleet. Seven backends behind one call, fused with hybrid and smart modes and RRF ranking — so a query spans the whole web, not the first page.
searching · autonomous
01What it does
Seven backends, one honest answer
Configurable search backends
Configured providers can be selected individually, or hybrid and smart modes can put available engines behind one web_search call. The runtime keeps the interface stable while deployments choose their providers.
hybrid mode (default)
A priority chain tries configured engines in order and uses the first non-empty result set. The runtime reports when no configured source answers instead of inventing coverage.
smart mode + RRF
All configured engines run in parallel via tokio::join!. Results are deduped by normalized URL and ranked with reciprocal rank fusion (k=60), so pages confirmed by several engines end up on top.
Social & news surfaces
search_social and search_news can use configured APIs or public web/RSS surfaces, with heuristic entity extraction and provenance retained for review.
Directories & registries
search_business_directory can query configured directories in parallel and merge records such as names, categories, addresses, phones, websites, and provenance when those fields are available.
Composable discovery pipeline
A worker can chain directory lookup, site parsing, social search, attribution, deduplication, and confidence ranking into a reusable workflow. Lead discovery is one downstream use, not the product identity.
02How it works
From query to ranked sources
Pick the surface
The worker chooses the right source for the mission: open web, repositories, documents, social profiles, press coverage, or configured directories.
Fan out to backends
Configured engines can run in parallel with per-source timeouts; unavailable providers are reported rather than hidden behind assumptions.
Guard & fetch
Each URL passes the SSRF guard — every resolved IP must be public. Redirects are followed manually, at most 5 hops, re-validated per hop.
Fuse & rank
smart mode normalizes URLs (case, fragment, trailing slash), dedupes and scores with RRF: each rank r contributes 1 / (60 + r) — multi-engine hits win.
Feed forward
Ranked results with source metadata go to fetching and extraction; the session’s fetch and MX caches (10-minute TTL) make follow-up reads instant.
03Under the hood
The runtime underneath
| [search] backend | hybrid (default) · smart · or a selected configured engine; provider availability is surfaced to the operator |
| hybrid order | Configured engines are tried in priority order; the first non-empty result set wins and unavailable sources remain visible |
| smart fusion | all configured engines in parallel; dedup by normalized URL; RRF score Σ 1 / (60 + rank) |
| results | snippets capped at 1,000 chars; limit default 10, max 50; structured source metadata for sources.md |
| directories | Configured directory sources are queried within their limits, then merged and deduplicated with provenance retained |
| news | Configured news APIs or public RSS; mentioned entities are extracted heuristically and kept with source provenance |
| social | Configured social APIs or public profile pages; blocked responses are reported and provenance remains attached |
| discovery pipeline | Configured sources can be composed for site parsing, role queries, attribution, and confidence scoring |
| SSRF guard | loopback / RFC1918 / link-local / cloud-metadata IPs blocked on every resolved address; ≤ 5 redirects, each hop re-validated |
| caches | session-shared fetch cache (64 entries) and MX cache (256 entries), TTL 10 minutes, inherited by every agent |
04Tools