AGENT 02 / 07

Searching

The discovery layer of the fleet. Seven backends behind one call, fused with hybrid and smart modes and RRF ranking — so a query spans the whole web, not the first page.

searchingfathom worker

searching · autonomous

01What it does

Seven backends, one honest answer

01

Configurable search backends

Configured providers can be selected individually, or hybrid and smart modes can put available engines behind one web_search call. The runtime keeps the interface stable while deployments choose their providers.

02

hybrid mode (default)

A priority chain tries configured engines in order and uses the first non-empty result set. The runtime reports when no configured source answers instead of inventing coverage.

03

smart mode + RRF

All configured engines run in parallel via tokio::join!. Results are deduped by normalized URL and ranked with reciprocal rank fusion (k=60), so pages confirmed by several engines end up on top.

04

Social & news surfaces

search_social and search_news can use configured APIs or public web/RSS surfaces, with heuristic entity extraction and provenance retained for review.

05

Directories & registries

search_business_directory can query configured directories in parallel and merge records such as names, categories, addresses, phones, websites, and provenance when those fields are available.

06

Composable discovery pipeline

A worker can chain directory lookup, site parsing, social search, attribution, deduplication, and confidence ranking into a reusable workflow. Lead discovery is one downstream use, not the product identity.

02How it works

From query to ranked sources

1

Pick the surface

The worker chooses the right source for the mission: open web, repositories, documents, social profiles, press coverage, or configured directories.

2

Fan out to backends

Configured engines can run in parallel with per-source timeouts; unavailable providers are reported rather than hidden behind assumptions.

3

Guard & fetch

Each URL passes the SSRF guard — every resolved IP must be public. Redirects are followed manually, at most 5 hops, re-validated per hop.

4

Fuse & rank

smart mode normalizes URLs (case, fragment, trailing slash), dedupes and scores with RRF: each rank r contributes 1 / (60 + r) — multi-engine hits win.

5

Feed forward

Ranked results with source metadata go to fetching and extraction; the session’s fetch and MX caches (10-minute TTL) make follow-up reads instant.

03Under the hood

The runtime underneath

[search] backendhybrid (default) · smart · or a selected configured engine; provider availability is surfaced to the operator
hybrid orderConfigured engines are tried in priority order; the first non-empty result set wins and unavailable sources remain visible
smart fusionall configured engines in parallel; dedup by normalized URL; RRF score Σ 1 / (60 + rank)
resultssnippets capped at 1,000 chars; limit default 10, max 50; structured source metadata for sources.md
directoriesConfigured directory sources are queried within their limits, then merged and deduplicated with provenance retained
newsConfigured news APIs or public RSS; mentioned entities are extracted heuristically and kept with source provenance
socialConfigured social APIs or public profile pages; blocked responses are reported and provenance remains attached
discovery pipelineConfigured sources can be composed for site parsing, role queries, attribution, and confidence scoring
SSRF guardloopback / RFC1918 / link-local / cloud-metadata IPs blocked on every resolved address; ≤ 5 redirects, each hop re-validated
cachessession-shared fetch cache (64 entries) and MX cache (256 entries), TTL 10 minutes, inherited by every agent

04Tools

What the Searcher can call

web_searchsearch_socialsearch_newssearch_business_directoryfind_leadsparse_corporate_siteweb_fetchweb_crawlweb_feed