Аутрич

Руководство по аутричу и генерации лидов

Превратите цель на обычном языке в проверенных, обогащённых, готовых к использованию лидов — а затем в персональные аутрич-черновики, синхронизированные с вашей CRM. Это сквозной воркфлоу, объединяющий все инструменты агента.

Turn a plain-language goal into verified, enriched, ready-to-use leads — then into personalized outreach drafts synced to your CRM. This is the end-to-end workflow that ties the agent’s tools together.

Полный конвейер

Исследовательский прогон — это один и тот же цикл на каждом уровне: план, разветвление, сбор в рамках бюджетов, проверка и передача более чистого сигнала дальше. Вот что происходит на самом деле, инструмент за инструментом.

A research run is the same loop at every level: plan, fan out, collect under budgets, verify, then hand cleaner signal to the next stage. Here is what actually happens, tool by tool.

01
Ask

Write the goal in plain language. The coordinator decomposes it into sub-tasks, each with its own token budget and a shared contact-target counter.

02
Find

find_leads searches business directories + parses up to 6 corporate sites per company + runs role-based social searches in parallel.

03
Extract

extract_contacts pulls emails, phones and social profiles from pages — including obfuscated patterns like "bob at acme dot com". Every record carries a confidence score and its source URL.

04
Verify

verify_email (syntax → MX → disposable → role → optional SMTP probe), verify_phone (libphonenumber), verify_social_profile (14 platforms). Only verified rows survive.

05
Enrich

enrich_company + enrich_person pull size, revenue, founding year, HQ, tech stack, and the person’s title, bio and direct contact.

06
Persist

Autosave writes every extracted lead to contacts.db (dedup + merge), absorbs findings into memory, and optionally pushes to your CRM with crm_id dedup.

Начните с цели на обычном языке

Никаких схем и настройки CRM заранее. Опишите, чего хотите, своими словами.

bash
fathom run "Find VP Engineering at Series B startups in SF, verify their emails, and save the leads"

The coordinator turns that into sub-tasks with per-task contact quotas. Each sub-task becomes a Researcher (or an OS process with use_multiprocess = true), runs until its budget is spent, then hands its findings forward.

Почему проверка важна

Every bounce costs sender reputation and budget. That is why email goes through five gates before it is saved. verify_email returns a confidence score in [0,1] and a verdict — you decide the cutoff that matches your risk appetite.

ЭтапЧто проверяет
SyntaxRFC 5322 subset: one @, ≤254 chars, local ≤64, no leading/double dots, non-numeric TLD.
MX lookupDNS-over-HTTPS (Google + Cloudflare fallback); falls back to A-record per RFC 5321 §5.1 when no MX.
DisposableBlocks ~50 known throwaway providers (mailinator, yopmail, guerrillamail…).
Role-basedFlags department mailboxes (info@, sales@, no-reply@, admin@…).
SMTP probeOptional (smtp_check=true): TCP:25 → HELO → MAIL FROM → RCPT TO. Accepts 250/251, rejects 5xx. Never sends DATA.

Оценка уверенности

  • База 0.5 — синтаксис, disposable и ролевые проверки корректируют её.
  • +0.4, если домен разрешается через MX/A.
  • SMTP принял → ≥ 0.95; отклонён → ≤ 0.2; без SMTP → максимум 0.9.
  • −0.3, если адрес похож на disposable. Итоговое значение ограничено [0,1].
💡

Don’t guess emails. suggest_emails generates up to 9 patterns (first.last, f.last, last.first…) and learns the company’s pattern from your known addresses — matching candidates get +0.15 confidence (capped 0.95).

Введите лидов в свой процесс

Всё автоматически попадает в contacts.db (дедуп + слияние). Дальше есть три пути.

Everything lands in contacts.db automatically (dedup + merge). From there you have three paths.

1 · Export

CSV, vCard, JSON or Excel via fathom contacts export — or richer PDF/HTML/DOCX session reports.

2 · Push to CRM

fathom contacts push-crm syncs to amoCRM, Bitrix24 or HubSpot. Dedup by crm_id means re-runs never double-create.

3 · Watch for new leads

watch mode re-runs the same query every N seconds, diffs the contact DB and alerts on new rows.

Контроль человека там, где это важно

Побочные эффекты ждут явного «да». save_contacts и git_push по умолчанию за воротами одобрения — подтвердите в TUI или через POST /api/v1/sessions/:id/approve в API.

Side effects wait for an explicit yes. save_contacts and git_push are gated behind an approval flow by default — approve in the TUI, or via POST /api/v1/sessions/:id/approve on the API.

OSINT Tools Deep Dive

The outreach pipeline is powered by five complementary data-gathering tools. Each one targets a different surface — business directories, social networks, news coverage, or individual corporate sites — and the coordinator orchestrates them in parallel to maximize coverage while staying within per-run budgets.

🔍

find_leads

Lead Generation Engine

The primary entry point for any outreach task. find_leads takes a natural-language goal and decomposes it into directory searches, corporate-site parses, and social-network probes. It runs up to 6 corporate-site parses per company in parallel and merges results on the fly.

Parameters

ParameterTypeDescription
industrystringSIC/NAICS category or free-text (e.g. “B2B SaaS”, “fintech”)
locationstringCity, region, or country — parsed into geo-coordinates for directory radius search
rolesstring[]Target titles: [“VP Engineering”, “CTO”, “Head of Platform”]
limitnumberMax companies to return (default 50, max 500)
enrichbooleanWhen true, auto-runs enrich_company on each hit

Example invocation

python
result = await agent.call_tool("find_leads", {
  industry: "B2B SaaS",
  location: "San Francisco Bay Area",
  roles: ["VP Engineering", "CTO"],
  limit: 25,
  enrich: true
})

Example output (one record)

json
{
"company": "Acme Analytics",
"website": "https://acme-analytics.com",
"industry": "B2B SaaS",
"location": "San Francisco, CA",
"employees": "80-120",
"funding_stage": "Series B",
"contacts": [
  {
    "name": "Jordan Lee",
    "title": "VP Engineering",
    "email": "[email protected]",
    "email_confidence": 0.97,
    "linkedin": "https://linkedin.com/in/jordanlee",
    "source_url": "https://acme-analytics.com/team"
  }
],
"confidence": 0.92,
"sources": [
  "https://2gis.com/san-francisco/firm/acme-analytics",
  "https://acme-analytics.com/team",
  "https://linkedin.com/company/acme-analytics"
]
}
📂

search_business_directory

Multi-directory Search

Searches across 2GIS, Google Maps, Yandex Maps, and Yellow Pages in a single call. The tool fans out to all available directories for the target geography, deduplicates by normalized name

  • address, and returns a unified result set.

Parameters

ParameterTypeDescription
querystringSearch query (e.g. “software company”, “coworking space”)
citystringTarget city name
countrystringCountry name or ISO-3166-1 alpha-2 code
limitnumberMax results per directory (default 20, max 100)

Returned fields per record

  • name — Normalized business name
  • category — Business category from the directory
  • address — Full street address
  • phone — Phone number in E.164 when available
  • website — Company website URL
  • email — Public email if listed in the directory
  • rating — Average star rating (Google Maps, 2GIS)
  • source — Which directory the record came from

Example output

json
{
"name": "NovaTech Solutions",
"category": "Software Development",
"address": "45 Market St, Suite 200, San Francisco, CA 94105",
"phone": "+14155550123",
"website": "https://novatech.io",
"email": "[email protected]",
"rating": 4.6,
"reviews_count": 38,
"source": "google_maps"
}
🌐

search_social

Social Network Search

Searches Twitter/X, Telegram, and LinkedIn for people and companies matching a query. Useful for finding decision-makers by title+company, discovering niche community participants, and cross-referencing names found through directory searches.

Parameters

ParameterTypeDescription
querystringSearch string (e.g. “VP Engineering Stripe”, “CTO fintech”)
platformsstring[]Platforms to search: [“twitter”, “telegram”, “linkedin”] (default: all)
limitnumberMax results per platform (default 10, max 50)

Returned fields per record

  • profile_url — Direct link to the profile
  • name — Display name
  • bio — Profile bio / description text
  • followers_count — Follower count (Twitter, LinkedIn approximate)
  • location — Self-reported location
  • platform — Source platform identifier

Example output

json
{
"profile_url": "https://twitter.com/jordanlee_eng",
"name": "Jordan Lee",
"bio": "VP Engineering @AcmeAnalytics · Building data pipelines at scale",
"followers_count": 4200,
"location": "San Francisco, CA",
"platform": "twitter"
}
📰

search_news

News & Press Search

Searches news articles, press releases, and blog coverage. Especially useful for identifying recently funded companies, new executive hires, product launches, and partnership announcements — all high-signal events for outreach timing.

Parameters

ParameterTypeDescription
querystringSearch terms (e.g. “Series B funding SaaS”, “CTO appointment fintech”)
limitnumberMax articles to return (default 20, max 100)
from_datestringStart date in ISO 8601 (e.g. “2025-01-01”)
to_datestringEnd date in ISO 8601 (e.g. “2025-06-30”)

Returned fields per record

  • title — Article headline
  • url — Canonical URL
  • source — Publication name
  • date — Publication date
  • snippet — 200-character excerpt
  • mentioned_persons — Extracted named entities (people)
  • mentioned_companies — Extracted named entities (companies)

Example output

json
{
"title": "Acme Analytics Raises $40M Series B Led by Sequoia",
"url": "https://techcrunch.com/2025/03/15/acme-series-b",
"source": "TechCrunch",
"date": "2025-03-15",
"snippet": "Acme Analytics, the San Francisco-based B2B data platform, has closed a $40M Series B round...",
"mentioned_persons": ["Jordan Lee", "Maria Chen"],
"mentioned_companies": ["Acme Analytics", "Sequoia Capital"]
}
🏢

parse_corporate_site

Website Analyzer

Analyzes a corporate website to extract structured company intelligence. Crawls the homepage, /about, /team, and /contact pages (following robots.txt). Resolves obfuscated emails and phone numbers embedded in JavaScript or rendered via anti-scraping techniques.

What gets extracted

  • Company name — From <title>, OG tags, or structured data (JSON-LD)
  • Description — Meta description or first meaningful paragraph
  • Industry — Inferred from content keywords + JSON-LD categories
  • Size — Employee count range from careers page or LinkedIn widget
  • HQ — Address from footer, contact page, or structured data
  • Contact emails — Extracted from pages, mailto links, and JS-rendered content
  • Contact phones — Parsed from all pages, including international formats
  • Social profiles — Twitter, LinkedIn, GitHub, Telegram links from headers/footers
  • Team members — Names and titles from /team, /about, or leadership pages

Example output

json
{
"company_name": "Acme Analytics",
"description": "Real-time data pipelines for modern B2B teams",
"industry": "B2B SaaS / Data Infrastructure",
"size": "80-120 employees",
"hq": "45 Market St, Suite 200, San Francisco, CA 94105",
"emails": ["[email protected]", "[email protected]"],
"phones": ["+14155550100"],
"social": {
  "twitter": "https://twitter.com/acmeanalytics",
  "linkedin": "https://linkedin.com/company/acme-analytics",
  "github": "https://github.com/acme-analytics"
},
"team": [
  { "name": "Jordan Lee", "title": "VP Engineering" },
  { "name": "Maria Chen", "title": "CEO" },
  { "name": "Alex Rivera", "title": "CTO" }
],
"technologies": ["React", "Go", "Kubernetes", "PostgreSQL"]
}

Verification Tools Deep Dive

Raw data is noisy. Verification tools clean, normalize, and score every contact record before it enters your pipeline. Here is exactly what each tool does and how to interpret its output.

✉️

verify_email

5-Stage Pipeline

The most critical verification tool. Every email passes through five sequential stages — each stage can short-circuit the pipeline if a hard failure is detected (e.g., invalid syntax). The final confidence score is a composite of all stage results.

The 5-stage pipeline in detail

StageWhat happensScore impact
1. Syntax

Validates against RFC 5322 subset: exactly one @, total length ≤254, local part ≤64, no leading/double dots, TLD is non-numeric and at least 2 chars.

Fail → 0.0 (hard stop)
2. MX lookup

DNS-over-HTTPS via Google (primary) and Cloudflare (fallback). If no MX records exist, falls back to A-record lookup per RFC 5321 §5.1.

Resolved → +0.4; NXDOMAIN → 0.1 (hard stop)
3. Disposable

Checks the domain against ~50 known throwaway providers: mailinator, yopmail, guerrillamail, tempmail, 10minutemail, etc.

Match → −0.3
4. Role-based

Flags departmental / noreply addresses: info@, sales@, support@, admin@, no-reply@, billing@, etc.

Match → −0.15 (still usable, but lower priority)
5. SMTP probe

Optional (smtp_check=true). Opens TCP:25, sends HELO, MAIL FROM, RCPT TO. Accepts 250/251 as deliverable, 5xx as rejected. Never sends DATA — no actual email is delivered.

Accepted → ≥0.95; Rejected → ≤0.2; Timeout → capped at 0.9

Confidence scoring algorithm

python
score = 0.5                                    # base
if mx_resolved:     score += 0.4               # domain is real
if disposable:      score -= 0.3               # throwaway domain
if role_based:      score -= 0.15              # department mailbox
if smtp_accepted:   score = max(score, 0.95)   # server confirmed
if smtp_rejected:   score = min(score, 0.2)    # server denied
if no_smtp_check:   score = min(score, 0.9)    # can't fully confirm
score = clamp(score, 0.0, 1.0)
if pattern_matched: score = min(score + 0.15, 0.95)  # suggest_emails bonus

Example responses

json — Valid corporate email
{
"email": "[email protected]",
"verdict": "valid",
"confidence": 0.97,
"stages": {
  "syntax": "pass",
  "mx": "pass",
  "disposable": "clean",
  "role_based": false,
  "smtp": "accepted (250 OK)"
},
"mx_records": ["alt1.gmail-smtp-in.l.google.com", "gmail-smtp-in.l.google.com"]
}
json — Disposable email
{
"email": "[email protected]",
"verdict": "disposable",
"confidence": 0.05,
"stages": {
  "syntax": "pass",
  "mx": "pass",
  "disposable": "flagged (mailinator)",
  "role_based": false,
  "smtp": "skipped"
}
}
json — Role-based, SMTP accepted
{
"email": "[email protected]",
"verdict": "role_based",
"confidence": 0.78,
"stages": {
  "syntax": "pass",
  "mx": "pass",
  "disposable": "clean",
  "role_based": "flagged (info@)",
  "smtp": "accepted (250 OK)"
}
}
📞

verify_phone

Phone Normalization

Normalizes phone numbers using libphonenumber (the same library behind Google’s Android dialer). Parses raw input into E.164 format, detects the country, and classifies the line type.

What it returns

  • e164 — E.164 canonical format (e.g. +14155550123)
  • country — ISO-3166-1 alpha-2 country code
  • line_type — mobile, fixed_line, voip, toll_free, or unknown
  • valid — Boolean: does the number match the numbering plan for the detected country?
  • carrier — Carrier name when available (mobile numbers only)

Example

json
{
"input": "8 (495) 123-45-67",
"e164": "+74951234567",
"country": "RU",
"line_type": "fixed_line",
"valid": true,
"carrier": null
}
👤

verify_social_profile

14 Platforms

Validates and enriches social profile URLs. Returns structured metadata specific to each platform — follower counts, bios, verification status, and more.

Supported platforms (14)

Twitter / XLinkedInGitHubTelegramInstagramFacebookYouTubeTikTokRedditMastodonThreadsBlueskyHacker NewsProduct Hunt

Example response

json
{
"url": "https://twitter.com/jordanlee_eng",
"platform": "twitter",
"valid": true,
"metadata": {
  "name": "Jordan Lee",
  "handle": "@jordanlee_eng",
  "bio": "VP Engineering @AcmeAnalytics",
  "followers": 4200,
  "following": 890,
  "verified": false,
  "joined": "2019-04-22",
  "location": "San Francisco, CA"
}
}
🧩

suggest_emails

Pattern Inference

Given a person’s full name and a company domain, generates candidate email addresses using the 9 most common corporate patterns. When known addresses for the same domain are available, it infers the company’s actual pattern and boosts matching candidates.

The 9 standard patterns

#PatternExample (Jordan Lee)
1first.last[email protected]
2firstlast[email protected]
3f.last[email protected]
4first[email protected]
5last[email protected]
6first_last[email protected]
7last.first[email protected]
8flast[email protected]
9first-last[email protected]

Confidence scoring for suggestions

  • Each candidate gets a base confidence of 0.3 — it’s just a pattern guess.
  • If the inferred pattern from known addresses matches → +0.15 (capped at 0.95).
  • Running verify_email on the suggestion can push the score up to 0.97.
  • Suggestions that fail MX lookup drop to 0.05 immediately.
💡

Pattern inference is cumulative. Each time you save a verified email for a domain, suggest_emails learns that domain’s pattern. After 3+ known addresses for the same domain, the pattern lock is reliable enough to boost candidates to +0.20.

Enrichment Tools Deep Dive

Once you have a verified contact, enrichment fills in the context that makes outreach personal and relevant. These two tools turn bare email addresses into full profiles.

🏗️

enrich_company

Company Intelligence

Takes a company name or domain and returns structured intelligence gathered from public sources: corporate filings, about pages, job postings, technology stacks, and press coverage.

What gets discovered

  • Industry — Primary and secondary SIC/NAICS categories
  • Company size — Employee range (e.g. 80-120)
  • Revenue — Estimated annual revenue when publicly available
  • Founding year — Year of incorporation or launch
  • HQ — Headquarters address
  • Description — One-paragraph company description
  • Technologies — Detected tech stack from job postings, headers, and scripts
  • Funding — Latest funding round, amount, and investors

Example response

json
{
"domain": "acme-analytics.com",
"company_name": "Acme Analytics",
"industry": "B2B SaaS — Data Infrastructure",
"employee_range": "80-120",
"estimated_revenue": "$10M-$25M",
"founded": 2019,
"hq": "San Francisco, CA, USA",
"description": "Real-time data pipeline platform for B2B analytics teams.",
"technologies": ["React", "Go", "Kubernetes", "PostgreSQL", "Snowflake"],
"funding": {
  "latest_round": "Series B",
  "amount": "$40M",
  "date": "2025-03-15",
  "investors": ["Sequoia Capital", "Y Combinator"]
}
}
🧑‍💼

enrich_person

Person Intelligence

Takes a name and company (or email) and discovers publicly available professional information. Cross-references LinkedIn, Twitter, GitHub, press mentions, and conference speaker lists.

What gets discovered

  • Title — Current job title
  • Company — Current employer
  • LinkedIn — Profile URL
  • Twitter / X — Profile URL and handle
  • GitHub — Profile URL and notable repos
  • Public contact details — Emails and phones found in public directories
  • Location — City / metro area
  • Bio — Summary from LinkedIn or conference bios

Example response

json
{
"name": "Jordan Lee",
"title": "VP Engineering",
"company": "Acme Analytics",
"location": "San Francisco, CA",
"bio": "Engineering leader with 12 years in data infrastructure. Previously at Stripe and Google.",
"linkedin": "https://linkedin.com/in/jordanlee",
"twitter": "https://twitter.com/jordanlee_eng",
"github": "https://github.com/jordanlee",
"emails": ["[email protected]"],
"phones": ["+14155550199"],
"recent_mentions": [
  {
    "title": "Acme Analytics Raises $40M Series B",
    "source": "TechCrunch",
    "date": "2025-03-15"
  }
]
}

Contact Management Workflow

Every contact discovered by the pipeline flows through a unified management layer backed by SQLite. The system handles deduplication, merge, export, and CRM sync — so you never lose data or create duplicates.

How contacts.db works

All contacts are stored in contacts.db, a SQLite database in your project’s data directory. The schema is designed for fast dedup and flexible metadata.

sql
CREATE TABLE contacts (
  id          TEXT PRIMARY KEY,   -- UUID v4
  name        TEXT,
  email       TEXT,
  phone       TEXT,
  title       TEXT,
  company     TEXT,
  linkedin    TEXT,
  twitter     TEXT,
  github      TEXT,
  website     TEXT,
  location    TEXT,
  bio         TEXT,
  confidence  REAL DEFAULT 0.0,
  source_url  TEXT,
  crm_id      TEXT,               -- set after CRM push
  created_at  DATETIME DEFAULT CURRENT_TIMESTAMP,
  updated_at  DATETIME DEFAULT CURRENT_TIMESTAMP
);

CREATE UNIQUE INDEX idx_dedup
  ON contacts(email, phone, linkedin);

Dedup logic

  • Primary key: email + phone + LinkedIn composite. If any two of the three match, the record is merged.
  • Merge strategy: non-empty fields from the newer record overwrite older empty fields; confidence takes the higher value.
  • Name normalization: diacritics stripped, whitespace collapsed, case-insensitive comparison.

extract_contacts

Extracts emails, phone numbers, and social profile URLs from free text, raw HTML, or a live URL. Handles obfuscated patterns like bob [at] acme [dot] com and JavaScript-rendered content.

python
result = await agent.call_tool("extract_contacts", {
  url: "https://acme-analytics.com/team"
})

result = await agent.call_tool("extract_contacts", {
  text: "Reach Jordan at [email protected] or +1 (415) 555-0199"
})

save_contacts

Saves extracted contacts to contacts.db with automatic dedup and merge. When CRM integration is configured, an optional push happens after save.

python
# Save with auto-push to CRM
await agent.call_tool("save_contacts", { push_crm: true })

# Save without CRM push (default)
await agent.call_tool("save_contacts", {})

CLI commands

CommandDescription
contacts listList all contacts in the database with optional –limit filter
contacts export –format csvExport contacts to CSV, vCard (.vcf), JSON, or XLSX
contacts export –format vcardExport as vCard 3.0 — compatible with Apple Contacts, Google Contacts, Outlook
contacts export –format xlsxExport as Excel workbook with separate sheets for contacts and companies
contacts dedupRun dedup pass on existing database; reports merge count and conflicts
contacts push-crmSync all unsent contacts to the configured CRM

CRM integration

amoCRM v4

CRM API integration. Pushes verified contacts to the configured provider and tracks the remote crm_id for idempotent re-runs, so repeated syncs do not create duplicate contacts. Provider credentials and account settings come from the [crm] config section.

Bitrix24 REST

Webhook-based integration. Creates contacts + leads + companies in a single batch call. Activity timeline auto-populated with outreach events. Requires BITRIX24_WEBHOOK_URL env variable.

HubSpot v3

Private app token integration via the CRM v3 Objects API. Creates contacts with associations to companies. Supports batch upsert (up to 100 records per call). Custom properties mapped via [crm].hubspot_property_map.

toml — config.toml CRM section
[crm]
provider = "amocrm"          # amocrm | bitrix24 | hubspot
domain = "mycompany.amocrm.com"
api_key = "..."

# Or for HubSpot:
# provider = "hubspot"
# domain = "api.hubapi.com"
# api_key = "pat-xxx-..."

Approval Flow Deep Dive

Not every tool should run without human review. The approval flow gates side-effect-causing operations — so the agent can plan freely, but it needs your explicit OK before it writes to your CRM, saves contacts, or pushes code.

Gated tools (by default)

ToolWhy it’s gated
save_contactsWrites to contacts.db and optionally pushes to CRM — real side effects
git_pushPushes code to a remote repository — irreversible once published

How to customize

Add or remove gated tools in config.toml under the [agent] section:

toml
[agent]
# Default: save_contacts and git_push require approval
approval_tools = ["save_contacts", "git_push"]

# Add more tools to the gate:
approval_tools = ["save_contacts", "git_push", "send_email", "crm_push"]

# Remove all gates (use with caution):
approval_tools = []

TUI approval workflow

  • The agent pauses and displays a summary of the pending action in the terminal UI.
  • You see: tool name, parameters, expected outcome, and any warnings.
  • Press y to approve, n to reject, or e to edit the parameters before approving.
  • Editing opens the parameter JSON in your $EDITOR for inline modification.

API approval workflow

When running headless or integrated into CI/CD, approve via the REST API:

bash
# Approve a pending action
curl -X POST http://localhost:8080/api/v1/sessions/sess_abc123/approve \\
-H "Content-Type: application/json" \\
-d '{"action_id": "act_789", "approved": true}'

# Reject a pending action
curl -X POST http://localhost:8080/api/v1/sessions/sess_abc123/approve \\
-H "Content-Type: application/json" \\
-d '{"action_id": "act_789", "approved": false}'

Timeout and fallback behavior

If no approval arrives within the configured window, the agent uses the fallback strategy:

toml
[agent]
# What to do when approval times out:
# "deny"  — skip the action, continue the pipeline
# "allow" — auto-approve (for fully automated runs)
# "pause" — pause the session, resume later
approval_fallback = "deny"

# Timeout in seconds (default: 300 = 5 minutes)
approval_timeout_seconds = 300
⚠️

Never set approval_fallback = “allow” in production unless you are running a fully automated pipeline with strict parameter validation. Auto-approve bypasses human review for CRM writes and code pushes.

Watch Mode & Monitoring

Watch mode transforms a one-shot outreach task into a recurring monitor. Re-run the same query on a schedule, diff the results, and get notified only when new contacts appear.

The –repeat flag

Add –repeat N to re-run every N seconds. The agent runs the full pipeline on each cycle — but only saves contacts that are new or changed since the last run.

bash
# Re-run every 24 hours (86400 seconds)
fathom run "Find new VP Engineering hires at Series B startups in SF" \\
--repeat 86400

How diff detection works

  • Each run snapshots the contact database state (hash of email + phone + LinkedIn).
  • Before saving, new records are compared against the previous snapshot.
  • Only genuinely new or updated contacts trigger notifications.
  • Removed contacts (present in snapshot, absent in new run) are logged but not deleted.

Notification channels

Webhook

POST JSON payload to any URL. Configurable via NOTIFY_WEBHOOK_URL env variable. Payload includes new contact count, contact details, and run metadata.

Email

Send a digest email via SMTP. Configure SMTP_HOST, SMTP_PORT, SMTP_USER, SMTP_PASS, and NOTIFY_EMAIL_TO.

Telegram

Send a message to a Telegram chat via Bot API. Set TELEGRAM_BOT_TOKEN and TELEGRAM_CHAT_ID env variables.

Example: daily lead monitoring setup

bash
# Full monitoring setup
export NOTIFY_WEBHOOK_URL="https://hooks.slack.com/services/T00/B00/xxx"
export TELEGRAM_BOT_TOKEN="123456:ABC-..."
export TELEGRAM_CHAT_ID="-1001234567890"

fathom run \\
"Find newly funded startups in Berlin, extract CTO contacts, verify emails" \\
--repeat 86400

Webhook payload example

json
{
"event": "watch_diff",
"run_id": "run_20250615_001",
"timestamp": "2025-06-15T08:00:00Z",
"new_contacts": 7,
"updated_contacts": 2,
"total_contacts": 156,
"contacts": [
  {
    "name": "New CTO at StartupX",
    "email": "[email protected]",
    "title": "CTO",
    "company": "StartupX",
    "confidence": 0.94
  }
],
"query": "newly funded startups in Berlin"
}

Real-World Workflow Examples

These end-to-end examples show how the pipeline works in practice — from a plain-language goal to verified, enriched leads ready for outreach.

Example 1: Build a targeted engineering leadership list

Goal

“Build a list of VPs of Engineering at B2B SaaS companies in the Bay Area with verified emails and LinkedIn profiles.”

Command

bash
fathom run \\
"Find VP Engineering at B2B SaaS companies in the Bay Area. \\
 Verify their emails and LinkedIn profiles. \\
 Save enriched contacts with confidence >= 0.85."

What happens step by step

StepToolWhat it does
1find_leadsSearches 2GIS + Google Maps for “B2B SaaS” companies in SF Bay Area. Parses up to 6 corporate sites per company. Searches LinkedIn for VP Engineering titles.
2extract_contactsPulls emails, phones, and LinkedIn URLs from all discovered pages. Handles obfuscated patterns.
3verify_email5-stage pipeline on each email. Filters out anything below 0.85 confidence.
4verify_social_profileValidates each LinkedIn URL and extracts profile metadata.
5enrich_personAdds bio, title confirmation, Twitter/GitHub profiles, and recent mentions.
6save_contactsSaves to contacts.db (dedup + merge). Waits for approval before write.

Expected output structure

json
{
"session_id": "sess_abc123",
"query": "VP Engineering at B2B SaaS companies in the Bay Area",
"total_companies_found": 87,
"total_contacts_extracted": 134,
"verified_contacts": 92,
"filtered_contacts": 41,
"contacts": [
  {
    "name": "Jordan Lee",
    "title": "VP Engineering",
    "company": "Acme Analytics",
    "email": "[email protected]",
    "email_confidence": 0.97,
    "linkedin": "https://linkedin.com/in/jordanlee",
    "linkedin_verified": true,
    "twitter": "https://twitter.com/jordanlee_eng",
    "location": "San Francisco, CA",
    "bio": "Engineering leader with 12 years in data infrastructure.",
    "company_size": "80-120 employees",
    "company_funding": "Series B ($40M)"
  }
]
}

Example 2: Monitor competitor hiring

Goal

“Monitor competitor hiring and alert me when they post new senior engineering roles.”

Command

bash
fathom run \\
"Monitor career pages and job boards for senior engineering roles \\
 at CompetitorA, CompetitorB, CompetitorC. \\
 Alert on new VP/Director/Staff-level postings." \\
--repeat 43200

What happens step by step

StepToolWhat it does
1parse_corporate_siteCrawls /careers and /jobs pages of each competitor. Extracts job titles, locations, and posting dates.
2search_newsSearches for recent hiring announcements and press coverage mentioning the competitors.
3Diff engineCompares current findings against the previous run’s snapshot. Identifies new postings.
4NotifySends a Telegram message with each new senior role: title, location, posting URL, and date.

Expected notification payload

json
{
"event": "watch_diff",
"new_postings": [
  {
    "company": "CompetitorA",
    "title": "Director of Platform Engineering",
    "location": "San Francisco, CA",
    "url": "https://competitora.com/careers/director-platform",
    "posted_date": "2025-06-14",
    "detected_at": "2025-06-15T08:00:00Z"
  }
],
"monitoring_since": "2025-06-01T00:00:00Z",
"total_runs": 30
}

Example 3: Enrich an existing CSV of contacts

Goal

“I have a CSV of 200 contacts with names and companies. Enrich them with verified emails and LinkedIn profiles.”

Command

bash
# Step 1: Import the CSV
fathom contacts list --limit 50

# Step 2: Enrich each contact
fathom run \\
"For each contact in the database that is missing an email, \\
 find their professional email and LinkedIn profile. \\
 Verify emails before saving." \\
--source-db contacts.db

# Step 3: Export enriched results
fathom contacts export --format csv -o ./enriched_contacts

What happens step by step

StepToolWhat it does
1ImportParses CSV, maps columns, and inserts into contacts.db with dedup.
2suggest_emailsFor each contact without an email, generates 9 pattern candidates using name + company domain.
3verify_emailRuns 5-stage verification on each candidate. Keeps the highest-confidence match above 0.8.
4search_socialSearches LinkedIn for the person’s name + company. Validates and enriches the profile URL.
5enrich_personFills in bio, title confirmation, location, and additional social profiles.
6save_contactsMerges enriched data into existing contact records. Updates only changed fields.
7ExportExports the enriched database to CSV with all new fields included.

Expected output structure

text — enriched_contacts.csv columns
name,email,email_confidence,title,company,linkedin,twitter,location,bio,source
Jordan Lee,[email protected],0.97,VP Engineering,Acme Analytics,https://linkedin.com/in/jordanlee,https://twitter.com/jordanlee_eng,"San Francisco, CA","Engineering leader with 12 years...",suggest_emails+verify
Maria Chen,[email protected],0.91,CTO,StartupX,https://linkedin.com/in/mariachen,,Berlin,"Former Google, building developer tools...",search_social+verify
💡

Pattern learning compounds. After enriching 200 contacts, suggest_emails has learned the email patterns for dozens of domains. Future enrichment runs for the same companies will be significantly faster and more accurate.