Outreach & Lead Generation Guide
Turn a plain-language goal into verified, enriched, ready-to-use leads — then into personalized outreach drafts synced to your CRM. This is the end-to-end workflow that ties the agent’s tools together.
The full pipeline
A research run is the same loop at every level: plan, fan out, collect under budgets, verify, then hand cleaner signal to the next stage. Here is what actually happens, tool by tool.
Write the goal in plain language. The coordinator decomposes it into sub-tasks, each with its own token budget and a shared contact-target counter.
find_leads searches business directories + parses up to 6 corporate sites per company + runs role-based social searches in parallel.
extract_contacts pulls emails, phones and social profiles from pages — including obfuscated patterns like "bob at acme dot com". Every record carries a confidence score and its source URL.
verify_email (syntax → MX → disposable → role → optional SMTP probe), verify_phone (libphonenumber), verify_social_profile (14 platforms). Only verified rows survive.
enrich_company + enrich_person pull size, revenue, founding year, HQ, tech stack, and the person’s title, bio and direct contact.
Autosave writes every extracted lead to contacts.db (dedup + merge), absorbs findings into memory, and optionally pushes to your CRM with crm_id dedup.
Start with a plain-language goal
No schemas, no CRM setup first. Ask for what you want in the terms that make sense to you.
fathom run "Find VP Engineering at Series B startups in SF, verify their emails, and save the leads"The coordinator turns that into sub-tasks with per-task contact quotas. Each sub-task becomes a
Researcher (or an OS process with use_multiprocess = true), runs until its budget is
spent, then hands its findings forward.
Why verification matters
Every bounce costs sender reputation and budget. That is why email goes through five gates
before it is saved. verify_email returns a confidence score in [0,1]
and a verdict — you decide the cutoff that matches your risk appetite.
| Stage | What it checks |
|---|---|
Syntax | RFC 5322 subset: one @, ≤254 chars, local ≤64, no leading/double dots, non-numeric TLD. |
MX lookup | DNS-over-HTTPS (Google + Cloudflare fallback); falls back to A-record per RFC 5321 §5.1 when no MX. |
Disposable | Blocks ~50 known throwaway providers (mailinator, yopmail, guerrillamail…). |
Role-based | Flags department mailboxes (info@, sales@, no-reply@, admin@…). |
SMTP probe | Optional (smtp_check=true): TCP:25 → HELO → MAIL FROM → RCPT TO. Accepts 250/251, rejects 5xx. Never sends DATA. |
Confidence scoring
- Base 0.5 — syntax, disposable and role checks adjust it.
- +0.4 if the domain resolves via MX/A.
- SMTP accepted → ≥ 0.95; rejected → ≤ 0.2; no SMTP → capped at 0.9.
- −0.3 if the address looks disposable. Final value clamped to [0,1].
Don’t guess emails. suggest_emails generates up to 9 patterns
(first.last, f.last, last.first…) and learns the company’s pattern from your known addresses —
matching candidates get +0.15 confidence (capped 0.95).
Get leads into your workflow
Everything lands in contacts.db automatically (dedup + merge). From there you have
three paths.
1 · Export
CSV, vCard, JSON or Excel via fathom contacts export — or richer PDF/HTML/DOCX session reports.
2 · Push to CRM
fathom contacts push-crm syncs to amoCRM, Bitrix24 or HubSpot. Dedup by crm_id means re-runs never double-create.
3 · Watch for new leads
watch mode re-runs the same query every N seconds, diffs the contact DB and alerts on new rows.
Human-in-the-loop where it counts
Side effects wait for an explicit yes. save_contacts and git_push are
gated behind an approval flow by default — approve in the TUI, or via
POST /api/v1/sessions/:id/approve on the API.
OSINT Tools Deep Dive
The outreach pipeline is powered by five complementary data-gathering tools. Each one targets a different surface — business directories, social networks, news coverage, or individual corporate sites — and the coordinator orchestrates them in parallel to maximize coverage while staying within per-run budgets.
find_leads
Lead Generation EngineThe primary entry point for any outreach task. find_leads takes a natural-language
goal and decomposes it into directory searches, corporate-site parses, and social-network probes.
It runs up to 6 corporate-site parses per company in parallel and merges results on the fly.
Parameters
| Parameter | Type | Description |
|---|---|---|
industry | string | SIC/NAICS category or free-text (e.g. “B2B SaaS”, “fintech”) |
location | string | City, region, or country — parsed into geo-coordinates for directory radius search |
roles | string[] | Target titles: [“VP Engineering”, “CTO”, “Head of Platform”] |
limit | number | Max companies to return (default 50, max 500) |
enrich | boolean | When true, auto-runs enrich_company on each hit |
Example invocation
result = await agent.call_tool("find_leads", {
industry: "B2B SaaS",
location: "San Francisco Bay Area",
roles: ["VP Engineering", "CTO"],
limit: 25,
enrich: true
})Example output (one record)
{
"company": "Acme Analytics",
"website": "https://acme-analytics.com",
"industry": "B2B SaaS",
"location": "San Francisco, CA",
"employees": "80-120",
"funding_stage": "Series B",
"contacts": [
{
"name": "Jordan Lee",
"title": "VP Engineering",
"email": "[email protected]",
"email_confidence": 0.97,
"linkedin": "https://linkedin.com/in/jordanlee",
"source_url": "https://acme-analytics.com/team"
}
],
"confidence": 0.92,
"sources": [
"https://2gis.com/san-francisco/firm/acme-analytics",
"https://acme-analytics.com/team",
"https://linkedin.com/company/acme-analytics"
]
}search_business_directory
Multi-directory SearchSearches across 2GIS, Google Maps, Yandex Maps, and Yellow Pages in a single call. The tool fans out to all available directories for the target geography, deduplicates by normalized name
- address, and returns a unified result set.
Parameters
| Parameter | Type | Description |
|---|---|---|
query | string | Search query (e.g. “software company”, “coworking space”) |
city | string | Target city name |
country | string | Country name or ISO-3166-1 alpha-2 code |
limit | number | Max results per directory (default 20, max 100) |
Returned fields per record
- name — Normalized business name
- category — Business category from the directory
- address — Full street address
- phone — Phone number in E.164 when available
- website — Company website URL
- email — Public email if listed in the directory
- rating — Average star rating (Google Maps, 2GIS)
- source — Which directory the record came from
Example output
{
"name": "NovaTech Solutions",
"category": "Software Development",
"address": "45 Market St, Suite 200, San Francisco, CA 94105",
"phone": "+14155550123",
"website": "https://novatech.io",
"email": "[email protected]",
"rating": 4.6,
"reviews_count": 38,
"source": "google_maps"
}search_social
Social Network SearchSearches Twitter/X, Telegram, and LinkedIn for people and companies matching a query. Useful for finding decision-makers by title+company, discovering niche community participants, and cross-referencing names found through directory searches.
Parameters
| Parameter | Type | Description |
|---|---|---|
query | string | Search string (e.g. “VP Engineering Stripe”, “CTO fintech”) |
platforms | string[] | Platforms to search: [“twitter”, “telegram”, “linkedin”] (default: all) |
limit | number | Max results per platform (default 10, max 50) |
Returned fields per record
- profile_url — Direct link to the profile
- name — Display name
- bio — Profile bio / description text
- followers_count — Follower count (Twitter, LinkedIn approximate)
- location — Self-reported location
- platform — Source platform identifier
Example output
{
"profile_url": "https://twitter.com/jordanlee_eng",
"name": "Jordan Lee",
"bio": "VP Engineering @AcmeAnalytics · Building data pipelines at scale",
"followers_count": 4200,
"location": "San Francisco, CA",
"platform": "twitter"
}search_news
News & Press SearchSearches news articles, press releases, and blog coverage. Especially useful for identifying recently funded companies, new executive hires, product launches, and partnership announcements — all high-signal events for outreach timing.
Parameters
| Parameter | Type | Description |
|---|---|---|
query | string | Search terms (e.g. “Series B funding SaaS”, “CTO appointment fintech”) |
limit | number | Max articles to return (default 20, max 100) |
from_date | string | Start date in ISO 8601 (e.g. “2025-01-01”) |
to_date | string | End date in ISO 8601 (e.g. “2025-06-30”) |
Returned fields per record
- title — Article headline
- url — Canonical URL
- source — Publication name
- date — Publication date
- snippet — 200-character excerpt
- mentioned_persons — Extracted named entities (people)
- mentioned_companies — Extracted named entities (companies)
Example output
{
"title": "Acme Analytics Raises $40M Series B Led by Sequoia",
"url": "https://techcrunch.com/2025/03/15/acme-series-b",
"source": "TechCrunch",
"date": "2025-03-15",
"snippet": "Acme Analytics, the San Francisco-based B2B data platform, has closed a $40M Series B round...",
"mentioned_persons": ["Jordan Lee", "Maria Chen"],
"mentioned_companies": ["Acme Analytics", "Sequoia Capital"]
}parse_corporate_site
Website AnalyzerAnalyzes a corporate website to extract structured company intelligence. Crawls the homepage,
/about, /team, and /contact pages (following
robots.txt). Resolves obfuscated emails and phone numbers embedded in JavaScript or
rendered via anti-scraping techniques.
What gets extracted
- Company name — From <title>, OG tags, or structured data (JSON-LD)
- Description — Meta description or first meaningful paragraph
- Industry — Inferred from content keywords + JSON-LD categories
- Size — Employee count range from careers page or LinkedIn widget
- HQ — Address from footer, contact page, or structured data
- Contact emails — Extracted from pages, mailto links, and JS-rendered content
- Contact phones — Parsed from all pages, including international formats
- Social profiles — Twitter, LinkedIn, GitHub, Telegram links from headers/footers
- Team members — Names and titles from
/team,/about, or leadership pages
Example output
{
"company_name": "Acme Analytics",
"description": "Real-time data pipelines for modern B2B teams",
"industry": "B2B SaaS / Data Infrastructure",
"size": "80-120 employees",
"hq": "45 Market St, Suite 200, San Francisco, CA 94105",
"emails": ["[email protected]", "[email protected]"],
"phones": ["+14155550100"],
"social": {
"twitter": "https://twitter.com/acmeanalytics",
"linkedin": "https://linkedin.com/company/acme-analytics",
"github": "https://github.com/acme-analytics"
},
"team": [
{ "name": "Jordan Lee", "title": "VP Engineering" },
{ "name": "Maria Chen", "title": "CEO" },
{ "name": "Alex Rivera", "title": "CTO" }
],
"technologies": ["React", "Go", "Kubernetes", "PostgreSQL"]
}Verification Tools Deep Dive
Raw data is noisy. Verification tools clean, normalize, and score every contact record before it enters your pipeline. Here is exactly what each tool does and how to interpret its output.
verify_email
5-Stage PipelineThe most critical verification tool. Every email passes through five sequential stages — each stage can short-circuit the pipeline if a hard failure is detected (e.g., invalid syntax). The final confidence score is a composite of all stage results.
The 5-stage pipeline in detail
| Stage | What happens | Score impact |
|---|---|---|
| 1. Syntax | Validates against RFC 5322 subset: exactly one | Fail → 0.0 (hard stop) |
| 2. MX lookup | DNS-over-HTTPS via Google (primary) and Cloudflare (fallback). If no MX records exist, falls back to A-record lookup per RFC 5321 §5.1. | Resolved → +0.4; NXDOMAIN → 0.1 (hard stop) |
| 3. Disposable | Checks the domain against ~50 known throwaway providers: mailinator, yopmail, guerrillamail, tempmail, 10minutemail, etc. | Match → −0.3 |
| 4. Role-based | Flags departmental / noreply addresses: | Match → −0.15 (still usable, but lower priority) |
| 5. SMTP probe | Optional ( | Accepted → ≥0.95; Rejected → ≤0.2; Timeout → capped at 0.9 |
Confidence scoring algorithm
score = 0.5 # base
if mx_resolved: score += 0.4 # domain is real
if disposable: score -= 0.3 # throwaway domain
if role_based: score -= 0.15 # department mailbox
if smtp_accepted: score = max(score, 0.95) # server confirmed
if smtp_rejected: score = min(score, 0.2) # server denied
if no_smtp_check: score = min(score, 0.9) # can't fully confirm
score = clamp(score, 0.0, 1.0)
if pattern_matched: score = min(score + 0.15, 0.95) # suggest_emails bonusExample responses
{
"email": "[email protected]",
"verdict": "valid",
"confidence": 0.97,
"stages": {
"syntax": "pass",
"mx": "pass",
"disposable": "clean",
"role_based": false,
"smtp": "accepted (250 OK)"
},
"mx_records": ["alt1.gmail-smtp-in.l.google.com", "gmail-smtp-in.l.google.com"]
}{
"email": "[email protected]",
"verdict": "disposable",
"confidence": 0.05,
"stages": {
"syntax": "pass",
"mx": "pass",
"disposable": "flagged (mailinator)",
"role_based": false,
"smtp": "skipped"
}
}{
"email": "[email protected]",
"verdict": "role_based",
"confidence": 0.78,
"stages": {
"syntax": "pass",
"mx": "pass",
"disposable": "clean",
"role_based": "flagged (info@)",
"smtp": "accepted (250 OK)"
}
}verify_phone
Phone NormalizationNormalizes phone numbers using libphonenumber (the same library behind Google’s
Android dialer). Parses raw input into E.164 format, detects the country, and classifies the
line type.
What it returns
- e164 — E.164 canonical format (e.g.
+14155550123) - country — ISO-3166-1 alpha-2 country code
- line_type —
mobile,fixed_line,voip,toll_free, orunknown - valid — Boolean: does the number match the numbering plan for the detected country?
- carrier — Carrier name when available (mobile numbers only)
Example
{
"input": "8 (495) 123-45-67",
"e164": "+74951234567",
"country": "RU",
"line_type": "fixed_line",
"valid": true,
"carrier": null
}verify_social_profile
14 PlatformsValidates and enriches social profile URLs. Returns structured metadata specific to each platform — follower counts, bios, verification status, and more.
Supported platforms (14)
Example response
{
"url": "https://twitter.com/jordanlee_eng",
"platform": "twitter",
"valid": true,
"metadata": {
"name": "Jordan Lee",
"handle": "@jordanlee_eng",
"bio": "VP Engineering @AcmeAnalytics",
"followers": 4200,
"following": 890,
"verified": false,
"joined": "2019-04-22",
"location": "San Francisco, CA"
}
}suggest_emails
Pattern InferenceGiven a person’s full name and a company domain, generates candidate email addresses using the 9 most common corporate patterns. When known addresses for the same domain are available, it infers the company’s actual pattern and boosts matching candidates.
The 9 standard patterns
| # | Pattern | Example (Jordan Lee) |
|---|---|---|
| 1 | first.last | [email protected] |
| 2 | firstlast | [email protected] |
| 3 | f.last | [email protected] |
| 4 | first | [email protected] |
| 5 | last | [email protected] |
| 6 | first_last | [email protected] |
| 7 | last.first | [email protected] |
| 8 | flast | [email protected] |
| 9 | first-last | [email protected] |
Confidence scoring for suggestions
- Each candidate gets a base confidence of 0.3 — it’s just a pattern guess.
- If the inferred pattern from known addresses matches → +0.15 (capped at 0.95).
- Running
verify_emailon the suggestion can push the score up to 0.97. - Suggestions that fail MX lookup drop to 0.05 immediately.
Pattern inference is cumulative. Each time you save a verified email for a
domain, suggest_emails learns that domain’s pattern. After 3+ known addresses
for the same domain, the pattern lock is reliable enough to boost candidates to
+0.20.
Enrichment Tools Deep Dive
Once you have a verified contact, enrichment fills in the context that makes outreach personal and relevant. These two tools turn bare email addresses into full profiles.
enrich_company
Company IntelligenceTakes a company name or domain and returns structured intelligence gathered from public sources: corporate filings, about pages, job postings, technology stacks, and press coverage.
What gets discovered
- Industry — Primary and secondary SIC/NAICS categories
- Company size — Employee range (e.g. 80-120)
- Revenue — Estimated annual revenue when publicly available
- Founding year — Year of incorporation or launch
- HQ — Headquarters address
- Description — One-paragraph company description
- Technologies — Detected tech stack from job postings, headers, and scripts
- Funding — Latest funding round, amount, and investors
Example response
{
"domain": "acme-analytics.com",
"company_name": "Acme Analytics",
"industry": "B2B SaaS — Data Infrastructure",
"employee_range": "80-120",
"estimated_revenue": "$10M-$25M",
"founded": 2019,
"hq": "San Francisco, CA, USA",
"description": "Real-time data pipeline platform for B2B analytics teams.",
"technologies": ["React", "Go", "Kubernetes", "PostgreSQL", "Snowflake"],
"funding": {
"latest_round": "Series B",
"amount": "$40M",
"date": "2025-03-15",
"investors": ["Sequoia Capital", "Y Combinator"]
}
}enrich_person
Person IntelligenceTakes a name and company (or email) and discovers publicly available professional information. Cross-references LinkedIn, Twitter, GitHub, press mentions, and conference speaker lists.
What gets discovered
- Title — Current job title
- Company — Current employer
- LinkedIn — Profile URL
- Twitter / X — Profile URL and handle
- GitHub — Profile URL and notable repos
- Public contact details — Emails and phones found in public directories
- Location — City / metro area
- Bio — Summary from LinkedIn or conference bios
Example response
{
"name": "Jordan Lee",
"title": "VP Engineering",
"company": "Acme Analytics",
"location": "San Francisco, CA",
"bio": "Engineering leader with 12 years in data infrastructure. Previously at Stripe and Google.",
"linkedin": "https://linkedin.com/in/jordanlee",
"twitter": "https://twitter.com/jordanlee_eng",
"github": "https://github.com/jordanlee",
"emails": ["[email protected]"],
"phones": ["+14155550199"],
"recent_mentions": [
{
"title": "Acme Analytics Raises $40M Series B",
"source": "TechCrunch",
"date": "2025-03-15"
}
]
}Contact Management Workflow
Every contact discovered by the pipeline flows through a unified management layer backed by SQLite. The system handles deduplication, merge, export, and CRM sync — so you never lose data or create duplicates.
How contacts.db works
All contacts are stored in contacts.db, a SQLite database in your project’s data
directory. The schema is designed for fast dedup and flexible metadata.
CREATE TABLE contacts (
id TEXT PRIMARY KEY, -- UUID v4
name TEXT,
email TEXT,
phone TEXT,
title TEXT,
company TEXT,
linkedin TEXT,
twitter TEXT,
github TEXT,
website TEXT,
location TEXT,
bio TEXT,
confidence REAL DEFAULT 0.0,
source_url TEXT,
crm_id TEXT, -- set after CRM push
created_at DATETIME DEFAULT CURRENT_TIMESTAMP,
updated_at DATETIME DEFAULT CURRENT_TIMESTAMP
);
CREATE UNIQUE INDEX idx_dedup
ON contacts(email, phone, linkedin);Dedup logic
- Primary key: email + phone + LinkedIn composite. If any two of the three match, the record is merged.
- Merge strategy: non-empty fields from the newer record overwrite older empty fields; confidence takes the higher value.
- Name normalization: diacritics stripped, whitespace collapsed, case-insensitive comparison.
extract_contacts
Extracts emails, phone numbers, and social profile URLs from free text, raw HTML, or a live URL.
Handles obfuscated patterns like bob [at] acme [dot] com and JavaScript-rendered
content.
result = await agent.call_tool("extract_contacts", {
url: "https://acme-analytics.com/team"
})
result = await agent.call_tool("extract_contacts", {
text: "Reach Jordan at [email protected] or +1 (415) 555-0199"
})save_contacts
Saves extracted contacts to contacts.db with automatic dedup and merge. When CRM
integration is configured, an optional push happens after save.
# Save with auto-push to CRM
await agent.call_tool("save_contacts", { push_crm: true })
# Save without CRM push (default)
await agent.call_tool("save_contacts", {})CLI commands
| Command | Description |
|---|---|
contacts list | List all contacts in the database with optional –limit filter |
contacts export –format csv | Export contacts to CSV, vCard (.vcf), JSON, or XLSX |
contacts export –format vcard | Export as vCard 3.0 — compatible with Apple Contacts, Google Contacts, Outlook |
contacts export –format xlsx | Export as Excel workbook with separate sheets for contacts and companies |
contacts dedup | Run dedup pass on existing database; reports merge count and conflicts |
contacts push-crm | Sync all unsent contacts to the configured CRM |
CRM integration
amoCRM v4
CRM API integration. Pushes verified contacts to the configured provider and tracks
the remote crm_id for idempotent re-runs, so repeated syncs do not create
duplicate contacts. Provider credentials and account settings come from the [crm] config section.
Bitrix24 REST
Webhook-based integration. Creates contacts + leads + companies in a single batch call.
Activity timeline auto-populated with outreach events. Requires
BITRIX24_WEBHOOK_URL env variable.
HubSpot v3
Private app token integration via the CRM v3 Objects API. Creates contacts with
associations to companies. Supports batch upsert (up to 100 records per call). Custom
properties mapped via [crm].hubspot_property_map.
[crm]
provider = "amocrm" # amocrm | bitrix24 | hubspot
domain = "mycompany.amocrm.com"
api_key = "..."
# Or for HubSpot:
# provider = "hubspot"
# domain = "api.hubapi.com"
# api_key = "pat-xxx-..."Approval Flow Deep Dive
Not every tool should run without human review. The approval flow gates side-effect-causing operations — so the agent can plan freely, but it needs your explicit OK before it writes to your CRM, saves contacts, or pushes code.
Gated tools (by default)
| Tool | Why it’s gated |
|---|---|
save_contacts | Writes to contacts.db and optionally pushes to CRM — real side effects |
git_push | Pushes code to a remote repository — irreversible once published |
How to customize
Add or remove gated tools in config.toml under the [agent] section:
[agent]
# Default: save_contacts and git_push require approval
approval_tools = ["save_contacts", "git_push"]
# Add more tools to the gate:
approval_tools = ["save_contacts", "git_push", "send_email", "crm_push"]
# Remove all gates (use with caution):
approval_tools = []TUI approval workflow
- The agent pauses and displays a summary of the pending action in the terminal UI.
- You see: tool name, parameters, expected outcome, and any warnings.
- Press
yto approve,nto reject, oreto edit the parameters before approving. - Editing opens the parameter JSON in your
$EDITORfor inline modification.
API approval workflow
When running headless or integrated into CI/CD, approve via the REST API:
# Approve a pending action
curl -X POST http://localhost:8080/api/v1/sessions/sess_abc123/approve \\
-H "Content-Type: application/json" \\
-d '{"action_id": "act_789", "approved": true}'
# Reject a pending action
curl -X POST http://localhost:8080/api/v1/sessions/sess_abc123/approve \\
-H "Content-Type: application/json" \\
-d '{"action_id": "act_789", "approved": false}'Timeout and fallback behavior
If no approval arrives within the configured window, the agent uses the fallback strategy:
[agent]
# What to do when approval times out:
# "deny" — skip the action, continue the pipeline
# "allow" — auto-approve (for fully automated runs)
# "pause" — pause the session, resume later
approval_fallback = "deny"
# Timeout in seconds (default: 300 = 5 minutes)
approval_timeout_seconds = 300Never set approval_fallback = “allow” in production unless you
are running a fully automated pipeline with strict parameter validation. Auto-approve bypasses
human review for CRM writes and code pushes.
Watch Mode & Monitoring
Watch mode transforms a one-shot outreach task into a recurring monitor. Re-run the same query on a schedule, diff the results, and get notified only when new contacts appear.
The –repeat flag
Add –repeat N to re-run every N seconds. The agent runs the full
pipeline on each cycle — but only saves contacts that are new or changed since the last run.
# Re-run every 24 hours (86400 seconds)
fathom run "Find new VP Engineering hires at Series B startups in SF" \\
--repeat 86400How diff detection works
- Each run snapshots the contact database state (hash of email + phone + LinkedIn).
- Before saving, new records are compared against the previous snapshot.
- Only genuinely new or updated contacts trigger notifications.
- Removed contacts (present in snapshot, absent in new run) are logged but not deleted.
Notification channels
Webhook
POST JSON payload to any URL. Configurable via NOTIFY_WEBHOOK_URL env
variable. Payload includes new contact count, contact details, and run metadata.
Send a digest email via SMTP. Configure SMTP_HOST, SMTP_PORT,
SMTP_USER, SMTP_PASS, and NOTIFY_EMAIL_TO.
Telegram
Send a message to a Telegram chat via Bot API. Set
TELEGRAM_BOT_TOKEN and TELEGRAM_CHAT_ID env variables.
Example: daily lead monitoring setup
# Full monitoring setup
export NOTIFY_WEBHOOK_URL="https://hooks.slack.com/services/T00/B00/xxx"
export TELEGRAM_BOT_TOKEN="123456:ABC-..."
export TELEGRAM_CHAT_ID="-1001234567890"
fathom run \\
"Find newly funded startups in Berlin, extract CTO contacts, verify emails" \\
--repeat 86400Webhook payload example
{
"event": "watch_diff",
"run_id": "run_20250615_001",
"timestamp": "2025-06-15T08:00:00Z",
"new_contacts": 7,
"updated_contacts": 2,
"total_contacts": 156,
"contacts": [
{
"name": "New CTO at StartupX",
"email": "[email protected]",
"title": "CTO",
"company": "StartupX",
"confidence": 0.94
}
],
"query": "newly funded startups in Berlin"
}Real-World Workflow Examples
These end-to-end examples show how the pipeline works in practice — from a plain-language goal to verified, enriched leads ready for outreach.
Example 1: Build a targeted engineering leadership list
“Build a list of VPs of Engineering at B2B SaaS companies in the Bay Area with verified emails and LinkedIn profiles.”
Command
fathom run \\
"Find VP Engineering at B2B SaaS companies in the Bay Area. \\
Verify their emails and LinkedIn profiles. \\
Save enriched contacts with confidence >= 0.85."What happens step by step
| Step | Tool | What it does |
|---|---|---|
| 1 | find_leads | Searches 2GIS + Google Maps for “B2B SaaS” companies in SF Bay Area. Parses up to 6 corporate sites per company. Searches LinkedIn for VP Engineering titles. |
| 2 | extract_contacts | Pulls emails, phones, and LinkedIn URLs from all discovered pages. Handles obfuscated patterns. |
| 3 | verify_email | 5-stage pipeline on each email. Filters out anything below 0.85 confidence. |
| 4 | verify_social_profile | Validates each LinkedIn URL and extracts profile metadata. |
| 5 | enrich_person | Adds bio, title confirmation, Twitter/GitHub profiles, and recent mentions. |
| 6 | save_contacts | Saves to contacts.db (dedup + merge). Waits for approval before write. |
Expected output structure
{
"session_id": "sess_abc123",
"query": "VP Engineering at B2B SaaS companies in the Bay Area",
"total_companies_found": 87,
"total_contacts_extracted": 134,
"verified_contacts": 92,
"filtered_contacts": 41,
"contacts": [
{
"name": "Jordan Lee",
"title": "VP Engineering",
"company": "Acme Analytics",
"email": "[email protected]",
"email_confidence": 0.97,
"linkedin": "https://linkedin.com/in/jordanlee",
"linkedin_verified": true,
"twitter": "https://twitter.com/jordanlee_eng",
"location": "San Francisco, CA",
"bio": "Engineering leader with 12 years in data infrastructure.",
"company_size": "80-120 employees",
"company_funding": "Series B ($40M)"
}
]
}Example 2: Monitor competitor hiring
“Monitor competitor hiring and alert me when they post new senior engineering roles.”
Command
fathom run \\
"Monitor career pages and job boards for senior engineering roles \\
at CompetitorA, CompetitorB, CompetitorC. \\
Alert on new VP/Director/Staff-level postings." \\
--repeat 43200What happens step by step
| Step | Tool | What it does |
|---|---|---|
| 1 | parse_corporate_site | Crawls /careers and /jobs pages of each competitor. Extracts job titles, locations, and posting dates. |
| 2 | search_news | Searches for recent hiring announcements and press coverage mentioning the competitors. |
| 3 | Diff engine | Compares current findings against the previous run’s snapshot. Identifies new postings. |
| 4 | Notify | Sends a Telegram message with each new senior role: title, location, posting URL, and date. |
Expected notification payload
{
"event": "watch_diff",
"new_postings": [
{
"company": "CompetitorA",
"title": "Director of Platform Engineering",
"location": "San Francisco, CA",
"url": "https://competitora.com/careers/director-platform",
"posted_date": "2025-06-14",
"detected_at": "2025-06-15T08:00:00Z"
}
],
"monitoring_since": "2025-06-01T00:00:00Z",
"total_runs": 30
}Example 3: Enrich an existing CSV of contacts
“I have a CSV of 200 contacts with names and companies. Enrich them with verified emails and LinkedIn profiles.”
Command
# Step 1: Import the CSV
fathom contacts list --limit 50
# Step 2: Enrich each contact
fathom run \\
"For each contact in the database that is missing an email, \\
find their professional email and LinkedIn profile. \\
Verify emails before saving." \\
--source-db contacts.db
# Step 3: Export enriched results
fathom contacts export --format csv -o ./enriched_contactsWhat happens step by step
| Step | Tool | What it does |
|---|---|---|
| 1 | Import | Parses CSV, maps columns, and inserts into contacts.db with dedup. |
| 2 | suggest_emails | For each contact without an email, generates 9 pattern candidates using name + company domain. |
| 3 | verify_email | Runs 5-stage verification on each candidate. Keeps the highest-confidence match above 0.8. |
| 4 | search_social | Searches LinkedIn for the person’s name + company. Validates and enriches the profile URL. |
| 5 | enrich_person | Fills in bio, title confirmation, location, and additional social profiles. |
| 6 | save_contacts | Merges enriched data into existing contact records. Updates only changed fields. |
| 7 | Export | Exports the enriched database to CSV with all new fields included. |
Expected output structure
name,email,email_confidence,title,company,linkedin,twitter,location,bio,source
Jordan Lee,[email protected],0.97,VP Engineering,Acme Analytics,https://linkedin.com/in/jordanlee,https://twitter.com/jordanlee_eng,"San Francisco, CA","Engineering leader with 12 years...",suggest_emails+verify
Maria Chen,[email protected],0.91,CTO,StartupX,https://linkedin.com/in/mariachen,,Berlin,"Former Google, building developer tools...",search_social+verifyPattern learning compounds. After enriching 200 contacts,
suggest_emails has learned the email patterns for dozens of domains. Future
enrichment runs for the same companies will be significantly faster and more accurate.