AGENT 04 / 07

Structuring

The structuring worker turns extracted fragments into canonical records — normalized, deduplicated, and linked in an entity graph so every worker can rely on the same context.

structuringfathom worker

structuring · autonomous

01What it does

Canonical records from scattered fragments

save_contacts01

Persist, never duplicate

Writes structured records into SQLite or PostgreSQL, deduplicating against existing data: the same normalized key means a merge, not a second row. Optional integrations can receive the canonical record.

deterministic autosave02

The model cannot forget

Prompt-only steering can lose a handoff when a model skips a save call. The runtime persists metadata from successful extraction and discovery tools itself, so structured results survive the next turn.

ContactDb schema03

Five tables, one dossier

contacts, social_profiles, companies, tags and notes — with source provenance, timestamps and a crm_id column for idempotent CRM sync. WAL mode, foreign keys, indexed lookups.

normalization04

One spelling per contact

normalize_email trims and lower-cases ("[email protected]" → [email protected]); normalize_phone keeps ASCII digits only, stored in the indexed phone_norm column. These are the dedup keys.

merge semantics05

Records grow, never split

The older record wins: blank fields are filled from the new data, social profiles move over with NOT EXISTS dedup, tags append via UNIQUE(contact_id, tag), notes accumulate — the duplicate row is deleted.

entity graph06

Entities and relationships

Entities deduplicate by (name, type); typed relations connect people, projects, systems, and organizations. Multi-hop BFS up to depth 4 supports context-aware downstream work.

02How it works

From tool output to canonical contact

1

Trigger

An extraction or discovery tool returns structured metadata with confidence and provenance for the current mission.

2

Autosave

The runtime persists structured items right after success — compaction cannot prune the handoff, and the model cannot forget the result.

3

Normalize

Emails lower-cased, phones reduced to digits, provenance recorded: extract_contacts:mailto link:https://acme.io.

4

Find-or-insert

One atomic transaction: look up by normalized email, then phone_norm. Match → merge; no match → insert. Concurrent agents cannot double-insert.

5

Graph

People, projects, systems, and organizations upsert as nodes; relations become typed edges with confidence for context-aware downstream work.

6

Sync

Contacts push to amoCRM, Bitrix24 or HubSpot four at a time with retry; the returned crm_id is stored so re-syncs update instead of duplicating.

03Under the hood

Schema, keys and merge rules

Schemacontacts(id, email, phone, phone_norm, name, title, company, source, crm_id, created_at, updated_at) · social_profiles · companies · tags · notes
StorageSQLite with WAL + busy_timeout 5000 out of the box · PostgreSQL behind the same interface via [contacts] pg_url
Dedup keysnormalized email first, then digits-only phone_norm · lowest id wins — the oldest record survives
Mergeblanks filled from new data · socials moved with NOT EXISTS · tags INSERT OR IGNORE · notes appended · temp row deleted, one transaction
Atomicityfind-or-insert runs in a single locked transaction — concurrent workers cannot both insert the same normalized record (TOCTOU-safe)
Autosavefires on extract_contacts / find_leads success regardless of the model · "[auto-persisted N contact(s): X new, Y merged]"
Provenancesource tags identify the producing tool and source URL · confidence is retained as record metadata for review
Lenient inputtags/notes accept a JSON array or a plain string; "; " and newlines split blobs — LLM shape drift never loses data
Graph dedupentity_nodes UNIQUE(name COLLATE NOCASE, entity_type) · edges UNIQUE(from, to, relation), default confidence 0.8
Multi-hopBFS depth clamped 1–4, simple paths only, inverse traversal marked ⁻¹
CRM pushbuffer_unordered(4) concurrency · one retry after 1 s · set_crm_id for idempotent re-pushes

04Tools

What the agent runs

save_contactsextract_contactsfind_leadsmemory

Feeds downstream workers

A canonical record is shared context

Structured records give every worker the same reliable subject, source links, tags, and notes. A coding, research, or operations task can reuse the canonical entity without re-parsing every source or creating duplicates.

save_contacts([{
  email:  "[email protected]",   # → [email protected]
  name:   "Ann Petrova",
  title:  "CTO",
  company:"Acme",
  socials:[{ platform: "linkedin", url: "…" }],
  tags:   ["saas", "decision-maker"]
}])
→ added: 0 · merged_with_existing: 1

# one canonical record =
# one consistent downstream handoff

Every downstream task needs canonical context