AGENT 04 / 07
Structuring
The structuring worker turns extracted fragments into canonical records — normalized, deduplicated, and linked in an entity graph so every worker can rely on the same context.
structuring · autonomous
01What it does
Canonical records from scattered fragments
Persist, never duplicate
Writes structured records into SQLite or PostgreSQL, deduplicating against existing data: the same normalized key means a merge, not a second row. Optional integrations can receive the canonical record.
The model cannot forget
Prompt-only steering can lose a handoff when a model skips a save call. The runtime persists metadata from successful extraction and discovery tools itself, so structured results survive the next turn.
Five tables, one dossier
contacts, social_profiles, companies, tags and notes — with source provenance, timestamps and a crm_id column for idempotent CRM sync. WAL mode, foreign keys, indexed lookups.
One spelling per contact
normalize_email trims and lower-cases ("[email protected]" → [email protected]); normalize_phone keeps ASCII digits only, stored in the indexed phone_norm column. These are the dedup keys.
Records grow, never split
The older record wins: blank fields are filled from the new data, social profiles move over with NOT EXISTS dedup, tags append via UNIQUE(contact_id, tag), notes accumulate — the duplicate row is deleted.
Entities and relationships
Entities deduplicate by (name, type); typed relations connect people, projects, systems, and organizations. Multi-hop BFS up to depth 4 supports context-aware downstream work.
02How it works
From tool output to canonical contact
Trigger
An extraction or discovery tool returns structured metadata with confidence and provenance for the current mission.
Autosave
The runtime persists structured items right after success — compaction cannot prune the handoff, and the model cannot forget the result.
Normalize
Emails lower-cased, phones reduced to digits, provenance recorded: extract_contacts:mailto link:https://acme.io.
Find-or-insert
One atomic transaction: look up by normalized email, then phone_norm. Match → merge; no match → insert. Concurrent agents cannot double-insert.
Graph
People, projects, systems, and organizations upsert as nodes; relations become typed edges with confidence for context-aware downstream work.
Sync
Contacts push to amoCRM, Bitrix24 or HubSpot four at a time with retry; the returned crm_id is stored so re-syncs update instead of duplicating.
03Under the hood
Schema, keys and merge rules
| Schema | contacts(id, email, phone, phone_norm, name, title, company, source, crm_id, created_at, updated_at) · social_profiles · companies · tags · notes |
| Storage | SQLite with WAL + busy_timeout 5000 out of the box · PostgreSQL behind the same interface via [contacts] pg_url |
| Dedup keys | normalized email first, then digits-only phone_norm · lowest id wins — the oldest record survives |
| Merge | blanks filled from new data · socials moved with NOT EXISTS · tags INSERT OR IGNORE · notes appended · temp row deleted, one transaction |
| Atomicity | find-or-insert runs in a single locked transaction — concurrent workers cannot both insert the same normalized record (TOCTOU-safe) |
| Autosave | fires on extract_contacts / find_leads success regardless of the model · "[auto-persisted N contact(s): X new, Y merged]" |
| Provenance | source tags identify the producing tool and source URL · confidence is retained as record metadata for review |
| Lenient input | tags/notes accept a JSON array or a plain string; "; " and newlines split blobs — LLM shape drift never loses data |
| Graph dedup | entity_nodes UNIQUE(name COLLATE NOCASE, entity_type) · edges UNIQUE(from, to, relation), default confidence 0.8 |
| Multi-hop | BFS depth clamped 1–4, simple paths only, inverse traversal marked ⁻¹ |
| CRM push | buffer_unordered(4) concurrency · one retry after 1 s · set_crm_id for idempotent re-pushes |
04Tools
What the agent runs
Feeds downstream workers
A canonical record is shared context
Structured records give every worker the same reliable subject, source links, tags, and notes. A coding, research, or operations task can reuse the canonical entity without re-parsing every source or creating duplicates.
save_contacts([{
email: "[email protected]", # → [email protected]
name: "Ann Petrova",
title: "CTO",
company:"Acme",
socials:[{ platform: "linkedin", url: "…" }],
tags: ["saas", "decision-maker"]
}])
→ added: 0 · merged_with_existing: 1
# one canonical record =
# one consistent downstream handoff