Memory that compounds
Facts are absorbed into a persistent, searchable knowledge base — so every autonomous worker session can start from what you already know instead of from zero.
Where facts live
| Scope | Purpose |
|---|---|
agent | Knowledge shared across all of the agent’s sessions — the persistent brain. |
user | Facts about the client you are researching or selling to. |
run | One-session episode facts, promoted to agent scope by distillation. |
Search that blends semantics and keywords
Memory search is a hybrid of semantic (vector) similarity and BM25 keyword matching:
score = 0.7·cosine_similarity + 0.3·bm25_normalized
final = score × max(0, 1 − 0.01·days_old)- Semantic weight 0.7 balances meaning over exact-match.
- Temporal decay 0.01/day slowly fades older facts — recent wins.
- top_k 5, min_score 0.25 keep results tight and relevant.
- Offline TF-IDF embeddings (512-dim) when no API key is set — no network required.
- Entity knowledge graph (memory_graph) for persons, companies, locations and typed relations
How facts get in
Every discovered fact passes through the absorb pipeline before it can enter memory.
Only self-contained facts (3–5000 chars) qualify. Very short, low-value fragments are rejected.
API keys, tokens and passwords are dropped before anything is stored.
Near-duplicate facts (cosine ≥ 0.85) merge into one canonical record.
The LLM (or heuristics) picks one of six outcomes: duplicate, supersede, contradict, coexist, related, or new.
Superseded rows are archived with a typed edge; contradictions both survive visibly.
Facts that change, tracked
Nothing is overwritten. When a fact is superseded, a new row is created and the old one is
archived behind a typed supersedes edge. contradicts keeps both
sides visible, and related_to links connected knowledge. The entity graph makes
role changes, funding events and tech migrations first-class facts your outreach can cite.
Memory shows up automatically
At depth 0 the top agent receives a memory digest in its system prompt: relevant memories, open TODOs, and recent records. Your outreach never repeats a touch, never asserts a stale role, and always builds on what you already gathered.
Grows, then tidies itself
An automated garbage collector archives stale run-scoped facts after 30 days and compacts
crowded (scope, key) groups of 200+ active rows into one without hard-deleting them.
The explicit fathom memory nuke –yes command is destructive and is not recoverable.
Distillation
Distillation promotes run-scoped episode facts into durable agent-scope knowledge. When a research
session completes, the distiller re-absorbs its run-facts through the full pipeline (validate, dedup,
classify) but targets the agent scope instead of the run scope. Originals are archived. This way,
facts discovered during a single session become permanent knowledge available to all future sessions
— without manual copying. Run fathom memory distill –session <key> or
let gc_auto = true run hourly memory maintenance in the background (including GC and distillation); this is separate from the destructive nuke command.
Memory Tools Reference
Seven specialized tools let agents read, write, search, link, and reason over memory.
File-backed memory is managed by the memory tool; semantic-memory operations have separate CLI and REST interfaces.
memory — File-backed CRUD
The memory tool manages file-backed context in MEMORY.md and USER.md.
It applies text operations to the selected target file; semantic records are managed separately by
memory_absorb, memory_search, and the fathom memory CLI.
| Action | Description | Example |
|---|---|---|
add | Append content to MEMORY.md or USER.md. | action=add target=memory content=… |
replace | Replace matching existing text. | action=replace old_text=… content=… |
remove | Remove matching existing text. | action=remove old_text=… |
batch | Apply a sequence of add, replace, or remove operations. | action=batch operations=[…] |
memory_absorb — Fact Absorption
The recommended way to store new knowledge. memory_absorb feeds raw text through the
full absorb pipeline — validation, secret scanning, consolidation, classification, and edge creation —
so you never store duplicates or secrets accidentally.
memory_absorb({
scope: "agent",
fact: "Stripe acquired Lemon Squeezy in March 2024 for merchant-of-record capabilities"
})The pipeline executes these steps in order:
Length must be 3–5000 characters. Fragments under 3 chars or gibberish are rejected with ERR_TOO_SHORT.
Regex patterns detect API keys, JWTs, SSH private keys, and passwords. Matched tokens are redacted before storage.
Embeddings are computed and compared against existing facts. Cosine similarity ≥ 0.85 triggers a merge into the canonical record.
LLM (or heuristic fallback) decides: new, duplicate, supersede, contradict, coexist, or related.
New rows are inserted; superseded rows get a supersedes edge; contradictions both survive with a contradicts edge.
Typical response:
{
"status": "absorbed",
"memory_id": "mem_8f3a2b",
"classification": "new",
"embedding_model": "tfidf-512",
"edges_created": 0,
"duplicates_merged": 0
}memory_search — Hybrid Search
Combines vector similarity (semantic) with BM25 keyword matching for robust retrieval.
Results are ranked by a weighted score and filtered by min_score.
memory_search({ scope: "agent", query: "Stripe pricing tiers", top_k: 5 })
memory_search({ scope: "user", query: "CEO name and background", top_k: 3 })
memory_search({ scope: "all", query: "competitor analysis fintech", min_score: 0.3 })Scoring formula in detail:
raw_score = semantic_weight × cosine(embedding_q, embedding_doc)
+ (1 − semantic_weight) × bm25_normalized(query, doc)
temporal = max(0, 1 − temporal_decay × days_since_last_update)
final_score = raw_score × temporal × importance_boostExample search result:
{
"results": [
{
"id": "mem_8f3a2b",
"content": "Stripe acquired Lemon Squeezy in March 2024",
"scope": "agent",
"score": 0.87,
"semantic_component": 0.91,
"bm25_component": 0.78,
"temporal_factor": 0.96,
"importance": 1.0,
"updated_at": "2024-03-15T10:30:00Z"
}
],
"total": 1,
"query_time_ms": 12
}memory_digest — Context Digest
Generates a condensed memory block suitable for injection into system prompts. The digest pulls the most relevant memories for the current context, open TODOs, and recent records.
memory_digest({
scope: "agent",
topic: "outreach to fintech companies",
max_tokens: 2000
})The digest includes three sections:
- Relevant memories — top-K facts matching the topic, ranked by score.
- Open TODOs — any records tagged as TODO that are still active.
- Recent records — facts absorbed in the last 48 hours, regardless of topic.
## Memory Digest — "outreach to fintech companies"
### Relevant Facts
- [agent] Stripe acquired Lemon Squeezy (MoR) in March 2024 (score: 0.87)
- [agent] Plaid raised Series D at $13.4B valuation (score: 0.72)
- [user] Target contact: Jane Doe, VP Engineering at Plaid (score: 0.68)
### Open TODOs
- [run] Draft personalized email for Jane Doe referencing Plaid's API-first approach
### Recent (last 48h)
- [run] Plaid announced new Transfer product on 2024-03-14memory_boost — Adjust Importance
Lets you increase or decrease the importance multiplier of a specific memory record.
Boosted facts rank higher in search results and digests. Importance is a float multiplier
applied during scoring — default is 1.0.
memory_boost({
id: "mem_8f3a2b",
importance: 2.0,
reason: "Critical acquisition — cite in all Stripe outreach"
})How importance affects ranking:
| Importance | Effect | Use Case |
|---|---|---|
0.5 | Half-weight — deprioritized in results | Background context, nice-to-know |
1.0 | Normal weight (default) | Standard facts |
1.5 | 50% boost in ranking | Important client preferences |
2.0 | Double weight — always surfaces first | Critical facts, must-cite knowledge |
3.0 | Triple weight — pinned to top | Deal-breaking constraints, legal requirements |
memory_link — Typed Edges
Creates a directed, typed edge between two memory records. Edges let you express relationships like “this fact supersedes that one” or “these two facts are related.”
| Edge Type | Description | Example |
|---|---|---|
related_to | General semantic connection between two facts. | “Stripe pricing” ↔ “Stripe MoR acquisition” |
supersedes | Newer fact replaces an older one. Old is archived. | “Series D at $13.4B” supersedes “Series C at $8B” |
contradicts | Two facts conflict — both kept visible for review. | “Q1 revenue $5M” contradicts “Q1 revenue $3M” |
implements | One fact is a concrete implementation of another. | “Used Stripe Checkout v2” implements “Migrated to Stripe” |
extends | One fact adds detail or context to another. | “Jane is ex-Google” extends “Jane is VP Engineering” |
references | One fact cites or references another as a source. | “Market size $50B” references “Gartner 2024 report” |
memory_link({
from_id: "mem_new_series_d",
to_id: "mem_old_series_c",
edge_type: "supersedes",
note: "Valuation updated after Series D announcement"
})memory_graph — Entity Knowledge Graph
Manages a lightweight entity-relationship graph alongside the memory store. Entities represent real-world things (people, companies, locations) and relations connect them with typed edges.
Three actions are available:
memory_graph({
action: "add",
entity: {
type: "person",
name: "Jane Doe",
attributes: { role: "VP Engineering", company: "Plaid" }
}
})
memory_graph({
action: "add",
entity: {
type: "company",
name: "Plaid",
attributes: { industry: "fintech", hq: "San Francisco" }
}
})
memory_graph({
action: "add",
entity: {
type: "location",
name: "San Francisco",
attributes: { country: "US", state: "CA" }
}
})memory_graph({
action: "add",
relation: {
from: "person:Jane Doe",
to: "company:Plaid",
type: "works_at",
attributes: { since: "2021" }
}
})
memory_graph({
action: "add",
relation: {
from: "company:Plaid",
to: "location:San Francisco",
type: "located_in"
}
})memory_graph({
action: "query",
start: "location:San Francisco",
relation: "located_in",
direction: "inbound",
hops: 2
})
// Returns: all people who work at companies located in San Francisco
memory_graph({
action: "list",
entity_type: "company",
limit: 20
})Entity Graph Deep Dive
The entity graph sits alongside the memory store and provides structured, queryable relationships between real-world entities. Unlike free-text memories, the graph enforces typed nodes and edges so you can traverse relationships programmatically.
Entity Types
| Type | Description | Common Attributes |
|---|---|---|
person | Individual humans — contacts, founders, executives. | role, company, email, linkedin |
company | Organizations — prospects, competitors, partners. | industry, hq, size, founded |
location | Geographic places — cities, countries, regions. | country, state, timezone |
Building a Relationship Map
Entities are connected through typed relations. The graph supports directional edges,
so person → works_at → company is distinct from company → employs → person.
person:Jane Doe ──works_at──▶ company:Plaid
company:Plaid ──located_in──▶ location:San Francisco
person:Jane Doe ──knows──▶ person:John Smith
company:Plaid ──competes_with──▶ company:Stripe
company:Stripe ──acquired──▶ company:Lemon SqueezyMulti-Hop Queries
The real power of the graph emerges with multi-hop traversal. You can ask questions
like “find all people who work at companies headquartered in Berlin” by following
the works_at and located_in edges across two hops.
// Hop 1: Find companies in Berlin
// location:Berlin ←──located_in── company:N26
// location:Berlin ←──located_in── company:Delivery Hero
// Hop 2: Find people at those companies
// company:N26 ←──works_at── person:Max Tayenthal
// company:Delivery Hero ←──works_at── person:Niklas Östberg
// Result: [Max Tayenthal, Niklas Östberg]Entity Deduplication
When the same entity is referenced from different sources (e.g., “Plaid” from a news article and “Plaid Inc.” from a CRM export), the graph merges them by normalized name and entity type. Review source context and conflicting attributes before export.
- Name normalization — Whitespace and case normalization identify matching names by entity type.
- Attribute merging — Conflicting attributes remain visible for review while non-conflicting values can be combined.
- Edge re-wiring — Relations attached to a reused node remain connected to the canonical entity.
Graph Visualization Concept
The graph can be visualized as a node-link diagram where each entity is a node and each relation is a typed, directed edge. Node size reflects connectivity (degree centrality) and edge thickness can encode relation strength or recency.
nodes = graph.list_entities()
edges = graph.list_relations()
for node in nodes:
draw_circle(
x=node.x, y=node.y,
radius=2 + log(node.degree),
color=type_colors[node.type], // person=blue, company=green, location=orange
label=node.name
)
for edge in edges:
draw_arrow(
from=edge.from, to=edge.to,
style=edge_type_styles[edge.type], // solid, dashed, dotted
label=edge.type
)Memory Configuration Reference
All memory behavior is controlled through the [memory] section of your
config.toml configuration file. Every field has a sensible default
so you only need to override what matters for your use case.
| Field | Type | Default | Description |
|---|---|---|---|
enabled | bool | true | Master switch. When false, all memory tools return stub responses and nothing is persisted. |
db_path | string | ~/.fathom/memory.db | Path to the SQLite database file. Created automatically on first run. |
embeddings | string | “auto” | Embedding backend: auto (OpenAI if key present, else TF-IDF), openai, or tfidf. |
semantic_weight | float | 0.7 | Weight for semantic similarity in hybrid search. BM25 gets 1 − semantic_weight. |
top_k | int | 5 | Number of results returned by memory_search and memory_digest. |
min_score | float | 0.25 | Minimum final score for a result to be included. Filters noise. |
temporal_decay | float | 0.01 | Daily decay rate applied to older facts. At 0.01, a 100-day-old fact scores at 0.0 of its original. |
auto_digest | bool | true | Automatically inject a memory digest into the top-level agent’s system prompt. |
llm_classify | bool | true | Use the LLM for absorb classification. When false, falls back to heuristic rules. |
rerank | bool | false | Enable a second-pass LLM reranker on search results. Adds latency but improves precision. |
gc_auto | bool | false | Run hourly background memory maintenance (GC and distillation) while sessions execute. |
gc_ttl_days | int | 30 | Run-scoped facts older than this are archived by the GC. |
gc_compact_above | int | 200 | When a (scope, key) group exceeds this many active rows, GC compacts them. |
Field Details
enabled — The master kill switch. Set to false to disable all memory features without removing configuration. Useful for testing or CI environments where persistent state is undesirable.
db_path — SQLite is used for zero-dependency local storage. The database contains tables for memories, embeddings, entity graph nodes, entity graph edges, and metadata. You can point this to a shared volume for multi-machine setups, though SQLite’s single-writer model may require WAL mode.
embeddings — Controls how text is vectorized for semantic search. auto checks for an OPENAI_API_KEY environment variable; if found, it uses text-embedding-3-small (1536-dim). Otherwise it falls back to a local TF-IDF model (512-dim) that works entirely offline. The tfidf and openai options force one backend regardless of environment.
semantic_weight — Balances meaning vs. exact keywords. A higher value (e.g., 0.85) favors semantic understanding — good for natural-language queries. A lower value (e.g., 0.5) favors keyword matching — better for exact product names, acronyms, or codes. The default 0.7 works well for mixed workloads.
top_k and min_score — Control result volume and quality. top_k caps how many results come back; min_score filters out anything below the threshold. Raise min_score to 0.4+ for precision-critical tasks like legal or compliance research.
temporal_decay — Makes freshness matter. The formula is max(0, 1 − decay × days_old). At the default 0.01, a fact is worth 90% after 10 days, 50% after 50 days, and 0% after 100 days. Set to 0.0 to disable decay entirely (all facts treated equally regardless of age).
auto_digest — When enabled, the system automatically generates and injects a memory digest into the top-level agent’s system prompt before each run. This means the agent always has relevant context without explicit tool calls.
llm_classify — Controls whether the absorb pipeline uses an LLM to classify incoming facts (new, duplicate, supersede, contradict, coexist, related) or falls back to embedding-distance heuristics. The LLM path is more accurate but adds one API call per absorb.
rerank — When enabled, search results go through a second pass where the LLM reorders them by relevance to the original query. This adds latency (one extra LLM call) but can significantly improve result quality for complex or ambiguous queries.
gc_* fields — The garbage collector handles two tasks: (1) archiving run-scoped facts older than gc_ttl_days, and (2) compacting (scope, key) groups that exceed gc_compact_above active rows into a single canonical record. Both are soft operations — archived data is always recoverable.
Example Configuration
[memory]
enabled = true
db_path = "~/.fathom/memory.db"
embeddings = "auto"
semantic_weight = 0.7
top_k = 5
min_score = 0.25
temporal_decay = 0.01
auto_digest = true
llm_classify = true
rerank = false
gc_ttl_days = 30
gc_compact_above = 200
gc_auto = falseCLI Memory Commands
All memory operations are accessible from the command line via fathom memory.
These commands are useful for debugging, manual curation, and automation scripts.
Search
fathom memory search "Stripe pricing tiers" --top-k 10 --scope agent
fathom memory search "competitor analysis" --scope all
fathom memory search "CEO background" --scope userThe search command runs hybrid vector + BM25 search and prints ranked results.
Use –format json for programmatic consumption or –format table (default)
for human-readable output.
List
fathom memory list --scope all --status active
fathom memory list --scope agent --status archived -n 50
fathom memory list --scope runLists records with optional filtering by scope, status (active, superseded, archived, all),
and a result limit. Results are ordered by most recently updated.
Get
fathom memory get mem_8f3a2b
fathom memory get mem_8f3a2b --follow full_history
fathom memory get mem_8f3a2b --follow latestFetches a single record by ID. The –follow flag controls how much related data
to include: none (default), edges (typed links), or
full_history (all superseded versions and edges).
Stats
fathom memory statsPrints a summary of the memory store: total records by scope and status, entity graph size, edge counts by type, embedding model in use, and database file size.
┌─────────────────────────────────────┐
│ Memory Store Statistics │
├─────────────────────────────────────┤
│ Active records: 1,247 │
│ agent scope: 892 │
│ user scope: 203 │
│ run scope: 152 │
│ Archived records: 389 │
│ Entity graph: │
│ nodes: 156 │
│ edges: 412 │
│ Embedding model: tfidf-512 │
│ DB size: 14.2 MB │
└─────────────────────────────────────┘Rebuild
fathom memory rebuildRebuilds the memory index from scratch. Useful after changing embedding models or if the index becomes corrupted. The command re-embeds all stored memories with the current model.
Distill
fathom memory distill --session run_abc123
fathom memory distill --session run_abc123 --dry-run
fathom memory distill --session run_abc123Promotes run-scoped facts from a completed session into the agent scope. Use –dry-run
to preview which facts would be promoted without actually writing. The destination is the persistent agent scope.
Garbage Collection
fathom memory gc
fathom memory gc --ttl-days 30 --dry-run
fathom memory gc --compact-above 200Manually triggers garbage collection. –dry-run shows what would be archived or
compacted without making changes. Use –ttl-days to override the default retention
period for run-scoped facts.
Nuke
fathom memory nuke --scope run --yes
fathom memory nuke --scope all --yesPermanently deletes all records in the specified scope. Requires –yes as a safety
confirmation. This is the destructive command; GC and archive operations are recoverable.
Use –scope all to wipe the entire memory store — useful for starting fresh
in development environments.
Memory API Endpoints
When the agent is running as a server, all memory operations are exposed via a REST API. Endpoints return JSON and follow standard HTTP status codes.
List & Search Memories
GET /api/v1/memories?scope=agent&q=stripe+pricing&top_k=5&min_score=0.25| Parameter | Type | Description |
|---|---|---|
scope | string | Filter by scope: agent, user, run, or all. |
q | string | Search query. When present, runs hybrid search. When absent, returns a plain list. |
top_k | int | Max results. Defaults to config top_k. |
min_score | float | Minimum score threshold. Defaults to config min_score. |
status | string | active (default), archived, or all. |
key | string | Filter by key prefix. |
{
"memories": [
{
"id": "mem_8f3a2b",
"scope": "agent",
"key": "stripe-acquisition",
"content": "Stripe acquired Lemon Squeezy in March 2024",
"importance": 1.0,
"score": 0.87,
"status": "active",
"created_at": "2024-03-15T10:30:00Z",
"updated_at": "2024-03-15T10:30:00Z"
}
],
"total": 1,
"page": 1
}Absorb Facts
POST /api/v1/memories/absorb
Content-Type: application/json
{
"scope": "agent",
"fact": "Stripe acquired Lemon Squeezy in March 2024 for merchant-of-record capabilities"
}{
"status": "absorbed",
"memory_id": "mem_8f3a2b",
"classification": "new",
"embedding_model": "tfidf-512",
"edges_created": 0,
"duplicates_merged": 0
}Store Statistics
GET /api/v1/memories/stats{
"active_count": 1247,
"archived_count": 389,
"by_scope": {
"agent": 892,
"user": 203,
"run": 152
},
"entity_graph": {
"nodes": 156,
"edges": 412
},
"embedding_model": "tfidf-512",
"db_size_bytes": 14889768
}Distill Run Facts
POST /api/v1/memories/distill
Content-Type: application/json
{
"session_key": "run_abc123",
"target_scope": "agent",
"dry_run": false
}{
"promoted": 12,
"archived_originals": 12,
"duplicates_found": 3,
"new_memories_created": 9
}Garbage Collection
POST /api/v1/memories/gc
Content-Type: application/json
{
"ttl_days": 30,
"compact_above": 200,
"dry_run": true
}{
"would_archive": 47,
"would_compact": 2,
"dry_run": true
}Get Single Memory
GET /api/v1/memories/mem_8f3a2b?follow=full_history{
"id": "mem_8f3a2b",
"scope": "agent",
"key": "stripe-acquisition",
"content": "Stripe acquired Lemon Squeezy in March 2024",
"importance": 1.0,
"status": "active",
"edges": [
{ "type": "related_to", "target_id": "mem_7c2d1e", "note": "Stripe ecosystem" }
],
"history": [
{
"version": 1,
"content": "Stripe reportedly acquiring Lemon Squeezy",
"superseded_by": "mem_8f3a2b",
"created_at": "2024-02-20T08:00:00Z"
}
],
"created_at": "2024-03-15T10:30:00Z",
"updated_at": "2024-03-15T10:30:00Z"
}Archive (Soft Delete)
DELETE /api/v1/memories/mem_8f3a2b{
"id": "mem_8f3a2b",
"status": "archived",
"archived_at": "2024-03-20T14:00:00Z"
}This is a soft delete — the record is moved to archived status and excluded from default
queries. The archived row remains available through the memory history and CLI maintenance commands; the HTTP API exposes archive, not restore.
Hard deletion is only available via the CLI nuke command.
Best Practices
File-backed vs. Semantic Memory
The memory tool (file-backed CRUD) is best for deterministic operations —
when you know the exact ID, need to update specific fields, or want to manage record
lifecycle directly. The semantic tools (memory_search, memory_absorb,
memory_digest) are best for discovery and knowledge ingestion — when you need
to find relevant context, absorb new information, or generate summaries.
Rule of thumb: Use memory for CRUD operations on known records.
Use memory_absorb to store new facts and memory_search to find
existing ones. Use memory_digest when you need a pre-assembled context block.
Structuring Facts for Search
Well-structured facts produce better search results. Follow these guidelines:
- Be self-contained — A fact should make sense without surrounding context. “The Series D was at $13.4B” is bad; “Plaid raised Series D at $13.4B valuation in April 2024” is good.
- Include key entities — Mention company names, people, and dates explicitly. Embeddings and BM25 both rely on these signals.
- One fact per record — Don’t bundle multiple unrelated facts. Each record should capture a single atomic piece of knowledge.
- Use descriptive keys — A key like
plaid-series-d-valuationis more useful thanfact-42for debugging and manual review. - Avoid duplicates — The absorb pipeline catches near-duplicates (cosine ≥ 0.85), but semantically different phrasings of the same fact may slip through. Rephrase consistently.
Distillation vs. Manual Management
Let distillation handle run-to-agent promotion automatically for most workloads. It ensures
only meaningful facts survive and prevents memory bloat from transient session data.
Use manual memory add when you need to seed the agent with baseline knowledge
before a campaign — things like company profiles, product catalogs, or competitive intelligence
that the agent should always know.
- Auto-distill — Enable
gc_auto = trueto run the hourly maintenance loop that performs GC and distillation. Best for continuous workflows. - Manual distill — Use
–dry-runfirst to review what would be promoted, then run without the flag. Best for high-stakes campaigns where you want control. - Seed knowledge — Use
memory_absorbwith scopeagentto pre-load facts before running research. This gives the agent a head start.
Tuning Garbage Collection
GC keeps the memory store healthy over time. The right settings depend on your usage pattern:
| Workload | gc_ttl_days | gc_compact_above | Rationale |
|---|---|---|---|
| High-volume research | 14 | 100 | Aggressive cleanup — lots of transient run data. |
| Standard outreach | 30 | 200 | Balanced — the default works for most teams. |
| Long-term accounts | 90 | 500 | Conservative — keep history for relationship building. |
| Compliance / audit | 365 | 1000 | Minimal cleanup — retain everything for audit trails. |
Memory Hygiene
Periodic maintenance keeps your memory store accurate and useful:
- Review contradictions — Run
fathom memory list –scope agentperiodically and look for records linked bycontradictsedges. Resolve them by boosting the correct one and archiving the stale one. - Audit importance scores — Facts with high importance that are no longer relevant will dominate digests. Use
memory_boostto dial them back when context changes. - Prune stale run facts — If
gc_autois disabled, runfathom memory gc –dry-runweekly to check for accumulation. - Rebuild after model changes — Switching embedding models invalidates existing vectors. Run
fathom memory rebuildafter changing theembeddingsconfig. - Monitor store size — Use
fathom memory statsto track growth. Store size depends on workload; usefathom memory statsto monitor growth in your deployment. - Back up the database — The SQLite file at
db_pathcontains everything. Include it in your backup rotation, especially before runningnuke.