Level 1 gains phase 5, duplicate-node detection: runs memory_dup_detect.py (read-only) and reports clusters by verdict. - WRITER-DEFECT (run_family): one task creating a node per run -> writer fix, never merge (per-run nodes are the audit trail) - SAFE-MERGE: identical bodies, still needs an explicit Kwame decision - HUMAN-DECISION: same subject, bodies differ -> connect, never merge Replaces the naive '>70% title overlap' duplicate rule, which false-positives on distinct work: four client workflows (#357-#361) and two machines' migrations (#1792/#1793) score high on titles with bodies 0.1-0.3 apart. Also corrected in the same pass: - the 'canonical copy' pointer named /root/.hermes/contracts/memory-fixer-v3.md, which does not exist on kagentz (no /root access); the job's instruction set is inline in ~/.hermes/cron/jobs.json - 'reports to Kwame via this Zulip DM' described the pre-2026-09-21 model; the report is DROPped to the gate (comms_drop.py) and exit 0 means QUEUED, not sent - added a Checks line so a skipped phase 5 is visible in the report
14 KiB
report_only_agents, kind, name, description, version
| report_only_agents | kind | name | description | version | |
|---|---|---|---|---|---|
|
pattern | memory-fixer | Auto-fix low-hanging fruit in the RA-H OS knowledge graph. No judgment calls — only deterministic Level 1 operations. Escalate anything that needs Kwame's input. Executes confirmed Kwame decisions to completion (state + updated_at). | 2.2.0 |
Memory Fixer
Executable copy: the
okyeame-memory-fixercron job on kagentz (hermes cron list) holds its instruction set inline in~/.hermes/cron/jobs.json(hermes cron edit <id> --prompt …; there is no--prompt-file, and~/.hermes/cron/memory-fixer-prompt.mdis a synced draft, not the live instruction). This file is the institutional record of the same contract; when the two diverge, the job prompt is what actually runs — diff it against this file before claiming a prompt change landed. ⚠️ Corrected 2026-09-26: the previous pointer (/root/.hermes/contracts/memory-fixer-v3.md) does not exist on kagentz — no/rootaccess from this container — and was verified unreachable, not merely stale.
Purpose
Auto-fix low-hanging fruit in the graph. No judgment calls — only deterministic Level 1 operations. Escalate anything that needs Kwame's input. When Kwame replies to an escalation, execute the decision to completion (update state and timestamps), never leaving a node in review-pending forever.
Type: Write-only (Level 1 fixes only) Scope: RA-H OS knowledge graph (192.168.68.65) Schedule: Daily at 8 AM ET Escalation: Level 2+ to Kwame as task items
Key Design Decision
The updateNode tool's metadata field performs a restricted merge — the state key only accepts 'processed' or 'not_processed'. Additionally, new metadata keys cannot be added via the merge.
Solution: Use the description field to tag stale nodes with review actions, since description is a simple string overwritable via updateNode.
Tag Format: [REVIEW: action] original description text...
Where action is one of:
archive— node is stale and should be archivedrefresh— node is stale and should be refreshed (infrastructure)keep— node has been confirmed as currentmerge— node is a duplicate candidate
Query for finding review-tagged nodes:
SELECT id, title, description
FROM nodes
WHERE description LIKE '[REVIEW:%';
Level 1 Auto-Fixes (No Kwame Decision Needed)
1. Missing type Auto-Classification
SELECT id, title,
CASE
WHEN title LIKE '%infrastructure%' OR title LIKE '%proxmox%' OR title LIKE '%setup%' THEN 'infrastructure'
WHEN title LIKE '%skill%' OR title LIKE '%how to%' OR title LIKE '%guide%' THEN 'skill'
WHEN title LIKE '%doc%' OR title LIKE '%template%' OR title LIKE '%brand%' THEN 'documentation'
WHEN title LIKE 'WAL:%' OR title LIKE 'TASK:%' THEN 'note'
WHEN title LIKE '%[LEARN]%' THEN 'documentation'
ELSE 'note'
END as auto_type
FROM nodes
WHERE json_extract(metadata, '$.type') IS NULL;
2. Missing namespace Auto-Population
SELECT id, title, json_extract(metadata, '$.tenant') as tenant
FROM nodes
WHERE json_extract(metadata, '$.namespace') IS NULL;
3. Staleness Review Tagging (refresh-suggested nodes only)
Using the type-based windows from the memory-monitor contract, tag nodes stale beyond their window. Only process a maximum of 10 nodes per run to avoid overwhelming Kwame. Prioritize infrastructure first, then dynamic, then ephemeral.
Archive-suggested nodes are NO LONGER tagged — they are archived outright (see Level 1 fix 4). Tagging with [REVIEW: refresh] applies only to living nodes (infrastructure, deployment, system, system-health, business, philosophy, research, learning, investigation, analysis, project, agent, registry, policy).
Exclusion Rules:
- Nodes with
state=review_pending,deprecated,archived, ornot_processedare NOT processed - Nodes whose
descriptionalready starts with[REVIEW:or[ARCHIVED]are NOT re-processed
SELECT id, title, json_extract(metadata, '$.type') as node_type,
CAST(julianday('now') - julianday(updated_at) AS INTEGER) as days_stale,
CASE
WHEN json_extract(metadata, '$.type') IN ('infrastructure', 'deployment', 'system', 'system-health') THEN 'refresh'
WHEN json_extract(metadata, '$.type') IN ('business', 'philosophy', 'research', 'learning', 'learn', 'investigation', 'analysis', 'project') THEN 'refresh'
ELSE 'archive'
END as suggested_action
FROM nodes
WHERE updated_at < datetime('now',
CASE
WHEN json_extract(metadata, '$.type') IN ('infrastructure', 'deployment', 'system', 'system-health') THEN '-14 days'
WHEN json_extract(metadata, '$.type') IN ('skill', 'documentation', 'template', 'protocol-enforcement', 'prd', 'architecture') THEN '-90 days'
WHEN json_extract(metadata, '$.type') IN ('note', 'wal', 'WAL', 'task', 'TASK', 'event') THEN '-30 days'
WHEN json_extract(metadata, '$.type') IN ('business', 'philosophy', 'research', 'learning', 'learn', 'investigation', 'analysis', 'project') THEN '-120 days'
WHEN json_extract(metadata, '$.type') IN ('deprecated-relay', 'audit', 'audit-report', 'incident', 'incident-report') THEN '-3650 days'
ELSE '-45 days'
END
)
AND json_extract(metadata, '$.state') NOT IN ('review_pending', 'deprecated', 'archived', 'not_processed')
AND (description IS NULL OR description NOT LIKE '[REVIEW:%')
ORDER BY days_stale ASC
LIMIT 10;
For each identified node, call updateNode(id, { description: "[REVIEW: action] " + originalDescription }).
4. Stale-Node Archiving (Level 1 — standing Kwame directive, 2026-09-11)
Kwame's standing directive: stale nodes CAN be archived by the fixer. No per-batch escalation, no [REVIEW: archive] tagging — archive them.
For every node whose suggested action is archive (i.e. its type is NOT one of the living types in fix 3), archive it in a single updateNode call:
updateNode(id, {
"description": "[ARCHIVED] " + originalDescriptionWithoutReviewTag,
"metadata": {"state": "archived"}
})
statetransitions DO work throughupdateNode(archived, and back toactive). The former "state only accepts processed/not_processed, use SSH" claim was wrong — verified 2026-09-11 by archiving 7 nodes (#61, #373, #388, #465, #475, #526, #1476) over the bridge withupdated_atauto-bumping. SSH to the bridge host is a fallback, not a requirement, and it is blocked from kagentz anyway.- Pass
descriptionandmetadatain the same call, and always keep theupdatesobject nested:{"id": N, "updates": {…}}. - Archiving is non-destructive: the node stays in the graph, marked
state: archived+[ARCHIVED]prefix. Living nodes (refresh-suggested) are NEVER archived without a specific Kwame decision — they are the cluster/agent/business canon.
Archive candidates are identified by the fix 3 query's suggested_action = 'archive' branch (the ELSE 'archive' case: anything not an infrastructure/skill/documentation/strategic/audit type).
5. Duplicate-Node Detection (Level 1 — read-only, every run)
The graph's duplicate problem is rarely an agent mistyping a title: it is recurring writers creating a new node per run instead of updating one. This phase detects that class and reports it. It is read-only and never merges.
python3 /home/hermes/.hermes/scripts/memory_dup_detect.py --json
Read-only, ~15s over the whole graph, exit 0. That script is the source of truth for the clustering logic — do not re-implement it in the prompt or hand-count "duplicates" from titles.
Consume each items[] entry's verdict field; do not invent your own:
verdict |
Meaning | Required action |
|---|---|---|
WRITER-DEFECT (run_family: true) |
ONE scheduled task writes a new node per run | Report the ids, the agents (the writer) and span_days. Never merge — each node is that run's audit record. If the family grew since the last report, say UNFIXED and name the writer. |
SAFE-MERGE |
Bodies identical | Still requires an explicit merge #A into #B decision from Kwame. |
HUMAN-DECISION |
Same subject, bodies differ | Propose connect (an edge), never merge. |
- Title overlap alone is not duplication. Four distinct client workflows of one family (#357-#361) and two different machines' migrations (#1792/#1793) both score high on title tokens while their bodies sit 0.1-0.3 apart. Confirm against body similarity before calling anything a duplicate.
- Report clusters as candidates for Kwame's decision, never as established duplicates — a wrong auto-merge destroys distinct content irrecoverably.
- Per-run history nodes are kept deliberately. Bulk-merging a run family destroys the audit trail the family exists for.
Level 2 Escalations (Kwame Decision Required)
- Refresh-suggested stale nodes flagged with
[REVIEW: refresh]— refresh or keep? (Archive-suggested nodes are auto-archived under fix 4 and are not escalated.) - Duplicate Nodes — as detected by fix 5, by
verdict, never by raw title overlap.WRITER-DEFECTis a writer fix (update one canonical node), not a merge decision;SAFE-MERGEandHUMAN-DECISIONclusters are escalated for merge-or-connect. - Orphan Nodes >90 days old — Archive or connect?
Reporting Format
The fixer does not send anything. Under the single-egress model (2026-09-21) every report leaves the node
through Mumuni's gate (comms_drop.py for the queue, comms_gate.py to release and read-back verify), so
exit 0 means QUEUED, never delivered. A report body is written to a file and handed to the outbox helper:
🦅 Memory Fixer — [HH:MM UTC]
Level 1 fixes applied:
- Missing type: X nodes classified
- Missing namespace: Y nodes populated
Stale nodes needing review (max 10):
1. [Node #XXX] Title — X days stale, SUGGEST: refresh
2. [Node #YYY] Title — Y days stale, SUGGEST: archive
...
Duplicate clusters (candidates — Kwame decides; the fixer never merges unilaterally):
1. [WRITER-DEFECT] #AAA/#BBB/#CCC — writer <agent>, N nodes, span Nd (UNFIXED if it grew since the last report)
2. [HUMAN-DECISION] #DDD/#EEE — same subject, bodies differ, SUGGEST: connect
3. "none" when the scan returned no clusters
Orphans >90 days:
1. [Node #EEE] Title — X days stale, orphaned
Reply with:
- "archive #XXX, #YYY" to mark for archive
- "archive all" to archive all stale nodes listed
- "keep #XXX" to confirm a node is current
- "merge #AAA into #BBB" to merge duplicates
- "refresh #XXX" to mark as current
Execution on Next Run
The fixer reads Kwame's previous response and executes the decision to completion — it must not leave a node in review-pending forever. Tagging alone is NOT enough; each confirmed decision must also update state and updated_at so the node drops out of the stale window on the next run.
⚠️ Corrected 2026-09-11:
updateNodeDOES acceptstatechanges —{"updates": {"description": …, "metadata": {"state": "archived"}}}works over the bridge, andupdated_atbumps automatically. The old "use direct SSH + SQLite for state transitions" instruction was based on a wrong assumption; SSH is a fallback only (and is blocked from kagentz). Use oneupdateNodecall for both the tag and the state.ssh root@192.168.68.65 "sqlite3 /root/.local/share/RA-H/db/rah.sqlite \"UPDATE nodes SET metadata = json_set(metadata, '$.state', '<state>'), updated_at = datetime('now') WHERE id = <id>;\""Use
updateNodeonly for description/source/title/link edits.
Decision → completed action mapping:
| Kwame reply | Description change | State | updated_at |
|---|---|---|---|
archive #XXX |
replace [REVIEW: archive] → [ARCHIVED] prefix |
archived |
bumped to now |
keep #XXX / refresh #XXX |
clear the [REVIEW: …] tag entirely |
active |
bumped to now |
merge #AAA into #BBB |
set [REVIEW: merge_into #BBB] on #AAA, then follow manual merge workflow |
handled manually | bumped to now |
archive all |
apply the archive row to every node listed in the prior report | archived |
bumped to now |
Why updated_at must be bumped (critical): the Level-1 staleness query keys off updated_at < now - window. If the fixer clears the tag but leaves a stale updated_at, the node is immediately re-flagged on the very next run and the cycle repeats forever. Bumping updated_at to now pushes the node back to the front of the window.
Exclusion after action: once an action is applied, the node's description no longer starts with [REVIEW: (archive → [ARCHIVED], refresh/keep → original text), so it is not re-processed.
After all actions are applied, verify with:
SELECT id, json_extract(metadata, '$.state') FROM nodes WHERE description LIKE '[REVIEW:%';
The result must be 0 rows when all decisions are executed. Report what was done.
Checks
- State integrity: archived nodes have
state: archived+[ARCHIVED]prefix; kept nodes arestate: activewithout a[REVIEW:]tag. - Auto-archive applied: no node should ever be left tagged
[REVIEW: archive]— that tag is retired. Any[REVIEW: archive]found means fix 4 was skipped; archive it and report. - No review-pending forever: after executing Kwame's decisions,
[REVIEW:%node count must be 0. - Duplicate scan ran: every report carries the fix 5 block (
nonewhen there were no clusters). A report with no duplicate section means phase 5 was skipped — a silently skipped detection phase is the failure this phase exists to prevent. - Timestamps: every executed decision (and every auto-archive) bumps
updated_at, so the node exits the stale window on the next run.
Logging
Every Level 1 fix logged to ~/.hermes/logs/memory-fixer/YYYY-MM-DD.md
Every Level 2 escalation logged and delivered to Kwame.