Compare commits

...
Author SHA1 Message Date
root e42b970dec fix(audit-hermes): handle fallback_providers as list or dict
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 17s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s
The audit assumed fallback_providers was always a dict (single provider).
Two live agents (koby, koonimo) carry it as a LIST of dicts (one entry per
fallback), so the script crashed with:

    File "audit-hermes-config.py", line 211, in audit
        fb.get("provider") == "deepseek",
    AttributeError: 'list' object has no attribute 'get'

Both are REAL agent configs, so this is not a malformed-input case — the
script simply could not audit two of the four agents it exists to audit.

Fix:
- Normalize fallback_providers to a list of entries (dict → [dict], list → list)
- Apply the existing checks to each entry
- A malformed entry (not a mapping) produces a reported VIOLATION naming the
  offending entry, NOT an uncaught exception

Adds regression test using the real failing shape (list of dicts) and proves
it bites against the pre-fix revision.

Real audit results after fix:
- mumuni: FAIL — 7 violations
- tanko: FAIL — 21 violations
- koby: FAIL — 16 violations (previously crashed)
- koonimo: FAIL — 10 violations (previously crashed)

No agent configs were changed. No existing rules were relaxed.
2026-09-27 11:17:18 +00:00
abiba-bot eadb927ec1 Merge pull request 'fix(hermes): clarify auxiliary model policy to match audit' (#140) from fix/hermes-aux-model-policy-20260927 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 9s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 1s
2026-09-27 10:05:44 +00:00
mumuni-bot 73d5097555 Merge pull request 'feat(memory-fixer): add duplicate-detection phase; correct stale canonical pointer + reporting model (v2.2.0)' (#139) from fix/memory-fixer-duplicate-detection-20260926 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 8s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 11s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 11s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 2s
Merge PR #139 — memory-fixer duplicate-detection phase (v2.2.0). All CI green; authorized by AGENTS.md step 4 (all green -> merge) and the authorization table (memory-fixer = normal sensitivity, any registered agent).
2026-09-27 02:16:43 +00:00
Mumuni c460ef905c feat(memory-fixer): add duplicate-detection phase; correct stale canonical pointer + reporting model (v2.2.0)
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 8s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 9s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s
Level 1 gains phase 5, duplicate-node detection: runs
memory_dup_detect.py (read-only) and reports clusters by verdict.

- WRITER-DEFECT (run_family): one task creating a node per run -> writer
  fix, never merge (per-run nodes are the audit trail)
- SAFE-MERGE: identical bodies, still needs an explicit Kwame decision
- HUMAN-DECISION: same subject, bodies differ -> connect, never merge

Replaces the naive '>70% title overlap' duplicate rule, which false-positives
on distinct work: four client workflows (#357-#361) and two machines'
migrations (#1792/#1793) score high on titles with bodies 0.1-0.3 apart.

Also corrected in the same pass:
- the 'canonical copy' pointer named /root/.hermes/contracts/memory-fixer-v3.md,
  which does not exist on kagentz (no /root access); the job's instruction set is
  inline in ~/.hermes/cron/jobs.json
- 'reports to Kwame via this Zulip DM' described the pre-2026-09-21 model; the
  report is DROPped to the gate (comms_drop.py) and exit 0 means QUEUED, not sent
- added a Checks line so a skipped phase 5 is visible in the report
2026-09-26 19:32:36 +00:00
3 changed files with 316 additions and 23 deletions
+36 -17
View File
@@ -94,7 +94,16 @@ def audit(path):
cfg = yaml.safe_load(f)
model = cfg.get("model", {})
fb = cfg.get("fallback_providers", {})
fb_raw = cfg.get("fallback_providers", {})
# Normalize: fallback_providers may be a dict (single provider) or a list of dicts
# (one entry per fallback). Both shapes are valid; we must handle both without crashing.
if isinstance(fb_raw, dict):
fb_entries = [fb_raw]
elif isinstance(fb_raw, list):
fb_entries = fb_raw
else:
fb_entries = [fb_raw] # Let it fail the check below as malformed
fb = fb_entries[0] if fb_entries else {}
comp = cfg.get("compression", {})
aux = cfg.get("auxiliary", {})
deleg = cfg.get("delegation", {})
@@ -207,22 +216,32 @@ def audit(path):
"Rule 14",
f"delegation.provider must be 'harness' (got {deleg.get('provider')!r})",
)
check(
fb.get("provider") == "deepseek",
"Rule 14",
f"fallback_providers.provider must be 'deepseek' (got {fb.get('provider')!r}) — "
f"true fallback diversity, not same endpoint as primary",
)
check(
fb.get("model") == "deepseek-v4-flash",
"Rule 14",
f"fallback_providers.model must be 'deepseek-v4-flash' (got {fb.get('model')!r})",
)
check(
fb.get("api_key_env") == "DEEPSEEK_API_KEY",
"Rule 14",
f"fallback_providers.api_key_env must be DEEPSEEK_API_KEY (got {fb.get('api_key_env')!r})",
)
# Check each fallback entry. A malformed entry (not a mapping) is a VIOLATION, not a crash.
for idx, entry in enumerate(fb_entries):
prefix = f"fallback_providers[{idx}]"
if not isinstance(entry, dict):
check(
False,
"Rule 14",
f"{prefix} must be a mapping (got {type(entry).__name__})",
)
continue
check(
entry.get("provider") == "deepseek",
"Rule 14",
f"{prefix}.provider must be 'deepseek' (got {entry.get('provider')!r}) — "
f"true fallback diversity, not same endpoint as primary",
)
check(
entry.get("model") == "deepseek-v4-flash",
"Rule 14",
f"{prefix}.model must be 'deepseek-v4-flash' (got {entry.get('model')!r})",
)
check(
entry.get("api_key_env") == "DEEPSEEK_API_KEY",
"Rule 14",
f"{prefix}.api_key_env must be DEEPSEEK_API_KEY (got {entry.get('api_key_env')!r})",
)
# --- custom_providers sanity ---
check(
+44 -6
View File
@@ -6,13 +6,19 @@ name: memory-fixer
description: >
Auto-fix low-hanging fruit in the RA-H OS knowledge graph. No judgment calls — only deterministic Level 1 operations.
Escalate anything that needs Kwame's input. Executes confirmed Kwame decisions to completion (state + updated_at).
version: 2.1.0
version: 2.2.0
---
---
# Memory Fixer
> **Canonical copy:** `/root/.hermes/contracts/memory-fixer-v3.md` (used by the `memory-fixer-daily` cron job). This file is the institutional record of the same contract. When the two diverge, treat the v3 source in `/root/.hermes/contracts/` as executable truth.
> **Executable copy:** the `okyeame-memory-fixer` cron job on kagentz (`hermes cron list`) holds its instruction
> set **inline in `~/.hermes/cron/jobs.json`** (`hermes cron edit <id> --prompt …`; there is no `--prompt-file`, and
> `~/.hermes/cron/memory-fixer-prompt.md` is a synced draft, not the live instruction). This file is the institutional
> record of the same contract; when the two diverge, the job prompt is what actually runs — diff it against this file
> before claiming a prompt change landed.
> ⚠️ Corrected 2026-09-26: the previous pointer (`/root/.hermes/contracts/memory-fixer-v3.md`) does not exist on
> kagentz — no `/root` access from this container — and was verified unreachable, not merely stale.
## Purpose
Auto-fix low-hanging fruit in the graph. No judgment calls — only deterministic Level 1 operations. Escalate anything that needs Kwame's input. When Kwame replies to an escalation, **execute the decision to completion** (update state and timestamps), never leaving a node in review-pending forever.
@@ -125,15 +131,44 @@ updateNode(id, {
**Archive candidates are identified by the fix 3 query's `suggested_action = 'archive'` branch** (the `ELSE 'archive'` case: anything not an infrastructure/skill/documentation/strategic/audit type).
### 5. Duplicate-Node Detection (Level 1 — read-only, every run)
The graph's duplicate problem is rarely an agent mistyping a title: it is **recurring writers creating a new
node per run instead of updating one**. This phase detects that class and reports it. It is read-only and
**never merges**.
```bash
python3 /home/hermes/.hermes/scripts/memory_dup_detect.py --json
```
Read-only, ~15s over the whole graph, exit 0. That script is the source of truth for the clustering logic —
do not re-implement it in the prompt or hand-count "duplicates" from titles.
Consume each `items[]` entry's `verdict` field; do not invent your own:
| `verdict` | Meaning | Required action |
|---|---|---|
| `WRITER-DEFECT` (`run_family: true`) | ONE scheduled task writes a new node per run | Report the ids, the `agents` (the writer) and `span_days`. **Never merge** — each node is that run's audit record. If the family grew since the last report, say `UNFIXED` and name the writer. |
| `SAFE-MERGE` | Bodies identical | Still requires an explicit `merge #A into #B` decision from Kwame. |
| `HUMAN-DECISION` | Same subject, bodies differ | Propose **connect (an edge)**, never merge. |
- **Title overlap alone is not duplication.** Four distinct client workflows of one family (#357-#361) and two
different machines' migrations (#1792/#1793) both score high on title tokens while their bodies sit 0.1-0.3
apart. Confirm against body similarity before calling anything a duplicate.
- Report clusters as **candidates for Kwame's decision**, never as established duplicates — a wrong auto-merge
destroys distinct content irrecoverably.
- Per-run history nodes are kept deliberately. Bulk-merging a run family destroys the audit trail the family exists for.
## Level 2 Escalations (Kwame Decision Required)
1. **Refresh-suggested stale nodes** flagged with `[REVIEW: refresh]` — refresh or keep? (Archive-suggested nodes are auto-archived under fix 4 and are not escalated.)
2. **Duplicate Nodes** (same title or >70% title overlap) — Merge or keep?
2. **Duplicate Nodes** — as detected by fix 5, by `verdict`, never by raw title overlap. `WRITER-DEFECT` is a writer fix (update one canonical node), not a merge decision; `SAFE-MERGE` and `HUMAN-DECISION` clusters are escalated for merge-or-connect.
3. **Orphan Nodes >90 days old** — Archive or connect?
## Reporting Format
The fixer reports to Kwame via this Zulip DM:
The fixer does **not** send anything. Under the single-egress model (2026-09-21) every report leaves the node
through Mumuni's gate (`comms_drop.py` for the queue, `comms_gate.py` to release and read-back verify), so
exit 0 means QUEUED, never delivered. A report body is written to a file and handed to the outbox helper:
```
🦅 Memory Fixer — [HH:MM UTC]
@@ -147,8 +182,10 @@ Stale nodes needing review (max 10):
2. [Node #YYY] Title — Y days stale, SUGGEST: archive
...
Duplicates needing decision:
1. [Node #AAA] vs [Node #BBB] — Same title
Duplicate clusters (candidates — Kwame decides; the fixer never merges unilaterally):
1. [WRITER-DEFECT] #AAA/#BBB/#CCC — writer <agent>, N nodes, span Nd (UNFIXED if it grew since the last report)
2. [HUMAN-DECISION] #DDD/#EEE — same subject, bodies differ, SUGGEST: connect
3. "none" when the scan returned no clusters
Orphans >90 days:
1. [Node #EEE] Title — X days stale, orphaned
@@ -195,6 +232,7 @@ The result must be 0 rows when all decisions are executed. Report what was done.
- **State integrity:** archived nodes have `state: archived` + `[ARCHIVED]` prefix; kept nodes are `state: active` without a `[REVIEW:]` tag.
- **Auto-archive applied:** no node should ever be left tagged `[REVIEW: archive]` — that tag is retired. Any `[REVIEW: archive]` found means fix 4 was skipped; archive it and report.
- **No review-pending forever:** after executing Kwame's decisions, `[REVIEW:%` node count must be 0.
- **Duplicate scan ran:** every report carries the fix 5 block (`none` when there were no clusters). A report with no duplicate section means phase 5 was skipped — a silently skipped detection phase is the failure this phase exists to prevent.
- **Timestamps:** every executed decision (and every auto-archive) bumps `updated_at`, so the node exits the stale window on the next run.
## Logging
@@ -0,0 +1,236 @@
"""Regression test for the fallback_providers list-shape crash in audit-hermes-config.py.
WHY THIS FILE EXISTS: audit-hermes-config.py assumed `fallback_providers` was always a dict
(single provider). Two live agents (koby, koonimo) carry it as a LIST of dicts (one entry per
fallback), so the script crashed with:
File "audit-hermes-config.py", line 211, in audit
fb.get("provider") == "deepseek",
AttributeError: 'list' object has no attribute 'get'
Both are REAL agent configs, so this is not a malformed-input case — the script simply could not
audit two of the four agents it exists to audit. Until fixed, the key-hygiene check had no
coverage for half the fleet while appearing to run.
These tests execute the real CLI (`python3 audit-hermes-config.py <config>`) and assert:
1. A config whose `fallback_providers` is a LIST of valid dicts does NOT crash (exit code is 0 or 1,
never a traceback/AttributeError).
2. A config whose `fallback_providers` contains a MALFORMED entry (a list element that is not a
mapping) reports a VIOLATION naming the offending entry, NOT an uncaught exception.
3. The dict shape still works (existing tests must stay green).
No network, vault, or SSH access is required.
"""
from __future__ import annotations
import pathlib
import subprocess
import sys
ROOT = pathlib.Path(__file__).resolve().parent.parent
AUDIT = ROOT / "audit-hermes-config.py"
# A valid config where fallback_providers is a LIST of dicts (the real koby/koonimo shape).
# One entry, well-formed: provider=deepseek, model=deepseek-v4-flash, api_key_env=DEEPSEEK_API_KEY.
# This must produce a real verdict (PASS or FAIL) without crashing.
LIST_SHAPE_VALID = """
model:
api_key: ""
api_key_env: LITELLM_API_KEY
base_url: http://192.168.68.116/litellm/v1
max_tokens: 4096
default: syslog-auto
provider: harness
fallback_providers:
- provider: deepseek
model: deepseek-v4-flash
api_key_env: DEEPSEEK_API_KEY
compression:
model: syslog-auto
provider: harness
threshold: 0.65
max_context_window: 131072
auxiliary:
vision:
model: gpu-vision
provider: harness
web_extract:
model: gpu-vision
provider: harness
compression:
model: syslog-auto
provider: harness
delegation:
provider: harness
custom_providers:
- name: harness
key_env: LITELLM_API_KEY
base_url: http://192.168.68.116/litellm/v1
"""
# A valid config where fallback_providers is a LIST with TWO entries (multiple fallbacks).
# Both entries well-formed. Must not crash and should produce a real verdict.
LIST_SHAPE_MULTI = """
model:
api_key: ""
api_key_env: LITELLM_API_KEY
base_url: http://192.168.68.116/litellm/v1
max_tokens: 4096
default: syslog-auto
provider: harness
fallback_providers:
- provider: deepseek
model: deepseek-v4-flash
api_key_env: DEEPSEEK_API_KEY
- provider: deepseek
model: deepseek-v4-flash
api_key_env: DEEPSEEK_API_KEY
compression:
model: syslog-auto
provider: harness
threshold: 0.65
max_context_window: 131072
auxiliary:
vision:
model: gpu-vision
provider: harness
web_extract:
model: gpu-vision
provider: harness
compression:
model: syslog-auto
provider: harness
delegation:
provider: harness
custom_providers:
- name: harness
key_env: LITELLM_API_KEY
base_url: http://192.168.68.116/litellm/v1
"""
# A config where fallback_providers is a LIST containing a MALFORMED entry:
# one element is a plain string, not a mapping. The checker must report a VIOLATION
# naming the offending entry (fallback_providers[1]) and NOT crash.
LIST_SHAPE_MALFORMED = """
model:
api_key: ""
api_key_env: LITELLM_API_KEY
base_url: http://192.168.68.116/litellm/v1
max_tokens: 4096
default: syslog-auto
provider: harness
fallback_providers:
- provider: deepseek
model: deepseek-v4-flash
api_key_env: DEEPSEEK_API_KEY
- "not-a-mapping"
compression:
model: syslog-auto
provider: harness
threshold: 0.65
max_context_window: 131072
auxiliary:
vision:
model: gpu-vision
provider: harness
web_extract:
model: gpu-vision
provider: harness
compression:
model: syslog-auto
provider: harness
delegation:
provider: harness
custom_providers:
- name: harness
key_env: LITELLM_API_KEY
base_url: http://192.168.68.116/litellm/v1
"""
# The original DICT shape (single provider) must still work — existing behaviour preserved.
DICT_SHAPE_VALID = """
model:
api_key: ""
api_key_env: LITELLM_API_KEY
base_url: http://192.168.68.116/litellm/v1
max_tokens: 4096
default: syslog-auto
provider: harness
fallback_providers:
provider: deepseek
model: deepseek-v4-flash
api_key_env: DEEPSEEK_API_KEY
compression:
model: syslog-auto
provider: harness
threshold: 0.65
max_context_window: 131072
auxiliary:
vision:
model: gpu-vision
provider: harness
web_extract:
model: gpu-vision
provider: harness
compression:
model: syslog-auto
provider: harness
delegation:
provider: harness
custom_providers:
- name: harness
key_env: LITELLM_API_KEY
base_url: http://192.168.68.116/litellm/v1
"""
def _run_config(tmp_path, name, text):
cfg = tmp_path / name
cfg.write_text(text)
proc = subprocess.run(
[sys.executable, str(AUDIT), str(cfg)],
capture_output=True, text=True,
)
return proc.returncode, proc.stdout, proc.stderr
def test_list_shape_single_entry_does_not_crash(tmp_path):
"""A LIST with one valid dict must not raise AttributeError; exit 0 (PASS)."""
code, out, err = _run_config(tmp_path, "list-single.yaml", LIST_SHAPE_VALID)
# Must NOT be a crash (traceback). A clean run exits 0 (PASS) or 1 (FAIL), never 2+ (exception).
assert code in (0, 1), f"Expected clean exit 0 or 1, got {code}\nSTDOUT:\n{out}\nSTDERR:\n{err}"
assert "AttributeError" not in err, f"Crashed with AttributeError:\n{err}"
assert "Traceback" not in err, f"Crashed with uncaught exception:\n{err}"
# The valid single-entry list should PASS (all rules satisfied).
assert code == 0, f"Expected PASS but got {code}\n{out}"
assert "RESULT: PASS" in out
def test_list_shape_multiple_entries_does_not_crash(tmp_path):
"""A LIST with two valid dicts must not raise AttributeError; exit 0 (PASS)."""
code, out, err = _run_config(tmp_path, "list-multi.yaml", LIST_SHAPE_MULTI)
assert code in (0, 1), f"Expected clean exit 0 or 1, got {code}\nSTDOUT:\n{out}\nSTDERR:\n{err}"
assert "AttributeError" not in err, f"Crashed with AttributeError:\n{err}"
assert "Traceback" not in err, f"Crashed with uncaught exception:\n{err}"
assert code == 0, f"Expected PASS but got {code}\n{out}"
assert "RESULT: PASS" in out
def test_list_shape_malformed_entry_reports_violation_not_crash(tmp_path):
"""A LIST containing a non-mapping element must be a reported VIOLATION, not a crash."""
code, out, err = _run_config(tmp_path, "list-malformed.yaml", LIST_SHAPE_MALFORMED)
# Must NOT be a crash.
assert "AttributeError" not in err, f"Crashed with AttributeError:\n{err}"
assert "Traceback" not in err, f"Crashed with uncaught exception:\n{err}"
# Should be a FAIL (exit 1) because the malformed entry is a violation.
assert code == 1, f"Expected FAIL (exit 1) but got {code}\n{out}"
assert "RESULT: FAIL" in out
# The violation must name the offending entry (fallback_providers[1]).
assert "fallback_providers[1]" in out, f"Violation did not name the offending entry:\n{out}"
def test_dict_shape_still_passes(tmp_path):
"""The original DICT shape (single provider) must still PASS — existing behaviour preserved."""
code, out, err = _run_config(tmp_path, "dict-valid.yaml", DICT_SHAPE_VALID)
assert code == 0, f"Expected PASS but got {code}\n{out}\nSTDERR:\n{err}"
assert "RESULT: PASS" in out