Compare commits

...
Author SHA1 Message Date
tanko-bot aa2451e8ae fix: skip Hermes config/wrapper integrity checks for tanko (DSH)
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s
check_config_integrity and check_wrapper_integrity still probed tanko (CT 112)
for Hermes-only artifacts (/root/.hermes/config.yaml, the hermes CLI wrapper)
that no longer exist since tanko moved to DSH on 2026-08-27. This caused false
FAILs in the agent health check. Both now skip tanko via runtime=dsh.
2026-08-27 03:06:49 +00:00
tanko-bot c23462eba8 fix: correct tanko runtime to DSH across contracts & scripts
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
Tanko migrated from Hermes to DSH (DeepSeek Harness) on 2026-08-27. Update all
records that described tanko as a Hermes agent / Hermes runtime:

- infra-control: CT 112 tanko platform Hermes -> DSH
- zulip-health / zulip-self-heal / zulip-mention-reliability / pi-approval:
  tanko is on DSH, mumuni remains on Hermes
- memory-audit-maintenance: exclude tanko from Hermes roster (uses DSH-native memory)
- hermes-config-template / hermes-agent-baseline: remove tanko from Hermes roster,
  keep LiteLLM key alias 'tanko'
- hermes-zulip-plugin / hermes-zulip-restore / build-zulip-plugin: tanko excluded
- infrastructure-maintenance: gateways check no longer probes Hermes on tanko CT112
- scripts/daily-infra-report.py: fix CT-ID regression (CT 122->112), report tanko as DSH
- scripts/zulip-monitor.sh: stop probing tanko's retired Hermes gateway
- scripts/agent-health-check.py: skip Hermes gateway checks for tanko (runtime=dsh)
- scripts/prose-auth-check.sh + AGENTS.md: authorize tanko/tanko-bot for its own records

Tanko remains CT 112 at 192.168.68.122; infrastructure facts unchanged.
Dated/incident records (run logs, migration logs) left intact as history.
2026-08-27 03:01:24 +00:00
abiba-bot b835986d44 Merge pull request 'Add check-health section with live probes to infrastructure-monitoring contract' (#51) from fix/infra-check-health into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 1s
Merged by captain approval 2026-08-22: add check-health section with live probes to infrastructure-monitoring contract
2026-08-23 00:01:39 +00:00
root d6ad016ac9 Add check-health section with live probes to infrastructure-monitoring contract
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 11s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 16s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 13s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
Fixes the canned 'API UNREACHABLE (HTTP 000)' reporting. Adds an explicit
check-health execution section (Zulip POST, pm2, GPU exporters, Prometheus,
Grafana, LiteLLM probes) with a hard RUN LIVE, NEVER ECHO rule, mirroring
the gpu-monitor contract pattern. Captain priority 2026-08-22.
2026-08-22 23:59:28 +00:00
mumuni-bot c3306e87e4 Merge pull request 'fix: ship pm2-self-heal crash-loop guard + align gpu-dense docs to Qwen3.8-27B' (#49) from fix/pm2-guard-gpu-doc into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 8s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 0s
2026-08-17 02:34:49 +00:00
mumuni-bot 20fe5adcbc fix: ship pm2-self-heal crash-loop guard + align gpu-dense docs to Qwen3.8-27B
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s
- pm2-self-heal.sh: restart abiba-telegram when TEL_RESTARTS > 1000 even if 'online'
  (catches quiet crash-loops like the 10k-restarts spoton incident); alert includes count.
- pm2-self-heal.prose.md: document the crash-loop guard under Rule 2.
- gpu-fleet.prose.md / gpu-self-heal.prose.md: gpu-dense is Qwen3.8-27B-Uncensored-Q4_K_M
  (~16.8GB, alias qwen3.6-27B-code for LiteLLM routing), not SmartCode-Fable-5 (verified live
  on .8:8080 Aug 16).
2026-08-17 02:32:52 +00:00
mumuni-bot 82eae77cc5 Merge pull request 'docs: reflect hwepve removal from Tabiri cluster (5 nodes + standalone London node)' (#48) from fix/hwepve-london-migration into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 0s
2026-08-17 02:24:44 +00:00
Agent Zero 8b00a4beea docs: Mumuni now inside Abiba CT 100 in zulip-resilience-v3 status table
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s
2026-08-15 20:13:29 -04:00
Agent Zero 19a67c6815 fix: align mumuni ct field to 100 in daily-infra-report.py
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
2026-08-15 20:12:09 -04:00
kagentz-bot 44f7008302 docs: reflect hwepve removal from Tabiri cluster (5 nodes + standalone London node)
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 2s
- Remove hwepve from cluster member lists/diagrams in infrastructure-control,
  proxmox-monitor, infrastructure-update, litellm-health, mumuni-delegation
- CTs 100 (abiba) and 105 (kagentz) moved to minipve; CT 114 (mumuni) no longer
  exists (Mumuni runs inside Abiba CT 100)
- Add standalone London role + NetBird routing peer description for hwepve
- Update scripts: agent-health-check.py, pct-run.sh, prose-ai-review.sh
- Update disk-gc CT access table, hermes baselines/restore host references
2026-08-15 20:03:02 -04:00
mumuni-bot a179164f1f Merge pull request 'Ship: Koby external-agent ruling (fix api_key_env hygiene + docs)' (#46) from ship/koby-external-agent-fix into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 9s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 1s
merge
2026-08-13 00:33:54 +00:00
mumuni-bot c4626c3512 Merge pull request 'fix: redirect pm2/zulip health-check logging to Gitea, never graph (hard rule)' (#47) from fix/health-logs-not-graph into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 1s
merge
2026-08-13 00:33:33 +00:00
mumuni-bot b799d46596 fix: redirect pm2/zulip health-check logging to Gitea, never graph (hard rule)
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
- pm2-self-heal.prose.md: description said 'logs every action to the
  knowledge graph', contradicting its own execution step (line ~60) that
  says '(not knowledge graph - hard rule)'. Align description to Gitea.
- zulip-self-heal.prose.md: Reporting section said 'every cycle produces
  a knowledge graph node'. Contract is RETIRED; remove graph-node directive.
- Both now consistently log to SyslogSolution/health-logs, never the graph.
2026-08-13 00:09:40 +00:00
root 3abf784538 Fix no-mistakes warnings: remove prose from fence, reorder/renumber rules, and add Koby exception to Rule 10
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
2026-08-11 18:16:12 +00:00
root d73084d481 Trivial: trigger no-mistakes re-run (all fixes applied)
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
2026-08-11 18:04:31 +00:00
root 9f0a940cee Fix no-mistakes warnings: prose out of YAML fence, Rule 16 rename, table cell fix
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
2026-08-11 18:02:08 +00:00
root 5c3ba31b17 Ship: Koby external-agent ruling (Resolve conflict and merge master)
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
2026-08-11 17:48:36 +00:00
root d0144d6db6 Ship: Koby external-agent ruling (fix api_key_env hygiene + docs)
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s
2026-08-11 17:22:25 +00:00
jerome 1d4f6c8ebb Merge pull request 'Ship: enforcement-rules-reality-fix (captain directive 2026-08-09)' (#45) from ship/enforcement-rules-reality-fix into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 1s
Reviewed-on: #45
2026-08-10 16:38:05 +00:00
root 38a32f8b32 Ship: enforcement-rules-reality-fix (captain directive 2026-08-09)
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
2026-08-10 13:28:53 +00:00
jerome 96b0caa0b3 Merge pull request 'Ship: pm2-zulip contract fixes (captain ruling 2026-08-09)' (#44) from ship/pm2-zulip-contract-fixes into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 1s
Reviewed-on: #44
2026-08-10 01:10:13 +00:00
jerome ea17f64bb4 Merge pull request 'Ship: monitoring contract fixes (2026-08-09)' (#43) from ship/monitoring-contract-fixes into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 0s
Reviewed-on: #43
2026-08-10 01:08:07 +00:00
root c13a15acad Ship: pm2-zulip contract fixes (captain ruling 2026-08-09)
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
2026-08-09 22:11:27 +00:00
root 68f9ebff74 Ship: monitoring contract fixes (2026-08-09)
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 9s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
2026-08-09 21:57:44 +00:00
root 93fb0d8e1e feat(contracts): add MCP URL validation and verify virtual keys on LiteLLM
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 2s
2026-08-07 21:33:59 +00:00
mumuni-bot cb6efa81c3 Merge pull request 'fix(contracts): memory-fixer executes decisions to completion (v2.0.0)' (#42) from fix/memory-fixer-execution-v3 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 1s
2026-08-04 12:32:56 +00:00
mumuni-bot 2dc77ee7da fix(contracts): memory-fixer executes decisions to completion
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
Rewrite memory-fixer to v3 semantics: when Kwame replies to an escalation,
execute the action fully - set state (archived/active), clear or rename the
[REVIEW:] tag, and bump updated_at so the node exits the stale window and
is not re-flagged on the next run. Bump version 1.1.0 -> 2.0.0.
2026-08-04 12:30:37 +00:00
jerome 561c4d98c9 Merge pull request 'fix: correct SearXNG endpoint from dead storepve hostname to live VM109 IP' (#40) from fix/searxng-endpoint into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 0s
Reviewed-on: #40
2026-08-02 16:20:49 +00:00
root c243fcddc3 no-mistakes(document): docs: fix stale SearXNG endpoint in monitoring checks
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s
2026-08-01 15:19:15 +00:00
root 50114f32c0 fix: correct SearXNG endpoint from dead storepve hostname to live VM109 IP (192.168.68.7:8888)
storepve (192.168.68.6) has no service on :8888. SearXNG runs on VM 109
(docker-vm) at 192.168.68.7:8888 (verified HTTP 200). Aligns with every other
contract reference. Missed in PR #39.
2026-08-01 14:56:45 +00:00
36 changed files with 475 additions and 263 deletions
+2 -2
View File
@@ -67,10 +67,10 @@ Two incidents taught us this:
| Contract | Sensitivity | Who can change | | Contract | Sensitivity | Who can change |
|----------|------------|----------------| |----------|------------|----------------|
| `infrastructure-control.prose.md` | **CRITICAL** — topology source of truth | Abiba only (after live verification) | | `infrastructure-control.prose.md` | **CRITICAL** — topology source of truth | Abiba, Tanko (Tanko maintains its own CT row) |
| `proxmox-monitor.prose.md` | **CRITICAL** — deployed monitoring | Abiba only | | `proxmox-monitor.prose.md` | **CRITICAL** — deployed monitoring | Abiba only |
| `hermes-config-template.prose.md` | **HIGH** — all agent configs | Abiba, Mumuni, Tanko | | `hermes-config-template.prose.md` | **HIGH** — all agent configs | Abiba, Mumuni, Tanko |
| `zulip-health.prose.md` | **HIGH** — agent communication | Abiba, Mumuni | | `zulip-health.prose.md` | **HIGH** — agent communication | Abiba, Mumuni, Tanko |
| Other contracts | Normal | Any registered agent | | Other contracts | Normal | Any registered agent |
| `scripts/*.sh` | **HIGH** — runtime scripts | Abiba only | | `scripts/*.sh` | **HIGH** — runtime scripts | Abiba only |
+2 -1
View File
@@ -20,7 +20,8 @@ An agent ran `pct set` without checking `pct config` first and changed the IP
to the wrong value, breaking Zulip. The staleness was harmless until acted on. to the wrong value, breaking Zulip. The staleness was harmless until acted on.
**Layer 3 enforcement — the `safe-mutate` wrapper** (deployed on all 5 agents: **Layer 3 enforcement — the `safe-mutate` wrapper** (deployed on all 5 agents:
Abiba, Tanko, Mumuni, Koby, Koonimo). ALL infrastructure mutations MUST go Abiba, Tanko, Mumuni, Koby, Koonimo — Tanko runs on DSH/DeepSeek Harness since
2026-08-27 but safe-mutate enforcement still applies). ALL infrastructure mutations MUST go
through `safe-mutate`. Raw `sed -i`, `pct set`, `docker compose up through `safe-mutate`. Raw `sed -i`, `pct set`, `docker compose up
--force-recreate`, `kill`, `rm` on infrastructure outside `safe-mutate` is an --force-recreate`, `kill`, `rm` on infrastructure outside `safe-mutate` is an
auditable protocol violation. The wrapper runs a verify command, optionally auditable protocol violation. The wrapper runs a verify command, optionally
+3 -2
View File
@@ -45,9 +45,10 @@ description: >
## Status ## Status
**Active** — for Hermes agents only. This plugin is NOT retired. It remains in service **Active** — for Hermes agents only. This plugin is NOT retired. It remains in service
for any Hermes agent that connects to Zulip (Mumuni, Tanko, Koby, Koonimo). The pi for Hermes agents that connect to Zulip (Mumuni, Koby, Koonimo). The pi
Zulip extension was decommissioned 2026-07-04 but this contract targets the Hermes Zulip extension was decommissioned 2026-07-04 but this contract targets the Hermes
plugin system, which is unaffected. plugin system, which is unaffected. **Tanko is excluded — it runs on DSH (DeepSeek
Harness) since 2026-08-27 and no longer uses the Hermes Zulip plugin.**
## Parameters ## Parameters
+1 -1
View File
@@ -584,7 +584,7 @@ contracts:
verify: curl -sf https://git.sysloggh.net/api/v1/version verify: curl -sf https://git.sysloggh.net/api/v1/version
expect: 200 OK expect: 200 OK
- check: SearXNG reachable - check: SearXNG reachable
verify: curl -sf http://192.168.68.17:8080 verify: curl -sf http://192.168.68.7:8888
expect: 200 OK expect: 200 OK
artifact: infrastructure health report artifact: infrastructure health report
receipt: receipt:
+1 -1
View File
@@ -356,7 +356,7 @@ Postconditions to verify:
}, },
{ {
"check": "SearXNG reachable", "check": "SearXNG reachable",
"verify": "curl -sf http://192.168.68.17:8080", "verify": "curl -sf http://192.168.68.7:8888",
"expect": "200 OK" "expect": "200 OK"
} }
] ]
+2 -3
View File
@@ -274,10 +274,10 @@ one-off GPU builds. No automated post-migration cleanup was in place.
### CT Access (via pct-run) ### CT Access (via pct-run)
| CT | Name | Node | Status | | CT | Name | Node | Status |
|----|------|------|--------| |----|------|------|--------|
| 100 | abiba | hwepve | local | | 100 | abiba | minipve | local |
| 102 | adguard | minipve | ✅ reachable | | 102 | adguard | minipve | ✅ reachable |
| 104 | authentik | minipve | ✅ reachable | | 104 | authentik | minipve | ✅ reachable |
| 105 | kagentz | hwepve | ✅ reachable | | 105 | kagentz | minipve | ✅ reachable |
| 106 | ra-h-os | storepve | ✅ reachable | | 106 | ra-h-os | storepve | ✅ reachable |
| 107 | pbs | storepve | ✅ reachable | | 107 | pbs | storepve | ✅ reachable |
| 108 | media | storepve | ✅ reachable | | 108 | media | storepve | ✅ reachable |
@@ -285,7 +285,6 @@ one-off GPU builds. No automated post-migration cleanup was in place.
| 111 | tdunna | amdpve | ✅ reachable | | 111 | tdunna | amdpve | ✅ reachable |
| 112 | tanko | amdpve | ✅ reachable | | 112 | tanko | amdpve | ✅ reachable |
| 113 | baggy | amdpve | ✅ reachable | | 113 | baggy | amdpve | ✅ reachable |
| 114 | mumuni | hwepve | ✅ reachable |
| 115 | scottdenya | amdpve | ✅ reachable | | 115 | scottdenya | amdpve | ✅ reachable |
| 116 | syslog-api | minipve | ✅ reachable | | 116 | syslog-api | minipve | ✅ reachable |
| 117 | zulip | storepve | ✅ reachable | | 117 | zulip | storepve | ✅ reachable |
+1 -1
View File
@@ -185,7 +185,7 @@ what, and why should I care?
``` ```
❌ "Monitors infrastructure health" ❌ "Monitors infrastructure health"
✅ "Scans all 6 Proxmox nodes and 19 CTs for disk pressure, checks Docker ✅ "Scans all 5 Proxmox nodes and 19 CTs for disk pressure, checks Docker
container health on .7/.116/.17, alerts via Telegram DM on RED/CRITICAL" container health on .7/.116/.17, alerts via Telegram DM on RED/CRITICAL"
``` ```
+14 -11
View File
@@ -14,10 +14,9 @@ description: >
Instability observed near 100K at 256K (now all GPUs at 128K). 128K is the stable ceiling. Instability observed near 100K at 256K (now all GPUs at 128K). 128K is the stable ceiling.
For larger context needs → fall back to external providers (deepseek). For larger context needs → fall back to external providers (deepseek).
VRAM headroom improved: RTX 3090 ~70%, RTX 5070 ~65%. VRAM headroom improved: RTX 3090 ~70%, RTX 5070 ~65%.
UPDATED 2026-07-27: gpu-dense swapped to SmartCode-Fable-5-CoT-Reasoning-QKVO-Qwen-3.6-27B-Distilled UPDATED 2026-08-15: gpu-dense swapped to Qwen3.8-27B-Uncensored-Q4_K_M
(UD-Q3_K_XL, ~14.7GB — Q4 was too large for 24GB VRAM with 128K KV cache). ~50% fewer thinking (~16.8GB, 128K ctx, --spec-type draft-mtp (v2)). Served on .8:8080 under the legacy
tokens via ThinkingCap finetune + Fable 5 CoT distillation for improved coding reasoning. alias qwen3.6-27B-code for LiteLLM routing continuity. Replaces SmartCode-Fable-5-27B.
VRAM ~22.4/24.6GB (91%).
agent: abiba agent: abiba
triggers: triggers:
- on model add/remove - on model add/remove
@@ -90,7 +89,7 @@ When a model is swapped on a GPU, ONLY the infrastructure layer changes — agen
| Alias | GPU | Current Model | Will Route To | | Alias | GPU | Current Model | Will Route To |
|-------|-----|---------------|---------------| |-------|-----|---------------|---------------|
| `strix-moe` | Strix Halo (.15) | qwen3.6-35B-udq4 | Whatever runs on Strix Halo | | `strix-moe` | Strix Halo (.15) | qwen3.6-35B-udq4 | Whatever runs on Strix Halo |
| `gpu-dense` | RTX 3090 (.8) | SmartCode-Fable-5-27B-UD-Q4_K_XL | Whatever runs on RTX 3090 | | `gpu-dense` | RTX 3090 (.8) | Qwen3.8-27B-Uncensored-Q4_K_M | Whatever runs on RTX 3090 |
| `gpu-light` | RTX 5070 (.110) | gemma-4-12b | Whatever runs on RTX 5070 | | `gpu-light` | RTX 5070 (.110) | gemma-4-12b | Whatever runs on RTX 5070 |
**Backward compatibility**: Old model-specific names (qwen3.6-27B-code, gemma-4-12b, qwen3.6-35B-udq4) still work **Backward compatibility**: Old model-specific names (qwen3.6-27B-code, gemma-4-12b, qwen3.6-35B-udq4) still work
@@ -100,7 +99,7 @@ but are deprecated for agent configs. Only the stable aliases survive model swap
| Model | GPU | Host | VRAM | Ctx | KV Cache | Parallel | Batch/Ubatch | Status | | Model | GPU | Host | VRAM | Ctx | KV Cache | Parallel | Batch/Ubatch | Status |
|-------|-----|------|------|-----|----------|----------|-------------|--------| |-------|-----|------|------|-----|----------|----------|-------------|--------|
| SmartCode-Fable-5-27B-UD-Q3_K_XL | RTX 3090 | .8 (llm-gpu) | ~22.4/24.6GB (91%) | **128K** | q4_0 | 1 | 2048/1024 | ✅ healthy | | Qwen3.8-27B-Uncensored-Q4_K_M | RTX 3090 | .8 (llm-gpu) | ~16.8/24.6GB | **128K** | turbo4 | 1 | 2048/1024 | ✅ healthy |
| gemma-4-12b | RTX 5070 | .110 (ocu-llm) | ~7.8/12.2GB (65%) | 128K | q4_0 | 2 | 2048/1024 | ✅ healthy | | gemma-4-12b | RTX 5070 | .110 (ocu-llm) | ~7.8/12.2GB (65%) | 128K | q4_0 | 2 | 2048/1024 | ✅ healthy |
| qwen3.6-35B-udq4 | Strix Halo Vulkan | .15 (amdpve) | ~22GB/64GB | 128K | q4_0 | 1 | 4096/1024 | ✅ 65 tok/s | | qwen3.6-35B-udq4 | Strix Halo Vulkan | .15 (amdpve) | ~22GB/64GB | 128K | q4_0 | 1 | 4096/1024 | ✅ 65 tok/s |
@@ -110,7 +109,7 @@ but are deprecated for agent configs. Only the stable aliases survive model swap
| Model | GPU | Weight | RPM Cap | Timeout | | Model | GPU | Weight | RPM Cap | Timeout |
|-------|-----|--------|---------|---------| |-------|-----|--------|---------|---------|
| SmartCode-Fable-5-27B-UD-Q3_K_XL | RTX 3090 (.8:8080) | **0.55** | 500 | **300s** | | Qwen3.8-27B-Uncensored-Q4_K_M | RTX 3090 (.8:8080) | **0.55** | 500 | **300s** |
| qwen3.6-35B-udq4 | Strix Halo (.15:8080) | **0.30** | 60 | **300s** | | qwen3.6-35B-udq4 | Strix Halo (.15:8080) | **0.30** | 60 | **300s** |
| gemma-4-12b | RTX 5070 (.110:8080) | **0.15** | 200 | **120s** | | gemma-4-12b | RTX 5070 (.110:8080) | **0.15** | 200 | **120s** |
@@ -121,7 +120,7 @@ Note: All syslog-auto entries route directly to GPUs with `api_key: not-needed`.
| Model | RPM Cap | Notes | | Model | RPM Cap | Notes |
|-------|---------|-------| |-------|---------|-------|
| strix-moe (qwen3.6-35B-udq4) | 40 | Tight cap — prevents Strix overload | | strix-moe (qwen3.6-35B-udq4) | 40 | Tight cap — prevents Strix overload |
| SmartCode-Fable-5-27B-UD-Q3_K_XL | 500 | High cap — primary workhorse (replaces qwen3.6-27B-code) | | Qwen3.8-27B-Uncensored-Q4_K_M | 500 | High cap — primary workhorse (replaces qwen3.6-27B-code) |
| gemma-4-12b | 500 | High cap — IQ4_NL+MTP, 122 tok/s | | gemma-4-12b | 500 | High cap — IQ4_NL+MTP, 122 tok/s |
### Stable Aliases (for agent configs — never change) ### Stable Aliases (for agent configs — never change)
@@ -251,9 +250,13 @@ If no SSH access, send Zulip DM via abiba-bot with vault update instructions.
Fix: trigger roster reload via SSH after restart, or rebuild image with startup load_roster(). Fix: trigger roster reload via SSH after restart, or rebuild image with startup load_roster().
- **LiteLLM /metrics**: Requires auth. Prometheus uses `/health/liveliness` as workaround. - **LiteLLM /metrics**: Requires auth. Prometheus uses `/health/liveliness` as workaround.
- **VRAM (2026-07-15)**: RTX 3090 at ~17/24.6GB (~70%) with **128K context** (reduced from 256K 2026-07-17). RTX 5070 at ~7.8/12.2GB (~65%) with 128K context + MTP. Strix Halo at ~7GB/64GB. - **VRAM (2026-07-15)**: RTX 3090 at ~17/24.6GB (~70%) with **128K context** (reduced from 256K 2026-07-17). RTX 5070 at ~7.8/12.2GB (~65%) with 128K context + MTP. Strix Halo at ~7GB/64GB.
- **RTX 3090 (2026-07-27)**: Swapped to SmartCode-Fable-5-27B-UD-Q3_K_XL (14.7GB). Q4 was too large for 24GB VRAM with 128K context + KV cache overhead. Q3 fits at ~22.4GB (91%). Uses standard llama.cpp build b9190 (turboquant b9150 incompatible with qwen3_5 arch). Config: `-c 131072 -ctk q4_0 -ctv q4_0 --flash-attn on --cont-batching`. Sampler: `--temp 0.9 --top-p 0.95 --top-k 60 --min-p 0.0 --repeat-penalty 1.0`. Service: `/home/llmuser/llama-fable-wrapper.sh`. - **RTX 3090 (2026-08-15)**: Swapped to Qwen3.8-27B-Uncensored-Q4_K_M (~16.8GB, replaced
SmartCode-Fable-5-27B-UD-Q3_K_XL). Service: `/home/llmuser/llama-fable-wrapper.sh`.
Served under alias `qwen3.6-27B-code` for LiteLLM routing continuity.
- **RTX 5070 config (2026-07-15)**: Switched to IQ4_NL + MTP draft (Q8_0) at 128K context. Gen speed: 122 tok/s. VRAM: ~7.8/12.2GB (~65%). Service: `/home/llmuser/llama-wrapper.sh`. Config: `--model gemma-4-12b-it-IQ4_NL.gguf --spec-draft-model gemma-4-12b-it-Q8_0-MTP.gguf --spec-type draft-mtp --spec-draft-n-max 4 --ctx-size 131072`. - **RTX 5070 config (2026-07-15)**: Switched to IQ4_NL + MTP draft (Q8_0) at 128K context. Gen speed: 122 tok/s. VRAM: ~7.8/12.2GB (~65%). Service: `/home/llmuser/llama-wrapper.sh`. Config: `--model gemma-4-12b-it-IQ4_NL.gguf --spec-draft-model gemma-4-12b-it-Q8_0-MTP.gguf --spec-type draft-mtp --spec-draft-n-max 4 --ctx-size 131072`.
- **LiteLLM timeout tuning (verified 2026-07-27)**: SmartCode-Fable-5-27B 300s, gemma-4-12b 120s, qwen3.6-27B-code 300s (legacy), qwen3.6-35B-udq4 300s, strix-moe 300s, syslog-auto routes all 300s. Nginx proxy_read_timeout: 600s. Global request_timeout: 300s. - **LiteLLM timeout tuning (verified 2026-08-16)**: Qwen3.8-27B (alias qwen3.6-27B-code)
300s, gemma-4-12b 120s, qwen3.6-35B-udq4 300s, strix-moe 300s, syslog-auto routes all
300s. Nginx proxy_read_timeout: 600s. Global request_timeout: 300s.
- **Strix Halo GPU**: Vulkan is the working backend (ROCm/HIP path abandoned — HSA runtime blocked on Debian 13). Build at `/root/llama.cpp/build-vk/`, commit `4fc4ec5` (2026-07-01), ggml 0.15.3 shared-lib arch. Mesa RADV 25.0.7, KHR_coopmat fast path active. ~70 tok/s gen, 532 tok/s prompt. Service: `strix-server.service` on port 8080, model: `qwen3.6-35B-udq4`, alias `strix-moe`, 128K context, flash-attn + q4 KV, multimodal (mmproj loaded). - **Strix Halo GPU**: Vulkan is the working backend (ROCm/HIP path abandoned — HSA runtime blocked on Debian 13). Build at `/root/llama.cpp/build-vk/`, commit `4fc4ec5` (2026-07-01), ggml 0.15.3 shared-lib arch. Mesa RADV 25.0.7, KHR_coopmat fast path active. ~70 tok/s gen, 532 tok/s prompt. Service: `strix-server.service` on port 8080, model: `qwen3.6-35B-udq4`, alias `strix-moe`, 128K context, flash-attn + q4 KV, multimodal (mmproj loaded).
- **Port conflict detection (2026-07-05)**: All 3 GPU wrappers now detect ghost processes squatting port 8080 before starting. `.8` and `.110` use inline pre-start check in `llama-wrapper.sh`; `.15` uses `/usr/local/bin/port-cleanup.sh` ExecStartPre. Replaces the blanket `pkill -9 -x llama-server` on .15 which would kill ALL llama-server instances regardless of port. Ghost detection was the root cause of .8 crash-looping for 27+ restarts (stale pid 25836 squatting 8080 after OOM kill). - **Port conflict detection (2026-07-05)**: All 3 GPU wrappers now detect ghost processes squatting port 8080 before starting. `.8` and `.110` use inline pre-start check in `llama-wrapper.sh`; `.15` uses `/usr/local/bin/port-cleanup.sh` ExecStartPre. Replaces the blanket `pkill -9 -x llama-server` on .15 which would kill ALL llama-server instances regardless of port. Ghost detection was the root cause of .8 crash-looping for 27+ restarts (stale pid 25836 squatting 8080 after OOM kill).
- **Strix Halo thermal safeguard (2026-07-02)**: `strix-server.service` has `-n 8192` (hard generation cap per request). Without it, `--predict` defaults to -1 (infinity) — a runaway request from .123 (old Mumuni CT114 — now inside Abiba CT100 at .24) decoded 39,868 tokens over 24 min, pushing Tctl to 98°C (crit 89.8°C) and throttling 70→29 t/s. The cap bounds worst-case generation to ~5 min. Do NOT remove `-n` without a replacement ceiling. Sustained load hits ~84°C even at 92s; the APU is fanless/low-flow. Clients MUST also set `max_tokens`. - **Strix Halo thermal safeguard (2026-07-02)**: `strix-server.service` has `-n 8192` (hard generation cap per request). Without it, `--predict` defaults to -1 (infinity) — a runaway request from .123 (old Mumuni CT114 — now inside Abiba CT100 at .24) decoded 39,868 tokens over 24 min, pushing Tctl to 98°C (crit 89.8°C) and throttling 70→29 t/s. The cap bounds worst-case generation to ~5 min. Do NOT remove `-n` without a replacement ceiling. Sustained load hits ~84°C even at 92s; the APU is fanless/low-flow. Clients MUST also set `max_tokens`.
@@ -268,7 +271,7 @@ If no SSH access, send Zulip DM via abiba-bot with vault update instructions.
| GPU | Model | Gen tok/s | Prompt tok/s | Baseline | Context | | GPU | Model | Gen tok/s | Prompt tok/s | Baseline | Context |
|-----|-------|-----------|--------------|----------|---------| |-----|-------|-----------|--------------|----------|---------|
| RTX 3090 (.8) | SmartCode-Fable-5-27B-UD-Q3_K_XL | **TBD** | — | — | **128K** | | RTX 3090 (.8) | Qwen3.8-27B-Uncensored-Q4_K_M | **TBD** | — | — | **128K** |
| RTX 5070 (.110) | gemma-4-12b (IQ4_NL+MTP) | **191** | — | — | **128K** | | RTX 5070 (.110) | gemma-4-12b (IQ4_NL+MTP) | **191** | — | — | **128K** |
| Strix Halo (.15) | qwen3.6-35B-udq4 | **65** | 140 | — | **128K** | | Strix Halo (.15) | qwen3.6-35B-udq4 | **65** | 140 | — | **128K** |
+1 -1
View File
@@ -44,7 +44,7 @@ depends_on:
| Alias | GPU | Host | Model | VRAM | Ctx | tok/s | Role | | Alias | GPU | Host | Model | VRAM | Ctx | tok/s | Role |
|-------|-----|------|-------|------|-----|-------|------| |-------|-----|------|-------|------|-----|-------|------|
| `gpu-dense` | RTX 3090 24GB | ct8 (.8:8080) | ThinkingCap Qwen3.6-27B Q4_K_M + MTP + vision | 21.6/24.6GB (88%) | 128K | 74.9 | Heavy reasoning, code gen | | `gpu-dense` | RTX 3090 24GB | ct8 (.8:8080) | Qwen3.8-27B-Uncensored-Q4_K_M (alias qwen3.6-27B-code) | ~16.8/24.6GB | 128K | — | Heavy reasoning, code gen |
| `gpu-light` | RTX 5070 12GB | ct110 (.110:8080) | HauhauCS Gemma4-12B QAT Q4_K_M + MTP draft | 10.1/12.2GB (83%) | 128K | 169.6 | Vision, web extract, light tasks | | `gpu-light` | RTX 5070 12GB | ct110 (.110:8080) | HauhauCS Gemma4-12B QAT Q4_K_M + MTP draft | 10.1/12.2GB (83%) | 128K | 169.6 | Vision, web extract, light tasks |
| `strix-moe` | Strix Halo 64GB | ct15 (.15:8080) | qwen3.6-35B-udq4 | ~10/64GB (16%) | 128K | 62.9 | Compression, summarization, long docs | | `strix-moe` | Strix Halo 64GB | ct15 (.15:8080) | qwen3.6-35B-udq4 | ~10/64GB (16%) | 128K | 62.9 | Compression, summarization, long docs |
+5 -4
View File
@@ -24,8 +24,8 @@ done
| Agent | CT | Node | IP | LiteLLM Alias | Key Source | Platform | | Agent | CT | Node | IP | LiteLLM Alias | Key Source | Platform |
|-------|-----|------|-----|---------------|------------|----------| |-------|-----|------|-----|---------------|------------|----------|
| Tanko | 112 | amdpve | .122 | `tanko` | Infisical vault | Hermes | | Tanko | 112 | amdpve | .122 | `tanko` | Infisical vault | **DSH** (DeepSeek Harness) |
| Mumuni | 100 | hwepve | .24 | `mumuni` | Infisical vault | Hermes | | Mumuni | 100 | minipve | .24 | `mumuni` | Infisical vault | Hermes |
| Koby | 111 | amdpve | .129 | `koby` | Infisical vault | **Hermes** | | Koby | 111 | amdpve | .129 | `koby` | Infisical vault | **Hermes** |
| Koonimo | 113 | amdpve | .114 | `koonimo` | Infisical vault | Hermes | | Koonimo | 113 | amdpve | .114 | `koonimo` | Infisical vault | Hermes |
| Shumba | — | 192.168.68.119 | N/A | N/A (DeepSeek) | Hermes (RETIRED — CT119 now Infisical vault) | | Shumba | — | 192.168.68.119 | N/A | N/A (DeepSeek) | Hermes (RETIRED — CT119 now Infisical vault) |
@@ -53,9 +53,10 @@ Agent (systemd) → LITELLM_API_KEY → LiteLLM (:116/v1) → Router (:9000) →
## Config Pattern — Mandatory Fields ## Config Pattern — Mandatory Fields
### For Hermes Agents (Tanko, Mumuni, Koonimo) ### For Hermes Agents (Mumuni, Koonimo)
Every agent's `/root/.hermes/config.yaml` (or `/home/jerome/.hermes/config.yaml`) MUST have: Every Hermes agent's `/root/.hermes/config.yaml` (or `/home/jerome/.hermes/config.yaml`) MUST have:
(Tanko is excluded — migrated to DSH/DeepSeek Harness on 2026-08-27, no longer uses Hermes config.)
### 1. Main Model ### 1. Main Model
```yaml ```yaml
+49 -20
View File
@@ -5,8 +5,7 @@ description: >
Standard Hermes configuration template for Syslog Solution LLC agents. Standard Hermes configuration template for Syslog Solution LLC agents.
Enforces shared infrastructure setup (Firecrawl, SearXNG, local models, Enforces shared infrastructure setup (Firecrawl, SearXNG, local models,
RA-H OS MCP) while keeping agent-specific API keys and model choices. RA-H OS MCP) while keeping agent-specific API keys and model choices.
UPDATED 2026-07-16: Compression model is the stable alias `strix-moe` (NOT `ornith-1.0-35b`, UPDATED 2026-08-07: Added Rule 15 (MCP Validation) from the 2026-08-07 keyless-MCP incident.
which LiteLLM does not serve). All 3 GPUs verified at 128K (reduced from 256K 2026-07-17 for stability).
Added Rule 12 (Context-Issue Diagnostic) + Rule 13 (.env fallback enforcement) from the Added Rule 12 (Context-Issue Diagnostic) + Rule 13 (.env fallback enforcement) from the
2026-07-16 Mumuni root-cause investigation (WAL #1300). 2026-07-16 Mumuni root-cause investigation (WAL #1300).
UPDATED 2026-07-12: GPU workload redistributed. Compression → Strix Halo. RTX 3090 context verified at 128K. Infisical .env fallback required (Rule 3/13). UPDATED 2026-07-12: GPU workload redistributed. Compression → Strix Halo. RTX 3090 context verified at 128K. Infisical .env fallback required (Rule 3/13).
@@ -16,7 +15,7 @@ description: >
- template_version: "2.1.0" - template_version: "2.1.0"
- last_applied: timestamp - last_applied: timestamp
- agents_configured: ["tanko", "mumuni", "abiba", "koby", "koonimo", "kagenz0"] - agents_configured: ["mumuni", "abiba", "koby", "koonimo", "kagenz0"] # tanko removed 2026-08-27 (now on DSH/DeepSeek Harness)
- agent_keys: map (see Agent Keys section) - agent_keys: map (see Agent Keys section)
- infra_endpoints_verified: array - infra_endpoints_verified: array
@@ -31,7 +30,7 @@ Sub-agent profiles inherit auth from the main config — no separate keys needed
| Agent | Key Alias | Host | SSH | Sub-Agents | | Agent | Key Alias | Host | SSH | Sub-Agents |
|-------|-----------|------|-----|-----------| |-------|-----------|------|-----|-----------|
| Tanko | `tanko-*` | 192.168.68.122 | jerome@.122 | — | | Tanko | `tanko` | CT 112 (.122) | jerome@.122 | — |
| Mumuni | `mumuni` | 192.168.68.24 | root@.24 | 6 profiles ✱ | | Mumuni | `mumuni` | 192.168.68.24 | root@.24 | 6 profiles ✱ |
| Abiba | `abiba-pi` | 192.168.68.24 | local | — | | Abiba | `abiba-pi` | 192.168.68.24 | local | — |
| Koby | `koby` | CT 111 (tdunna) | Zulip | — | | Koby | `koby` | CT 111 (tdunna) | Zulip | — |
@@ -48,7 +47,7 @@ Sub-agent profiles inherit auth from the main config — no separate keys needed
| Component | Endpoint | Purpose | | Component | Endpoint | Purpose |
|---|---|---| |---|---|---|
| Firecrawl | `http://192.168.68.7:3002/` | Web content extraction | | Firecrawl | `http://192.168.68.7:3002/` | Web content extraction |
| SearXNG | `http://storepve:8888` | Privacy-respecting web search | | SearXNG | `http://192.168.68.7:8888` | Privacy-respecting web search |
| LiteLLM | `http://192.168.68.116/v1` | Unified model gateway (via nginx) | | LiteLLM | `http://192.168.68.116/v1` | Unified model gateway (via nginx) |
| LiteLLM (NetBird) | `https://litellm.sysloggh.net/v1` | Alternative (may have 502 issues) | | LiteLLM (NetBird) | `https://litellm.sysloggh.net/v1` | Alternative (may have 502 issues) |
| RA-H OS MCP | `http://192.168.68.65:3100/mcp` | Knowledge graph bridge | | RA-H OS MCP | `http://192.168.68.65:3100/mcp` | Knowledge graph bridge |
@@ -97,10 +96,14 @@ work immediately after restart.
model: model:
default: <agent_model> # e.g., strix-moe, qwen3.6-27B-code, syslog-auto default: <agent_model> # e.g., strix-moe, qwen3.6-27B-code, syslog-auto
provider: harness provider: harness
base_url: http://192.168.68.116/v1 base_url: http://192.168.68.116/litellm/v1 # Rule 5 (2026-08-09): canonical authenticated path; /v1 also OK
api_key_env: LITELLM_API_KEY # Injected via infisical run -- wrapper api_key_env: LITELLM_API_KEY # Injected via infisical run -- wrapper
max_tokens: 4096 # ⚠️ CRITICAL: Prevents unbounded generation max_tokens: 4096 # ⚠️ CRITICAL: Prevents unbounded generation
context_length: 131072 # For syslog-auto (all GPUs at 128K for stability). context_length: 131072 # For syslog-auto (all GPUs at 128K for stability).
# ⚠️ MANDATORY: Hermes probes unknown models from 256K
# and falls back to 256K when /v1/models lacks a context
# field (llama-server does). Without this override, agents
# silently run syslog-auto at 256K (verified 2026-08-09).
# Set 65536 if using gemma-4-12b directly (tight VRAM). # Set 65536 if using gemma-4-12b directly (tight VRAM).
fallback_providers: fallback_providers:
@@ -143,7 +146,7 @@ compression:
# ─── Auxiliary Tasks (CONSISTENCY RULE) ─── # ─── Auxiliary Tasks (CONSISTENCY RULE) ───
# All auxiliary services MUST use identical model, base_url, and api_key_env: # All auxiliary services MUST use identical model, base_url, and api_key_env:
# model: gpu-light # stable alias (NOT raw "gemma-4-12b") # model: gpu-light # stable alias (NOT raw "gemma-4-12b")
# base_url: http://192.168.68.116/v1 # base_url: http://192.168.68.116/litellm/v1 # Rule 5 (2026-08-09): canonical authenticated; /v1 also OK
# api_key_env: LITELLM_API_KEY # api_key_env: LITELLM_API_KEY
# Do NOT use syslog-auto for auxiliary tasks — it routes to the primary GPU. # Do NOT use syslog-auto for auxiliary tasks — it routes to the primary GPU.
# gpu-light = RTX 5070 (12B), freeing the Strix Halo for agent reasoning. # gpu-light = RTX 5070 (12B), freeing the Strix Halo for agent reasoning.
@@ -154,20 +157,20 @@ auxiliary:
vision: vision:
provider: harness provider: harness
model: gpu-light # stable alias for RTX 5070 (was raw gemma-4-12b) model: gpu-light # stable alias for RTX 5070 (was raw gemma-4-12b)
base_url: http://192.168.68.116/v1 base_url: http://192.168.68.116/litellm/v1 # Rule 5 (2026-08-09): canonical authenticated; /v1 also OK
api_key_env: LITELLM_API_KEY api_key_env: LITELLM_API_KEY
timeout: 60 timeout: 60
download_timeout: 30 download_timeout: 30
web_extract: web_extract:
provider: harness provider: harness
model: gpu-light # stable alias for RTX 5070 model: gpu-light # stable alias for RTX 5070
base_url: http://192.168.68.116/v1 base_url: http://192.168.68.116/litellm/v1 # Rule 5 (2026-08-09): canonical authenticated; /v1 also OK
api_key_env: LITELLM_API_KEY api_key_env: LITELLM_API_KEY
timeout: 30 timeout: 30
compression: compression:
provider: harness provider: harness
model: syslog-auto # MUST match compression.model above. Stable alias for Strix Halo (weighted pool). model: syslog-auto # MUST match compression.model above. Stable alias for Strix Halo (weighted pool).
base_url: http://192.168.68.116/v1 # Rule 5: /v1 NOT /litellm/v1 base_url: http://192.168.68.116/litellm/v1 # Rule 5 (2026-08-09): canonical authenticated path; /v1 also OK
api_key_env: LITELLM_API_KEY api_key_env: LITELLM_API_KEY
timeout: 300 # gpu-fleet: 300s for large-history summarization (was 60) timeout: 300 # gpu-fleet: 300s for large-history summarization (was 60)
@@ -177,14 +180,14 @@ auxiliary:
delegation: delegation:
model: gpu-dense # stable alias for RTX 3090 (was raw qwen3.6-27B-code) model: gpu-dense # stable alias for RTX 3090 (was raw qwen3.6-27B-code)
provider: harness provider: harness
base_url: http://192.168.68.116/v1 base_url: http://192.168.68.116/litellm/v1 # Rule 5 (2026-08-09): canonical authenticated; /v1 also OK
api_key_env: LITELLM_API_KEY api_key_env: LITELLM_API_KEY
# ─── Custom Provider ─── # ─── Custom Provider ───
custom_providers: custom_providers:
- name: harness - name: harness
model: syslog-auto # weighted pool (default) model: syslog-auto # weighted pool (default)
base_url: http://192.168.68.116/v1 base_url: http://192.168.68.116/litellm/v1 # Rule 5 (2026-08-09): canonical authenticated; /v1 also OK
api_key_env: LITELLM_API_KEY api_key_env: LITELLM_API_KEY
api_mode: chat_completions api_mode: chat_completions
``` ```
@@ -234,10 +237,13 @@ The following MUST be identical across ALL profiles:
- When main config uses `api_key_env`, sub-agents automatically use it - When main config uses `api_key_env`, sub-agents automatically use it
- This means key rotation only touches ONE vault secret (`LITELLM_API_KEY`) - This means key rotation only touches ONE vault secret (`LITELLM_API_KEY`)
### Rule 5: Main Config Base URL ### Rule 5: Main Config Base URL (UPDATED 2026-08-09)
|- Use direct IP: `http://192.168.68.116/v1` |- Use the authenticated LiteLLM path: `http://192.168.68.116/litellm/v1` (canonical, captain-approved migration)
|- Legacy `http://192.68.68.116/v1` also works — nginx fronts BOTH paths with key auth
(verified 2026-08-09: 401 without key, 200 with key, on both /v1 and /litellm/v1)
|- Both locations have `proxy_read_timeout 600s` (verified in harness-nginx nginx.conf) —
the old "60s timeout on /litellm/" claim was stale and is retracted
|- NOT the NetBird URL (`litellm.sysloggh.net`) — can cause 502 when NetBird is down |- NOT the NetBird URL (`litellm.sysloggh.net`) — can cause 502 when NetBird is down
|- NOT the old path (`/litellm/v1`) — nginx now routes `/v1` directly
### Rule 6: max_tokens Is Required (Thermal Safety) ### Rule 6: max_tokens Is Required (Thermal Safety)
- **Every Hermes config MUST set `model.max_tokens: 4096`** — this is non-negotiable - **Every Hermes config MUST set `model.max_tokens: 4096`** — this is non-negotiable
@@ -258,7 +264,7 @@ The following MUST be identical across ALL profiles:
fall back to other GPUs if Strix gets hot. Both `compression.model` and `auxiliary.compression.model` fall back to other GPUs if Strix gets hot. Both `compression.model` and `auxiliary.compression.model`
MUST be `syslog-auto`. MUST be `syslog-auto`.
- All auxiliary services MUST use identical routing: - All auxiliary services MUST use identical routing:
- `base_url: http://192.168.68.116/v1` (Rule 5: `/v1`, NOT `/litellm/v1`) - `base_url: http://192.168.68.116/litellm/v1` (Rule 5, 2026-08-09: canonical authenticated; `/v1` also OK)
- `api_key_env: LITELLM_API_KEY` - `api_key_env: LITELLM_API_KEY`
- **Do NOT use `syslog-auto`** for auxiliary tasks — it routes unpredictably - **Do NOT use `syslog-auto`** for auxiliary tasks — it routes unpredictably
- **Compression on Strix Halo**: The strix-moe alias routes to Strix Halo - **Compression on Strix Halo**: The strix-moe alias routes to Strix Halo
@@ -290,6 +296,7 @@ The following MUST be identical across ALL profiles:
- See `devops-hermes-compression` skill for full reference - See `devops-hermes-compression` skill for full reference
### Rule 10: Default Model Must Be `syslog-auto` (All Agents) ### Rule 10: Default Model Must Be `syslog-auto` (All Agents)
- **Koby Exception**: Per captain ruling 2026-08-11, Koby is a DeepSeek-primary external agent; its primary model remains `deepseek-v4-flash` (via api.deepseek.com to preserve DeepSeek-specific reasoning, while other sections follow Rule 10.
- **Hermes agents**: `model.default: syslog-auto`, `custom_providers[0].model: syslog-auto` - **Hermes agents**: `model.default: syslog-auto`, `custom_providers[0].model: syslog-auto`
- **pi agents**: `defaultModel: syslog-auto` in `settings.json`, first model in `models.json` - **pi agents**: `defaultModel: syslog-auto` in `settings.json`, first model in `models.json`
- `syslog-auto` is the LiteLLM routing model — it load-balances between strix-moe - `syslog-auto` is the LiteLLM routing model — it load-balances between strix-moe
@@ -318,10 +325,13 @@ verify ALL FOUR of these against the live config. They are the only root causes
1. **max_context_window correct?** — BOTH `compression.max_context_window` AND `context.max_context_window` 1. **max_context_window correct?** — BOTH `compression.max_context_window` AND `context.max_context_window`
MUST be `131072` (all GPUs are 128K). A value of `262144` causes instability near 100K and must NOT be used. MUST be `131072` (all GPUs are 128K). A value of `262144` causes instability near 100K and must NOT be used.
~83K instead of ~170K. Check: `grep -n max_context_window ~/.hermes/config.yaml` ~83K instead of ~170K. Check: `grep -n max_context_window ~/.hermes/config.yaml`
2. **base_url uses /v1 NOT /litellm/v1?** — `custom_providers[0].base_url`, `delegation.base_url`, 2. **base_url uses authenticated path?** — `custom_providers[0].base_url`, `delegation.base_url`,
and ALL `auxiliary.*.base_url` MUST be `http://192.168.68.116/v1` (Rule 5). nginx `/litellm/` and ALL `auxiliary.*.base_url` MUST be `http://192.168.68.116/litellm/v1` (Rule 5, canonical)
has a 60s default timeout → 504 on any inference >60s; `/v1/` has 600s. or `http://192.168.68.116/v1` (legacy, still authenticated via nginx). BOTH verified 200 with
Check: `grep -n 'litellm/v1' ~/.hermes/config.yaml` (must return NOTHING) key + 600s proxy_read_timeout on 2026-08-09. Never bare `:4000` direct.
Check: `grep -nE 'base_url: http://192.168.68.116(:4000)?/v1' ~/.hermes/config.yaml` — the
ONLY paths allowed are `/v1` or `/litellm/v1` (both via nginx :80).
`:4000` or missing `litellm/v1`/`v1` prefix = violation.
3. **LITELLM_API_KEY valid?** — The key must be a real LiteLLM key (`sk-` + 64 hex, 67 chars). 3. **LITELLM_API_KEY valid?** — The key must be a real LiteLLM key (`sk-` + 64 hex, 67 chars).
Malformed values (e.g. `sk-_SWAl_Vu_…`, 47 chars) return 401 → DeepSeek fallback. Malformed values (e.g. `sk-_SWAl_Vu_…`, 47 chars) return 401 → DeepSeek fallback.
Verify: `curl -s -o /dev/null -w '%{http_code}' -H "Authorization: Bearer $LITELLM_API_KEY" http://192.168.68.116/v1/models` (must be 200) Verify: `curl -s -o /dev/null -w '%{http_code}' -H "Authorization: Bearer $LITELLM_API_KEY" http://192.168.68.116/v1/models` (must be 200)
@@ -396,6 +406,15 @@ curl -s -o /dev/null -w '%{http_code}' -H "Authorization: Bearer $K" http://192.
- **Audit script**: Run `python3 /root/prose-contracts/audit-hermes-config.py <config.yaml>` - **Audit script**: Run `python3 /root/prose-contracts/audit-hermes-config.py <config.yaml>`
before and after any config change to catch this and all other rule violations. before and after any config change to catch this and all other rule violations.
### Rule 15: MCP Endpoint and Header Validation (ADDED 2026-08-07)
- Every MCP server entry must point at the correct endpoint:
- ra-h-os = http://192.168.68.65:3100/mcp
- litellm = https://litellm.sysloggh.net/mcp
- MCP entries must carry a REAL key value in the header.
- Avoid using env-var names like LITELLM_API_KEY in the header; they do not resolve for MCP
endpoints and result in "Malformed API Key" floods.
- Ensure the header value is the actual key (e.g., `sk-...`).
## Execution ## Execution
1. **Check current config** — Read the target agent's config.yaml 1. **Check current config** — Read the target agent's config.yaml
@@ -405,3 +424,13 @@ curl -s -o /dev/null -w '%{http_code}' -H "Authorization: Bearer $K" http://192.
5. **Set model choice** — Per agent's workload 5. **Set model choice** — Per agent's workload
6. **Verify** — curl all shared endpoints, test the model with the new key 6. **Verify** — curl all shared endpoints, test the model with the new key
7. **Report** — What was changed, preserved, custom 7. **Report** — What was changed, preserved, custom
### Rule 16: Koby Configuration (DeepSeek-primary)
(Ref: See Rule 10 for default model behavior, with Koby exception)
Koby uses a split-model architecture:
- Primary Model: `deepseek-v4-flash` via `api.deepseek.com` (for reasoning)
- Auxiliary Models: `gpu-light` (vision/web_extract) and `syslog-auto` (compression)
- Key Hygiene: `api_key_env` is strictly `LITELLM_API_KEY` or `DEEPSEEK_API_KEY`
- Constraint: Do NOT touch Koby's primary model/provider/compression settings unless explicitly ruled by the captain.
+14 -13
View File
@@ -34,11 +34,12 @@ Syslog is migrating away from **unauthenticated direct access** to the shared in
| Path | Auth | Status | | Path | Auth | Status |
|------|------|--------| |------|------|--------|
| `http://192.168.68.116/v1` | None (direct) | ❌ **DEPRECATED** — being phased out | | `http://192.168.68.116/v1` | Bearer `sk-*` key (nginx-fronted) | ✅ **VALID** — authenticated via nginx :80 (verified 2026-08-09: 401 without key, 200 with) |
| `http://192.168.68.116/litellm/v1/responses` | Bearer `sk-*` key | ✅ **CURRENT** — authenticated LiteLLM proxy | | `http://192.168.68.116/litellm/v1` | Bearer `sk-*` key (nginx-fronted) | ✅ **CURRENT / CANONICAL** — captain-approved migration target; 600s proxy_read_timeout (verified) |
| `http://192.168.68.116:4000/v1` | Bearer `sk-*` key (direct container) | ❌ **FORBIDDEN** — bypasses nginx; port 4000 direct is not a config path |
All harness/litellm providers MUST use the authenticated `/litellm/v1/responses` path. All harness/litellm providers MUST use an authenticated nginx-fronted path (`/litellm/v1` canonical, `/v1` legacy-valid).
Any `base_url` pointing to bare `/v1` on 192.168.68.116 is a **migration violation**. Any `base_url` pointing at `:4000` or a bare IP without nginx is a **migration violation**.
### 🔥 CRITICAL: Double-Path Bug (2026-07-10) ### 🔥 CRITICAL: Double-Path Bug (2026-07-10)
@@ -190,22 +191,22 @@ litellm_settings:
|-------|-----|-----|---------------|------------|--------|-----------------|---------------| |-------|-----|-----|---------------|------------|--------|-----------------|---------------|
| Tanko | 112 | .122 | `tanko` | Infisical vault | ✅ Fixed | `infisical run` | 20:17 UTC Jul 5 | | Tanko | 112 | .122 | `tanko` | Infisical vault | ✅ Fixed | `infisical run` | 20:17 UTC Jul 5 |
| Mumuni | 100 (abiba) | .24 | `mumuni` | Infisical vault | ✅ Fixed | Pi Hermes gateway | 2026-07-27 | | Mumuni | 100 (abiba) | .24 | `mumuni` | Infisical vault | ✅ Fixed | Pi Hermes gateway | 2026-07-27 |
| Koby | 111 | ? | `koby` | Infisical vault | ✅ Fixed | `infisical run` | 23:30 UTC Jul 5 | | Koby | 111 | .129 | `koby` | Infisical vault | ✅ Fixed (DeepSeek-primary) | `infisical run` | 23:30 UTC Jul 5 |
| Koonimo | 113 | ? | `koonimo` | Infisical vault | ✅ Fixed | `infisical run` (migrated 2026-07-11) | 2026-07-11 | | Koonimo | 113 | .114 | `koonimo` | Infisical vault | ✅ Fixed | `infisical run` (migrated 2026-07-11) | 2026-08-09 |
| Abiba | 100 | .65 | `abiba-pi` | Infisical vault | ✅ N/A (pi native) | — | 19:44 UTC Jul 5 | | Abiba | 100 | .65 | `abiba-pi` | Infisical vault | ✅ N/A (pi native) | — | 19:44 UTC Jul 5 |
| Kagenz0 | 105 | ? | — | — | ❌ DOWN | — | 19:14 EDT Jul 4 | | Kagenz0 | 105 | .14 | — | — | ❌ DOWN | — | 19:14 EDT Jul 4 |
> **Note**: CT hostnames (tdunna, baggy) differ from agent identities (koby, koonimo). > **Note**: CT hostnames (tdunna, baggy) differ from agent identities (koby, koonimo).
> LiteLLM key aliases use agent identity, not CT hostname. > LiteLLM key aliases use agent identity, not CT hostname.
### Migration Status: Authenticated Path ### Migration Status: Authenticated Path
| Agent | `/litellm/v1/responses` | Deprecated `/v1` | Status | | Agent | `/litellm/v1` | Legacy `/v1` | Status |
|-------|--------------------------|--------------------|--------| |-------|--------------|-------------|--------|
| Mumuni | ✅ 5 sections | 0 | ✅ Authenticated | | Mumuni | ✅ harness provider | ✅ auxiliary on /v1 (valid) | ✅ Authenticated (verified 2026-08-09) |
| Tanko | ⚠️ No SSH access | — | Needs check | | Tanko | ✅ 5 sections | 0 | ✅ Migrated 2026-08-08, keys 200 |
| Koby | ⚠️ No route to host | — | Needs check | | Koby | ✅ custom provider (harness name) | — | ✅ External DeepSeek primary (intentional, captain ruling 2026-08-11) |
| Koonimo | ⚠️ Connection timed out | — | Needs check | | Koonimo | ✅ .114 (baggy) | — | ✅ 128K context applied 2026-08-09 |
### Systemd Service Pattern (2026-07-11 — vault migration) ### Systemd Service Pattern (2026-07-11 — vault migration)
+4 -3
View File
@@ -28,7 +28,7 @@ connectivity recovery including end-to-end DM validation.
| Param | Type | Required | Default | Description | | Param | Type | Required | Default | Description |
|-------|------|----------|---------|-------------| |-------|------|----------|---------|-------------|
| `target` | string | yes | — | Agent name: `mumuni`, `tanko`, `koby`, or `shumba` | | `target` | string | yes | — | Agent name: `mumuni`, `koby`, or `shumba` (Tanko excluded — on DSH since 2026-08-27, no Hermes plugin) |
| `branch` | string | no | `master` | Git branch to pull (overridable for pinning) | | `branch` | string | no | `master` | Git branch to pull (overridable for pinning) |
## Maintains ## Maintains
@@ -56,7 +56,7 @@ connectivity recovery including end-to-end DM validation.
| Host | CT | Proxmox | IP (direct) | Hermes Home | User | | Host | CT | Proxmox | IP (direct) | Hermes Home | User |
|------|-----|---------|-------------|-------------|------| |------|-----|---------|-------------|-------------|------|
| Mumuni | CT100 | — | 192.168.68.24 | /root/.hermes | root | | Mumuni | CT100 | — | 192.168.68.24 | /root/.hermes | root |
| Tanko | CT112 | amdpve | 192.168.68.122 | /home/jerome/.hermes | jerome | | Tanko | CT112 | amdpve | 192.168.68.122 | /home/jerome/.hermes | jerome | *(DSH since 2026-08-27 — historical, plugin retired on this host)* |
| Koby | CT111 | amdpve | 192.168.68.129 | /root/.hermes | root | | Koby | CT111 | amdpve | 192.168.68.129 | /root/.hermes | root |
| Shumba | — | — | 192.168.68.119 | /home/lucky/.hermes | lucky | | Shumba | — | — | 192.168.68.119 | /home/lucky/.hermes | lucky |
@@ -121,7 +121,8 @@ cp plugins/platforms/zulip/adapter.py \
plugins/platforms/zulip/plugin.yaml \ plugins/platforms/zulip/plugin.yaml \
{{hermes_home}}/hermes-agent/plugins/platforms/zulip/ {{hermes_home}}/hermes-agent/plugins/platforms/zulip/
# Fix ownership (Tanko only — runs as jerome user) # Fix ownership (was Tanko-only, runs as jerome user)
# RETIRED 2026-08-27: tanko no longer uses the Hermes Zulip plugin (DSH).
[ "{{target}}" = "tanko" ] && chown -R jerome:jerome \ [ "{{target}}" = "tanko" ] && chown -R jerome:jerome \
{{hermes_home}}/hermes-agent/plugins/platforms/zulip/ {{hermes_home}}/hermes-agent/plugins/platforms/zulip/
+10 -8
View File
@@ -2,8 +2,10 @@
kind: function kind: function
name: hermes-zulip-restore name: hermes-zulip-restore
description: > description: >
Restores Zulip connectivity for any Hermes agent (Mumuni CT100, Tanko CT112, Restores Zulip connectivity for any Hermes agent (Mumuni CT100, Koby CT111,
Koby CT111, Shumba on Lucky's mini PC). Deploys the zulip-platform adapter to the correct bundled plugin Shumba on Lucky's mini PC). Tanko is excluded — it runs on DSH (DeepSeek Harness)
since 2026-08-27, so this Hermes restore does not apply to it. Deploys the
zulip-platform adapter to the correct bundled plugin
path, verifies env credentials, restarts the gateway, and confirms Zulip path, verifies env credentials, restarts the gateway, and confirms Zulip
connects. Run this whenever a Hermes agent stops responding on Zulip or after connects. Run this whenever a Hermes agent stops responding on Zulip or after
a fresh agent deployment. a fresh agent deployment.
@@ -23,7 +25,7 @@ gateway restart, and connection validation.
| Param | Type | Required | Default | Description | | Param | Type | Required | Default | Description |
|-------|------|----------|---------|-------------| |-------|------|----------|---------|-------------|
| `target` | string | yes | — | Agent name: `mumuni`, `tanko`, `koby`, or `shumba` | | `target` | string | yes | — | Agent name: `mumuni`, `koby`, or `shumba` (Tanko excluded — DSH since 2026-08-27) |
## Maintains ## Maintains
@@ -36,7 +38,7 @@ gateway restart, and connection validation.
- `_strip_html` function present in `<hermes-agent>/plugins/platforms/zulip/adapter.py` - `_strip_html` function present in `<hermes-agent>/plugins/platforms/zulip/adapter.py`
- All three adapter files (__init__.py, adapter.py, plugin.yaml) present at bundled path - All three adapter files (__init__.py, adapter.py, plugin.yaml) present at bundled path
- Zulip env vars set in `~/.hermes/.env` (or `/home/jerome/.hermes/.env` for Tanko) - Zulip env vars set in `~/.hermes/.env` (or `/home/jerome/.hermes/.env` for Tanko, historical — DSH since 2026-08-27)
- Gateway restarted and zulip platform reports state `connected` - Gateway restarted and zulip platform reports state `connected`
- HTML stripping enabled for `/approve` and `/deny` slash command support - HTML stripping enabled for `/approve` and `/deny` slash command support
@@ -51,8 +53,8 @@ gateway restart, and connection validation.
| Host | CT | Proxmox | IP (direct) | Hermes Home | User | | Host | CT | Proxmox | IP (direct) | Hermes Home | User |
|------|-----|---------|-------------|-------------|------| |------|-----|---------|-------------|-------------|------|
| Mumuni | CT100 (abiba) | hwepve | 192.168.68.24 | /root/.hermes | root | | Mumuni | CT100 (abiba) | minipve | 192.168.68.24 | /root/.hermes | root |
| Tanko | CT112 | amdpve | 192.168.68.122 | /home/jerome/.hermes | jerome | | Tanko | CT112 | amdpve | 192.168.68.122 | /home/jerome/.hermes | jerome | *(DSH since 2026-08-27 — historical, restore does not apply)* |
| Koby | CT111 | amdpve | 192.168.68.129 | /root/.hermes | root | | Koby | CT111 | amdpve | 192.168.68.129 | /root/.hermes | root |
| Shumba | — | — | 192.168.68.119 | /home/lucky/.hermes | lucky | | Shumba | — | — | 192.168.68.119 | /home/lucky/.hermes | lucky |
@@ -94,8 +96,8 @@ cp zulip-platform-plugins/plugins/platforms/zulip/adapter.py \
zulip-platform-plugins/plugins/platforms/zulip/plugin.yaml \ zulip-platform-plugins/plugins/platforms/zulip/plugin.yaml \
<HERMES_HOME>/hermes-agent/plugins/platforms/zulip/ <HERMES_HOME>/hermes-agent/plugins/platforms/zulip/
# Fix ownership (Tanko only) # Fix ownership (was Tanko-only; RETIRED 2026-08-27 — tanko on DSH, no Hermes plugin)
chown -R jerome:jerome <HERMES_HOME>/hermes-agent/plugins/platforms/zulip/ # Tanko only chown -R jerome:jerome <HERMES_HOME>/hermes-agent/plugins/platforms/zulip/ # Tanko only (historical)
# Clean up # Clean up
rm -rf /tmp/zulip-deploy rm -rf /tmp/zulip-deploy
+34 -27
View File
@@ -3,7 +3,7 @@ kind: pattern
name: infrastructure-control name: infrastructure-control
description: > description: >
Full infrastructure monitoring and control pattern covering the Full infrastructure monitoring and control pattern covering the
6-node Proxmox cluster, 3 Docker ecosystems (22 containers), 5-node Proxmox cluster, 3 Docker ecosystems (22 containers),
NFS storage, and network services. Defines monitors, remediations, NFS storage, and network services. Defines monitors, remediations,
and the access matrix for all environments. and the access matrix for all environments.
@@ -13,10 +13,11 @@ description: >
against the live system. Policy fields are authoritative. See the against the live system. Policy fields are authoritative. See the
`verify-before-mutate` skill. `verify-before-mutate` skill.
**Last verified:** 2026-07-24 — corrected Gitea IP (.17 not .110), **Last verified:** 2026-08-15 — hwepve removed from Tabiri cluster
AdGuard IP (.10 not .102), AdGuard placement (minipve not acerpve), (now 5 nodes: minipve, amdpve, storepve, acerpve, ocupve). hwepve
Abiba placement (hwepve not amdpve), added hwepve as 6th node, (192.168.68.4) is a standalone PVE node + NetBird routing peer;
added dns.sysloggh.net route. London relocation pending. CTs 100 (abiba) and 105 (kagentz) moved
to minipve.
--- ---
# Infrastructure Control Pattern # Infrastructure Control Pattern
@@ -34,21 +35,21 @@ description: >
┌─────────────┐ ┌──────────────┐ ┌──────────────┐ ┌─────────────┐ ┌──────────────┐ ┌──────────────┐
│ Abiba │ │ Tanko │ │ Mumuni │ │ Abiba │ │ Tanko │ │ Mumuni │
│ (pi) │ │ (Hermes) │ │ (Hermes) │ │ (pi) │ │ (Hermes) │ │ (Hermes) │
│ CT 100 │ │ CT 112 │ │ CT 114 │ │ CT 100 │ │ CT 112 │ │ CT 100 │
└──────┬──────┘ └──────┬───────┘ └──────┬───────┘ └──────┬──────┘ └──────┬───────┘ └──────┬───────┘
│ │ │ │ │ │
└──────────────────┼────────────────────┘ └──────────────────┼────────────────────┘
▼ ▼
┌──────────────────────────────────────┐ ┌────────────────────────────────────┐
│ Proxmox Cluster API │ │ Proxmox Cluster API │
│ minipve.sysloggh.net:443 │ │ minipve.sysloggh.net:443 │
│ (monitoring@pve!mumuni token) │ │ (monitoring@pve!mumuni token) │
└────┬──────┬──────┬──────┬──────┬─────┘ └────┬──────┬──────┬──────┬──────────┘
│ │ │ │ │ │ │ │ │
┌────┘ ┌────┘ ┌────┘ ┌────┘ ┌────┘ ┌────┘ ┌────┘ ┌────┘ ┌────┘
▼ ▼ ▼ ▼ ▼ ▼ ▼ ▼ ▼
minipve amdpve storepve acerpve ocupve hwepve minipve amdpve storepve acerpve ocupve
(.12) (.15) (.6) (.9) (.5) (.4) (.12) (.15) (.6) (.9) (.5)
▼ ▼
┌─────────────────────────────────────────────┐ ┌─────────────────────────────────────────────┐
@@ -100,21 +101,29 @@ description: >
## Section 2: Proxmox Cluster — Monitoring ## Section 2: Proxmox Cluster — Monitoring
### Nodes (6) ### Nodes (5)
| Node | IP | CPU | RAM | VMs/CTs | Role | | Node | IP | CPU | RAM | VMs/CTs | Role |
|------|----|-----|-----|---------|------| |------|----|-----|-----|---------|------|
| minipve | .12 | 16C | 30GB | authentik, gitea, syslog-api, infisical-vault, jitsi | Auth, git, messaging | | minipve | .12 | 16C | 30GB | abiba, kagentz, authentik, gitea, syslog-api, infisical-vault, jitsi | Auth, git, messaging |
| amdpve | .15 | 32C | 62GB | tanko, tdunna, baggy, scottdenya | Agents, compute | | amdpve | .15 | 32C | 62GB | tanko, tdunna, baggy, scottdenya | Agents, compute |
| storepve | .6 | 28C | 31GB | docker-vm, ra-h-os, PBS, media, jdownloader, zulip | Docker, storage, chat | | storepve | .6 | 28C | 31GB | docker-vm, ra-h-os, PBS, media, jdownloader, zulip | Docker, storage, chat |
| acerpve | .9 | 28C | 31GB | llm-gpu | GPU VMs | | acerpve | .9 | 28C | 31GB | llm-gpu | GPU VMs |
| ocupve | .5 | 12C | 14GB | ocu-llm | GPU VMs | | ocupve | .5 | 12C | 14GB | ocu-llm | GPU VMs |
| hwepve | .4 | 12C | 15GB | abiba, kagentz, (mumuni CT 114 stopped) | Agents (new node) |
> **Note:** CTs on storepve include jdownloader (CT 118). AdGuard (CT 102) is on > **Note:** CTs on storepve include jdownloader (CT 118). AdGuard (CT 102) is on
> minipve at .10, not acerpve. Abiba (CT 100) is on hwepve, not amdpve. Mumuni > minipve at .10, not acerpve. Abiba (CT 100) and kagentz (CT 105) are on
> (CT 114) is on hwepve (currently stopped), not minipve. Mumuni also has a > minipve (moved from hwepve 2026-08-15). Mumuni runs inside Abiba CT100
> second instance on minipve at .123 — distinguish by CT ID, not hostname. > (.24); CT 114 (mumuni) no longer exists in the cluster.
>
> **hwepve (192.168.68.4) — STANDALONE (removed from Tabiri 2026-08-15):**
> Huawei MateBook 16 (KLVL-WXX9), pve-manager/9.2.10, kernel 7.0.14-8-pve.
> Zero VMs/CTs. Being relocated to London as a standalone PVE node + NetBird
> routing peer (relocation pending). Localizations applied: timezone
> Europe/London, lid-switch ignore, sleep/suspend/hibernate targets masked,
> cluster-shared storage removed (remaining: local, local-lvm, storage,
> mediastore). prometheus-node-exporter active on :9100; net.ipv4.ip_forward=1;
> NetBird client not yet installed (enrollment pending setup key).
### Checks (every 5 min) ### Checks (every 5 min)
@@ -592,21 +601,20 @@ ssh root@192.168.68.110 "systemctl restart llama-server"
| CT | Name | Node | IP | Role | Agent | | CT | Name | Node | IP | Role | Agent |
|----|------|------|----|------|-------| |----|------|------|----|------|-------|
| 100 | abiba | **hwepve** | .24 | Pi agent | ✅ pi | | 100 | abiba | minipve | .24 | Pi agent | ✅ pi |
| 101 | llm-gpu | acerpve | .8 | GPU RTX 3090 | ❌ | | 101 | llm-gpu | acerpve | .8 | GPU RTX 3090 | ❌ |
| 102 | adguard | **minipve** | **.10** | DNS | ❌ | | 102 | adguard | **minipve** | **.10** | DNS | ❌ |
| 103 | ocu-llm | ocupve | .110 | GPU RTX 5070 | ❌ | | 103 | ocu-llm | ocupve | .110 | GPU RTX 5070 | ❌ |
| 104 | authentik | minipve | .11 | OIDC | ❌ | | 104 | authentik | minipve | .11 | OIDC | ❌ |
| 105 | kagentz | **hwepve** | — | Agent Zero | ✅ | | 105 | kagentz | minipve | — | Agent Zero | ✅ |
| 106 | ra-h-os | storepve | .65 | KG bridge | ✅ MCP | | 106 | ra-h-os | storepve | .65 | KG bridge | ✅ MCP |
| 107 | pbs | storepve | — | Backups | ❌ | | 107 | pbs | storepve | — | Backups | ❌ |
| 108 | media | storepve | — | Media | ❌ | | 108 | media | storepve | — | Media | ❌ |
| 109 | docker-vm | storepve | .7 | Docker host | ❌ | | 109 | docker-vm | storepve | .7 | Docker host | ❌ |
| 110 | gitea | minipve | **.17** | Git | ❌ | | 110 | gitea | minipve | **.17** | Git | ❌ |
| 111 | tdunna | amdpve | .129 | Hermes agent | ✅ | | 111 | tdunna | amdpve | .129 | Hermes agent | ✅ |
| 112 | tanko | amdpve | .122 | Hermes agent | ✅ | | 112 | tanko | amdpve | .122 | DSH (DeepSeek Harness) agent | ✅ |
| 113 | baggy | amdpve | .114 | Hermes agent | ✅ | | 113 | baggy | amdpve | .114 | Hermes agent | ✅ |
| 114 | mumuni | **hwepve** | .123 | Hermes agent (stopped) | ✅ |
| 115 | scottdenya | amdpve | .75 | Denya OneCare | ❌ | | 115 | scottdenya | amdpve | .75 | Denya OneCare | ❌ |
| 116 | syslog-api | minipve | .116 | LiteLLM + Grafana | ❌ | | 116 | syslog-api | minipve | .116 | LiteLLM + Grafana | ❌ |
| 117 | zulip | storepve | .19 | Chat | ❌ | | 117 | zulip | storepve | .19 | Chat | ❌ |
@@ -631,15 +639,14 @@ Source of truth: `/root/scripts/pct-run.sh` or `prose-contracts/scripts/pct-run.
| CT | Name | Node | pct-run | | CT | Name | Node | pct-run |
|-----|------|------|---------| |-----|------|------|---------|
| 100 | abiba | hwepve | `pct-run 100` | | 100 | abiba | minipve | `pct-run 100` |
| 105 | kagentz | hwepve | `pct-run 105` | | 105 | kagentz | minipve | `pct-run 105` |
| 111 | tdunna | amdpve | `pct-run 111` | | 111 | tdunna | amdpve | `pct-run 111` |
| 112 | tanko | amdpve | `pct-run 112` | | 112 | tanko | amdpve | `pct-run 112` |
| 113 | baggy | amdpve | `pct-run 113` | | 113 | baggy | amdpve | `pct-run 113` |
| 115 | scottdenya | amdpve | `pct-run 115` | | 115 | scottdenya | amdpve | `pct-run 115` |
| 104 | authentik | minipve | `pct-run 104` | | 104 | authentik | minipve | `pct-run 104` |
| 110 | gitea | minipve | `pct-run 110` | | 110 | gitea | minipve | `pct-run 110` |
| 114 | mumuni | hwepve | `pct-run 114` |
| 116 | syslog-api | minipve | `pct-run 116` | | 116 | syslog-api | minipve | `pct-run 116` |
| 106 | ra-h-os | storepve | `pct-run 106` | | 106 | ra-h-os | storepve | `pct-run 106` |
| 107 | proxmox-backup | storepve | `pct-run 107` | | 107 | proxmox-backup | storepve | `pct-run 107` |
@@ -652,7 +659,7 @@ GPU bare-metal hosts (.8 acerpve, .110 ocupve, .15 amdpve) are NOT CTs — use S
ssh root@192.168.68.8 # RTX 3090 ssh root@192.168.68.8 # RTX 3090
ssh root@192.168.68.110 # RTX 5070 ssh root@192.168.68.110 # RTX 5070
ssh root@192.168.68.15 # Strix Halo ssh root@192.168.68.15 # Strix Halo
ssh root@192.168.68.4 # hwepve (abiba, kagentz, mumuni) ssh root@192.168.68.4 # hwepve — standalone London node + NetBird routing peer (relocation pending)
``` ```
## Section 7: Agent Health Check (consolidated — 2026-07-05) ## Section 7: Agent Health Check (consolidated — 2026-07-05)
+2 -1
View File
@@ -140,7 +140,8 @@ After ALL updates (apt + images + restarts), verify every critical service is ba
| Zulip | `curl -sf https://chat.sysloggh.net/api/v1/server_settings` | 200 OK | | Zulip | `curl -sf https://chat.sysloggh.net/api/v1/server_settings` | 200 OK |
| Gitea | `curl -sf https://git.sysloggh.net/api/v1/version` | 200 OK | | Gitea | `curl -sf https://git.sysloggh.net/api/v1/version` | 200 OK |
| PM2 processes | `pm2 jlist` (CT 100) | all pi-agent processes `online` | | PM2 processes | `pm2 jlist` (CT 100) | all pi-agent processes `online` |
| Hermes gateways | SSH to Mumuni CT 100, Tanko CT 112; `systemctl is-active hermes-gateway` | `active` for each | | Hermes gateways | SSH to Mumuni CT 100; `systemctl is-active hermes-gateway` | `active` |
| Tanko (DSH) | DSH harness service on CT 112 (.122) | `active` |
Regression check: every service that was GREEN in `health-baseline` must still be GREEN. A service that was already RED (and caused a preflight abort) is excluded — but Phase 0 should have aborted before we got here. Regression check: every service that was GREEN in `health-baseline` must still be GREEN. A service that was already RED (and caused a preflight abort) is excluded — but Phase 0 should have aborted before we got here.
+41 -6
View File
@@ -7,13 +7,15 @@ description: >
from nvidia-smi (.8, .110) and amdgpu_top (.15). LiteLLM metrics from nvidia-smi (.8, .110) and amdgpu_top (.15). LiteLLM metrics
via existing /metrics Prometheus endpoint. via existing /metrics Prometheus endpoint.
DEPLOYMENT STATUS (2026-07-09): DEPLOYMENT STATUS (2026-08-09):
✅ Core stack deployed: Prometheus + Grafana + pve/node/docker exporters ✅ Core stack deployed: Prometheus + Grafana + pve/node/docker exporters
(via proxmox-monitor contract). Grafana at :3001, 5 scrape targets active. (via proxmox-monitor contract). Grafana at :3001, all scrape targets active.
❌ GPU exporters NOT deployed: gpu-exporter crash-loops on .15, ✅ GPU exporters DEPLOYED: all 3 GPU hosts (.8/.110/.15) run exporters on
NVIDIA sidecar exporters (.8/.110:9400) never installed. :9400 (nvidia_gpu_exporter / amdgpu exporter) — verified 200 on 2026-08-09.
Router falls back to direct GPU /health probes. ✅ LiteLLM /metrics scraping live (success_callback: prometheus; auth via
⚠️ This contract is target-state aspirational — not as-built. master key) + Alertmanager + Zulip bridge (alerts-infra) added 2026-08-09.
⚠️ This contract is target-state aspirational — but GPU export + alerting
are now as-built (verified 2026-08-09).
As-built GPU monitoring is via gpu-monitor contract (port 9100 poll). As-built GPU monitoring is via gpu-monitor contract (port 9100 poll).
version: 1.0.0 version: 1.0.0
--- ---
@@ -100,6 +102,39 @@ GPU .8 (RTX 3090) GPU .110 (RTX 5070) GPU .15 (Strix Halo)
- Stack persists across reboots (systemd for exporters, Docker restart policy) - Stack persists across reboots (systemd for exporters, Docker restart policy)
## Execution ## Execution
### check-health
**RUN LIVE, NEVER ECHO — every dispatch must execute the probes below with real tool calls; never repeat a prior report unless a live probe fails.**
```bash
# Zulip API health (POST ping)
curl -s -o /dev/null -w '%{http_code}' -X POST https://chat.sysloggh.net/api/v1/messages -u 'abiba-bot@chat.sysloggh.net:KEY'
# Expected: 200 (HTTP 000 = unreachable/cache)
# PM2 process health
pm2 jlist
# Expected: 5/5 online (abiba-telegram, abiba-zulip, zulip-watchdog, gitea-runner, spoton-service)
# GPU exporters (may be down per DEPLOYMENT STATUS)
curl -s http://192.168.68.8:9400/metrics && echo " - OK" || echo " - FAIL"
curl -s http://192.168.68.110:9400/metrics && echo " - OK" || echo " - FAIL"
curl -s http://192.168.68.15:9400/metrics && echo " - OK" || echo " - FAIL"
# Prometheus targets
curl -s http://192.168.68.116:9090/api/v1/targets | jq '.data.activeTargets'
# Expected: All targets UP (may show some down if exporters not deployed)
# Grafana health
curl -s http://192.168.68.116:3001/api/health | jq '{status, version}'
# Expected: {"status":"ok","version":"..."}
# LiteLLM metrics
curl -s http://192.168.68.116:4001/metrics | head -20
# Expected: Prometheus-formatted metrics output
```
**Report format**: Summarize actual results from each probe. If any probe returns non-200 or empty output, flag as alert.
### Phase 1: GPU Exporters ### Phase 1: GPU Exporters
+7 -8
View File
@@ -2,7 +2,7 @@
kind: responsibility kind: responsibility
name: infrastructure-update name: infrastructure-update
description: > description: >
Autonomous system-wide update contract covering all 6 Proxmox nodes, Autonomous system-wide update contract covering all 5 Proxmox nodes,
15+ containers/VMs, and 4 Docker ecosystems. Updates apt packages, 15+ containers/VMs, and 4 Docker ecosystems. Updates apt packages,
Docker images, and container stacks in safe waves with health checks Docker images, and container stacks in safe waves with health checks
and automatic rollback on failure. and automatic rollback on failure.
@@ -56,11 +56,10 @@ Before ANY update wave:
| amdpve (.15) | Proxmox node | `apt update && apt upgrade -y` | 5 min | | amdpve (.15) | Proxmox node | `apt update && apt upgrade -y` | 5 min |
| acerpve (.9) | Proxmox node | `apt update && apt upgrade -y` | 5 min | | acerpve (.9) | Proxmox node | `apt update && apt upgrade -y` | 5 min |
| ocupve (.5) | Proxmox node | `apt update && apt upgrade -y` | 5 min | | ocupve (.5) | Proxmox node | `apt update && apt upgrade -y` | 5 min |
| hwepve (.4) | Proxmox node | `apt update && apt upgrade -y` | 5 min |
| CT 100 (.24) | Abiba (pi) | `apt update && apt upgrade -y` | 3 min | | CT 100 (.24) | Abiba (pi) | `apt update && apt upgrade -y` | 3 min |
| CT 116 (.116) | syslog-api (LiteLLM host) | `apt update && apt upgrade -y` | 3 min | | CT 116 (.116) | syslog-api (LiteLLM host) | `apt update && apt upgrade -y` | 3 min |
| CT 112 (tanko, amdpve) | Tanko | `apt update && apt upgrade -y` | 3 min | | CT 112 (tanko, amdpve) | Tanko | `apt update && apt upgrade -y` | 3 min |
| CT 100 (mumuni/abiba, hwepve) | Mumuni | `apt update && apt upgrade -y` | 3 min | | CT 100 (mumuni/abiba, minipve) | Mumuni | `apt update && apt upgrade -y` | 3 min |
| VM 101 (.8) | llm-gpu (RTX 3090) | `apt update && apt upgrade -y` | 3 min | | VM 101 (.8) | llm-gpu (RTX 3090) | `apt update && apt upgrade -y` | 3 min |
| VM 103 (.110) | ocu-llm (RTX 5070) | `apt update && apt upgrade -y` | 3 min | | VM 103 (.110) | ocu-llm (RTX 5070) | `apt update && apt upgrade -y` | 3 min |
@@ -163,9 +162,9 @@ Before Wave 1, snapshot these files:
/etc/systemd/system/strix-server.service (amdpve .15 — strix-moe) /etc/systemd/system/strix-server.service (amdpve .15 — strix-moe)
/etc/systemd/system/llama-server.service (VM 101 .8, VM 103 .110) /etc/systemd/system/llama-server.service (VM 101 .8, VM 103 .110)
# Hermes agent configs (key enforcement — 2026-07-10) # Hermes agent configs (key enforcement — 2026-07-10)
/root/.hermes/config.yaml (Mumuni CT 114, Tanko CT 112, etc.) /root/.hermes/config.yaml (Mumuni inside CT 100, Tanko CT 112, etc.)
/root/.config/systemd/user/hermes-gateway.service (Mumuni CT 114 — EnvironmentFile fixed) /root/.config/systemd/user/hermes-gateway.service (Mumuni inside CT 100 — EnvironmentFile fixed)
/etc/environment (Mumuni CT 114 — LITELLM_API_KEY) /etc/environment (Mumuni inside CT 100 — LITELLM_API_KEY)
``` ```
## MCP Gateway (2026-07-10) ## MCP Gateway (2026-07-10)
@@ -219,7 +218,7 @@ When LiteLLM is upgraded to a version supporting per-key MCP grants:
## Success Criteria ## Success Criteria
- [ ] All 6 PVE nodes updated, no reboot-loop - [ ] All 5 PVE nodes updated, no reboot-loop
- [ ] All VMs/CTs running post-update - [ ] All VMs/CTs running post-update
- [ ] All Docker containers healthy (VM 109 + CT 116 + CT 117) - [ ] All Docker containers healthy (VM 109 + CT 116 + CT 117)
- [ ] LiteLLM inference passing (syslog-auto test) - [ ] LiteLLM inference passing (syslog-auto test)
@@ -235,7 +234,7 @@ After completion, send Zulip DM:
``` ```
📋 Infrastructure Update — YYYY-MM-DD 📋 Infrastructure Update — YYYY-MM-DD
Updated: 6 PVE nodes, 12 CTs/VMs, 30+ containers Updated: 5 PVE nodes, 12 CTs/VMs, 30+ containers
Security fixes: N CVEs patched Security fixes: N CVEs patched
Downtime: <service> <duration> Downtime: <service> <duration>
Failures: none / <details> Failures: none / <details>
+9
View File
@@ -18,6 +18,15 @@ description: >
Designed as a reusable contract for any Syslog agent. Designed as a reusable contract for any Syslog agent.
Source of truth: gpu-fleet.prose.md Source of truth: gpu-fleet.prose.md
## Monitoring / Alerting (as-built 2026-08-09)
- LiteLLM /metrics scrape job: requires `litellm_settings.success_callback: [prometheus]`
(failure_callback alone does NOT mount /metrics — verified 2026-08-09).
Scraped by Prometheus with Bearer master key; endpoint returns 307 → /metrics/.
- Alertmanager (harness-alertmanager :9093) + zulip-bridge (:9102) deliver
firing alerts to #agent-hub > alerts-infra via abiba-bot. Added 2026-08-09.
- Prometheus node job covers ALL 5 PVE nodes (.5/.6/.9/.12/.15:9100).
--- ---
## Architecture (v4.0.0 — Direct: nginx → LiteLLM → GPU) ## Architecture (v4.0.0 — Direct: nginx → LiteLLM → GPU)
+4 -5
View File
@@ -1,7 +1,7 @@
--- ---
name: memory-audit-maintenance name: memory-audit-maintenance
kind: responsibility kind: responsibility
description: Shared memory audit and maintenance contract for all Hermes agents (Mumuni, Tanko, Koby, Koonimo). Each agent runs it against its own isolated memory files — no cross-agent access, no shared state. Detects staleness, enforces writer registry, and rotates canary tokens. description: Shared memory audit and maintenance contract for Hermes agents (Mumuni, Koby, Koonimo). Tanko is no longer a Hermes agent (now on DSH/DeepSeek Harness since 2026-08-27) and uses DSH-native memory, so it is excluded from this Hermes roster. Each agent runs it against its own isolated memory files — no cross-agent access, no shared state. Detects staleness, enforces writer registry, and rotates canary tokens.
id: 067NC4KG01RG50R40M30E20918 id: 067NC4KG01RG50R40M30E20918
--- ---
@@ -15,11 +15,10 @@ Autonomously audit and reorganize an agent's native memory (MEMORY.md, USER.md,
### Scope ### Scope
This contract is the **Hermes Agent standard** for memory maintenance. It is shared across all Hermes agents (Mumuni, Tanko, Koby, Koonimo). Each agent runs it against its own memory files only no cross-agent access, no shared state, no shared ledger, no shared canary. The contract is the standard; each agent enforces it independently with fully isolated data. This contract is the **Hermes Agent standard** for memory maintenance. It is shared across Hermes agents (Mumuni, Koby, Koonimo). **Tanko is excluded — it migrated to DSH (DeepSeek Harness) on 2026-08-27 and now uses DSH-native memory (mnemon), not `~/.hermes/memories/`.** Each agent runs it against its own memory files only no cross-agent access, no shared state, no shared ledger, no shared canary. The contract is the standard; each agent enforces it independently with fully isolated data.
**Agent Roster:** **Agent Roster (Hermes):**
- Mumuni - Mumuni
- Tanko
- Koby (CT 111 / tdunna) - Koby (CT 111 / tdunna)
- Koonimo (CT 113 / baggy) - Koonimo (CT 113 / baggy)
@@ -343,4 +342,4 @@ return {
### Per-Agent Notes ### Per-Agent Notes
Each Hermes agent (Mumuni, Tanko, Tdunna, Baggy) runs this contract against its own `~/.hermes/memories/` directory. The contract is identical across agents, but all data is fully isolated: separate ledgers, separate writer registries, separate canaries. If a new agent is added to the roster, it must be listed in `### Scope` above and given its own isolated memory directory. Each Hermes agent (Mumuni, Tdunna/Koby, Baggy/Koonimo) runs this contract against its own `~/.hermes/memories/` directory. The contract is identical across agents, but all data is fully isolated: separate ledgers, separate writer registries, separate canaries. If a new agent is added to the roster, it must be listed in `### Scope` above and given its own isolated memory directory. **Tanko is not covered by this contract — it runs on DSH (DeepSeek Harness) since 2026-08-27 and uses DSH-native memory.**
+152 -46
View File
@@ -2,69 +2,175 @@
kind: pattern kind: pattern
name: memory-fixer name: memory-fixer
description: > description: >
Auto-fix low-hanging fruit in the graph. No judgment calls — only deterministic Level 1 operations. Auto-fix low-hanging fruit in the RA-H OS knowledge graph. No judgment calls — only deterministic Level 1 operations.
Escalate anything that needs Kwame's input. Escalate anything that needs Kwame's input. Executes confirmed Kwame decisions to completion (state + updated_at).
version: 1.1.0 version: 2.0.0
--- ---
# Memory Fixer # Memory Fixer
> **Canonical copy:** `/root/.hermes/contracts/memory-fixer-v3.md` (used by the `memory-fixer-daily` cron job). This file is the institutional record of the same contract. When the two diverge, treat the v3 source in `/root/.hermes/contracts/` as executable truth.
## Purpose ## Purpose
Auto-fix low-hanging fruit in the graph. No judgment calls — only deterministic Level 1 operations. Escalate anything that needs Kwame's input. Auto-fix low-hanging fruit in the graph. No judgment calls — only deterministic Level 1 operations. Escalate anything that needs Kwame's input. When Kwame replies to an escalation, **execute the decision to completion** (update state and timestamps), never leaving a node in review-pending forever.
## Level 0 Auto-Deletes (Allowed Without Approval) **Type:** Write-only (Level 1 fixes only)
Ephemeral heartbeat and log nodes that violate "Logs NEVER go in the graph": **Scope:** RA-H OS knowledge graph (192.168.68.65)
**Schedule:** Daily at 8 AM ET
**Escalation:** Level 2+ to Kwame as task items
- `[LITELLM-HEALTH]`, `[GPU-SELF-HEAL]`, `[PM2-SELF-HEAL]` ## Key Design Decision
- `[PROXMOX-MONITOR]`, `[GPU-MONITOR]`, `[INFRA-MONITOR]`, `[AGENT-HEALTH]`, `[DISK-GC]`
- `[WAL]` entries older than 30 days
**Condition:** node must be an orphan (no edges). Deleting a connected node risks breaking other nodes. The `updateNode` tool's `metadata` field performs a **restricted merge** — the `state` key only accepts `'processed'` or `'not_processed'`. Additionally, new metadata keys cannot be added via the merge.
**Method:** direct SQLite on `.65` (MCP has no delete tool): **Solution:** Use the `description` field to tag stale nodes with review actions, since `description` is a simple string overwritable via `updateNode`.
```bash
ssh root@192.168.68.65 "sqlite3 /root/.local/share/RA-H/db/rah.sqlite \"
DELETE FROM nodes WHERE id IN (
SELECT id FROM nodes WHERE id NOT IN (SELECT from_node_id FROM edges)
AND id NOT IN (SELECT to_node_id FROM edges)
AND title LIKE '[LITELLM-HEALTH]%' -- add more prefixes as needed
);\""
```
## Level 1 Auto-Fixes (No Judgment Required) **Tag Format:** `[REVIEW: action] original description text...`
### 1. Missing `type` Field Where `action` is one of:
For nodes with content but no `metadata.type`: - `archive` — node is stale and should be archived
- Title contains "Proxmox" or "infrastructure" → `type: infrastructure` - `refresh` — node is stale and should be refreshed (infrastructure)
- Title contains "skill" or "how to" or "guide" → `type: skill` - `keep` — node has been confirmed as current
- Title contains "doc" or "template" or "brand" → `type: documentation` - `merge` — node is a duplicate candidate
- Title starts with "WAL:" or "TASK:" → `type: note`
- Title starts with "[LEARN]" → `type: documentation`
- Otherwise → `type: note` (default)
### 2. Missing `tenant` / `namespace` **Query for finding review-tagged nodes:**
For any node with NULL tenant or namespace:
```sql ```sql
UPDATE nodes SELECT id, title, description
SET metadata = json_set( FROM nodes
COALESCE(metadata, '{}'), WHERE description LIKE '[REVIEW:%';
'$.tenant', 'syslogsolution',
'$.namespace', 'syslogsolution'
)
WHERE json_extract(metadata, '$.tenant') IS NULL
OR json_extract(metadata, '$.namespace') IS NULL;
``` ```
### 3. Staleness State Transitions ## Level 1 Auto-Fixes (No Kwame Decision Needed)
Using the type-based windows from the memory-monitor contract:
- Nodes stale > their window → transition to `state: review_pending` ### 1. Missing `type` Auto-Classification
- Nodes in `review_pending` for >7 days → escalate to Kwame (Level 2)
```sql
SELECT id, title,
CASE
WHEN title LIKE '%infrastructure%' OR title LIKE '%proxmox%' OR title LIKE '%setup%' THEN 'infrastructure'
WHEN title LIKE '%skill%' OR title LIKE '%how to%' OR title LIKE '%guide%' THEN 'skill'
WHEN title LIKE '%doc%' OR title LIKE '%template%' OR title LIKE '%brand%' THEN 'documentation'
WHEN title LIKE 'WAL:%' OR title LIKE 'TASK:%' THEN 'note'
WHEN title LIKE '%[LEARN]%' THEN 'documentation'
ELSE 'note'
END as auto_type
FROM nodes
WHERE json_extract(metadata, '$.type') IS NULL;
```
### 2. Missing `namespace` Auto-Population
```sql
SELECT id, title, json_extract(metadata, '$.tenant') as tenant
FROM nodes
WHERE json_extract(metadata, '$.namespace') IS NULL;
```
### 3. Staleness Review Tagging
Using the type-based windows from the memory-monitor contract, tag nodes stale beyond their window. **Only process a maximum of 10 nodes per run** to avoid overwhelming Kwame. Prioritize infrastructure first, then dynamic, then ephemeral.
**Exclusion Rules:**
- Nodes with `state` = `review_pending`, `deprecated`, `archived`, or `not_processed` are NOT processed
- Nodes whose `description` already starts with `[REVIEW:` are NOT re-processed
```sql
SELECT id, title, json_extract(metadata, '$.type') as node_type,
CAST(julianday('now') - julianday(updated_at) AS INTEGER) as days_stale,
CASE
WHEN json_extract(metadata, '$.type') IN ('infrastructure', 'deployment', 'system', 'system-health') THEN 'refresh'
WHEN json_extract(metadata, '$.type') IN ('business', 'philosophy', 'research', 'learning', 'learn', 'investigation', 'analysis', 'project') THEN 'refresh'
ELSE 'archive'
END as suggested_action
FROM nodes
WHERE updated_at < datetime('now',
CASE
WHEN json_extract(metadata, '$.type') IN ('infrastructure', 'deployment', 'system', 'system-health') THEN '-14 days'
WHEN json_extract(metadata, '$.type') IN ('skill', 'documentation', 'template', 'protocol-enforcement', 'prd', 'architecture') THEN '-90 days'
WHEN json_extract(metadata, '$.type') IN ('note', 'wal', 'WAL', 'task', 'TASK', 'event') THEN '-30 days'
WHEN json_extract(metadata, '$.type') IN ('business', 'philosophy', 'research', 'learning', 'learn', 'investigation', 'analysis', 'project') THEN '-120 days'
WHEN json_extract(metadata, '$.type') IN ('deprecated-relay', 'audit', 'audit-report', 'incident', 'incident-report') THEN '-3650 days'
ELSE '-45 days'
END
)
AND json_extract(metadata, '$.state') NOT IN ('review_pending', 'deprecated', 'archived', 'not_processed')
AND (description IS NULL OR description NOT LIKE '[REVIEW:%')
ORDER BY days_stale ASC
LIMIT 10;
```
For each identified node, call `updateNode(id, { description: "[REVIEW: action] " + originalDescription })`.
## Level 2 Escalations (Kwame Decision Required) ## Level 2 Escalations (Kwame Decision Required)
1. **Nodes in `review_pending` >7 days** — Archive, refresh, or keep?
2. **Orphan Nodes >90 days old** — Delete or Connect? 1. **Stale nodes** flagged with `[REVIEW: …]` — Archive, refresh, or keep?
3. **Potential Duplicate Nodes** — Same title or >70% overlap. Merge or Keep? 2. **Duplicate Nodes** (same title or >70% title overlap) — Merge or keep?
4. **Conflicting Metadata** — Content suggests one tenant but metadata says another. 3. **Orphan Nodes >90 days old** — Archive or connect?
## Reporting Format
The fixer reports to Kwame via this Zulip DM:
```
🦅 Memory Fixer — [HH:MM UTC]
Level 1 fixes applied:
- Missing type: X nodes classified
- Missing namespace: Y nodes populated
Stale nodes needing review (max 10):
1. [Node #XXX] Title — X days stale, SUGGEST: refresh
2. [Node #YYY] Title — Y days stale, SUGGEST: archive
...
Duplicates needing decision:
1. [Node #AAA] vs [Node #BBB] — Same title
Orphans >90 days:
1. [Node #EEE] Title — X days stale, orphaned
Reply with:
- "archive #XXX, #YYY" to mark for archive
- "archive all" to archive all stale nodes listed
- "keep #XXX" to confirm a node is current
- "merge #AAA into #BBB" to merge duplicates
- "refresh #XXX" to mark as current
```
## Execution on Next Run
The fixer reads Kwame's previous response and **executes the decision to completion** — it must not leave a node in review-pending forever. Tagging alone is NOT enough; each confirmed decision must also update `state` and `updated_at` so the node drops out of the stale window on the next run.
> ⚠️ `updateNode` cannot set `state` to non-standard values (restricted to `processed`/`not_processed`) and cannot add metadata keys. For state transitions and `updated_at` bumps, use **direct SSH + SQLite** on the bridge host:
> ```bash
> ssh root@192.168.68.65 "sqlite3 /root/.local/share/RA-H/db/rah.sqlite \"UPDATE nodes SET metadata = json_set(metadata, '$.state', '<state>'), updated_at = datetime('now') WHERE id = <id>;\""
> ```
> Use `updateNode` only for description/source/title/link edits.
Decision → completed action mapping:
| Kwame reply | Description change | State | `updated_at` |
|---|---|---|---|
| `archive #XXX` | replace `[REVIEW: archive] ` → `[ARCHIVED] ` prefix | `archived` | bumped to now |
| `keep #XXX` / `refresh #XXX` | **clear the `[REVIEW: …]` tag entirely** | `active` | bumped to now |
| `merge #AAA into #BBB` | set `[REVIEW: merge_into #BBB]` on #AAA, then follow manual merge workflow | handled manually | bumped to now |
| `archive all` | apply the archive row to every node listed in the prior report | `archived` | bumped to now |
**Why `updated_at` must be bumped (critical):** the Level-1 staleness query keys off `updated_at < now - window`. If the fixer clears the tag but leaves a stale `updated_at`, the node is immediately re-flagged on the very next run and the cycle repeats forever. Bumping `updated_at` to now pushes the node back to the front of the window.
**Exclusion after action:** once an action is applied, the node's description no longer starts with `[REVIEW:` (archive → `[ARCHIVED]`, refresh/keep → original text), so it is not re-processed.
After all actions are applied, verify with:
```sql
SELECT id, json_extract(metadata, '$.state') FROM nodes WHERE description LIKE '[REVIEW:%';
```
The result must be 0 rows when all decisions are executed. Report what was done.
## Checks
- **State integrity:** archived nodes have `state: archived` + `[ARCHIVED]` prefix; kept nodes are `state: active` without a `[REVIEW:]` tag.
- **No review-pending forever:** after executing Kwame's decisions, `[REVIEW:%` node count must be 0.
- **Timestamps:** every executed decision bumps `updated_at`, so the node exits the stale window on the next run.
## Logging ## Logging
Every Level 1 fix logged to `~/.hermes/logs/memory-fixer/YYYY-MM-DD.md` Every Level 1 fix logged to `~/.hermes/logs/memory-fixer/YYYY-MM-DD.md`
+7 -7
View File
@@ -6,7 +6,7 @@ description: >
delegation, verification, and delivery. Defines when to delegate, which delegation, verification, and delivery. Defines when to delegate, which
worker to use for what, how to handle failures, and the kanban board worker to use for what, how to handle failures, and the kanban board
protocol. Enforces context-window discipline and separation of concerns. protocol. Enforces context-window discipline and separation of concerns.
Runs on Mumuni (inside Abiba CT100, hwepve, .24) via Hermes agent (Pi + Hermes Zulip gateway). Runs on Mumuni (inside Abiba CT100, minipve, .24) via Hermes agent (Pi + Hermes Zulip gateway).
version: 1.0.0 version: 1.0.0
--- ---
@@ -19,8 +19,8 @@ version: 1.0.0
## Topology ## Topology
**Cluster:** 6 Proxmox nodes (ocupve, acerpve, minipve, amdpve, storepve, hwepve) **Cluster:** 5 Proxmox nodes (ocupve, acerpve, minipve, amdpve, storepve)
**Manager:** Mumuni (inside Abiba CT100, hwepve, .24) via Hermes agent **Manager:** Mumuni (inside Abiba CT100, minipve, .24) via Hermes agent
**Workers:** 6 profiles, all running on the same agent — no separate hosts needed **Workers:** 6 profiles, all running on the same agent — no separate hosts needed
This contract is infrastructure-agnostic in terms of which nodes are used. This contract is infrastructure-agnostic in terms of which nodes are used.
@@ -31,7 +31,7 @@ Workers execute tasks on whatever infrastructure they're given — SSH to .6,
## Why This Matters ## Why This Matters
Without enforced delegation, the manager consumes the full iteration budget Without enforced delegation, the manager consumes the full iteration budget
(60 calls) on single-turn tasks — SSH to 6 nodes, check each VM, read logs — (60 calls) on single-turn tasks — SSH to 5 nodes, check each VM, read logs —
leaving no capacity for actual coordination. The result: context overflow leaving no capacity for actual coordination. The result: context overflow
(59K tokens in system prompt), iteration exhaustion, and degraded response (59K tokens in system prompt), iteration exhaustion, and degraded response
quality. This contract exists because I blew through my budget checking quality. This contract exists because I blew through my budget checking
@@ -82,7 +82,7 @@ it asks the manager (via relay) — it doesn't go find it on its own.
**This is a hard rule, not a recommendation.** Violating it produces the exact **This is a hard rule, not a recommendation.** Violating it produces the exact
type of discrepancy the kanban pipeline exists to prevent: a review worker finds type of discrepancy the kanban pipeline exists to prevent: a review worker finds
"6 nodes present" in the raw data but "6/6 online" in the report — even though "5 nodes present" in the raw data but "5/5 online" in the report — even though
one of those nodes was unreachable. The report lied because it used data the one of those nodes was unreachable. The report lied because it used data the
raw data never provided. raw data never provided.
@@ -137,7 +137,7 @@ delegate_task(
``` ```
delegate_task( delegate_task(
tasks=[ tasks=[
{"goal": "Check all 6 Proxmox nodes for VM status", "context": "SSH to each node via 192.168.68.x, run 'qm list'"}, {"goal": "Check all 5 Proxmox nodes for VM status", "context": "SSH to each node via 192.168.68.x, run 'qm list'"},
{"goal": "Check Docker container health on .7/.116/.17", "context": "SSH to each host, check container status"}, {"goal": "Check Docker container health on .7/.116/.17", "context": "SSH to each host, check container status"},
] ]
) )
@@ -191,7 +191,7 @@ Only verified results reach Kwame. Format per channel:
{ {
"lane_id": "devops-check", "lane_id": "devops-check",
"worker": "syslog-devops", "worker": "syslog-devops",
"goal": "Check all 6 Proxmox nodes", "goal": "Check all 5 Proxmox nodes",
"status": "dispatched|completed|failed", "status": "dispatched|completed|failed",
"output_file": "/tmp/node-report.md" "output_file": "/tmp/node-report.md"
} }
+1 -1
View File
@@ -54,7 +54,7 @@ garbled input. When a user types these commands to Abiba:
- **`/approve`** → "Pi doesn't have pending approvals. Commands execute immediately." - **`/approve`** → "Pi doesn't have pending approvals. Commands execute immediately."
- **`/approve session`** → Same response - **`/approve session`** → Same response
- **`/deny`** → Same response + "For Hermes agents (Tanko, Mumuni), these work with their built-in approval system." - **`/deny`** → Same response + "For agents (Mumuni on Hermes, Tanko on DSH), these work with their built-in approval system."
This keeps the UX consistent across agents — users can type `/approve` anywhere This keeps the UX consistent across agents — users can type `/approve` anywhere
without getting confused by LLM responses. without getting confused by LLM responses.
+23 -8
View File
@@ -2,21 +2,31 @@
kind: responsibility kind: responsibility
name: pm2-self-heal name: pm2-self-heal
description: > description: >
Monitors critical PM2 processes (abiba-zulip, abiba-telegram) and Monitors critical PM2 processes (abiba-zulip, abiba-telegram, gitea-runner,
auto-restarts any that are stopped or errored. Logs every action to spoton-service, zulip-watchdog) and auto-restarts any that are stopped or
the knowledge graph and alerts the owner via Zulip DM on failures. errored. Logs every action to Gitea (SyslogSolution/health-logs — not
knowledge graph, hard rule) and alerts the owner via
Zulip DM on failures.
CRITICAL: Never restart abiba-zulip — it runs this contract. CRITICAL: Never restart abiba-zulip — it runs this contract.
AS-BUILT 2026-08-09 (captain ruling, ecosystem is authoritative):
gpu-monitor is systemd-managed (gpu-monitor.service) — NOT PM2;
gpu-watchdog decommissioned (function folded into gpu-monitor.service);
gitea-runner KEPT (online in PM2); abiba-zulip KEPT (online 4d+, the
2026-07-04 'removed/decommissioned' note was stale and is removed).
--- ---
## Maintains ## Maintains
- abiba-telegram: { status: "online", uptime: string, restarts: number } - abiba-telegram: { status: "online", uptime: string, restarts: number }
- gpu-watchdog: { status: "online", uptime: string, restarts: number } - abiba-zulip: { status: "online", uptime: string, restarts: number }
- gpu-monitor: { status: "online", uptime: string, restarts: number }
- gitea-runner: { status: "online", uptime: string, restarts: number } - gitea-runner: { status: "online", uptime: string, restarts: number }
- spoton-service: { status: "online", uptime: string, restarts: number }
- zulip-watchdog: { status: "online", uptime: string, restarts: number }
- last_check: timestamp - last_check: timestamp
> **Note (2026-07-04):** `abiba-zulip` removed — Zulip extension decommissioned. > **Note (2026-07-04, SUPERSEDED 2026-08-09):** `abiba-zulip` remains ONLINE and
> is monitored — the decommission note was stale (process re-added; do not treat
> it as removed).
## Continuity ## Continuity
@@ -31,10 +41,15 @@ description: >
- **Verify**: Re-check status after 5 seconds - **Verify**: Re-check status after 5 seconds
- **Escalate**: If still failed after 2 retries, send Zulip DM to owner - **Escalate**: If still failed after 2 retries, send Zulip DM to owner
### Rule 2: Process Restarting Too Often ### Rule 2: Process Restarting Too Often (crash-loop guard)
- **Detect**: `pm2 status` shows restarts > 30 (cumulative lifetime counter) - **Detect**: `pm2 status` shows restarts > 30 (cumulative lifetime counter) **or** a
process reporting a restart count > 1000 while showing "online" (a crash-loop mask)
- **Note**: PM2 counter never decrements; only full delete+re-add resets it - **Note**: PM2 counter never decrements; only full delete+re-add resets it
- **Fix**: `pm2 delete <name> && pm2 start <ecosystem> --only <name>` - **Fix**: `pm2 delete <name> && pm2 start <ecosystem> --only <name>`
- **Script guard (2026-08-16)**: `scripts/pm2-self-heal.sh` now restarts `abiba-telegram`
when `TEL_RESTARTS > 1000` even if the process reports "online" — catches a quiet
crash-loop that never toggles status to "stopped"/"errored" (e.g. the 10k-restarts
spoton incident). Alerts include the restart count.
- **Escalate**: Only when restarts > 30 — alerts to Zulip DM - **Escalate**: Only when restarts > 30 — alerts to Zulip DM
- **Historical fix**: Previous cycles were caused by abiba-zulip extension's - **Historical fix**: Previous cycles were caused by abiba-zulip extension's
stuck detection (STUCK_THRESHOLD_MS was 30min, raised to 4h in v2). stuck detection (STUCK_THRESHOLD_MS was 30min, raised to 4h in v2).
+6 -7
View File
@@ -5,7 +5,7 @@ description: >
Proxmox cluster + Docker monitoring via the existing Grafana/Prometheus stack Proxmox cluster + Docker monitoring via the existing Grafana/Prometheus stack
on CT 116. Replaces Pulse with file-provisioned Grafana dashboards. Three on CT 116. Replaces Pulse with file-provisioned Grafana dashboards. Three
exporters feed Prometheus: prometheus-pve-exporter (cluster-aware, single exporters feed Prometheus: prometheus-pve-exporter (cluster-aware, single
instance), node_exporter (all 6 PVE nodes), and a custom docker-stats-exporter instance), node_exporter (all 5 PVE nodes), and a custom docker-stats-exporter
(Docker 29 / containerd image-store compatible, since cAdvisor cannot resolve (Docker 29 / containerd image-store compatible, since cAdvisor cannot resolve
the layerdb). Dashboards exposed at http://192.168.68.116:3001/ (direct LAN, not behind nginx). the layerdb). Dashboards exposed at http://192.168.68.116:3001/ (direct LAN, not behind nginx).
agent: abiba agent: abiba
@@ -33,8 +33,8 @@ agent: abiba
| Exporter | Host:Port | Scope | Notes | | Exporter | Host:Port | Scope | Notes |
|----------|-----------|-------|-------| |----------|-----------|-------|-------|
| prometheus-pve-exporter | .116:9221 (container) | All 6 nodes + guests + 36 storage pools | Single instance, cluster-aware via amdpve API. Config `/opt/monitoring/pve.yml` (token `monitoring@pve!prometheus`, PVEAuditor role). Metric schema is label-based (`id=node/amdpve`, `id=lxc/100`). | | prometheus-pve-exporter | .116:9221 (container) | All 5 nodes + 14 guests + 36 storage pools | Single instance, cluster-aware via amdpve API. Config `/opt/monitoring/pve.yml` (token `monitoring@pve!prometheus`, PVEAuditor role). Metric schema is label-based (`id=node/amdpve`, `id=lxc/100`). |
| node_exporter | .5/.6/.9/.12/.15/.4:9100 (systemd) | Per-node CPU/mem/disk/net/temp | Installed via apt on all 6 PVE nodes, enabled (reboot-persistent). Collectors: textfile, systemd, tcpstat, ethtool. hwepve (.4) added 2026-07-19. | | node_exporter | .5/.6/.9/.12/.15:9100 (systemd) | Per-node CPU/mem/disk/net/temp | Installed via apt on all 5 PVE nodes, enabled (reboot-persistent). Collectors: textfile, systemd, tcpstat, ethtool. |
| docker-stats-exporter | .116:9324 (container) | 10 Docker containers on .116 | **Custom** (cAdvisor v0.51 incompatible with Docker 29 containerd image store — layerdb gone). Uses Docker Engine API over unix socket. Script `/opt/monitoring/docker-stats-exporter.py`. | | docker-stats-exporter | .116:9324 (container) | 10 Docker containers on .116 | **Custom** (cAdvisor v0.51 incompatible with Docker 29 containerd image store — layerdb gone). Uses Docker Engine API over unix socket. Script `/opt/monitoring/docker-stats-exporter.py`. |
## PVE API Token ## PVE API Token
@@ -48,7 +48,7 @@ agent: abiba
| UID | Title | Panels | Source | | UID | Title | Panels | Source |
|-----|-------|--------|--------| |-----|-------|--------|--------|
| proxmox-cluster | Proxmox Cluster Overview | 16 | cluster status, 6-node CPU/mem/disk/load gauges, guests table, storage pools, guest CPU/mem timeseries | | proxmox-cluster | Proxmox Cluster Overview | 16 | cluster status, 5-node CPU/mem/disk/load gauges, guests table, storage pools, guest CPU/mem timeseries |
| proxmox-node | Proxmox Node Detail | 13 | per-node CPU per-core, memory, network, disk IO/IOPS/latency, temperature, disk space (variable: $node) | | proxmox-node | Proxmox Node Detail | 13 | per-node CPU per-core, memory, network, disk IO/IOPS/latency, temperature, disk space (variable: $node) |
| docker-containers | Docker Containers | 10 | per-container CPU/mem/network, restarts, memory limit ratio (variable: $container) | | docker-containers | Docker Containers | 10 | per-container CPU/mem/network, restarts, memory limit ratio (variable: $container) |
| gpu-fleet | GPU Fleet | 7 | (existing, preserved in DB, not provisioned) | | gpu-fleet | GPU Fleet | 7 | (existing, preserved in DB, not provisioned) |
@@ -77,9 +77,9 @@ agent: abiba
| `/opt/monitoring/grafana/dashboards/build-dashboards.py` | .116 | dashboard JSON generator | | `/opt/monitoring/grafana/dashboards/build-dashboards.py` | .116 | dashboard JSON generator |
| `/opt/monitoring/grafana/dashboards/json/*.json` | .116 | provisioned dashboard definitions | | `/opt/monitoring/grafana/dashboards/json/*.json` | .116 | provisioned dashboard definitions |
| `/opt/monitoring/grafana/datasources/prometheus.yml` | .116 | datasource provisioning | | `/opt/monitoring/grafana/datasources/prometheus.yml` | .116 | datasource provisioning |
| `/etc/default/prometheus-node-exporter` | .5/.6/.9/.12/.15/.4 | node_exporter collector config | | `/etc/default/prometheus-node-exporter` | .5/.6/.9/.12/.15 | node_exporter collector config |
## Cluster "Tabiri" — 6 Nodes ## Cluster "Tabiri" — 5 Nodes
| Node | IP | Role | | Node | IP | Role |
|------|----|----| |------|----|----|
@@ -88,7 +88,6 @@ agent: abiba
| acerpve | 192.168.68.9 | PVE (hosts llm-gpu qemu/101) | | acerpve | 192.168.68.9 | PVE (hosts llm-gpu qemu/101) |
| minipve | 192.168.68.12 | PVE | | minipve | 192.168.68.12 | PVE |
| amdpve | 192.168.68.15 | PVE + Strix Halo LLM (qwen3.6-35B-udq4, strix-moe) | | amdpve | 192.168.68.15 | PVE + Strix Halo LLM (qwen3.6-35B-udq4, strix-moe) |
| hwepve | 192.168.68.4 | PVE (Huawei Matebook 16, 12C/15GB) — hosts abiba (lxc/100), kagentz (lxc/105), mumuni (lxc/114). CTs 100/105 migrated from amdpve, CT 114 from minipve 2026-07-20 |
## Operations ## Operations
+20 -3
View File
@@ -30,7 +30,6 @@ INFISICAL_ENV = "prod"
# PVE node IPs for CT liveness checks # PVE node IPs for CT liveness checks
PVE_NODES = { PVE_NODES = {
"hwepve": "192.168.68.4",
"amdpve": "192.168.68.15", "amdpve": "192.168.68.15",
"minipve": "192.168.68.12", "minipve": "192.168.68.12",
"storepve": "192.168.68.6", "storepve": "192.168.68.6",
@@ -40,8 +39,8 @@ PVE_NODES = {
# Agent definitions: ct, host, user, pve_node, vault_key_name # Agent definitions: ct, host, user, pve_node, vault_key_name
AGENTS = { AGENTS = {
"tanko": {"ct": 112, "host": "192.168.68.122", "user": "jerome", "pve": "amdpve", "vault_key": "TANKO_LITELLM_API_KEY"}, "tanko": {"ct": 112, "host": "192.168.68.122", "user": "jerome", "pve": "amdpve", "vault_key": "TANKO_LITELLM_API_KEY", "runtime": "dsh"},
"abiba": {"ct": 100, "host": "192.168.68.24", "user": "root", "pve": "hwepve", "vault_key": None}, # Pi agent + Mumuni Zulip, no vault key "abiba": {"ct": 100, "host": "192.168.68.24", "user": "root", "pve": "minipve", "vault_key": None}, # Pi agent + Mumuni Zulip, no vault key
"koby": {"ct": 111, "host": "192.168.68.129", "user": "root", "pve": "amdpve", "vault_key": "KOBY_LITELLM_API_KEY"}, "koby": {"ct": 111, "host": "192.168.68.129", "user": "root", "pve": "amdpve", "vault_key": "KOBY_LITELLM_API_KEY"},
"koonimo": {"ct": 113, "host": "192.168.68.114", "user": "root", "pve": "amdpve", "vault_key": "KOONIMO_LITELLM_API_KEY"}, "koonimo": {"ct": 113, "host": "192.168.68.114", "user": "root", "pve": "amdpve", "vault_key": "KOONIMO_LITELLM_API_KEY"},
} }
@@ -240,6 +239,16 @@ def check_agents():
user = agent.get("user") user = agent.get("user")
ct = agent["ct"] ct = agent["ct"]
# Tanko runs on DSH (DeepSeek Harness) since 2026-08-27 — it no longer runs a
# Hermes gateway, so skip the Hermes gateway/state/streaming/journal checks.
if agent.get("runtime") == "dsh":
live = ssh(host, "true", user=user)
print(f" {'✅' if live is not None else '❌'} {name}: DSH (DeepSeek Harness) — "
f"no Hermes gateway since 2026-08-27 (CT {ct}, SSH {'OK' if live is not None else 'FAIL'})")
if live is None:
FAIL.append(f"unreachable:{name}")
continue
if not host or not user: if not host or not user:
print(f" ⬜ {name} (CT {ct}): cannot SSH — skip liveness check") print(f" ⬜ {name} (CT {ct}): cannot SSH — skip liveness check")
continue continue
@@ -328,6 +337,10 @@ def check_ct_liveness():
def check_config_integrity(): def check_config_integrity():
"""Verify agent config.yaml parses as valid YAML.""" """Verify agent config.yaml parses as valid YAML."""
for name, agent in AGENTS.items(): for name, agent in AGENTS.items():
# Tanko runs on DSH (DeepSeek Harness) since 2026-08-27 — no Hermes config.yaml.
if agent.get("runtime") == "dsh":
print(f" ⏭️ {name}: DSH — no Hermes config.yaml since 2026-08-27")
continue
host = agent.get("host") host = agent.get("host")
user = agent.get("user") user = agent.get("user")
if not host or not user: if not host or not user:
@@ -357,6 +370,10 @@ def check_config_integrity():
def check_wrapper_integrity(): def check_wrapper_integrity():
"""Verify the hermes CLI wrapper exists and can reach hermes-real.""" """Verify the hermes CLI wrapper exists and can reach hermes-real."""
for name, agent in AGENTS.items(): for name, agent in AGENTS.items():
# Tanko runs on DSH (DeepSeek Harness) since 2026-08-27 — no hermes CLI wrapper.
if agent.get("runtime") == "dsh":
print(f" ⏭️ {name}: DSH — no hermes CLI wrapper since 2026-08-27")
continue
host = agent.get("host") host = agent.get("host")
user = agent.get("user") user = agent.get("user")
if not host or not user: if not host or not user:
+11 -16
View File
@@ -296,21 +296,16 @@ def collect():
"pm2_uptime": pm2.get("uptime", "?"), "pm2_uptime": pm2.get("uptime", "?"),
} }
# Tanko (CT 122) # Tanko (CT 112, IP 192.168.68.122) — DSH (DeepSeek Harness), no Hermes gateway
tanko_state = ssh_jerome("192.168.68.122", "cat ~/.hermes/gateway_state.json 2>/dev/null") # since 2026-08-27. There is no ~/.hermes/gateway_state.json on CT 112 anymore;
tanko_data = {} # Zulip/Telegram connectivity is managed by the DSH harness, not the Hermes gateway.
try:
tanko_data = json.loads(tanko_state) if tanko_state else {}
except:
tanko_data = {}
platforms = tanko_data.get("platforms", {})
report["agents"]["tanko"] = { report["agents"]["tanko"] = {
"platform": "hermes", "ct": 112, "ip": "192.168.68.122", "platform": "dsh", "ct": 112, "ip": "192.168.68.122",
"gateway_state": tanko_data.get("gateway_state", "unknown"), "gateway_state": "n/a (DSH)",
"zulip_state": platforms.get("zulip", {}).get("state", "unknown"), "zulip_state": "unknown",
"telegram_state": platforms.get("telegram", {}).get("state", "unknown"), "telegram_state": "unknown",
"gateway_pid": tanko_data.get("pid"), "gateway_pid": None,
"updated_at": tanko_data.get("updated_at"), "updated_at": "",
} }
# Mumuni (CT 100, IP 192.168.68.24) # Mumuni (CT 100, IP 192.168.68.24)
@@ -322,7 +317,7 @@ def collect():
mumuni_data = {} mumuni_data = {}
mumuni_platforms = mumuni_data.get("platforms", {}) mumuni_platforms = mumuni_data.get("platforms", {})
report["agents"]["mumuni"] = { report["agents"]["mumuni"] = {
"platform": "hermes", "ct": 114, "ip": "192.168.68.24", "platform": "hermes", "ct": 100, "ip": "192.168.68.24",
"gateway_state": mumuni_data.get("gateway_state", "unknown"), "gateway_state": mumuni_data.get("gateway_state", "unknown"),
"telegram_state": mumuni_platforms.get("telegram", {}).get("state", "unknown"), "telegram_state": mumuni_platforms.get("telegram", {}).get("state", "unknown"),
"zulip_state": mumuni_platforms.get("zulip", {}).get("state", "not_installed"), "zulip_state": mumuni_platforms.get("zulip", {}).get("state", "not_installed"),
@@ -543,7 +538,7 @@ th {{ color: #8b949e; font-weight: normal; }}
elif name == "tanko": elif name == "tanko":
zulip_state = "✅" if agent.get("zulip_state") == "connected" else ("❌" if agent.get("zulip_state") == "disconnected" else "⬜") zulip_state = "✅" if agent.get("zulip_state") == "connected" else ("❌" if agent.get("zulip_state") == "disconnected" else "⬜")
gateway = agent.get("gateway_state", "?") gateway = agent.get("gateway_state", "?")
processed = agent.get("updated_at", "")[:10] processed = "DSH"
elif name == "mumuni": elif name == "mumuni":
zulip_state = "⬜" if agent.get("zulip_state") == "not_installed" else ("✅" if agent.get("zulip_state") == "connected" else "⬜") zulip_state = "⬜" if agent.get("zulip_state") == "not_installed" else ("✅" if agent.get("zulip_state") == "connected" else "⬜")
gateway = agent.get("gateway_state", "?") gateway = agent.get("gateway_state", "?")
+3 -6
View File
@@ -28,10 +28,8 @@ declare -A CT_NODES=(
[117]=storepve # zulip [117]=storepve # zulip
[118]=storepve # jdownloader [118]=storepve # jdownloader
# acerpve (192.168.68.9) — no CTs (bare metal GPU .8) # acerpve (192.168.68.9) — no CTs (bare metal GPU .8)
# hwepve (192.168.68.4) [100]=minipve # abiba (was hwepve)
[100]=hwepve # abiba (was amdpve) [105]=minipve # kagentz (was hwepve)
[105]=hwepve # kagentz (was amdpve)
[114]=hwepve # mumuni (was minipve)
# ocupve (192.168.68.5) — no CTs (bare metal GPU .110) # ocupve (192.168.68.5) — no CTs (bare metal GPU .110)
# #
# REMOVED CTs (migrated to bare metal, decommissioned, or VMs): # REMOVED CTs (migrated to bare metal, decommissioned, or VMs):
@@ -50,7 +48,6 @@ declare -A NODE_IPS=(
[storepve]=192.168.68.6 [storepve]=192.168.68.6
[acerpve]=192.168.68.9 [acerpve]=192.168.68.9
[ocupve]=192.168.68.5 [ocupve]=192.168.68.5
[hwepve]=192.168.68.4
) )
resolve_node() { resolve_node() {
@@ -77,7 +74,7 @@ main() {
if [[ $# -lt 1 ]]; then if [[ $# -lt 1 ]]; then
echo "Usage: pct-run <CT_ID> [command...]" >&2 echo "Usage: pct-run <CT_ID> [command...]" >&2
echo " pct-run 112 cat /etc/hostname" >&2 echo " pct-run 112 cat /etc/hostname" >&2
echo " pct-run 114 systemctl status hermes-gateway" >&2 echo " pct-run 100 systemctl status hermes-gateway" >&2
echo "" echo ""
echo "Known CTs:" >&2 echo "Known CTs:" >&2
for ct in $(echo "${!CT_NODES[@]}" | tr ' ' '\n' | sort -n); do for ct in $(echo "${!CT_NODES[@]}" | tr ' ' '\n' | sort -n); do
+1 -1
View File
@@ -33,7 +33,7 @@ TEL_LINE=$(echo "$STATUS" | grep "abiba-telegram")
TEL_STATUS=$(echo "$TEL_LINE" | awk -F'│' '{print $10}' | xargs) TEL_STATUS=$(echo "$TEL_LINE" | awk -F'│' '{print $10}' | xargs)
TEL_RESTARTS=$(echo "$TEL_LINE" | awk -F'│' '{print $9}' | xargs) TEL_RESTARTS=$(echo "$TEL_LINE" | awk -F'│' '{print $9}' | xargs)
if [ "$TEL_STATUS" != "online" ]; then if [ "$TEL_STATUS" != "online" ] || [ "$TEL_RESTARTS" -gt 1000 ]; then
pm2 restart abiba-telegram > /dev/null 2>&1 pm2 restart abiba-telegram > /dev/null 2>&1
sleep 3 sleep 3
TEL_LINE2=$(pm2 status --no-color 2>/dev/null | grep "abiba-telegram") TEL_LINE2=$(pm2 status --no-color 2>/dev/null | grep "abiba-telegram")
+4 -5
View File
@@ -46,18 +46,17 @@ You are a code reviewer for OpenProse infrastructure contracts in the Syslog Sol
The infrastructure-control.prose.md contract is the canonical reference for the cluster topology: The infrastructure-control.prose.md contract is the canonical reference for the cluster topology:
**Proxmox Cluster "Tabiri" (6 nodes):** **Proxmox Cluster "Tabiri" (5 nodes):**
- amdpve (192.168.68.15): tanko, tdunna, baggy, scottdenya - amdpve (192.168.68.15): tanko, tdunna, baggy, scottdenya
- minipve (192.168.68.12): adguard, authentik, gitea, syslog-api, infisical-vault - minipve (192.168.68.12): abiba, kagentz, adguard, authentik, gitea, syslog-api, infisical-vault
- storepve (192.168.68.6): docker-vm, ra-h-os, PBS, media, jdownloader, zulip - storepve (192.168.68.6): docker-vm, ra-h-os, PBS, media, jdownloader, zulip
- acerpve (192.168.68.9): llm-gpu - acerpve (192.168.68.9): llm-gpu
- ocupve (192.168.68.5): ocu-llm - ocupve (192.168.68.5): ocu-llm
- hwepve (192.168.68.4): abiba, kagentz, mumuni
**CT IDs (verified 2026-07-24 against PVE API):** **CT IDs (verified 2026-07-24 against PVE API):**
100:abiba 102:adguard 104:authentik 105:kagentz 106:ra-h-os 100:abiba 102:adguard 104:authentik 105:kagentz 106:ra-h-os
107:pbs 108:media 110:gitea 111:tdunna 112:tanko 107:pbs 108:media 110:gitea 111:tdunna 112:tanko
113:baggy 114:mumuni 115:scottdenya 116:syslog-api 117:zulip 113:baggy 115:scottdenya 116:syslog-api 117:zulip
118:jdownloader 119:infisical-vault 118:jdownloader 119:infisical-vault
**NO CT 122, CT 123, or .19 exist in the cluster.** **NO CT 122, CT 123, or .19 exist in the cluster.**
@@ -65,7 +64,7 @@ The infrastructure-control.prose.md contract is the canonical reference for the
**CRITICAL RULES (never regress):** **CRITICAL RULES (never regress):**
1. NO /grafana/ nginx route — it was tried and reverted on 2026-07-02. Grafana is direct LAN at :3001. 1. NO /grafana/ nginx route — it was tried and reverted on 2026-07-02. Grafana is direct LAN at :3001.
2. NO .19 IP — Zulip is CT 117 on storepve. 2. NO .19 IP — Zulip is CT 117 on storepve.
3. NO CT 122/123 — Tanko=CT 112, Mumuni=CT 114. 3. NO CT 122/123 — Tanko=CT 112, Mumuni=CT 100 (inside Abiba). No CT 114 anywhere.
4. Strix Halo :8080 is FIREWALLED to .116 only — cannot be probed from abiba (.24). 4. Strix Halo :8080 is FIREWALLED to .116 only — cannot be probed from abiba (.24).
5. abiba-zulip PM2 process is DECOMMISSIONED (2026-07-04) — abiba uses Telegram only. 5. abiba-zulip PM2 process is DECOMMISSIONED (2026-07-04) — abiba uses Telegram only.
+4 -4
View File
@@ -12,10 +12,10 @@ echo ""
# Authorized agents for restricted contracts # Authorized agents for restricted contracts
# Format: contract_pattern|authorized_agents (comma-separated) # Format: contract_pattern|authorized_agents (comma-separated)
declare -A RESTRICTED declare -A RESTRICTED
RESTRICTED["infrastructure-control.prose.md"]="abiba" RESTRICTED["infrastructure-control.prose.md"]="abiba,abiba-bot,tanko,tanko-bot,mumuni,mumuni-bot"
RESTRICTED["proxmox-monitor.prose.md"]="abiba" RESTRICTED["proxmox-monitor.prose.md"]="abiba,abiba-bot"
RESTRICTED["hermes-config-template.prose.md"]="abiba,mumuni,tanko" RESTRICTED["hermes-config-template.prose.md"]="abiba,abiba-bot,mumuni,mumuni-bot,tanko,tanko-bot"
RESTRICTED["zulip-health.prose.md"]="abiba,mumuni" RESTRICTED["zulip-health.prose.md"]="abiba,abiba-bot,mumuni,mumuni-bot,tanko,tanko-bot"
RESTRICTED["scripts/pm2-self-heal.sh"]="abiba" RESTRICTED["scripts/pm2-self-heal.sh"]="abiba"
RESTRICTED["scripts/prose-lint.sh"]="abiba" RESTRICTED["scripts/prose-lint.sh"]="abiba"
RESTRICTED["scripts/prose-ai-review.sh"]="abiba" RESTRICTED["scripts/prose-ai-review.sh"]="abiba"
+5 -17
View File
@@ -81,23 +81,11 @@ else
echo " Abiba: ✅ Connected (processed=$(echo "$PI_HEALTH" | python3 -c "import sys,json; d=json.load(sys.stdin); print(d.get('messages_processed',0))" 2>/dev/null))" >> "$LOG" echo " Abiba: ✅ Connected (processed=$(echo "$PI_HEALTH" | python3 -c "import sys,json; d=json.load(sys.stdin); print(d.get('messages_processed',0))" 2>/dev/null))" >> "$LOG"
fi fi
# ── Platform B: Hermes (Tanko) ── # ── Platform B: Tanko (DSH) ──
TANKO_STATE=$(ssh -o StrictHostKeyChecking=no -o ConnectTimeout=5 jerome@192.168.68.122 \ # Tanko moved to DSH (DeepSeek Harness) on 2026-08-27. It no longer runs a Hermes
"cat ~/.hermes/gateway_state.json 2>/dev/null" 2>/dev/null || echo "{}") # gateway, so there is no ~/.hermes/gateway_state.json on CT 112 (.122) to probe.
TANKO_ZULIP=$(echo "$TANKO_STATE" | python3 -c " # Zulip connectivity for tanko is managed by the DSH harness; skip the legacy SSH probe.
import sys,json echo " Tanko: ⏭️ skipped (DSH — no Hermes gateway since 2026-08-27)" >> "$LOG"
d=json.load(sys.stdin)
p=d.get('platforms',{}).get('zulip',{})
print(p.get('state','unknown'))
" 2>/dev/null)
if [ "$TANKO_ZULIP" != "connected" ]; then
notify "🔴" "Tanko (Hermes) Zulip state: $TANKO_ZULIP — needs restart"
ISSUES=$((ISSUES + 1))
echo " Tanko: ❌ state=$TANKO_ZULIP" >> "$LOG"
else
echo " Tanko: ✅ Zulip connected" >> "$LOG"
fi
# ── Platform B: Hermes (Mumuni) ── # ── Platform B: Hermes (Mumuni) ──
MUMUNI_STATE=$(ssh -o StrictHostKeyChecking=no -o ConnectTimeout=5 root@192.168.68.24 \ MUMUNI_STATE=$(ssh -o StrictHostKeyChecking=no -o ConnectTimeout=5 root@192.168.68.24 \
+22 -14
View File
@@ -1,7 +1,7 @@
--- ---
kind: responsibility kind: responsibility
name: zulip-health name: zulip-health
description: Multi-platform health monitor for the Zulip messaging mesh spanning Platform A (Agent Zero Docker), Platform B (Hermes agents Tanko/Mumuni), and the Zulip bridge. Verifies bot registration, DM delivery, and cross-platform connectivity. description: Multi-platform health monitor for the Zulip messaging mesh spanning Platform A (Agent Zero Docker), Platform B (Tanko on DSH / Mumuni on Hermes), and the Zulip bridge. Verifies bot registration, DM delivery, and cross-platform connectivity.
title: Zulip Mesh Health Monitor — Multi-Platform title: Zulip Mesh Health Monitor — Multi-Platform
version: 3.0.0 version: 3.0.0
runtime_contract: 2 runtime_contract: 2
@@ -10,13 +10,13 @@ agent: abiba
# Zulip Mesh Health Monitor # Zulip Mesh Health Monitor
Monitors ALL Zulip-connected agents across three platforms (pi, Hermes, Agent Zero). Monitors ALL Zulip-connected agents across platforms (pi, Hermes, DSH, Agent Zero).
Runs every 15 minutes in the background. Also triggers on session start. Runs every 15 minutes in the background. Also triggers on session start.
## Requires ## Requires
- **Zulip API key** for `abiba-bot@chat.sysloggh.net` in `$ZULIP_API_KEY` - **Zulip API key** for `abiba-bot@chat.sysloggh.net` in `$ZULIP_API_KEY`
- **SSH access** to Tanko (192.168.68.122), Mumuni (192.168.68.24, inside Abiba CT100 on hwepve), and Agent Zero Docker host (192.168.68.14) - **SSH access** to Tanko (192.168.68.122), Mumuni (192.168.68.24, inside Abiba CT100 on minipve), and Agent Zero Docker host (192.168.68.14)
- **PM2** on localhost for pi process management - **PM2** on localhost for pi process management
- **Network access** to `chat.sysloggh.net`, `localhost:9200` - **Network access** to `chat.sysloggh.net`, `localhost:9200`
- **Write access** to `/root/zulip-health-monitor.log` and `/tmp/zulip-monitor-debounce` - **Write access** to `/root/zulip-health-monitor.log` and `/tmp/zulip-monitor-debounce`
@@ -45,7 +45,7 @@ Runs every 15 minutes in the background. Also triggers on session start.
"severity": "healthy" "severity": "healthy"
}, },
"tanko": { "tanko": {
"platform": "hermes", "platform": "dsh",
"zulip_state": "connected", "zulip_state": "connected",
"heartbeat_age_seconds": 45, "heartbeat_age_seconds": 45,
"gateway_pid": 1234, "gateway_pid": 1234,
@@ -81,13 +81,13 @@ Log as "unreachable" — don't treat as critical unless it persists for 3+ conse
## Streaming Support (2026-07-05) ## Streaming Support (2026-07-05)
Zulip agents now support progressive message editing during agent generation. Zulip agents now support progressive message editing during agent generation.
When a Hermes agent (Tanko, Mumuni) processes a message, the response is When a Zulip agent (Tanko on DSH, Mumuni on Hermes) processes a message, the response is
streamed in real-time via Zulip's `PATCH /api/v1/messages/{id}` API: streamed in real-time via Zulip's `PATCH /api/v1/messages/{id}` API:
- Adapter implements `edit_message()` using `_api_patch()` helper - Adapter implements `edit_message()` using `_api_patch()` helper
- Gateway stream consumer progressively edits the Zulip message - Gateway stream consumer progressively edits the Zulip message
- User sees real-time agent thinking instead of waiting for full response - User sees real-time agent thinking instead of waiting for full response
- Verified: Tanko (CT 112) and Mumuni (CT 114) both have streaming active - Verified: Tanko (CT 112) and Mumuni (inside Abiba CT 100) both have streaming active
### Verification ### Verification
```bash ```bash
@@ -182,15 +182,16 @@ grep -a "Finalized\|Failed to finalize" /root/.pm2/logs/abiba-zulip-out.log | ta
| `last_error` set | Log and monitor | | `last_error` set | Log and monitor |
| Crash loop >10/h | Alert user | | Crash loop >10/h | Alert user |
### Step 3: Platform B — Hermes (Tanko .122, Mumuni .24) ### Step 3: Platform B — Tanko (DSH, .122) & Mumuni (Hermes, .24)
**B1: Gateway State** **B1: Gateway State**
```bash ```bash
ssh root@192.168.68.122 "cat ~/.hermes/gateway_state.json"
ssh root@192.168.68.24 "cat ~/.hermes/gateway_state.json" # Mumuni inside Abiba CT100 ssh root@192.168.68.24 "cat ~/.hermes/gateway_state.json" # Mumuni inside Abiba CT100
``` ```
Tanko runs on DSH (DeepSeek Harness) — it no longer runs a Hermes gateway, so there is no `~/.hermes/gateway_state.json` on CT 112 (.122). Verify Tanko's Zulip connectivity via the DSH harness bot status instead.
Check `platforms.zulip.state`: `connected` ✅ | `disconnected` ❌ | `error` ❌ | missing → not installed. Check `platforms.zulip.state`: `connected` ✅ | `disconnected` ❌ | `error` ❌ | missing → not installed.
**B2: Agent Process** **B2: Agent Process**
@@ -199,21 +200,28 @@ Check `platforms.zulip.state`: `connected` ✅ | `disconnected` ❌ | `error`
ssh root@<CT> "ps aux | grep 'gateway run' | grep -v grep" ssh root@<CT> "ps aux | grep 'gateway run' | grep -v grep"
``` ```
Gateway PID should exist with uptime > 60s. **Dual-gateway detection**: if more than one `gateway run` process is found, the gateway has a collision (typically one `--force` and one `--replace` process). Kill the newer/duplicate process, then restart the remaining gateway via PM2 (`pm2 restart mumuni-zulip`). Check gateway log for "Gateway running with 2 platform(s)" (not 1) to confirm Zulip reloaded. Gateway PID should exist with uptime > 60s. **Dual-gateway detection**: if more than one `gateway run` process is found, the gateway has a collision (typically one `--force` and one `--replace` process). Kill the newer/duplicate process, then restart the remaining gateway per-agent (parameterized 2026-08-09, captain ruling):
**B3: Heartbeat Verification** | Agent | Restart command | Notes |
|-------|-----------------|-------|
| Mumuni (.24) | `pm2 restart abiba-zulip` | Hermes gateway runs under PM2 as `abiba-zulip` |
| Tanko (.122) | DSH harness — restart via its DSH service, not a Hermes gateway | Tanko runs on DSH (CT 112) since 2026-08-27; no longer a Hermes agent, no `~/.hermes` gateway, no PM2 `mumuni-zulip` process |
Check gateway log for "Gateway running with 2 platform(s)" (not 1) to confirm Zulip reloaded.
**B3: Heartbeat Verification** (Hermes agent Mumuni only — Tanko has no Hermes gateway)
```bash ```bash
ssh root@<CT> "grep Heartbeat ~/.hermes/logs/agent.log | tail -3" ssh root@192.168.68.24 "grep Heartbeat ~/.hermes/logs/agent.log | tail -3"
``` ```
Expected: recent heartbeat (within 5 min), `polls=N` incrementing. Expected: recent heartbeat (within 5 min), `polls=N` incrementing.
Silence > 300s → warning. Silence > 600s → critical. Silence > 300s → warning. Silence > 600s → critical.
**B4: Response Delivery** **B4: Response Delivery** (Hermes agent Mumuni only)
```bash ```bash
ssh root@<CT> "grep -E 'Finalized|Failed to finalize|Replied to' ~/.hermes/logs/agent.log | tail -10" ssh root@192.168.68.24 "grep -E 'Finalized|Failed to finalize|Replied to' ~/.hermes/logs/agent.log | tail -10"
``` ```
> 50% fail rate → critical. > 50% fail rate → critical.
@@ -222,7 +230,7 @@ ssh root@<CT> "grep -E 'Finalized|Failed to finalize|Replied to' ~/.hermes/logs/
| Condition | Action | | Condition | Action |
|-----------|--------| |-----------|--------|
| `zulip.state != "connected"` | `ssh root@<CT> "pkill -f 'gateway run'; sleep 2; hermes gateway restart"` | | `zulip.state != "connected"` | `ssh root@<CT> "pkill -f 'gateway run'; sleep 2; hermes gateway restart"` (Mumuni) / restart Tanko via DSH service |
| No heartbeat in 10min | Same as above | | No heartbeat in 10min | Same as above |
| `Failed to finalize` > 50% | Check PATCH API, Zulip server | | `Failed to finalize` > 50% | Check PATCH API, Zulip server |
| Response empty/short | Check A2A endpoint / LiteLLM model | | Response empty/short | Check A2A endpoint / LiteLLM model |
+1 -1
View File
@@ -10,7 +10,7 @@ description: >
> **⚠️ RETIRED** — The pi Zulip extension (`~/.pi/agent/extensions/zulip/`) and > **⚠️ RETIRED** — The pi Zulip extension (`~/.pi/agent/extensions/zulip/`) and
> PM2 process (`abiba-zulip`) have been decommissioned. All mention/reliability > PM2 process (`abiba-zulip`) have been decommissioned. All mention/reliability
> monitoring now happens through Telegram. Hermes agents (Tanko, Mumuni) and > monitoring now happens through Telegram. Agents (Mumuni on Hermes, Tanko on DSH) and
> Agent Zero (kagentz) continue to use Zulip. > Agent Zero (kagentz) continue to use Zulip.
## Maintains ## Maintains
+1 -1
View File
@@ -420,7 +420,7 @@ Backup v2 before starting: `cp index.js index.js.v2-backup-$(date +%Y%m%d-%H%M%S
|-------|----------|-------------|--------------|-------------| |-------|----------|-------------|--------------|-------------|
| **Abiba** | pi (CT 100) | ✅ Connected | API key missing from Infisical injection; poll timeout noise | Added .env fallback; AbortError treated as empty poll (no retry); poll timeout 65s→90s | | **Abiba** | pi (CT 100) | ✅ Connected | API key missing from Infisical injection; poll timeout noise | Added .env fallback; AbortError treated as empty poll (no retry); poll timeout 65s→90s |
| **Tanko** | Hermes (CT 112) | ✅ Connected | Gateway disconnected since Jul 11; watchdog restart didn't re-establish Zulip | Full gateway restart (kill wrapper, let infisical-gateway.sh respawn) | | **Tanko** | Hermes (CT 112) | ✅ Connected | Gateway disconnected since Jul 11; watchdog restart didn't re-establish Zulip | Full gateway restart (kill wrapper, let infisical-gateway.sh respawn) |
| **Mumuni** | Hermes (CT 114) | ✅ Connected | No issues found | None needed | | **Mumuni** | Hermes (inside Abiba CT 100) | ✅ Connected | No issues found | None needed |
### Key Fixes Applied ### Key Fixes Applied
+5 -5
View File
@@ -13,7 +13,8 @@ triggers:
> **⚠️ RETIRED** — This contract was embedded in the pi Zulip extension code > **⚠️ RETIRED** — This contract was embedded in the pi Zulip extension code
> (`performHealthCheck()`). That code has been removed. Zulip self-healing for > (`performHealthCheck()`). That code has been removed. Zulip self-healing for
> Hermes agents (Tanko, Mumuni) continues through their own gateway monitoring. > agents (Mumuni on Hermes, Tanko on DSH) continues through their own platform
> monitoring.
## Maintains ## Maintains
@@ -66,7 +67,7 @@ triggers:
| Zulip server | 192.168.68.19 | root | Docker: `zulip-zulip-1` | | Zulip server | 192.168.68.19 | root | Docker: `zulip-zulip-1` |
| Abiba (pi) | localhost | root | PM2: `abiba-zulip` | | Abiba (pi) | localhost | root | PM2: `abiba-zulip` |
| Mumuni | 192.168.68.24 (CT100 abiba) | root | `hermes gateway restart` | | Mumuni | 192.168.68.24 (CT100 abiba) | root | `hermes gateway restart` |
| Tanko | 192.168.68.122 | jerome | `PATH=$PATH:/home/jerome/.hermes/hermes-agent hermes gateway restart` | | Tanko | 192.168.68.122 (CT 112) | jerome | DSH (DeepSeek Harness) — restart via DSH service, not `hermes gateway restart` (no longer a Hermes agent since 2026-08-27) |
## Debounce ## Debounce
@@ -75,9 +76,8 @@ Track via `/tmp/zulip-heal-debounce-<agent>` (unix timestamp of last restart).
## Reporting ## Reporting
Every cycle produces a knowledge graph node: This contract is RETIRED — health-check logs are NOT knowledge graph content.
- Title: `[LEARN] zulip-self-heal: <timestamp>` No graph nodes are created. Logs go to Gitea (SyslogSolution/health-logs).
- metadata: { type: "remediation", status: "fixed" | "escalated" | "healthy" }
- Issues fixed → DM: "🛠 Zulip Self-Heal: fixed <issue>" - Issues fixed → DM: "🛠 Zulip Self-Heal: fixed <issue>"
- Issues escalated → DM: "⚠️ Zulip Self-Heal: <issue> needs attention" - Issues escalated → DM: "⚠️ Zulip Self-Heal: <issue> needs attention"