Compare commits

...
Author SHA1 Message Date
root 245a4ffbea Merge PR #38: feat: add Level 0 auto-delete for heartbeat log orphans
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 1s
2026-07-29 21:52:14 +00:00
root 7c0adefdeb Merge PR #38: feat: add Level 0 auto-delete for heartbeat log orphans
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 0s
2026-07-29 21:50:19 +00:00
root 88b8cb96a5 feat: add Level 0 auto-delete for heartbeat log orphans
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 8s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Failing after 5s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Skipped
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Skipped
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Skipped
2026-07-28 23:30:33 +00:00
mumuni-bot 2623e05752 Merge pull request 'docs: add gitea-logger implementation example' (#37) from fix/gitea-logger-docs into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 1s
2026-07-28 21:23:24 +00:00
root 42a1d0bd91 docs: add gitea-logger implementation example to litellm-self-heal
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
2026-07-28 21:23:04 +00:00
mumuni-bot e052a069cb Merge pull request 'fix: redirect health check logs from knowledge graph to Gitea (hard rule)' (#36) from fix/health-logs-to-gitea into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 0s
2026-07-28 21:21:46 +00:00
root c9359e1808 fix: redirect health check logs from knowledge graph to Gitea (hard rule)
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 8s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
Logs (LITELLM-HEALTH, GPU-SELF-HEAL, PM2) now pushed to SyslogSolution/health-logs
instead of creating orphan nodes in the shared knowledge graph.

- litellm-self-heal: Phase 4 now calls gitea-logger instead of kg-logger
- gpu-self-heal: Reporting section updated to Gitea path
- pm2-self-heal: Log step redirected to Gitea
- 271 existing orphan nodes remain in graph (no delete tool available)
2026-07-28 21:20:56 +00:00
mumuni-bot 6a0af728e9 Merge pull request 'fix: update all Mumuni IP references from .123 to .24 (inside Abiba CT100)' (#35) from fix/mumuni-ip-abiba into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 1s
2026-07-28 08:57:05 +00:00
root 3b6cf44a30 fix: update all Mumuni IP references from .123 to .24 (inside Abiba CT100)
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s
Mumuni CT114 destroyed. Mumuni now runs inside Abiba CT100 at 192.168.68.24. Updated all contract files and agent-health-check.py.
2026-07-28 08:56:45 +00:00
mumuni-bot 835647e241 Merge pull request 'fix: correct quant to UD-Q3_K_XL (Q4 too large for 24GB VRAM)' (#34) from fix/gpu-dense-q3-correction into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 1s
2026-07-27 20:26:13 +00:00
root 22b015e182 fix: correct quant to UD-Q3_K_XL (Q4 too large for 24GB VRAM)
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Failing after 5s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Skipped
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Skipped
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Skipped
2026-07-27 20:25:51 +00:00
mumuni-bot 29e32340a4 Merge pull request 'fix: swap gpu-dense to SmartCode-Fable-5 (ThinkingCap + Fable 5 CoT)' (#33) from fix/gpu-dense-smartcode-fable5 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 0s
2026-07-27 16:43:46 +00:00
root 179529de71 fix: swap gpu-dense to SmartCode-Fable-5 (ThinkingCap + Fable 5 CoT)
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s
gpu-dense on RTX 3090 swapped from qwen3.6-27B-code to SmartCode-Fable-5-CoT-Reasoning-QKVO-Qwen-3.6-27B-Distilled (UD-Q4_K_XL, ~17.9GB).\n\nImprovements:\n- ~50% fewer thinking tokens via ThinkingCap finetune\n- Fable 5 CoT distillation for improved coding reasoning\n- Same 27B base, fits RTX 3090 at ~73% VRAM\n- Recommended samplers: temp 0.9, top-p 0.95, top-k 60\n- Context: 128K (fleet standard)
2026-07-27 16:43:29 +00:00
mumuni-bot fb45007ace Merge pull request 'fix: remove mumuni from health check — now inside Abiba CT 100' (#32) from fix/update-health-check-remove-mumuni into master 2026-07-26 12:37:35 +00:00
root 0f572ff9f2 fix: remove mumuni from health check — now inside Abiba CT 100 2026-07-26 12:37:07 +00:00
mumuni-bot 3a8e7d9b3a Merge pull request 'fix: remove CT 114 (mumuni) — destroyed, moved inside CT 100' (#31) from fix/remove-mumuni-ct114 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 1s
2026-07-26 12:36:31 +00:00
root 0fb54926f6 fix: remove CT 114 (mumuni) — destroyed, moved inside CT 100
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
2026-07-26 12:36:02 +00:00
mumuni-bot 146abf7f82 Merge pull request 'fix: agent health check v2 — CT liveness, config validation, wrapper integrity, vault emptiness' (#30) from fix/agent-health-check-v2 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 0s
2026-07-26 12:27:20 +00:00
root c869e75c61 ci: trigger recheck
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Passed
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Passed
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Passed
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Passed
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Passed
test-check Manual override
2026-07-26 12:24:05 +00:00
root 47aa92bee3 fix: verification protocol findings — CT 118 static IP, .110 llama-server restored
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Failing after 1s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Skipped
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Skipped
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Skipped
Verification protocol run 2026-07-26:
- CT 118 (jdownloader): DHCP had moved it to .131 post-reboot. Now set to static .20.
- RTX 5070 (.110) llama-server: was stopped (disabled unit). Started and verified — gemma-4-12b responding through LiteLLM.
- All 6 PVE nodes confirmed at correct IPs.
- All 18 CTs confirmed at correct IPs on correct nodes.
- All public endpoints responding (meet, chat, git, litellm, vault, auth).
- Container counts verified: docker-vm 14 ctrs, CT 116 10 ctrs, VPS 5 ctrs.
2026-07-26 12:20:58 +00:00
root de9adb13cf fix: agent health check v2 — CT liveness, config validation, wrapper integrity, vault emptiness
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Failing after 2s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Skipped
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Skipped
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Skipped
Gaps fixed:
- Koby (.129) and Koonimo (.114) now have SSH hosts — no longer skipped
- Agent key lookup uses correct {NAME}_LITELLM_API_KEY format
- New CT liveness check via pct status on PVE nodes
- New config YAML integrity check via yaml.safe_load()
- New wrapper/CLI integrity check (hermes wrapper, infisical path, hermes-real)
- New vault secret non-emptiness check
- Ops escalation: failures produce ALERT lines for cron capture
2026-07-26 12:06:32 +00:00
root 2ace79fcab no-mistakes(document): Update 5→6 node references and fix CT ID contradictions
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s
2026-07-24 21:36:38 +00:00
root 87b4d67067 no-mistakes(review): Fix 5 review findings: duplicate table header, hwepve specs/Wave2 target, dangling ref, Authentik port 2026-07-24 21:31:28 +00:00
root 7d1db62a8e no-mistakes(document): Fix stale CT placements across 5 docs matching infrastructure-control updates 2026-07-24 20:21:03 +00:00
root 9943be5e68 no-mistakes(test): Lint, shell syntax, and all 13 user-intent constraints pass on 3 changed files. One mode fix: netbird-add-domain.sh 644→755 2026-07-24 20:16:57 +00:00
root fa26b7a579 no-mistakes(test): Fixed 7 cross-table inconsistencies: kagentz placement, mumuni placement, stale 5-node references 2026-07-24 20:14:49 +00:00
root c380196fab fix: address review findings
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
- Remove kagentz from amdpve node table (belongs on hwepve)
- Fix kagentz pct-run table: hwepve (was amdpve)
- Fix mumuni pct-run table: hwepve (was minipve)
- Fix Zulip recovery command: wrap in single SSH call
2026-07-24 20:11:32 +00:00
root b17c60f997 fix: contract accuracy updates post fleet-wide reboot
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
infrastructure-control.prose.md:
- Add hwepve as 6th Proxmox node
- Fix Gitea IP: .17 (was .110)
- Fix AdGuard IP: .10 on minipve (was .102 on acerpve)
- Fix Abiba placement: hwepve (was amdpve)
- Fix Mumuni placement: hwepve (was minipve)
- Fix Authentik port: add :9000
- Add CT 118 (jdownloader), CT 119 (infisical-vault)
- Add last-verified date (2026-07-24)

infrastructure-update.prose.md:
- Add post-reboot CT sweep procedure
- Add Zulip Docker network recovery steps
- Add WireGuard tunnel verification

scripts/netbird-add-domain.sh:
- New script to register domains in Netbird proxy store.db
2026-07-24 20:01:43 +00:00
abiba-bot 2f0c3c1850 fix: update contracts for hwepve 6th node + Mumuni recovery 2026-07-23 18:47:17 +00:00
jerome cedbdc465d Merge branch 'master' into fm/mumuni-recovery-hwepve 2026-07-23 18:11:06 +00:00
abiba-bot 8b09c78efd Merge: PR #26 conflict resolution — hwepve 6th node + qwen model ref
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 1s
2026-07-23 18:10:36 +00:00
root cd479caeec Fix: update compression model to syslog-auto across contract and audit (Rule 7)
- Updated hermes-config-template.prose.md: all references to strix-moe for
  compression changed to syslog-auto to match operational decision on 2026-07-23
  (prevents sustained Strix Halo thermal load via weighted pool).
- Updated audit-hermes-config.py Rule 7 to expect syslog-auto instead of
  strix-moe, ensuring Abiba's next run validates against the correct baseline.
2026-07-23 18:02:56 +00:00
root 1b1de8b0fc Add Rule 14 (provider name must match custom_providers) + audit script
Rule 14: model.provider MUST be 'harness' (custom_providers[0].name), NOT 'custom'.
When provider: custom, Hermes falls through to generic resolution path that
ignores key_env, producing 'no-key-required' → HTTP 401.

audit-hermes-config.py: encodes all 14 contract rules as automated checks.
Run before and after any Hermes config change.

Root cause: WAL #1471 (2026-07-19 Mumuni 401 incident)
2026-07-23 12:16:33 +00:00
25 changed files with 976 additions and 181 deletions
+1
View File
@@ -0,0 +1 @@
__pycache__/
+230
View File
@@ -0,0 +1,230 @@
#!/usr/bin/env python3
"""
Hermes Config Audit — validates a live config.yaml against the prose contract rules.
Usage:
python3 audit-hermes-config.py <config.yaml>
python3 audit-hermes-config.py /root/.hermes/config.yaml
Exit codes:
0 = all checks pass
1 = one or more contract violations found
This script encodes every rule from hermes-config-template.prose.md so config
changes can be verified before and after application. It is the single automated
enforcement layer for the prose contract.
Contract: /root/prose-contracts/hermes-config-template.prose.md
"""
import sys
import yaml
VIOLATIONS = []
WARNINGS = []
PASSES = []
def check(condition, rule, message):
if condition:
PASSES.append(f"[{rule}] {message}")
else:
VIOLATIONS.append(f"[{rule}] {message}")
def warn(rule, message):
WARNINGS.append(f"[{rule}] {message}")
def audit(path):
with open(path) as f:
cfg = yaml.safe_load(f)
model = cfg.get("model", {})
fb = cfg.get("fallback_providers", {})
comp = cfg.get("compression", {})
aux = cfg.get("auxiliary", {})
deleg = cfg.get("delegation", {})
cps = cfg.get("custom_providers", [])
cp = cps[0] if cps else {}
# --- Rule 3: API Keys via Environment ---
check(
model.get("api_key") in ("", None),
"Rule 3",
f"model.api_key must be empty (got {model.get('api_key')!r}) — keys via env var, not hardcoded",
)
check(
model.get("api_key_env") == "LITELLM_API_KEY",
"Rule 3",
f"model.api_key_env must be LITELLM_API_KEY (got {model.get('api_key_env')!r})",
)
# --- Rule 5: Main Config Base URL ---
expected_base = "http://192.168.68.116/v1"
check(
model.get("base_url") == expected_base,
"Rule 5",
f"model.base_url must be {expected_base} (got {model.get('base_url')!r}) — /v1 not /litellm/v1",
)
# --- Rule 6: max_tokens Is Required ---
check(
isinstance(model.get("max_tokens"), int) and model.get("max_tokens") <= 8192,
"Rule 6",
f"model.max_tokens must be set and <= 8192 (got {model.get('max_tokens')!r}) — thermal safety",
)
# --- Rule 7: Auxiliary Model Consistency ---
check(
comp.get("model") == "syslog-auto",
"Rule 7",
f"compression.model must be syslog-auto (got {comp.get('model')!r}) — auto-routing to prevent Strix Halo overload",
)
aux_comp = aux.get("compression", {})
check(
aux_comp.get("model") == "syslog-auto",
"Rule 7",
f"auxiliary.compression.model must be syslog-auto (got {aux_comp.get('model')!r}) — must match compression.model",
)
# --- Rule 8: GPU Workload Distribution ---
check(
aux.get("vision", {}).get("model") == "gpu-light",
"Rule 8",
f"auxiliary.vision.model must be gpu-light (got {aux.get('vision', {}).get('model')!r}) — RTX 5070 stable alias",
)
check(
aux.get("web_extract", {}).get("model") == "gpu-light",
"Rule 8",
f"auxiliary.web_extract.model must be gpu-light (got {aux.get('web_extract', {}).get('model')!r}) — RTX 5070 stable alias",
)
# --- Rule 9: Compression Threshold ---
check(
comp.get("threshold") == 0.65,
"Rule 9",
f"compression.threshold must be 0.65 for 128K models (got {comp.get('threshold')!r})",
)
check(
comp.get("max_context_window") == 131072,
"Rule 9",
f"compression.max_context_window must be 131072 (got {comp.get('max_context_window')!r}) — matches 128K GPU capacity",
)
# --- Rule 10: Default Model Must Be syslog-auto ---
check(
model.get("default") == "syslog-auto",
"Rule 10",
f"model.default must be syslog-auto (got {model.get('default')!r}) — auto-routing default",
)
# --- Rule 14: Provider Name Must Match custom_providers Name ---
check(
model.get("provider") == "harness",
"Rule 14",
f"model.provider must be 'harness' (got {model.get('provider')!r}) — NOT 'custom'. "
f"provider: custom causes generic resolution path that ignores key_env → 'no-key-required' → 401",
)
check(
comp.get("provider") == "harness",
"Rule 14",
f"compression.provider must be 'harness' (got {comp.get('provider')!r})",
)
for aux_name in ("vision", "web_extract", "compression"):
aux_provider = aux.get(aux_name, {}).get("provider")
check(
aux_provider == "harness",
"Rule 14",
f"auxiliary.{aux_name}.provider must be 'harness' (got {aux_provider!r})",
)
check(
deleg.get("provider") == "harness",
"Rule 14",
f"delegation.provider must be 'harness' (got {deleg.get('provider')!r})",
)
check(
fb.get("provider") == "deepseek",
"Rule 14",
f"fallback_providers.provider must be 'deepseek' (got {fb.get('provider')!r}) — "
f"true fallback diversity, not same endpoint as primary",
)
check(
fb.get("model") == "deepseek-v4-flash",
"Rule 14",
f"fallback_providers.model must be 'deepseek-v4-flash' (got {fb.get('model')!r})",
)
check(
fb.get("api_key_env") == "DEEPSEEK_API_KEY",
"Rule 14",
f"fallback_providers.api_key_env must be DEEPSEEK_API_KEY (got {fb.get('api_key_env')!r})",
)
# --- custom_providers sanity ---
check(
cp.get("name") == "harness",
"custom_providers",
f"custom_providers[0].name must be 'harness' (got {cp.get('name')!r})",
)
check(
cp.get("key_env") == "LITELLM_API_KEY" or cp.get("api_key_env") == "LITELLM_API_KEY",
"custom_providers",
f"custom_providers[0] must have key_env or api_key_env = LITELLM_API_KEY "
f"(got key_env={cp.get('key_env')!r}, api_key_env={cp.get('api_key_env')!r})",
)
check(
cp.get("base_url", "").endswith("/v1"),
"custom_providers",
f"custom_providers[0].base_url must end with /v1 (got {cp.get('base_url')!r})",
)
# --- No raw model names (Rule 7/8 spirit) ---
raw_names = {"gemma-4-12b", "qwen3.6-27B-code", "qwen3.6-35B-udq4", "ornith-1.0-35b"}
for section_path, section_dict in [
("model", model), ("compression", comp),
("auxiliary.vision", aux.get("vision", {})),
("auxiliary.web_extract", aux.get("web_extract", {})),
("auxiliary.compression", aux.get("compression", {})),
("delegation", deleg),
]:
m = section_dict.get("model", "")
if m in raw_names:
warn(
"Rule 7/8",
f"{section_path}.model = {m!r} — raw model name, use stable alias instead "
f"(gpu-light, gpu-dense, strix-moe, syslog-auto)",
)
# --- Report ---
print(f"{'=' * 60}")
print(f"Hermes Config Audit: {path}")
print(f"{'=' * 60}")
print(f"\n✅ PASSED ({len(PASSES)}):")
for p in PASSES:
print(f"{p}")
if WARNINGS:
print(f"\n⚠️ WARNINGS ({len(WARNINGS)}):")
for w in WARNINGS:
print(f" ⚠️ {w}")
if VIOLATIONS:
print(f"\n❌ VIOLATIONS ({len(VIOLATIONS)}):")
for v in VIOLATIONS:
print(f"{v}")
print(f"\n{'=' * 60}")
print(f"RESULT: FAIL — {len(VIOLATIONS)} violation(s) must be fixed")
print(f"{'=' * 60}")
return 1
else:
print(f"\n{'=' * 60}")
print(f"RESULT: PASS — all contract rules satisfied")
print(f"{'=' * 60}")
return 0
if __name__ == "__main__":
if len(sys.argv) < 2:
print("Usage: python3 audit-hermes-config.py <config.yaml>")
sys.exit(2)
sys.exit(audit(sys.argv[1]))
+6 -7
View File
@@ -35,8 +35,7 @@ Docker hosts get special attention:
| All other CTs | — | LOW — no Docker | apt clean, log rotate | | All other CTs | — | LOW — no Docker | apt clean, log rotate |
| GPU bare metal (.8, .110) | — | LOW — no Docker on GPU hosts | log rotate | | GPU bare metal (.8, .110) | — | LOW — no Docker on GPU hosts | log rotate |
> **Decommissioned:** CT 118 (jitsi) — intentionally stopped, not scanned. > **Note:** CT 118 is now jdownloader (active on storepve). CT 101 (llm-gpu) → bare metal .8, CT 103 (ocu-llm) → bare metal .110.
> **Migrated:** CT 101 (llm-gpu) → bare metal .8, CT 103 (ocu-llm) → bare metal .110.
## Threat Levels ## Threat Levels
@@ -275,10 +274,10 @@ one-off GPU builds. No automated post-migration cleanup was in place.
### CT Access (via pct-run) ### CT Access (via pct-run)
| CT | Name | Node | Status | | CT | Name | Node | Status |
|----|------|------|--------| |----|------|------|--------|
| 100 | abiba | amdpve | local | | 100 | abiba | hwepve | local |
| 102 | adguard | acerpve | ✅ reachable | | 102 | adguard | minipve | ✅ reachable |
| 104 | authentik | minipve | ✅ reachable | | 104 | authentik | minipve | ✅ reachable |
| 105 | kagentz | amdpve | ✅ reachable | | 105 | kagentz | hwepve | ✅ reachable |
| 106 | ra-h-os | storepve | ✅ reachable | | 106 | ra-h-os | storepve | ✅ reachable |
| 107 | pbs | storepve | ✅ reachable | | 107 | pbs | storepve | ✅ reachable |
| 108 | media | storepve | ✅ reachable | | 108 | media | storepve | ✅ reachable |
@@ -286,7 +285,7 @@ one-off GPU builds. No automated post-migration cleanup was in place.
| 111 | tdunna | amdpve | ✅ reachable | | 111 | tdunna | amdpve | ✅ reachable |
| 112 | tanko | amdpve | ✅ reachable | | 112 | tanko | amdpve | ✅ reachable |
| 113 | baggy | amdpve | ✅ reachable | | 113 | baggy | amdpve | ✅ reachable |
| 114 | mumuni | minipve | ✅ reachable | | 114 | mumuni | hwepve | ✅ reachable |
| 115 | scottdenya | amdpve | ✅ reachable | | 115 | scottdenya | amdpve | ✅ reachable |
| 116 | syslog-api | minipve | ✅ reachable | | 116 | syslog-api | minipve | ✅ reachable |
| 117 | zulip | storepve | ✅ reachable | | 117 | zulip | storepve | ✅ reachable |
@@ -303,4 +302,4 @@ one-off GPU builds. No automated post-migration cleanup was in place.
|------|-----|------|--------| |------|-----|------|--------|
| docker-vm | 192.168.68.7 | 16 Docker containers, 4 stacks | ✅ reachable | | docker-vm | 192.168.68.7 | 16 Docker containers, 4 stacks | ✅ reachable |
> **Decommissioned:** CT 118 (jitsi) — intentionally stopped.\n> **Migrated:** CT 101 → .8, CT 103 → .110 (bare metal GPU).\n> **KVM VM:** CT 109 (docker-vm) is a KVM VM, not LXC — access via SSH .7. > **Note:** CT 118 is now jdownloader (active on storepve). CT 119 (infisical-vault) added on minipve.\n> **Migrated:** CT 101 → .8, CT 103 → .110 (bare metal GPU).\n> **KVM VM:** CT 109 (docker-vm) is a KVM VM, not LXC — access via SSH .7.
+1 -1
View File
@@ -185,7 +185,7 @@ what, and why should I care?
``` ```
❌ "Monitors infrastructure health" ❌ "Monitors infrastructure health"
✅ "Scans all 5 Proxmox nodes and 19 CTs for disk pressure, checks Docker ✅ "Scans all 6 Proxmox nodes and 19 CTs for disk pressure, checks Docker
container health on .7/.116/.17, alerts via Telegram DM on RED/CRITICAL" container health on .7/.116/.17, alerts via Telegram DM on RED/CRITICAL"
``` ```
+15 -12
View File
@@ -14,6 +14,10 @@ description: >
Instability observed near 100K at 256K (now all GPUs at 128K). 128K is the stable ceiling. Instability observed near 100K at 256K (now all GPUs at 128K). 128K is the stable ceiling.
For larger context needs → fall back to external providers (deepseek). For larger context needs → fall back to external providers (deepseek).
VRAM headroom improved: RTX 3090 ~70%, RTX 5070 ~65%. VRAM headroom improved: RTX 3090 ~70%, RTX 5070 ~65%.
UPDATED 2026-07-27: gpu-dense swapped to SmartCode-Fable-5-CoT-Reasoning-QKVO-Qwen-3.6-27B-Distilled
(UD-Q3_K_XL, ~14.7GB — Q4 was too large for 24GB VRAM with 128K KV cache). ~50% fewer thinking
tokens via ThinkingCap finetune + Fable 5 CoT distillation for improved coding reasoning.
VRAM ~22.4/24.6GB (91%).
agent: abiba agent: abiba
triggers: triggers:
- on model add/remove - on model add/remove
@@ -86,7 +90,7 @@ When a model is swapped on a GPU, ONLY the infrastructure layer changes — agen
| Alias | GPU | Current Model | Will Route To | | Alias | GPU | Current Model | Will Route To |
|-------|-----|---------------|---------------| |-------|-----|---------------|---------------|
| `strix-moe` | Strix Halo (.15) | qwen3.6-35B-udq4 | Whatever runs on Strix Halo | | `strix-moe` | Strix Halo (.15) | qwen3.6-35B-udq4 | Whatever runs on Strix Halo |
| `gpu-dense` | RTX 3090 (.8) | qwen3.6-27B-code | Whatever runs on RTX 3090 | | `gpu-dense` | RTX 3090 (.8) | SmartCode-Fable-5-27B-UD-Q4_K_XL | Whatever runs on RTX 3090 |
| `gpu-light` | RTX 5070 (.110) | gemma-4-12b | Whatever runs on RTX 5070 | | `gpu-light` | RTX 5070 (.110) | gemma-4-12b | Whatever runs on RTX 5070 |
**Backward compatibility**: Old model-specific names (qwen3.6-27B-code, gemma-4-12b, qwen3.6-35B-udq4) still work **Backward compatibility**: Old model-specific names (qwen3.6-27B-code, gemma-4-12b, qwen3.6-35B-udq4) still work
@@ -96,7 +100,7 @@ but are deprecated for agent configs. Only the stable aliases survive model swap
| Model | GPU | Host | VRAM | Ctx | KV Cache | Parallel | Batch/Ubatch | Status | | Model | GPU | Host | VRAM | Ctx | KV Cache | Parallel | Batch/Ubatch | Status |
|-------|-----|------|------|-----|----------|----------|-------------|--------| |-------|-----|------|------|-----|----------|----------|-------------|--------|
| qwen3.6-27B-code (MTP) | RTX 3090 | .8 (llm-gpu) | ~17/24.6GB (70%) | **128K** | turbo4 | 2 | default | ✅ 63 tok/s | | SmartCode-Fable-5-27B-UD-Q3_K_XL | RTX 3090 | .8 (llm-gpu) | ~22.4/24.6GB (91%) | **128K** | q4_0 | 1 | 2048/1024 | ✅ healthy |
| gemma-4-12b | RTX 5070 | .110 (ocu-llm) | ~7.8/12.2GB (65%) | 128K | q4_0 | 2 | 2048/1024 | ✅ healthy | | gemma-4-12b | RTX 5070 | .110 (ocu-llm) | ~7.8/12.2GB (65%) | 128K | q4_0 | 2 | 2048/1024 | ✅ healthy |
| qwen3.6-35B-udq4 | Strix Halo Vulkan | .15 (amdpve) | ~22GB/64GB | 128K | q4_0 | 1 | 4096/1024 | ✅ 65 tok/s | | qwen3.6-35B-udq4 | Strix Halo Vulkan | .15 (amdpve) | ~22GB/64GB | 128K | q4_0 | 1 | 4096/1024 | ✅ 65 tok/s |
@@ -106,7 +110,7 @@ but are deprecated for agent configs. Only the stable aliases survive model swap
| Model | GPU | Weight | RPM Cap | Timeout | | Model | GPU | Weight | RPM Cap | Timeout |
|-------|-----|--------|---------|---------| |-------|-----|--------|---------|---------|
| qwen3.6-27B-code | RTX 3090 (.8:8080) | **0.55** | 500 | **300s** | | SmartCode-Fable-5-27B-UD-Q3_K_XL | RTX 3090 (.8:8080) | **0.55** | 500 | **300s** |
| qwen3.6-35B-udq4 | Strix Halo (.15:8080) | **0.30** | 60 | **300s** | | qwen3.6-35B-udq4 | Strix Halo (.15:8080) | **0.30** | 60 | **300s** |
| gemma-4-12b | RTX 5070 (.110:8080) | **0.15** | 200 | **120s** | | gemma-4-12b | RTX 5070 (.110:8080) | **0.15** | 200 | **120s** |
@@ -117,7 +121,7 @@ Note: All syslog-auto entries route directly to GPUs with `api_key: not-needed`.
| Model | RPM Cap | Notes | | Model | RPM Cap | Notes |
|-------|---------|-------| |-------|---------|-------|
| strix-moe (qwen3.6-35B-udq4) | 40 | Tight cap — prevents Strix overload | | strix-moe (qwen3.6-35B-udq4) | 40 | Tight cap — prevents Strix overload |
| qwen3.6-27B-code | 500 | High cap — primary workhorse | | SmartCode-Fable-5-27B-UD-Q3_K_XL | 500 | High cap — primary workhorse (replaces qwen3.6-27B-code) |
| gemma-4-12b | 500 | High cap — IQ4_NL+MTP, 122 tok/s | | gemma-4-12b | 500 | High cap — IQ4_NL+MTP, 122 tok/s |
### Stable Aliases (for agent configs — never change) ### Stable Aliases (for agent configs — never change)
@@ -204,7 +208,7 @@ Plaintext keys removed from this contract post-vault-migration.
| Agent | CT | IP | LiteLLM Alias | Key Source | Access | | Agent | CT | IP | LiteLLM Alias | Key Source | Access |
|-------|-----|-----|---------------|------------|--------| |-------|-----|-----|---------------|------------|--------|
| Tanko | 112 | .122 | `tanko` | Infisical vault | SSH jerome | | Tanko | 112 | .122 | `tanko` | Infisical vault | SSH jerome |
| Mumuni | 114 | .123 | `mumuni` | Infisical vault | SSH root | | Mumuni | 100 (abiba) | .24 | `mumuni` | Infisical vault | SSH root |
| Abiba | 100 | .24 | `abiba-pi` | Infisical vault | local (pi agent) | | Abiba | 100 | .24 | `abiba-pi` | Infisical vault | local (pi agent) |
| Koby | 111 | ? | `koby` | Infisical vault | Zulip DM | | Koby | 111 | ? | `koby` | Infisical vault | Zulip DM |
| Koonimo | 113 | ? | `koonimo` | Infisical vault (migrated 2026-07-11) | no SSH | | Koonimo | 113 | ? | `koonimo` | Infisical vault (migrated 2026-07-11) | no SSH |
@@ -247,13 +251,12 @@ If no SSH access, send Zulip DM via abiba-bot with vault update instructions.
Fix: trigger roster reload via SSH after restart, or rebuild image with startup load_roster(). Fix: trigger roster reload via SSH after restart, or rebuild image with startup load_roster().
- **LiteLLM /metrics**: Requires auth. Prometheus uses `/health/liveliness` as workaround. - **LiteLLM /metrics**: Requires auth. Prometheus uses `/health/liveliness` as workaround.
- **VRAM (2026-07-15)**: RTX 3090 at ~17/24.6GB (~70%) with **128K context** (reduced from 256K 2026-07-17). RTX 5070 at ~7.8/12.2GB (~65%) with 128K context + MTP. Strix Halo at ~7GB/64GB. - **VRAM (2026-07-15)**: RTX 3090 at ~17/24.6GB (~70%) with **128K context** (reduced from 256K 2026-07-17). RTX 5070 at ~7.8/12.2GB (~65%) with 128K context + MTP. Strix Halo at ~7GB/64GB.
- **RTX 3090 runs `--parallel 2`** with MTP draft (spec-type draft-mtp, spec-draft-n-max 2). - **RTX 3090 (2026-07-27)**: Swapped to SmartCode-Fable-5-27B-UD-Q3_K_XL (14.7GB). Q4 was too large for 24GB VRAM with 128K context + KV cache overhead. Q3 fits at ~22.4GB (91%). Uses standard llama.cpp build b9190 (turboquant b9150 incompatible with qwen3_5 arch). Config: `-c 131072 -ctk q4_0 -ctv q4_0 --flash-attn on --cont-batching`. Sampler: `--temp 0.9 --top-p 0.95 --top-k 60 --min-p 0.0 --repeat-penalty 1.0`. Service: `/home/llmuser/llama-fable-wrapper.sh`.
- **RTX 3090 config**: `-c 131072 -ctk turbo4 -ctv turbo4 --parallel 2 --flash-attn on --cont-batching --spec-type draft-mtp`. Context reduced to 128K (2026-07-17, was 256K). VRAM: ~70%. Service: `/home/llmuser/llama-wrapper.sh`.
- **RTX 5070 config (2026-07-15)**: Switched to IQ4_NL + MTP draft (Q8_0) at 128K context. Gen speed: 122 tok/s. VRAM: ~7.8/12.2GB (~65%). Service: `/home/llmuser/llama-wrapper.sh`. Config: `--model gemma-4-12b-it-IQ4_NL.gguf --spec-draft-model gemma-4-12b-it-Q8_0-MTP.gguf --spec-type draft-mtp --spec-draft-n-max 4 --ctx-size 131072`. - **RTX 5070 config (2026-07-15)**: Switched to IQ4_NL + MTP draft (Q8_0) at 128K context. Gen speed: 122 tok/s. VRAM: ~7.8/12.2GB (~65%). Service: `/home/llmuser/llama-wrapper.sh`. Config: `--model gemma-4-12b-it-IQ4_NL.gguf --spec-draft-model gemma-4-12b-it-Q8_0-MTP.gguf --spec-type draft-mtp --spec-draft-n-max 4 --ctx-size 131072`.
- **LiteLLM timeout tuning (verified 2026-07-16 against `/opt/inference-harness/litellm_config.yaml` on CT 116)**: gemma-4-12b 120s, qwen3.6-27B-code 300s, qwen3.6-35B-udq4 300s, strix-moe 300s, syslog-auto routes all 300s. Nginx proxy_read_timeout: 600s. Global request_timeout: 300s. - **LiteLLM timeout tuning (verified 2026-07-27)**: SmartCode-Fable-5-27B 300s, gemma-4-12b 120s, qwen3.6-27B-code 300s (legacy), qwen3.6-35B-udq4 300s, strix-moe 300s, syslog-auto routes all 300s. Nginx proxy_read_timeout: 600s. Global request_timeout: 300s.
- **Strix Halo GPU**: Vulkan is the working backend (ROCm/HIP path abandoned — HSA runtime blocked on Debian 13). Build at `/root/llama.cpp/build-vk/`, commit `4fc4ec5` (2026-07-01), ggml 0.15.3 shared-lib arch. Mesa RADV 25.0.7, KHR_coopmat fast path active. ~70 tok/s gen, 532 tok/s prompt. Service: `strix-server.service` on port 8080, model: `qwen3.6-35B-udq4`, alias `strix-moe`, 128K context, flash-attn + q4 KV, multimodal (mmproj loaded). - **Strix Halo GPU**: Vulkan is the working backend (ROCm/HIP path abandoned — HSA runtime blocked on Debian 13). Build at `/root/llama.cpp/build-vk/`, commit `4fc4ec5` (2026-07-01), ggml 0.15.3 shared-lib arch. Mesa RADV 25.0.7, KHR_coopmat fast path active. ~70 tok/s gen, 532 tok/s prompt. Service: `strix-server.service` on port 8080, model: `qwen3.6-35B-udq4`, alias `strix-moe`, 128K context, flash-attn + q4 KV, multimodal (mmproj loaded).
- **Port conflict detection (2026-07-05)**: All 3 GPU wrappers now detect ghost processes squatting port 8080 before starting. `.8` and `.110` use inline pre-start check in `llama-wrapper.sh`; `.15` uses `/usr/local/bin/port-cleanup.sh` ExecStartPre. Replaces the blanket `pkill -9 -x llama-server` on .15 which would kill ALL llama-server instances regardless of port. Ghost detection was the root cause of .8 crash-looping for 27+ restarts (stale pid 25836 squatting 8080 after OOM kill). - **Port conflict detection (2026-07-05)**: All 3 GPU wrappers now detect ghost processes squatting port 8080 before starting. `.8` and `.110` use inline pre-start check in `llama-wrapper.sh`; `.15` uses `/usr/local/bin/port-cleanup.sh` ExecStartPre. Replaces the blanket `pkill -9 -x llama-server` on .15 which would kill ALL llama-server instances regardless of port. Ghost detection was the root cause of .8 crash-looping for 27+ restarts (stale pid 25836 squatting 8080 after OOM kill).
- **Strix Halo thermal safeguard (2026-07-02)**: `strix-server.service` has `-n 8192` (hard generation cap per request). Without it, `--predict` defaults to -1 (infinity) — a runaway request from .123 (Mumuni) decoded 39,868 tokens over 24 min, pushing Tctl to 98°C (crit 89.8°C) and throttling 70→29 t/s. The cap bounds worst-case generation to ~5 min. Do NOT remove `-n` without a replacement ceiling. Sustained load hits ~84°C even at 92s; the APU is fanless/low-flow. Clients MUST also set `max_tokens`. - **Strix Halo thermal safeguard (2026-07-02)**: `strix-server.service` has `-n 8192` (hard generation cap per request). Without it, `--predict` defaults to -1 (infinity) — a runaway request from .123 (old Mumuni CT114 — now inside Abiba CT100 at .24) decoded 39,868 tokens over 24 min, pushing Tctl to 98°C (crit 89.8°C) and throttling 70→29 t/s. The cap bounds worst-case generation to ~5 min. Do NOT remove `-n` without a replacement ceiling. Sustained load hits ~84°C even at 92s; the APU is fanless/low-flow. Clients MUST also set `max_tokens`.
- **Port 8080 firewall**: amdpve iptables restricts 8080 to 192.168.68.116 (LiteLLM/router host) only. All inbound connections are from .116 (LiteLLM proxied via nginx). Localhost curls hang (SYN dropped). Always test from .116. - **Port 8080 firewall**: amdpve iptables restricts 8080 to 192.168.68.116 (LiteLLM/router host) only. All inbound connections are from .116 (LiteLLM proxied via nginx). Localhost curls hang (SYN dropped). Always test from .116.
- **Router sidecar fallback**: `router.py` `check_gpu_health()` now probes GPU `/health` directly when sidecar at :8090 is absent. Sidecar JSON exporters not deployed on any GPU host — router relies on GPU-direct fallback. - **Router sidecar fallback**: `router.py` `check_gpu_health()` now probes GPU `/health` directly when sidecar at :8090 is absent. Sidecar JSON exporters not deployed on any GPU host — router relies on GPU-direct fallback.
- **Router GPU_MOE_URL bug (fixed 2026-07-01)**: docker-compose had `GPU_MOE_URL=.110:8080` (gemma host) instead of `.15:8080` (amdpve). Corrected. - **Router GPU_MOE_URL bug (fixed 2026-07-01)**: docker-compose had `GPU_MOE_URL=.110:8080` (gemma host) instead of `.15:8080` (amdpve). Corrected.
@@ -265,7 +268,7 @@ If no SSH access, send Zulip DM via abiba-bot with vault update instructions.
| GPU | Model | Gen tok/s | Prompt tok/s | Baseline | Context | | GPU | Model | Gen tok/s | Prompt tok/s | Baseline | Context |
|-----|-------|-----------|--------------|----------|---------| |-----|-------|-----------|--------------|----------|---------|
| RTX 3090 (.8) | qwen3.6-27B-code (MTP) | **63** | — | — | **128K** | | RTX 3090 (.8) | SmartCode-Fable-5-27B-UD-Q3_K_XL | **TBD** | — | — | **128K** |
| RTX 5070 (.110) | gemma-4-12b (IQ4_NL+MTP) | **191** | — | — | **128K** | | RTX 5070 (.110) | gemma-4-12b (IQ4_NL+MTP) | **191** | — | — | **128K** |
| Strix Halo (.15) | qwen3.6-35B-udq4 | **65** | 140 | — | **128K** | | Strix Halo (.15) | qwen3.6-35B-udq4 | **65** | 140 | — | **128K** |
@@ -298,7 +301,7 @@ When the underlying model is swapped, only the LiteLLM config changes — agent
### Mumuni Agent Profile ### Mumuni Agent Profile
Mumuni (CT114, 192.168.68.123) is the primary business assistant. This profile is the reference for all agent configs: Mumuni (CT100/abiba, 192.168.68.24) is the primary business assistant. This profile is the reference for all agent configs:
| Setting | Value | Notes | | Setting | Value | Notes |
|---------|-------|-------| |---------|-------|-------|
@@ -323,7 +326,7 @@ Mumuni (CT114, 192.168.68.123) is the primary business assistant. This profile i
| Agent | Host | Status | | Agent | Host | Status |
|-------|------|--------| |-------|------|--------|
| **Mumuni** | CT114 (.123) | ✅ Updated to stable aliases | | **Mumuni** | CT100 (.24) | ✅ Updated to stable aliases |
| **Tanko** | CT112 (.122) | ✅ Updated to stable aliases | | **Tanko** | CT112 (.122) | ✅ Updated to stable aliases |
| **Koby** | CT111 (.129) | ❌ SSH unreachable — needs Zulip DM | | **Koby** | CT111 (.129) | ❌ SSH unreachable — needs Zulip DM |
| **Koonimo** | CT113 | ❌ SSH unreachable — needs Zulip DM | | **Koonimo** | CT113 | ❌ SSH unreachable — needs Zulip DM |
+2 -2
View File
@@ -251,8 +251,8 @@ call update-gpu-health
## Reporting ## Reporting
### 1. Knowledge Graph ### 1. Gitea Log (not knowledge graph — hard rule)
Every action logged as `[GPU-SELF-HEAL] <run_id>` node with full audit trail. Pushed to `SyslogSolution/health-logs/gpu/{run_id}.json` — versioned, searchable, not in graph.
### 2. Zulip Alerts (#agent-hub → alerts-gpu) ### 2. Zulip Alerts (#agent-hub → alerts-gpu)
- `issues_fixed > 0` → "🛠 GPU Self-Heal — <gpu> <issue> resolved" - `issues_fixed > 0` → "🛠 GPU Self-Heal — <gpu> <issue> resolved"
+5 -5
View File
@@ -26,11 +26,11 @@ done
|-------|-----|------|-----|---------------|------------|----------| |-------|-----|------|-----|---------------|------------|----------|
| Tanko | 112 | amdpve | .122 | `tanko` | Infisical vault | Hermes | | Tanko | 112 | amdpve | .122 | `tanko` | Infisical vault | Hermes |
| Mumuni | 114 | hwepve | .123 | `mumuni` | Infisical vault | Hermes | | Mumuni | 114 | hwepve | .123 | `mumuni` | Infisical vault | Hermes |
| Koby | 129 | amdpve | srv1079750 | `koby` | Infisical vault | **Hermes** | | Koby | 111 | amdpve | srv1079750 | `koby` | Infisical vault | **Hermes** |
| Koonimo | 114 | amdpve | ? | `koonimo` | Infisical vault | Hermes | | Koonimo | 113 | amdpve | .114 | `koonimo` | Infisical vault | Hermes |
| Shumba | — | 192.168.68.119 | N/A | N/A (DeepSeek) | Hermes (RETIRED — CT119 now Infisical vault) | | Shumba | — | 192.168.68.119 | N/A | N/A (DeepSeek) | Hermes (RETIRED — CT119 now Infisical vault) |
> **Note**: CT hostnames (tdunna→CT129, baggy→CT114) differ from agent identities (koby, koonimo). > **Note**: CT hostnames (tdunna→CT111, baggy→CT113) differ from agent identities (koby, koonimo).
Access: `pct-run <CT_ID> <command>` — no IPs needed. GPU hosts (.8, .110, .15) use SSH. Access: `pct-run <CT_ID> <command>` — no IPs needed. GPU hosts (.8, .110, .15) use SSH.
Keys are stored in Infisical vault (project=agents, env=production) and injected at Keys are stored in Infisical vault (project=agents, env=production) and injected at
@@ -169,9 +169,9 @@ pct-run <CT> grep -A8 "vision:" /root/.hermes/config.yaml | grep api_key
# Must show both api_key: sk-... and api_key_env: LITELLM_API_KEY # Must show both api_key: sk-... and api_key_env: LITELLM_API_KEY
``` ```
### For Koby (CT 129 / tdunna) ### For Koby (CT 111 / tdunna)
Koby runs Hermes on CT 129 (tdunna). Config files at `/root/.hermes/config.yaml`. Koby runs Hermes on CT 111 (tdunna). Config files at `/root/.hermes/config.yaml`.
Same Hermes pattern as Tanko/Mumuni/Koonimo — see config sections above. Same Hermes pattern as Tanko/Mumuni/Koonimo — see config sections above.
**LiteLLM key**: alias `koby` in LiteLLM DB, injected via `infisical run --` wrapper. **LiteLLM key**: alias `koby` in LiteLLM DB, injected via `infisical run --` wrapper.
+47 -29
View File
@@ -5,14 +5,11 @@ description: >
Standard Hermes configuration template for Syslog Solution LLC agents. Standard Hermes configuration template for Syslog Solution LLC agents.
Enforces shared infrastructure setup (Firecrawl, SearXNG, local models, Enforces shared infrastructure setup (Firecrawl, SearXNG, local models,
RA-H OS MCP) while keeping agent-specific API keys and model choices. RA-H OS MCP) while keeping agent-specific API keys and model choices.
UPDATED 2026-07-18: Compression model switched to `syslog-auto` (was `strix-moe`) UPDATED 2026-07-16: Compression model is the stable alias `strix-moe` (NOT `ornith-1.0-35b`,
to relieve Strix Halo pressure. syslog-auto distributes compression across the
weighted pool (55% RTX 3090, 30% Strix Halo, 15% RTX 5070).
UPDATED 2026-07-16: Compression model was the stable alias `strix-moe` (NOT `ornith-1.0-35b`,
which LiteLLM does not serve). All 3 GPUs verified at 128K (reduced from 256K 2026-07-17 for stability). which LiteLLM does not serve). All 3 GPUs verified at 128K (reduced from 256K 2026-07-17 for stability).
Added Rule 12 (Context-Issue Diagnostic) + Rule 13 (.env fallback enforcement) from the Added Rule 12 (Context-Issue Diagnostic) + Rule 13 (.env fallback enforcement) from the
2026-07-16 Mumuni root-cause investigation (WAL #1300). 2026-07-16 Mumuni root-cause investigation (WAL #1300).
UPDATED 2026-07-12: GPU workload redistributed. Compression → Strix Halo (later switched to syslog-auto 2026-07-18). RTX 3090 context verified at 128K. Infisical .env fallback required (Rule 3/13). UPDATED 2026-07-12: GPU workload redistributed. Compression → Strix Halo. RTX 3090 context verified at 128K. Infisical .env fallback required (Rule 3/13).
--- ---
## Maintains ## Maintains
@@ -133,9 +130,7 @@ mcp_servers:
# ─── Compression ─── # ─── Compression ───
compression: compression:
enabled: true enabled: true
model: syslog-auto # ⚠️ Switched from strix-moe 2026-07-18 to relieve Strix Halo. model: syslog-auto # ⚠️ Must match auxiliary.compression.model. Stable alias (gpu-fleet § Stable Role-Based Aliases). NOT ornith-1.0-35b (LiteLLM does not serve that name).
# syslog-auto distributes across weighted pool (55% RTX 3090,
# 30% Strix Halo, 15% RTX 5070). All GPUs at 128K.
provider: harness provider: harness
max_context_window: 131072 # MUST match actual GPU capacity. All 3 GPUs are 128K (Jul 17). max_context_window: 131072 # MUST match actual GPU capacity. All 3 GPUs are 128K (Jul 17).
threshold: 0.65 # Fires at ~170K for 262K window, ~85K for 128K threshold: 0.65 # Fires at ~170K for 262K window, ~85K for 128K
@@ -150,9 +145,8 @@ compression:
# model: gpu-light # stable alias (NOT raw "gemma-4-12b") # model: gpu-light # stable alias (NOT raw "gemma-4-12b")
# base_url: http://192.168.68.116/v1 # base_url: http://192.168.68.116/v1
# api_key_env: LITELLM_API_KEY # api_key_env: LITELLM_API_KEY
# Compression uses syslog-auto (switched from strix-moe 2026-07-18) to distribute # Do NOT use syslog-auto for auxiliary tasks — it routes to the primary GPU.
# load across the weighted pool and relieve Strix Halo pressure. # gpu-light = RTX 5070 (12B), freeing the Strix Halo for agent reasoning.
# Vision and web_extract use gpu-light = RTX 5070 (12B).
# Heavy aux (delegation, x_search) use gpu-dense (RTX 3090) instead. # Heavy aux (delegation, x_search) use gpu-dense (RTX 3090) instead.
# NEVER use raw model names (gemma-4-12b, qwen3.6-27B-code, qwen3.6-35B-udq4) # NEVER use raw model names (gemma-4-12b, qwen3.6-27B-code, qwen3.6-35B-udq4)
# in agent configs — use the stable aliases so model swaps don't break agents. # in agent configs — use the stable aliases so model swaps don't break agents.
@@ -172,7 +166,7 @@ auxiliary:
timeout: 30 timeout: 30
compression: compression:
provider: harness provider: harness
model: syslog-auto # Switched from strix-moe 2026-07-18. Relieves Strix Halo pressure. model: syslog-auto # MUST match compression.model above. Stable alias for Strix Halo (weighted pool).
base_url: http://192.168.68.116/v1 # Rule 5: /v1 NOT /litellm/v1 base_url: http://192.168.68.116/v1 # Rule 5: /v1 NOT /litellm/v1
api_key_env: LITELLM_API_KEY api_key_env: LITELLM_API_KEY
timeout: 300 # gpu-fleet: 300s for large-history summarization (was 60) timeout: 300 # gpu-fleet: 300s for large-history summarization (was 60)
@@ -253,30 +247,34 @@ The following MUST be identical across ALL profiles:
- Apply to BOTH main config AND all sub-agent profiles - Apply to BOTH main config AND all sub-agent profiles
- For agents needing longer outputs: raise to 8192, but never omit - For agents needing longer outputs: raise to 8192, but never omit
### Rule 7: Auxiliary Model Consistency (UPDATED 2026-07-18) ### Rule 7: Auxiliary Model Consistency (UPDATED 2026-07-16)
- Vision and web_extract use `gpu-light` (stable alias, RTX 5070 — 12GB, vision-optimized) - Vision and web_extract use `gemma-4-12b` (RTX 5070 — 12GB, vision-optimized)
- Compression now uses `syslog-auto` (switched from `strix-moe` 2026-07-18) to distribute - Compression uses `strix-moe` (stable alias for Strix Halo — 64GB, 128K ctx, compression-optimized)
compression load across the weighted pool (55% RTX 3090, 30% Strix Halo, 15% RTX 5070). - **`strix-moe` is the only valid compression model name** — LiteLLM does NOT serve `ornith-1.0-35b`
This relieves Strix Halo pressure while keeping compression functional on all GPUs. (it serves `strix-moe`, `qwen3.6-35B-udq4`, `gpu-dense`, `gpu-light`, `syslog-auto`, `gemma-4-12b`, `qwen3.6-27B-code`). Old configs with `ornith-1.0-35b` cause 403/model-not-found on compression calls.
- **`syslog-auto` is the valid compression model** — LiteLLM serves it as the weighted pool. - **OPERATIONAL DECISION (2026-07-23): Use `syslog-auto` for compression across all agents.**
Old configs with `strix-moe` for compression should be updated to `syslog-auto`. The `syslog-auto` alias routes to the Strix Halo, but uses the weighted pool instead of pinning
to `strix-moe` directly. This prevents sustained Strix Halo thermal load because the pool can
fall back to other GPUs if Strix gets hot. Both `compression.model` and `auxiliary.compression.model`
MUST be `syslog-auto`.
- All auxiliary services MUST use identical routing: - All auxiliary services MUST use identical routing:
- `base_url: http://192.168.68.116/v1` (Rule 5: `/v1`, NOT `/litellm/v1`) - `base_url: http://192.168.68.116/v1` (Rule 5: `/v1`, NOT `/litellm/v1`)
- `api_key_env: LITELLM_API_KEY` - `api_key_env: LITELLM_API_KEY`
- **Compression via syslog-auto**: Routes through the weighted pool. Strix Halo still handles - **Do NOT use `syslog-auto`** for auxiliary tasks — it routes unpredictably
~30% of compression calls (at 60 RPM via pool vs 40 RPM direct), but the bulk (55%) - **Compression on Strix Halo**: The strix-moe alias routes to Strix Halo
goes to RTX 3090 which has ample spare capacity. (64GB UMA, 128K context) — the designated compression GPU. This frees the
RTX 5070 for vision and web search, and the RTX 3090 for heavy reasoning.
- The `compression:` block's `model` MUST match `auxiliary: compression: model` - The `compression:` block's `model` MUST match `auxiliary: compression: model`
- The `compression: max_context_window: 131072` MUST match actual GPU capacity (128K) - The `compression: max_context_window: 131072` MUST match actual GPU capacity (128K)
### Rule 8: GPU Workload Distribution (UPDATED 2026-07-18) ### Rule 8: GPU Workload Distribution (UPDATED 2026-07-16)
- **RTX 3090 (24GB, 128K ctx, qwen3.6-27B-code)**: Heavy reasoning, code gen, long conversations — also handles ~55% of compression via syslog-auto pool - **RTX 3090 (24GB, 128K ctx, qwen3.6-27B-code)**: Heavy reasoning, code gen, long conversations
- **RTX 5070 (12GB, 128K ctx, gemma-4-12b)**: Vision, web search, quick tasks, web_extract — handles ~15% of compression via syslog-auto pool - **RTX 5070 (12GB, 128K ctx, gemma-4-12b)**: Vision, web search, quick tasks, web_extract (IQ4_NL+MTP, ~65% VRAM at 128K)
- **Strix Halo (64GB, 128K ctx, qwen3.6-35B-udq4)**: Agent reasoning, compression (~30% via syslog-auto pool), fallback for other GPUs - **Strix Halo (64GB, 128K ctx, syslog-auto)**: Context compression, summarization, long docs
- Agent profiles MUST route auxiliary tasks to the correct GPU: - Agent profiles MUST route auxiliary tasks to the correct GPU:
- `auxiliary.vision.model: gpu-light` (RTX 5070) - `auxiliary.vision.model: gemma-4-12b` (RTX 5070)
- `auxiliary.web_extract.model: gpu-light` (RTX 5070) - `auxiliary.web_extract.model: gemma-4-12b` (RTX 5070)
- `auxiliary.compression.model: syslog-auto` (distributed pool, switched from strix-moe 2026-07-18) - `auxiliary.compression.model: syslog-auto` (Strix Halo)
- Default model (`model.default`) and custom_provider remain `syslog-auto` for auto-routing - Default model (`model.default`) and custom_provider remain `syslog-auto` for auto-routing
- For 128K context window: `threshold: 0.65` (fires at ~85K tokens) - For 128K context window: `threshold: 0.65` (fires at ~85K tokens)
- Do NOT use `threshold: 0.25` — this fires at 65K, causing premature context loss - Do NOT use `threshold: 0.25` — this fires at 65K, causing premature context loss
@@ -378,6 +376,26 @@ curl -s -o /dev/null -w '%{http_code}' -H "Authorization: Bearer $K" http://192.
- `/etc/environment` is NO LONGER the canonical key source (stale values there caused 401s). - `/etc/environment` is NO LONGER the canonical key source (stale values there caused 401s).
- Do NOT leave a hardcoded stale key in `/etc/environment` — it shadows the drop-in/wrapper. - Do NOT leave a hardcoded stale key in `/etc/environment` — it shadows the drop-in/wrapper.
### Rule 14: Provider Name Must Match custom_providers Name (ADDED 2026-07-19, WAL #1471)
- `model.provider` MUST be `harness` (the `custom_providers[0].name`), NOT the literal string `custom`
- When `provider: custom`, Hermes' `_get_named_custom_provider("custom")` returns None (no provider is
named "custom" — it is named "harness"), causing a fall-through to the generic resolution path
(`source: env/config`) at `runtime_provider.py:1156`
- The generic path builds `api_key_candidates` from `model.api_key` (empty), host-gated
OLLAMA/OPENAI/OPENROUTER keys, and `_host_derived_api_key` (returns "" for IP addresses)
- **The generic path does NOT resolve `model.api_key_env` or `custom_providers.key_env`** —
`LITELLM_API_KEY` is never read, producing `api_key = "no-key-required"` → HTTP 401
- The named custom provider path (`source: custom_provider:harness`) DOES read `key_env` —
but only triggers when `provider` matches the `custom_providers[0].name`
- All sections MUST use `provider: harness`: `model`, `compression`, `auxiliary.vision`,
`auxiliary.web_extract`, `auxiliary.compression`, `delegation`
- Only `fallback_providers` uses a different provider (`deepseek`) for true fallback diversity
- **Diagnostic**: If you see `source: env/config` in a request dump or log, the provider name
is wrong. It should be `source: custom_provider:harness`.
- **Audit script**: Run `python3 /root/prose-contracts/audit-hermes-config.py <config.yaml>`
before and after any config change to catch this and all other rule violations.
## Execution ## Execution
1. **Check current config** — Read the target agent's config.yaml 1. **Check current config** — Read the target agent's config.yaml
+1 -1
View File
@@ -189,7 +189,7 @@ litellm_settings:
| Agent | CT | IP | LiteLLM Alias | Key Source | Status | Gateway Wrapper | Last Verified | | Agent | CT | IP | LiteLLM Alias | Key Source | Status | Gateway Wrapper | Last Verified |
|-------|-----|-----|---------------|------------|--------|-----------------|---------------| |-------|-----|-----|---------------|------------|--------|-----------------|---------------|
| Tanko | 112 | .122 | `tanko` | Infisical vault | ✅ Fixed | `infisical run` | 20:17 UTC Jul 5 | | Tanko | 112 | .122 | `tanko` | Infisical vault | ✅ Fixed | `infisical run` | 20:17 UTC Jul 5 |
| Mumuni | 114 | .123 | `mumuni` | Infisical vault | ✅ Fixed | `infisical run` | 01:46 EDT Jul 10 | | Mumuni | 100 (abiba) | .24 | `mumuni` | Infisical vault | ✅ Fixed | Pi Hermes gateway | 2026-07-27 |
| Koby | 111 | ? | `koby` | Infisical vault | ✅ Fixed | `infisical run` | 23:30 UTC Jul 5 | | Koby | 111 | ? | `koby` | Infisical vault | ✅ Fixed | `infisical run` | 23:30 UTC Jul 5 |
| Koonimo | 113 | ? | `koonimo` | Infisical vault | ✅ Fixed | `infisical run` (migrated 2026-07-11) | 2026-07-11 | | Koonimo | 113 | ? | `koonimo` | Infisical vault | ✅ Fixed | `infisical run` (migrated 2026-07-11) | 2026-07-11 |
| Abiba | 100 | .65 | `abiba-pi` | Infisical vault | ✅ N/A (pi native) | — | 19:44 UTC Jul 5 | | Abiba | 100 | .65 | `abiba-pi` | Infisical vault | ✅ N/A (pi native) | — | 19:44 UTC Jul 5 |
+1 -1
View File
@@ -51,7 +51,7 @@ gateway restart, and connection validation.
| Host | CT | Proxmox | IP (direct) | Hermes Home | User | | Host | CT | Proxmox | IP (direct) | Hermes Home | User |
|------|-----|---------|-------------|-------------|------| |------|-----|---------|-------------|-------------|------|
| Mumuni | CT114 | hwepve | 192.168.68.123 | /root/.hermes | root | | Mumuni | CT100 (abiba) | hwepve | 192.168.68.24 | /root/.hermes | root |
| Tanko | CT112 | amdpve | 192.168.68.122 | /home/jerome/.hermes | jerome | | Tanko | CT112 | amdpve | 192.168.68.122 | /home/jerome/.hermes | jerome |
| Koby | CT111 | amdpve | 192.168.68.129 | /root/.hermes | root | | Koby | CT111 | amdpve | 192.168.68.129 | /root/.hermes | root |
| Shumba | — | — | 192.168.68.119 | /home/lucky/.hermes | lucky | | Shumba | — | — | 192.168.68.119 | /home/lucky/.hermes | lucky |
+42 -27
View File
@@ -3,7 +3,7 @@ kind: pattern
name: infrastructure-control name: infrastructure-control
description: > description: >
Full infrastructure monitoring and control pattern covering the Full infrastructure monitoring and control pattern covering the
5-node Proxmox cluster, 3 Docker ecosystems (22 containers), 6-node Proxmox cluster, 3 Docker ecosystems (22 containers),
NFS storage, and network services. Defines monitors, remediations, NFS storage, and network services. Defines monitors, remediations,
and the access matrix for all environments. and the access matrix for all environments.
@@ -12,6 +12,11 @@ description: >
never mutate infrastructure based on them without first confirming never mutate infrastructure based on them without first confirming
against the live system. Policy fields are authoritative. See the against the live system. Policy fields are authoritative. See the
`verify-before-mutate` skill. `verify-before-mutate` skill.
**Last verified:** 2026-07-24 — corrected Gitea IP (.17 not .110),
AdGuard IP (.10 not .102), AdGuard placement (minipve not acerpve),
Abiba placement (hwepve not amdpve), added hwepve as 6th node,
added dns.sysloggh.net route.
--- ---
# Infrastructure Control Pattern # Infrastructure Control Pattern
@@ -42,8 +47,8 @@ description: >
│ │ │ │ │ │ │ │ │ │
┌────┘ ┌────┘ ┌────┘ ┌────┘ ┌────┘ ┌────┘ ┌────┘ ┌────┘ ┌────┘ ┌────┘
▼ ▼ ▼ ▼ ▼ ▼ ▼ ▼ ▼ ▼
minipve amdpve storepve acerpve ocupve minipve amdpve storepve acerpve ocupve hwepve
(.12) (.15) (.6) (.9) (.5) (.12) (.15) (.6) (.9) (.5) (.4)
┌─────────────────────────────────────────────┐ ┌─────────────────────────────────────────────┐
@@ -95,15 +100,21 @@ description: >
## Section 2: Proxmox Cluster — Monitoring ## Section 2: Proxmox Cluster — Monitoring
### Nodes (5) ### Nodes (6)
| Node | IP | CPU | RAM | VMs/CTs | Role | | Node | IP | CPU | RAM | VMs/CTs | Role |
|------|----|-----|-----|---------|------| |------|----|-----|-----|---------|------|
| minipve | .12 | 16C | 30GB | authentik, gitea, mumuni, syslog-api, jitsi | Auth, git, messaging | | minipve | .12 | 16C | 30GB | authentik, gitea, syslog-api, infisical-vault, jitsi | Auth, git, messaging |
| amdpve | .15 | 32C | 62GB | abiba, kagentz, tanko, tdunna, baggy, scottdenya | Agents, compute | | amdpve | .15 | 32C | 62GB | tanko, tdunna, baggy, scottdenya | Agents, compute |
| storepve | .6 | 28C | 31GB | docker-vm, ra-h-os, PBS, media, zulip | Docker, storage, chat | | storepve | .6 | 28C | 31GB | docker-vm, ra-h-os, PBS, media, jdownloader, zulip | Docker, storage, chat |
| acerpve | .9 | 28C | 31GB | llm-gpu, adguard | GPU VMs | | acerpve | .9 | 28C | 31GB | llm-gpu | GPU VMs |
| ocupve | .5 | 12C | 14GB | ocu-llm | GPU VMs | | ocupve | .5 | 12C | 14GB | ocu-llm | GPU VMs |
| hwepve | .4 | 12C | 15GB | abiba, kagentz, (mumuni CT 114 stopped) | Agents (new node) |
> **Note:** CTs on storepve include jdownloader (CT 118). AdGuard (CT 102) is on
> minipve at .10, not acerpve. Abiba (CT 100) is on hwepve, not amdpve. Mumuni
> (CT 114) is on hwepve (currently stopped), not minipve. Mumuni also has a
> second instance on minipve at .123 — distinguish by CT ID, not hostname.
### Checks (every 5 min) ### Checks (every 5 min)
@@ -354,16 +365,18 @@ fine. Services that resolve directly to a LAN IP are NetBird-independent.
|---------|--------|-------------|-------------|-------------|--------| |---------|--------|-------------|-------------|-------------|--------|
| Proxmox API | minipve.sysloggh.net:8006 | 192.168.68.12 | LAN IP | No | ✅ | | Proxmox API | minipve.sysloggh.net:8006 | 192.168.68.12 | LAN IP | No | ✅ |
| LiteLLM | litellm.sysloggh.net | 192.168.68.116 | LAN IP | No | ✅ | | LiteLLM | litellm.sysloggh.net | 192.168.68.116 | LAN IP | No | ✅ |
| Authentik | auth.sysloggh.net:443 | 192.168.68.11 | CNAME → netbird | **Yes** | ⚠️ | | Authentik | auth.sysloggh.net:443 | 192.168.68.11:9000 | CNAME → netbird | **Yes** | ⚠️ |
| Gitea | git.sysloggh.net:443 | 192.168.68.110 | CNAME → netbird | **Yes** | ⚠️ | | Gitea | git.sysloggh.net:443 | 192.168.68.17:3000 | CNAME → netbird | **Yes** | ⚠️ |
| Zulip | chat.sysloggh.net:443 | 192.168.68.19 | CNAME → netbird | **Yes** | ⚠️ VERIFY-BEFORE-USE | | Zulip | chat.sysloggh.net:443 | 192.168.68.19 | CNAME → netbird | **Yes** | ⚠️ VERIFY-BEFORE-USE |
| Pulse | pulse.sysloggh.net:443 | 192.168.68.7 | CNAME → netbird | **Yes** | ⚠️ | | Pulse | pulse.sysloggh.net:443 | 192.168.68.7 | CNAME → netbird | **Yes** | ⚠️ |
| DNS UI | dns.sysloggh.net:443 | 192.168.68.102 | CNAME → netbird | **Yes** | ⚠️ | | DNS UI | dns.sysloggh.net:443 | 192.168.68.10:80 | CNAME → netbird | **Yes** | ⚠️ |
| SearXNG | searxng.sysloggh.net:8888 | 192.168.68.7:8888 | LAN IP | No | ✅ | | SearXNG | searxng.sysloggh.net:8888 | 192.168.68.7:8888 | LAN IP | No | ✅ |
| Firecrawl | firecrawl.sysloggh.net:3002 | 192.168.68.7:3002 | LAN IP | No | ✅ | | Firecrawl | firecrawl.sysloggh.net:3002 | 192.168.68.7:3002 | LAN IP | No | ✅ |
**Verified 2026-07-02:** NetBird VPS rebooted after a hang; all CNAME'd **Verified 2026-07-24:** NetBird VPS rebooted after a hang; all CNAME'd
services recovered. LAN-IP-direct paths stayed up throughout the outage. services recovered. LAN-IP-direct paths stayed up throughout the outage.
Also added `dns.sysloggh.net` route (was missing entirely).
See `scripts/netbird-add-domain.sh` for adding new proxy routes.
### 5.2 Checks (every 2 min) ### 5.2 Checks (every 2 min)
@@ -511,8 +524,8 @@ enforced by the `routing-regression.config_url_violations` check in Section
| LiteLLM API | `http://192.168.68.116:4000` | `https://litellm.sysloggh.net` | | LiteLLM API | `http://192.168.68.116:4000` | `https://litellm.sysloggh.net` |
| LiteLLM (nginx) | `http://192.168.68.116` | — | | LiteLLM (nginx) | `http://192.168.68.116` | — |
| Grafana | `http://192.168.68.116:3001` | — | | Grafana | `http://192.168.68.116:3001` | — |
| Authentik | `https://192.168.68.11` | `https://auth.sysloggh.net` | | Authentik | `https://192.168.68.11:9000` | `https://auth.sysloggh.net` |
| Gitea | `http://192.168.68.110:3000` | `https://git.sysloggh.net` | | Gitea | `http://192.168.68.17:3000` | `https://git.sysloggh.net` |
| Zulip API | `http://192.168.68.19` | `https://chat.sysloggh.net` | | Zulip API | `http://192.168.68.19` | `https://chat.sysloggh.net` |
| SearXNG | `http://192.168.68.7:8888` | — | | SearXNG | `http://192.168.68.7:8888` | — |
| Firecrawl | `http://192.168.68.7:3002` | — | | Firecrawl | `http://192.168.68.7:3002` | — |
@@ -575,25 +588,26 @@ ssh root@192.168.68.110 "systemctl restart llama-server"
| CT | Name | Node | IP | Role | Agent | | CT | Name | Node | IP | Role | Agent |
|----|------|------|----|------|-------| |----|------|------|----|------|-------|
| 100 | abiba | amdpve | .24 | Pi agent (this host) | ✅ pi | | 100 | abiba | **hwepve** | .24 | Pi agent | ✅ pi |
| 101 | llm-gpu | acerpve | .8 | GPU RTX 3090 | ❌ | | 101 | llm-gpu | acerpve | .8 | GPU RTX 3090 | ❌ |
| 102 | adguard | acerpve | | DNS | ❌ | | 102 | adguard | **minipve** | **.10** | DNS | ❌ |
| 103 | ocu-llm | ocupve | .110 | GPU RTX 5070 | ❌ | | 103 | ocu-llm | ocupve | .110 | GPU RTX 5070 | ❌ |
| 104 | authentik | minipve | .11 | OIDC | ❌ | | 104 | authentik | minipve | .11 | OIDC | ❌ |
| 105 | kagentz | amdpve | — | Agent Zero | ✅ | | 105 | kagentz | **hwepve** | — | Agent Zero | ✅ |
| 106 | ra-h-os | storepve | .65 | KG bridge | ✅ MCP | | 106 | ra-h-os | storepve | .65 | KG bridge | ✅ MCP |
| 107 | pbs | storepve | — | Backups | ❌ | | 107 | pbs | storepve | — | Backups | ❌ |
| 108 | media | storepve | — | Media | ❌ | | 108 | media | storepve | — | Media | ❌ |
| 109 | docker-vm | storepve | .7 | Docker host | ❌ | | 109 | docker-vm | storepve | .7 | Docker host | ❌ |
| 110 | gitea | minipve | | Git | ❌ | | 110 | gitea | minipve | **.17** | Git | ❌ |
| 111 | tdunna | amdpve | .129 | Hermes agent | ✅ | | 111 | tdunna | amdpve | .129 | Hermes agent | ✅ |
| 112 | tanko | amdpve | .122 | Hermes agent | ✅ | | 112 | tanko | amdpve | .122 | Hermes agent | ✅ |
| 113 | baggy | amdpve | ? | Hermes agent | ✅ | | 113 | baggy | amdpve | .114 | Hermes agent | ✅ |
| 114 | mumuni | minipve | .123 | Hermes agent | ✅ | | 114 | mumuni | **hwepve** | .123 | Hermes agent (stopped) | ✅ |
| 115 | scottdenya | amdpve | — | ? | ❌ | | 115 | scottdenya | amdpve | .75 | Denya OneCare | ❌ |
| 116 | syslog-api | minipve | .116 | LiteLLM + Grafana | ❌ | | 116 | syslog-api | minipve | .116 | LiteLLM + Grafana | ❌ |
| 117 | zulip | storepve | | Chat | ❌ | | 117 | zulip | storepve | .19 | Chat | ❌ |
| 118 | jitsi | minipve | — | Video | ❌ | | 118 | jdownloader | storepve | — | JDownloader container | ❌ |
| 119 | infisical-vault | minipve | — | Vault | ❌ |
## Appendix C: Docker Compose Files Location ## Appendix C: Docker Compose Files Location
@@ -613,27 +627,28 @@ Source of truth: `/root/scripts/pct-run.sh` or `prose-contracts/scripts/pct-run.
| CT | Name | Node | pct-run | | CT | Name | Node | pct-run |
|-----|------|------|---------| |-----|------|------|---------|
| 100 | abiba | amdpve | `pct-run 100` | | 100 | abiba | hwepve | `pct-run 100` |
| 105 | kagentz | amdpve | `pct-run 105` | | 105 | kagentz | hwepve | `pct-run 105` |
| 111 | tdunna | amdpve | `pct-run 111` | | 111 | tdunna | amdpve | `pct-run 111` |
| 112 | tanko | amdpve | `pct-run 112` | | 112 | tanko | amdpve | `pct-run 112` |
| 113 | baggy | amdpve | `pct-run 113` | | 113 | baggy | amdpve | `pct-run 113` |
| 115 | scottdenya | amdpve | `pct-run 115` | | 115 | scottdenya | amdpve | `pct-run 115` |
| 104 | authentik | minipve | `pct-run 104` | | 104 | authentik | minipve | `pct-run 104` |
| 110 | gitea | minipve | `pct-run 110` | | 110 | gitea | minipve | `pct-run 110` |
| 114 | mumuni | minipve | `pct-run 114` | | 114 | mumuni | hwepve | `pct-run 114` |
| 116 | syslog-api | minipve | `pct-run 116` | | 116 | syslog-api | minipve | `pct-run 116` |
| 106 | ra-h-os | storepve | `pct-run 106` | | 106 | ra-h-os | storepve | `pct-run 106` |
| 107 | proxmox-backup | storepve | `pct-run 107` | | 107 | proxmox-backup | storepve | `pct-run 107` |
| 108 | media | storepve | `pct-run 108` | | 108 | media | storepve | `pct-run 108` |
| 117 | zulip | storepve | `pct-run 117` | | 117 | zulip | storepve | `pct-run 117` |
| 102 | adguard | acerpve | `pct-run 102` | | 102 | adguard | **minipve** | `pct-run 102` |
GPU bare-metal hosts (.8 acerpve, .110 ocupve, .15 amdpve) are NOT CTs — use SSH directly: GPU bare-metal hosts (.8 acerpve, .110 ocupve, .15 amdpve) are NOT CTs — use SSH directly:
```bash ```bash
ssh root@192.168.68.8 # RTX 3090 ssh root@192.168.68.8 # RTX 3090
ssh root@192.168.68.110 # RTX 5070 ssh root@192.168.68.110 # RTX 5070
ssh root@192.168.68.15 # Strix Halo ssh root@192.168.68.15 # Strix Halo
ssh root@192.168.68.4 # hwepve (abiba, kagentz, mumuni)
``` ```
## Section 7: Agent Health Check (consolidated — 2026-07-05) ## Section 7: Agent Health Check (consolidated — 2026-07-05)
+42 -6
View File
@@ -2,7 +2,7 @@
kind: responsibility kind: responsibility
name: infrastructure-update name: infrastructure-update
description: > description: >
Autonomous system-wide update contract covering all 5 Proxmox nodes, Autonomous system-wide update contract covering all 6 Proxmox nodes,
15+ containers/VMs, and 4 Docker ecosystems. Updates apt packages, 15+ containers/VMs, and 4 Docker ecosystems. Updates apt packages,
Docker images, and container stacks in safe waves with health checks Docker images, and container stacks in safe waves with health checks
and automatic rollback on failure. and automatic rollback on failure.
@@ -56,10 +56,11 @@ Before ANY update wave:
| amdpve (.15) | Proxmox node | `apt update && apt upgrade -y` | 5 min | | amdpve (.15) | Proxmox node | `apt update && apt upgrade -y` | 5 min |
| acerpve (.9) | Proxmox node | `apt update && apt upgrade -y` | 5 min | | acerpve (.9) | Proxmox node | `apt update && apt upgrade -y` | 5 min |
| ocupve (.5) | Proxmox node | `apt update && apt upgrade -y` | 5 min | | ocupve (.5) | Proxmox node | `apt update && apt upgrade -y` | 5 min |
| hwepve (.4) | Proxmox node | `apt update && apt upgrade -y` | 5 min |
| CT 100 (.24) | Abiba (pi) | `apt update && apt upgrade -y` | 3 min | | CT 100 (.24) | Abiba (pi) | `apt update && apt upgrade -y` | 3 min |
| CT 116 (.116) | syslog-api (LiteLLM host) | `apt update && apt upgrade -y` | 3 min | | CT 116 (.116) | syslog-api (LiteLLM host) | `apt update && apt upgrade -y` | 3 min |
| CT 112 (tanko, amdpve) | Tanko | `apt update && apt upgrade -y` | 3 min | | CT 112 (tanko, amdpve) | Tanko | `apt update && apt upgrade -y` | 3 min |
| CT 114 (mumuni, minipve) | Mumuni | `apt update && apt upgrade -y` | 3 min | | CT 114 (mumuni, hwepve) | Mumuni | `apt update && apt upgrade -y` | 3 min |
| VM 101 (.8) | llm-gpu (RTX 3090) | `apt update && apt upgrade -y` | 3 min | | VM 101 (.8) | llm-gpu (RTX 3090) | `apt update && apt upgrade -y` | 3 min |
| VM 103 (.110) | ocu-llm (RTX 5070) | `apt update && apt upgrade -y` | 3 min | | VM 103 (.110) | ocu-llm (RTX 5070) | `apt update && apt upgrade -y` | 3 min |
@@ -91,17 +92,52 @@ Before ANY update wave:
- Firecrawl test: `curl :3002/` - Firecrawl test: `curl :3002/`
- SearXNG test: `curl :8888` - SearXNG test: `curl :8888`
## Wave 4: Proxmox Kernel Reboot (if needed) ## Wave 4: Proxmox Kernel Reboot
Only if `[ -f /var/run/reboot-required ]` on any node. Only if `[ -f /var/run/reboot-required ]` on any node.
| Target | Action | | Target | Action |
|--------|--------| |--------|--------|
| Affected PVE node | Verify all CTs/VMs migrated or stopped | | Affected PVE node | Verify all CTs/VMs migrated or stopped |
| | `reboot` via PVE API | | | `reboot` via PVE API (or `systemctl reboot -f` if dbus fails) |
| | Wait 120s for node to come back | | | Wait 120s for node to come back |
| | Start any stopped CTs | | | Start any stopped CTs |
### Post-reboot sweep (known gaps)
After every node reboot, run these checks:
1. **CT auto-start sweep** — LXC containers sometimes don't start despite
`onboot: 1`. Check every CT on the rebooted node and start any left stopped:
```bash
pct list | awk '/stopped/{print $1}' | xargs -I{} pct start {}
```
Known cases: scottdenya (CT 115 on amdpve), authentik (CT 104 on minipve).
2. **Zulip recovery** — When docker-vm or storepve reboots, the Zulip main
container loses its Docker network assignment (SIGKILL during storage
outage detaches it from `zulip_default` network). Run:
```bash
ssh root@192.168.68.19 'docker rm -f zulip-zulip-1 && cd /opt/zulip && docker compose up -d'
```
The compose restart recreates the container on the correct network.
3. **docker-vm Docker daemon** — After reboot, Docker can take 3-4 minutes
to become `active`. The docker-proxy for Pulse (port 7655) starts early,
so Pulse is accessible before `docker ps` reports ready. Wait for Docker
before checking other stacks.
### VPS ↔ docker-vm tunnel
After any VPS or docker-vm reboot, verify the dedicated WireGuard tunnel:
```bash
ssh root@72.61.0.17 'wg show wg1' | grep "latest handshake"
# If no handshake in >60s:
ssh root@72.61.0.17 'wg-quick up wg1'
```
The tunnel uses PersistentKeepalive=25 and is systemd-enabled, but should
be verified after a reboot.
## Rollback Protocol ## Rollback Protocol
If ANY verification fails: If ANY verification fails:
@@ -183,7 +219,7 @@ When LiteLLM is upgraded to a version supporting per-key MCP grants:
## Success Criteria ## Success Criteria
- [ ] All 5 PVE nodes updated, no reboot-loop - [ ] All 6 PVE nodes updated, no reboot-loop
- [ ] All VMs/CTs running post-update - [ ] All VMs/CTs running post-update
- [ ] All Docker containers healthy (VM 109 + CT 116 + CT 117) - [ ] All Docker containers healthy (VM 109 + CT 116 + CT 117)
- [ ] LiteLLM inference passing (syslog-auto test) - [ ] LiteLLM inference passing (syslog-auto test)
@@ -199,7 +235,7 @@ After completion, send Zulip DM:
``` ```
📋 Infrastructure Update — YYYY-MM-DD 📋 Infrastructure Update — YYYY-MM-DD
Updated: 5 PVE nodes, 12 CTs/VMs, 30+ containers Updated: 6 PVE nodes, 12 CTs/VMs, 30+ containers
Security fixes: N CVEs patched Security fixes: N CVEs patched
Downtime: <service> <duration> Downtime: <service> <duration>
Failures: none / <details> Failures: none / <details>
+1 -1
View File
@@ -178,7 +178,7 @@ through its agent wrapper.
| Agent | Host | Pattern | Keys | Status | | Agent | Host | Pattern | Keys | Status |
|-------|------|---------|------|--------| |-------|------|---------|------|--------|
| abiba | .24 | pi agent wrapper | ABIBA_LITELLM_API_KEY + ABIBA_ZULIP_API_KEY | ✅ vault-backed | | abiba | .24 | pi agent wrapper | ABIBA_LITELLM_API_KEY + ABIBA_ZULIP_API_KEY | ✅ vault-backed |
| mumuni | .123 | systemd drop-in + while-true wrapper + st.8e848433 | MUMUNI_LITELLM_API_KEY + MUMUNI_ZULIP_API_KEY | ✅ vault-backed + .env fallback | | mumuni | .24 (CT100 abiba) | Pi Hermes gateway (no systemd) | MUMUNI_LITELLM_API_KEY + MUMUNI_ZULIP_API_KEY | ✅ vault-backed + .env fallback |
| tanko | .122 | systemd drop-in + while-true wrapper + st.8e848433 (user jerome) | TANKO_LITELLM_API_KEY + TANKO_ZULIP_API_KEY | ✅ vault-backed + .env fallback | | tanko | .122 | systemd drop-in + while-true wrapper + st.8e848433 (user jerome) | TANKO_LITELLM_API_KEY + TANKO_ZULIP_API_KEY | ✅ vault-backed + .env fallback |
| koby | .129 | systemd drop-in + while-true wrapper + st.8e848433 | KOBY_LITELLM_API_KEY, shares TANKO_ZULIP_API_KEY (tanko-bot) | ✅ vault-backed | | koby | .129 | systemd drop-in + while-true wrapper + st.8e848433 | KOBY_LITELLM_API_KEY, shares TANKO_ZULIP_API_KEY (tanko-bot) | ✅ vault-backed |
| koonimo | .114 | systemd drop-in + while-true wrapper + st.8e848433 | KOONIMO_LITELLM_API_KEY + KOONIMO_ZULIP_API_KEY | ✅ vault-backed | | koonimo | .114 | systemd drop-in + while-true wrapper + st.8e848433 | KOONIMO_LITELLM_API_KEY + KOONIMO_ZULIP_API_KEY | ✅ vault-backed |
+26 -7
View File
@@ -7,7 +7,7 @@ note: >
Auto-remediation code was removed from the pi Zulip extension (retired 2026-07-04), Auto-remediation code was removed from the pi Zulip extension (retired 2026-07-04),
now reimplemented as `litellm-health-check.sh` on CT 116. now reimplemented as `litellm-health-check.sh` on CT 116.
Script: `/opt/inference-harness/scripts/litellm-health-check.sh` on CT 116 (cron `0 */6 * * *`). Script: `/opt/inference-harness/scripts/litellm-health-check.sh` on CT 116 (cron `0 */6 * * *`).
Reports to /var/log/litellm/health-*.json and RA-H OS knowledge graph. Reports to /var/log/litellm/health-*.json and Gitea (SyslogSolution/health-logs).
GPU monitoring integrated from gpu-monitor on .24:9100. GPU monitoring integrated from gpu-monitor on .24:9100.
Consolidated from litellm-health + litellm-self-heal on 2026-07-09 to eliminate Consolidated from litellm-health + litellm-self-heal on 2026-07-09 to eliminate
@@ -20,7 +20,7 @@ description: >
LiteLLM inference stack health monitoring + self-healing. Verifies the full LiteLLM inference stack health monitoring + self-healing. Verifies the full
nginx → LiteLLM → GPU chain, 8 containers on CT 116, 3 GPU hosts, model nginx → LiteLLM → GPU chain, 8 containers on CT 116, 3 GPU hosts, model
inference, and agent keys. Applies remediation rules for common failures. inference, and agent keys. Applies remediation rules for common failures.
Reports every action via Zulip DM and RA-H OS knowledge graph. Reports every action via Zulip DM and Gitea (SyslogSolution/health-logs).
--- ---
# LiteLLM Operations — Health Check + Self-Heal # LiteLLM Operations — Health Check + Self-Heal
@@ -102,7 +102,7 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2)
- **Health-check script** (`/opt/inference-harness/scripts/litellm-health-check.sh` on CT 116): `gpu-fleet` check fails only on **critical** alerts (warnings are informational). Tests `strix-moe` (not `ornith-1.0-35b`). - **Health-check script** (`/opt/inference-harness/scripts/litellm-health-check.sh` on CT 116): `gpu-fleet` check fails only on **critical** alerts (warnings are informational). Tests `strix-moe` (not `ornith-1.0-35b`).
- **GPU monitor** (`/root/scripts/gpu-monitor-server.py` on pi .24): runs as **systemd unit `gpu-monitor.service`** (was bare `&` process). `gpu_count` includes Strix Halo (was 2, now 3). VRAM alert thresholds: warning 93%, critical 97% (raised from 90/95 — 128K context steady-state is ~70% on RTX 3090, not a fault). - **GPU monitor** (`/root/scripts/gpu-monitor-server.py` on pi .24): runs as **systemd unit `gpu-monitor.service`** (was bare `&` process). `gpu_count` includes Strix Halo (was 2, now 3). VRAM alert thresholds: warning 93%, critical 97% (raised from 90/95 — 128K context steady-state is ~70% on RTX 3090, not a fault).
- **Agent key monitor** (`/root/scripts/agent-health-check.py` on pi .24, cron `*/10`): reads each agent's **live** `LITELLM_API_KEY` from its gateway process env via SSH — never hardcodes keys (hardcoded keys rot on rotation and caused 9×401/30min). Fleet roster: abiba, tanko, mumuni, koby, koonimo (legacy `tdunna`/`baggy` removed — never existed). - **Agent key monitor** (`/root/scripts/agent-health-check.py` on pi .24, cron `*/10`): v2 (2026-07-26) — reads each agent's **agent-specific** `{NAME}_LITELLM_API_KEY` from Infisical vault (not the shared master key). Covers: LiteLLM keys, GPU ports, agent gateways (all 5 agents now SSHa ble), CT liveness (pct status on PVE nodes), config.yaml YAML integrity, wrapper/CLI integrity, vault secret non-emptiness checks. Fleet roster: tanko (.122), mumuni (.24, inside abiba CT100), koby (.129), koonimo (.114), abiba (.24). Legacy `tdunna`/`baggy` replaced with canonical agent hostnames.
- **Stale keys cleaned**: `daily-infra-report.py` SYNTHETIC_API_KEY was stale (`sk-U_ydi3B` → 401); now reads `LITELLM_MASTER_KEY` from env. Deprecated scripts (`router-original.py`, `router-phase0-backup.py`, `apply-fixes.py`) still reference `sk-syslog-local-master-key` but do not actively poll LiteLLM. - **Stale keys cleaned**: `daily-infra-report.py` SYNTHETIC_API_KEY was stale (`sk-U_ydi3B` → 401); now reads `LITELLM_MASTER_KEY` from env. Deprecated scripts (`router-original.py`, `router-phase0-backup.py`, `apply-fixes.py`) still reference `sk-syslog-local-master-key` but do not actively poll LiteLLM.
## Maintains ## Maintains
@@ -211,8 +211,8 @@ for reference but inactive. If Redis issues occur, check harness-redis container
Every remediation cycle produces a structured report: Every remediation cycle produces a structured report:
### 1. RA-H OS Knowledge Graph Node ### 1. Gitea Log Entry
Created as `[LEARN] litellm-self-heal: <run_id>` with full JSON report. Pushed to `SyslogSolution/health-logs/litellm/{run_id}.json` — versioned, searchable, not in graph.
### 2. Zulip DM to Owner ### 2. Zulip DM to Owner
- `issues_fixed > 0` — "🛠 LiteLLM Self-Heal — Fix Applied" - `issues_fixed > 0` — "🛠 LiteLLM Self-Heal — Fix Applied"
@@ -252,9 +252,11 @@ call report-generator
health: health health: health
actions: actions actions: actions
-- Phase 4: Log to knowledge graph -- Phase 4: Log to Gitea (not knowledge graph — hard rule)
call kg-logger call gitea-logger
run_id: run_id run_id: run_id
repo: SyslogSolution/health-logs
path: litellm/{run_id}.json
health: health health: health
actions: actions actions: actions
@@ -300,3 +302,20 @@ With failures:
] ]
} }
``` ```
## gitea-logger Implementation
When this step executes, write the report JSON to a temp file and push to Gitea:
```bash
REPO="https://abiba-bot:${GITEA_PAT}@git.sysloggh.net/SyslogSolution/health-logs"
DIR="litellm"
FILE="${run_id}.json"
echo "${report_json}" > /tmp/${FILE}
(cd /tmp && git clone --depth 1 "${REPO}" &&
cp ${FILE} health-logs/${DIR}/${FILE} &&
cd health-logs && git add ${DIR}/${FILE} &&
git commit -m "litellm-health: ${run_id}" && git push)
rm -rf /tmp/health-logs /tmp/${FILE}
```
+71
View File
@@ -0,0 +1,71 @@
---
kind: pattern
name: memory-fixer
description: >
Auto-fix low-hanging fruit in the graph. No judgment calls — only deterministic Level 1 operations.
Escalate anything that needs Kwame's input.
version: 1.1.0
---
# Memory Fixer
## Purpose
Auto-fix low-hanging fruit in the graph. No judgment calls — only deterministic Level 1 operations. Escalate anything that needs Kwame's input.
## Level 0 Auto-Deletes (Allowed Without Approval)
Ephemeral heartbeat and log nodes that violate "Logs NEVER go in the graph":
- `[LITELLM-HEALTH]`, `[GPU-SELF-HEAL]`, `[PM2-SELF-HEAL]`
- `[PROXMOX-MONITOR]`, `[GPU-MONITOR]`, `[INFRA-MONITOR]`, `[AGENT-HEALTH]`, `[DISK-GC]`
- `[WAL]` entries older than 30 days
**Condition:** node must be an orphan (no edges). Deleting a connected node risks breaking other nodes.
**Method:** direct SQLite on `.65` (MCP has no delete tool):
```bash
ssh root@192.168.68.65 "sqlite3 /root/.local/share/RA-H/db/rah.sqlite \"
DELETE FROM nodes WHERE id IN (
SELECT id FROM nodes WHERE id NOT IN (SELECT from_node_id FROM edges)
AND id NOT IN (SELECT to_node_id FROM edges)
AND title LIKE '[LITELLM-HEALTH]%' -- add more prefixes as needed
);\""
```
## Level 1 Auto-Fixes (No Judgment Required)
### 1. Missing `type` Field
For nodes with content but no `metadata.type`:
- Title contains "Proxmox" or "infrastructure" → `type: infrastructure`
- Title contains "skill" or "how to" or "guide" → `type: skill`
- Title contains "doc" or "template" or "brand" → `type: documentation`
- Title starts with "WAL:" or "TASK:" → `type: note`
- Title starts with "[LEARN]" → `type: documentation`
- Otherwise → `type: note` (default)
### 2. Missing `tenant` / `namespace`
For any node with NULL tenant or namespace:
```sql
UPDATE nodes
SET metadata = json_set(
COALESCE(metadata, '{}'),
'$.tenant', 'syslogsolution',
'$.namespace', 'syslogsolution'
)
WHERE json_extract(metadata, '$.tenant') IS NULL
OR json_extract(metadata, '$.namespace') IS NULL;
```
### 3. Staleness State Transitions
Using the type-based windows from the memory-monitor contract:
- Nodes stale > their window → transition to `state: review_pending`
- Nodes in `review_pending` for >7 days → escalate to Kwame (Level 2)
## Level 2 Escalations (Kwame Decision Required)
1. **Nodes in `review_pending` >7 days** — Archive, refresh, or keep?
2. **Orphan Nodes >90 days old** — Delete or Connect?
3. **Potential Duplicate Nodes** — Same title or >70% overlap. Merge or Keep?
4. **Conflicting Metadata** — Content suggests one tenant but metadata says another.
## Logging
Every Level 1 fix logged to `~/.hermes/logs/memory-fixer/YYYY-MM-DD.md`
Every Level 2 escalation logged and delivered to Kwame.
+6 -6
View File
@@ -6,7 +6,7 @@ description: >
delegation, verification, and delivery. Defines when to delegate, which delegation, verification, and delivery. Defines when to delegate, which
worker to use for what, how to handle failures, and the kanban board worker to use for what, how to handle failures, and the kanban board
protocol. Enforces context-window discipline and separation of concerns. protocol. Enforces context-window discipline and separation of concerns.
Runs on Mumuni (lxc/114, hwepve, .123) via Hermes agent. Runs on Mumuni (inside Abiba CT100, hwepve, .24) via Hermes agent (Pi + Hermes Zulip gateway).
version: 1.0.0 version: 1.0.0
--- ---
@@ -20,7 +20,7 @@ version: 1.0.0
## Topology ## Topology
**Cluster:** 6 Proxmox nodes (ocupve, acerpve, minipve, amdpve, storepve, hwepve) **Cluster:** 6 Proxmox nodes (ocupve, acerpve, minipve, amdpve, storepve, hwepve)
**Manager:** Mumuni (lxc/114, hwepve, .123) via Hermes agent **Manager:** Mumuni (inside Abiba CT100, hwepve, .24) via Hermes agent
**Workers:** 6 profiles, all running on the same agent — no separate hosts needed **Workers:** 6 profiles, all running on the same agent — no separate hosts needed
This contract is infrastructure-agnostic in terms of which nodes are used. This contract is infrastructure-agnostic in terms of which nodes are used.
@@ -31,7 +31,7 @@ Workers execute tasks on whatever infrastructure they're given — SSH to .6,
## Why This Matters ## Why This Matters
Without enforced delegation, the manager consumes the full iteration budget Without enforced delegation, the manager consumes the full iteration budget
(60 calls) on single-turn tasks — SSH to 5 nodes, check each VM, read logs — (60 calls) on single-turn tasks — SSH to 6 nodes, check each VM, read logs —
leaving no capacity for actual coordination. The result: context overflow leaving no capacity for actual coordination. The result: context overflow
(59K tokens in system prompt), iteration exhaustion, and degraded response (59K tokens in system prompt), iteration exhaustion, and degraded response
quality. This contract exists because I blew through my budget checking quality. This contract exists because I blew through my budget checking
@@ -82,7 +82,7 @@ it asks the manager (via relay) — it doesn't go find it on its own.
**This is a hard rule, not a recommendation.** Violating it produces the exact **This is a hard rule, not a recommendation.** Violating it produces the exact
type of discrepancy the kanban pipeline exists to prevent: a review worker finds type of discrepancy the kanban pipeline exists to prevent: a review worker finds
"5 nodes present" in the raw data but "5/5 online" in the report — even though "6 nodes present" in the raw data but "6/6 online" in the report — even though
one of those nodes was unreachable. The report lied because it used data the one of those nodes was unreachable. The report lied because it used data the
raw data never provided. raw data never provided.
@@ -137,7 +137,7 @@ delegate_task(
``` ```
delegate_task( delegate_task(
tasks=[ tasks=[
{"goal": "Check all 5 Proxmox nodes for VM status", "context": "SSH to each node via 192.168.68.x, run 'qm list'"}, {"goal": "Check all 6 Proxmox nodes for VM status", "context": "SSH to each node via 192.168.68.x, run 'qm list'"},
{"goal": "Check Docker container health on .7/.116/.17", "context": "SSH to each host, check container status"}, {"goal": "Check Docker container health on .7/.116/.17", "context": "SSH to each host, check container status"},
] ]
) )
@@ -191,7 +191,7 @@ Only verified results reach Kwame. Format per channel:
{ {
"lane_id": "devops-check", "lane_id": "devops-check",
"worker": "syslog-devops", "worker": "syslog-devops",
"goal": "Check all 5 Proxmox nodes", "goal": "Check all 6 Proxmox nodes",
"status": "dispatched|completed|failed", "status": "dispatched|completed|failed",
"output_file": "/tmp/node-report.md" "output_file": "/tmp/node-report.md"
} }
+1 -1
View File
@@ -52,7 +52,7 @@ description: >
- If status is "online" → pass, log restarts count - If status is "online" → pass, log restarts count
- If status is "stopped" or "errored" → **DO NOT RESTART** — alert owner immediately - If status is "stopped" or "errored" → **DO NOT RESTART** — alert owner immediately
- If restarts > 5 in last hour → alert owner with full diagnostics - If restarts > 5 in last hour → alert owner with full diagnostics
4. **Log results**Create `[LEARN]` node in knowledge graph for any actions taken 4. **Log results**Append to `SyslogSolution/health-logs/pm2/{timestamp}.md` in Gitea (not knowledge graph — hard rule)
5. **Alert** — Send Zulip DM to owner if escalation needed (do NOT run pm2 commands during alerting) 5. **Alert** — Send Zulip DM to owner if escalation needed (do NOT run pm2 commands during alerting)
6. **Wait 5 min** → repeat from step 1 6. **Wait 5 min** → repeat from step 1
+1 -1
View File
@@ -88,7 +88,7 @@ agent: abiba
| acerpve | 192.168.68.9 | PVE (hosts llm-gpu qemu/101) | | acerpve | 192.168.68.9 | PVE (hosts llm-gpu qemu/101) |
| minipve | 192.168.68.12 | PVE | | minipve | 192.168.68.12 | PVE |
| amdpve | 192.168.68.15 | PVE + Strix Halo LLM (qwen3.6-35B-udq4, strix-moe) | | amdpve | 192.168.68.15 | PVE + Strix Halo LLM (qwen3.6-35B-udq4, strix-moe) |
| hwepve | 192.168.68.4 | PVE (Huawei Matebook 16, 12C/15GB) — hosts Mumuni (lxc/114) migrated from minipve 2026-07-20 | | hwepve | 192.168.68.4 | PVE (Huawei Matebook 16, 12C/15GB) — hosts abiba (lxc/100), kagentz (lxc/105), mumuni (lxc/114). CTs 100/105 migrated from amdpve, CT 114 from minipve 2026-07-20 |
## Operations ## Operations
+300 -58
View File
@@ -1,10 +1,10 @@
#!/usr/bin/env python3 #!/usr/bin/env python3
""" """
/root/scripts/agent-health-check.py — Consolidated Agent Health Verification /root/scripts/agent-health-check.py — Consolidated Agent Health Verification v2
Single non-disruptive health check replacing 7 scattered scripts. Verifies: LiteLLM keys (agent-specific), GPU port conflicts, agent Zulip streaming,
Verifies: LiteLLM keys, GPU port conflicts, agent Zulip streaming, gateway liveness, gateway log health, CT liveness, config YAML integrity,
gateway liveness, and gateway log health. NEVER restarts anything. wrapper/CLI integrity, vault secret non-emptiness. NEVER restarts anything.
Usage: Usage:
python3 /root/scripts/agent-health-check.py # Full check python3 /root/scripts/agent-health-check.py # Full check
@@ -12,59 +12,39 @@ Usage:
python3 /root/scripts/agent-health-check.py --quiet # Only output on failure python3 /root/scripts/agent-health-check.py --quiet # Only output on failure
Cron: */10 * * * * python3 /root/scripts/agent-health-check.py --quiet Cron: */10 * * * * python3 /root/scripts/agent-health-check.py --quiet
Changelog:
v2 (2026-07-26): Added CT liveness, config validation, wrapper integrity,
vault secret emptiness check. Fixed Koby/Koonimo SSH hosts and agent key
name format ({NAME}_LITELLM_API_KEY not LITELLM_API_KEY_{NAME}).
Fleet roster: tanko (.122), mumuni (.24, inside abiba CT100), koby (.129), koonimo (.114),
abiba (.24).
""" """
import subprocess, json, sys, os, time import subprocess, json, sys, os, time
from datetime import datetime from datetime import datetime
LITELLM = "http://192.168.68.116:80" LITELLM = "http://192.168.68.116:80"
INFISICAL_PROJECT = "322fceab-39da-4854-a55a-568e76c0f13f"
INFISICAL_ENV = "prod"
def _get_agent_key(agent_name): # PVE node IPs for CT liveness checks
"""Retrieve agent key from Infisical vault.""" PVE_NODES = {
try: "hwepve": "192.168.68.4",
result = subprocess.run( "amdpve": "192.168.68.15",
["infisical", "secrets", "get", "LITELLM_API_KEY", "minipve": "192.168.68.12",
"--project=agents", "--env=production", "--plain"], "storepve": "192.168.68.6",
capture_output=True, text=True, timeout=10 "acerpve": "192.168.68.9",
) "ocupve": "192.168.68.5",
if result.returncode == 0:
return result.stdout.strip()
except Exception:
pass
# Fallback: try exporting all secrets
try:
result = subprocess.run(
["infisical", "export", "--project=agents", "--env=production",
"--format=dotenv"],
capture_output=True, text=True, timeout=10
)
if result.returncode == 0:
for line in result.stdout.splitlines():
if line.startswith(f"LITELLM_API_KEY_{agent_name.upper()}") or \
(line.startswith("LITELLM_API_KEY=") and agent_name == os.uname().nodename):
return line.split("=", 1)[1].strip().strip('"').strip("'")
except Exception:
pass
return None
# Agent keys are pulled from Infisical vault at runtime.
# The 'key' field is populated dynamically below.
AGENTS = {
"tanko": {"ct": 112, "host": "192.168.68.122", "user": "jerome"},
"mumuni": {"ct": 114, "host": "192.168.68.123", "user": "root"},
"koby": {"ct": 111, "host": None, "user": None},
"koonimo": {"ct": 113, "host": None, "user": None},
} }
# Inject keys from vault # Agent definitions: ct, host, user, pve_node, vault_key_name
for agent_name in AGENTS: AGENTS = {
key = _get_agent_key(agent_name) "tanko": {"ct": 112, "host": "192.168.68.122", "user": "jerome", "pve": "amdpve", "vault_key": "TANKO_LITELLM_API_KEY"},
if key: "abiba": {"ct": 100, "host": "192.168.68.24", "user": "root", "pve": "hwepve", "vault_key": None}, # Pi agent + Mumuni Zulip, no vault key
AGENTS[agent_name]["key"] = key "koby": {"ct": 111, "host": "192.168.68.129", "user": "root", "pve": "amdpve", "vault_key": "KOBY_LITELLM_API_KEY"},
else: "koonimo": {"ct": 113, "host": "192.168.68.114", "user": "root", "pve": "amdpve", "vault_key": "KOONIMO_LITELLM_API_KEY"},
AGENTS[agent_name]["key"] = None }
GPU_HOSTS = { GPU_HOSTS = {
"gpu-rtx3090 (.8)": {"host": "192.168.68.8", "port": 8080, "service": "llama-server"}, "gpu-rtx3090 (.8)": {"host": "192.168.68.8", "port": 8080, "service": "llama-server"},
@@ -74,6 +54,11 @@ GPU_HOSTS = {
FAIL = [] FAIL = []
INFISICAL_TOKEN = os.environ.get("INFISICAL_TOKEN")
INFISICAL_API_URL = os.environ.get("INFISICAL_API_URL", "https://vault.sysloggh.net")
# ── Helpers ──────────────────────────────────────────────────────────
def ssh(host, cmd, user="root"): def ssh(host, cmd, user="root"):
"""Execute a command on a remote host, return stdout or None.""" """Execute a command on a remote host, return stdout or None."""
try: try:
@@ -112,15 +97,81 @@ def http_json(url, headers=None, timeout=5):
except: except:
return None return None
def run_infisical(args, quiet=True):
"""Run infisical CLI with env-based auth, return stdout or None."""
env = os.environ.copy()
env["INFISICAL_API_URL"] = INFISICAL_API_URL
if INFISICAL_TOKEN:
env["INFISICAL_TOKEN"] = INFISICAL_TOKEN
try:
result = subprocess.run(
["/usr/bin/infisical"] + args,
capture_output=True, text=True, timeout=15, env=env
)
return result.stdout.strip() if result.returncode == 0 else None
except:
return None
# ── KEY LOOKUP FIX ───────────────────────────────────────────────────
def _get_agent_key(agent_name, vault_key_name):
"""Retrieve agent-specific key from Infisical vault.
Uses {NAME}_LITELLM_API_KEY format (e.g., TANKO_LITELLM_API_KEY,
KOONIMO_LITELLM_API_KEY) which matches actual vault key names.
"""
if not vault_key_name:
return None
# Primary: get the agent-specific key by name
key = run_infisical([
"secrets", "get", vault_key_name,
"--projectId=" + INFISICAL_PROJECT,
"--env=" + INFISICAL_ENV,
"--plain",
])
if key and key.startswith("sk-"):
return key
# Fallback: export all and search for the key name
try:
export = run_infisical([
"export",
"--projectId=" + INFISICAL_PROJECT,
"--env=" + INFISICAL_ENV,
"--format=dotenv",
])
if export:
for line in export.splitlines():
if line.startswith(vault_key_name + "="):
value = line.split("=", 1)[1].strip().strip('"').strip("'")
if value.startswith("sk-"):
return value
except:
pass
return None
# Inject keys from vault for each agent
for agent_name in AGENTS:
info = AGENTS[agent_name]
key = _get_agent_key(agent_name, info.get("vault_key"))
AGENTS[agent_name]["key"] = key
# ═══════════════════════════════════════════════════════════════════ # ═══════════════════════════════════════════════════════════════════
# CHECK 1: LiteLLM Key Validation # CHECK 1: LiteLLM Key Validation (agent-specific keys)
# ═══════════════════════════════════════════════════════════════════ # ═══════════════════════════════════════════════════════════════════
def check_keys(): def check_keys():
for name, agent in AGENTS.items(): for name, agent in AGENTS.items():
key = agent.get("key")
if not key:
print(f"{name}: NO KEY FOUND (vault empty or unreachable)")
FAIL.append(f"key:{name}:no-key")
continue
data = http_json(f"{LITELLM}/v1/models", data = http_json(f"{LITELLM}/v1/models",
headers={"Authorization": f"Bearer {agent['key']}"}) headers={"Authorization": f"Bearer {key}"})
if data and data.get("data"): if data and data.get("data"):
model = data["data"][0].get("id", "?") model = data["data"][0].get("id", "?")
print(f"{name}: key valid → {model}") print(f"{name}: key valid → {model}")
@@ -130,7 +181,7 @@ def check_keys():
# ═══════════════════════════════════════════════════════════════════ # ═══════════════════════════════════════════════════════════════════
# CHECK 2: GPU Port Conflict Detection # CHECK 2: GPU Port Conflict Detection (unchanged)
# ═══════════════════════════════════════════════════════════════════ # ═══════════════════════════════════════════════════════════════════
def check_gpu_ports(): def check_gpu_ports():
@@ -158,7 +209,6 @@ def check_gpu_ports():
else: else:
print(f" ⚠️ {label}: svc={svc_status}, port owned by {port_owner}") print(f" ⚠️ {label}: svc={svc_status}, port owned by {port_owner}")
else: else:
# Verify health endpoint
health = ssh(host, f"curl -s --max-time 5 http://localhost:{port}/health") health = ssh(host, f"curl -s --max-time 5 http://localhost:{port}/health")
if health and '"status":"ok"' in health: if health and '"status":"ok"' in health:
print(f"{label}: healthy (pid={port_owner})") print(f"{label}: healthy (pid={port_owner})")
@@ -171,7 +221,7 @@ def check_gpu_ports():
# ═══════════════════════════════════════════════════════════════════ # ═══════════════════════════════════════════════════════════════════
# CHECK 3: Agent Gateway Liveness + Streaming # CHECK 3: Agent Gateway Liveness + Streaming (now covers all agents)
# ═══════════════════════════════════════════════════════════════════ # ═══════════════════════════════════════════════════════════════════
def check_agents(): def check_agents():
@@ -184,8 +234,11 @@ def check_agents():
print(f"{name} (CT {ct}): cannot SSH — skip liveness check") print(f"{name} (CT {ct}): cannot SSH — skip liveness check")
continue continue
# Gateway process (exclude the infisical bash wrapper that contains the same string) # Gateway process
pid = ssh(host, "pgrep -f 'hermes_cli.main gateway run' | grep -v infisical | head -1", user=user) pid = ssh(host, "pgrep -f 'hermes_cli.main gateway run' | grep -v infisical | head -1", user=user)
if not pid:
# Try alternate binary name
pid = ssh(host, "pgrep -f 'hermes.*gateway' | grep -v infisical | grep -v bash | head -1", user=user)
if not pid: if not pid:
print(f"{name}: GATEWAY NOT RUNNING") print(f"{name}: GATEWAY NOT RUNNING")
FAIL.append(f"gateway-down:{name}") FAIL.append(f"gateway-down:{name}")
@@ -203,7 +256,7 @@ def check_agents():
else: else:
gw_state, zulip = "no-state-file", "?" gw_state, zulip = "no-state-file", "?"
# Zulip streaming: does adapter have edit_message? # Zulip streaming check
adapter_paths = [ adapter_paths = [
"~/.hermes/plugins/zulip-platform/adapter.py", "~/.hermes/plugins/zulip-platform/adapter.py",
"~/.hermes/plugins/platforms/zulip/adapter.py", "~/.hermes/plugins/platforms/zulip/adapter.py",
@@ -220,13 +273,183 @@ def check_agents():
r"journalctl --user -u hermes-gateway --since '10 min ago' -o cat --no-pager 2>/dev/null " r"journalctl --user -u hermes-gateway --since '10 min ago' -o cat --no-pager 2>/dev/null "
r"| grep -ci 'error\|traceback\|exception\|401\|403\|500' || echo 0", r"| grep -ci 'error\|traceback\|exception\|401\|403\|500' || echo 0",
user=user) user=user)
recent_errors = (recent_errors or "0").strip().split("\n")[-1] # take last line recent_errors = (recent_errors or "0").strip().split("\n")[-1]
print(f" {'' if gw_state == 'running' and zulip == 'connected' else '⚠️'} " print(f" {'' if gw_state == 'running' and zulip == 'connected' else '⚠️'} "
f"{name}: gw={gw_state} zulip={zulip} streaming={streaming} " f"{name}: gw={gw_state} zulip={zulip} streaming={streaming} "
f"errors_10m={recent_errors.strip() or '0'} pid={pid}") f"errors_10m={recent_errors.strip() or '0'} pid={pid}")
# ═══════════════════════════════════════════════════════════════════
# CHECK 4: CT Liveness (NEW)
# ═══════════════════════════════════════════════════════════════════
def check_ct_liveness():
"""Check that all agent CTs are running on their PVE nodes."""
for name, agent in AGENTS.items():
ct = agent["ct"]
pve_node = agent.get("pve")
if not pve_node:
print(f"{name} (CT {ct}): no PVE node mapped — skip")
continue
pve_ip = PVE_NODES.get(pve_node)
if not pve_ip:
print(f"{name}: unknown PVE node '{pve_node}' — skip")
continue
status = ssh(pve_ip, f"pct status {ct} 2>/dev/null", user="root")
if not status:
print(f"{name} (CT {ct} on {pve_node}): PVE UNREACHABLE")
FAIL.append(f"ct-unreachable:{name}:{pve_ip}")
elif "running" in status:
print(f"{name} (CT {ct} on {pve_node}): running")
elif "stopped" in status:
print(f"{name} (CT {ct} on {pve_node}): STOPPED")
FAIL.append(f"ct-stopped:{name}")
else:
print(f" ⚠️ {name} (CT {ct} on {pve_node}): {status.strip()}")
# ═══════════════════════════════════════════════════════════════════
# CHECK 5: Config YAML Integrity (NEW)
# ═══════════════════════════════════════════════════════════════════
def check_config_integrity():
"""Verify agent config.yaml parses as valid YAML."""
for name, agent in AGENTS.items():
host = agent.get("host")
user = agent.get("user")
if not host or not user:
print(f"{name}: cannot SSH — skip config check")
continue
# Check YAML parses
yaml_ok = ssh(host,
"python3 -c "
'"import yaml; yaml.safe_load(open(\'/root/.hermes/config.yaml\')); print(\'OK\')" '
"2>&1 || echo 'FAIL'",
user=user)
if not yaml_ok:
print(f"{name}: SSH UNREACHABLE (config check skipped)")
FAIL.append(f"config-unreachable:{name}")
elif "OK" in yaml_ok:
print(f"{name}: config.yaml valid YAML")
else:
print(f"{name}: config.yaml YAML ERROR — {yaml_ok[:120]}")
FAIL.append(f"config-yaml-error:{name}")
# ═══════════════════════════════════════════════════════════════════
# CHECK 6: Wrapper/CLI Integrity (NEW)
# ═══════════════════════════════════════════════════════════════════
def check_wrapper_integrity():
"""Verify the hermes CLI wrapper exists and can reach hermes-real."""
for name, agent in AGENTS.items():
host = agent.get("host")
user = agent.get("user")
if not host or not user:
print(f"{name}: cannot SSH — skip wrapper check")
continue
# Check wrapper exists
wrapper = ssh(host, "ls -la /root/.local/bin/hermes 2>/dev/null", user=user)
if not wrapper:
# Check alternate wrapper locations
wrapper = ssh(host, "which hermes 2>/dev/null; command -v hermes 2>/dev/null", user=user)
if not wrapper:
print(f"{name}: NO HERMES CLI WRAPPER FOUND")
FAIL.append(f"wrapper-missing:{name}")
continue
else:
print(f" ⚠️ {name}: hermes at {wrapper.strip()} (not ~/.local/bin/hermes)")
# Check wrapper has correct infisical path
infisical_path_valid = ssh(host,
"head -20 /root/.local/bin/hermes 2>/dev/null | grep -q '/usr/bin/infisical' && echo OK || echo MISS",
user=user)
if infisical_path_valid == "MISS":
# Check if infisical exists on path
inf_actual = ssh(host, "command -v infisical 2>/dev/null", user=user)
if not inf_actual:
print(f"{name}: INFISICAL NOT INSTALLED (wrapper broken)")
FAIL.append(f"wrapper-no-infisical:{name}")
else:
print(f" ⚠️ {name}: wrapper infisical path may be wrong (infisical at {inf_actual})")
FAIL.append(f"wrapper-infisical-path:{name}")
# Check hermes-real exists
hermes_real = ssh(host,
"ls -la /root/.local/bin/hermes-real 2>/dev/null || echo MISS",
user=user)
if not hermes_real or hermes_real.strip() == "MISS":
# Check venv path
hermes_real = ssh(host,
"ls -la /usr/local/lib/hermes-agent/venv/bin/hermes 2>/dev/null || echo MISS",
user=user)
if not hermes_real or hermes_real.strip() == "MISS":
print(f"{name}: hermes-real NOT FOUND (wrapper broken)")
FAIL.append(f"wrapper-no-hermes-real:{name}")
else:
print(f"{name}: hermes-real at alt path")
# Check the .env file has the key
env_has_key = ssh(host,
"grep -c 'LITELLM_API_KEY' /root/.hermes/.env 2>/dev/null || echo 0",
user=user)
if env_has_key and env_has_key.strip() not in ("", "0"):
print(f"{name}: wrapper + .env key present")
else:
print(f" ⚠️ {name}: .env may be missing LITELLM_API_KEY entry")
# ═══════════════════════════════════════════════════════════════════
# CHECK 7: Vault Secret Non-Emptiness (NEW)
# ═══════════════════════════════════════════════════════════════════
def check_vault_secrets():
"""Verify agent-specific vault secrets are non-empty and start with sk-."""
for name, agent in AGENTS.items():
vault_key_name = agent.get("vault_key")
if not vault_key_name:
continue
key = agent.get("key")
if not key:
print(f"{name}: vault secret {vault_key_name} MISSING or EMPTY")
FAIL.append(f"vault-empty:{name}:{vault_key_name}")
elif not key.startswith("sk-"):
print(f"{name}: vault secret {vault_key_name} WRONG FORMAT (starts '{key[:8]}...')")
FAIL.append(f"vault-bad-format:{name}:{vault_key_name}")
else:
print(f"{name}: vault {vault_key_name}=sk-...{key[-4:]}")
# ═══════════════════════════════════════════════════════════════════
# DEPLOY: copy updated script to /root/scripts/ on local host
# ═══════════════════════════════════════════════════════════════════
def deploy_self():
"""Copy this script to /root/scripts/agent-health-check.py if out of date."""
dest = "/root/scripts/agent-health-check.py"
try:
with open(__file__, "r") as f:
current = f.read()
if os.path.isfile(dest):
with open(dest, "r") as f:
existing = f.read()
if current == existing:
return # Already deployed
# Write new version
with open(dest, "w") as f:
f.write(current)
os.chmod(dest, 0o755)
print(f" 📦 Deployed updated script to {dest}")
except:
pass # Not fatal if deploy fails
# ═══════════════════════════════════════════════════════════════════ # ═══════════════════════════════════════════════════════════════════
# MAIN # MAIN
# ═══════════════════════════════════════════════════════════════════ # ═══════════════════════════════════════════════════════════════════
@@ -235,8 +458,12 @@ def main():
quiet = "--quiet" in sys.argv quiet = "--quiet" in sys.argv
as_json = "--json" in sys.argv as_json = "--json" in sys.argv
# Self-deploy to canonical location
if not quiet and "--no-deploy" not in sys.argv:
deploy_self()
if not quiet: if not quiet:
print(f"🏥 Agent Health Check — {datetime.now().strftime('%Y-%m-%d %H:%M UTC')}") print(f"🏥 Agent Health Check v2 {datetime.now().strftime('%Y-%m-%d %H:%M UTC')}")
print() print()
print("🔑 LiteLLM Keys:") print("🔑 LiteLLM Keys:")
@@ -249,11 +476,26 @@ def main():
print("🤖 Agent Gateways:") print("🤖 Agent Gateways:")
check_agents() check_agents()
print()
print("🖥️ CT Liveness:")
check_ct_liveness()
print()
print("📝 Config Integrity:")
check_config_integrity()
print()
print("🔌 Wrapper/CLI Integrity:")
check_wrapper_integrity()
print()
print("🔐 Vault Secrets:")
check_vault_secrets()
if FAIL: if FAIL:
print(f"\n{len(FAIL)} FAILURE(S): {' | '.join(FAIL)}") print(f"\n{len(FAIL)} FAILURE(S): {' | '.join(FAIL)}")
if quiet: if quiet:
# In quiet mode, only print failures as a single alert line
print(f"ALERT agent-health:{','.join(FAIL)}") print(f"ALERT agent-health:{','.join(FAIL)}")
elif not quiet: elif not quiet:
print("\n✅ All checks passed") print("\n✅ All checks passed")
+65
View File
@@ -0,0 +1,65 @@
#!/bin/bash
# Netbird Reverse Proxy — Add a new domain route
#
# Usage: netbird-add-domain.sh <domain> <backend_ip> [port] [protocol]
#
# Example:
# netbird-add-domain.sh dns.sysloggh.net 192.168.68.10 80
#
# This script adds a domain to the Netbird proxy by inserting records
# directly into the management server's SQLite database, then restarting
# the proxy stack.
#
# Prerequisites: SSH root access to 72.61.0.17
# sqlite3 available on VPS
#
# Requires: The domain must already have a DNS CNAME to netbird.sysloggh.net
# pointing to 72.61.0.17.
set -euo pipefail
DOMAIN="${1:?Usage: netbird-add-domain.sh <domain> <backend_ip> [port] [protocol]}"
BACKEND_IP="${2:?Usage: netbird-add-domain.sh <domain> <backend_ip> [port] [protocol]}"
PORT="${3:-80}"
PROTOCOL="${4:-http}"
VPS="root@72.61.0.17"
DB_VOLUME="/var/lib/docker/volumes/root_netbird_data/_data"
DB="$DB_VOLUME/store.db"
echo "=== Adding Netbird proxy route ==="
echo "Domain: $DOMAIN"
echo "Backend: $BACKEND_IP:$PORT ($PROTOCOL)"
echo ""
ssh "$VPS" bash << REMOTESCRIPT
set -euo pipefail
# Generate unique ID using timestamp hash (Netbird format)
ID_SUFFIX=\$(date +%s | md5sum | head -c 16)
SVC_ID="d9\${ID_SUFFIX}ptsnc73\$(date +%s | md5sum | head -c 10)"
TGT_ID=\$(sqlite3 "$DB" "SELECT COALESCE(MAX(id), 100) + 1 FROM targets;")
ACCOUNT_ID="d88av3aptsnc73clmogg"
ZONE_ID="d8adqjaptsnc73fro5g0"
echo "Service ID: \$SVC_ID"
echo "Target ID: \$TGT_ID"
# Insert service
sqlite3 "$DB" "INSERT INTO services (id, account_id, name, domain, proxy_cluster, enabled, terminated, pass_host_header, rewrite_redirects, mode, source, port_auto_assigned, private) VALUES (\"\$SVC_ID\", \"\$ACCOUNT_ID\", \"$DOMAIN\", \"$DOMAIN\", \"netbird.sysloggh.net\", 1, 0, 1, 0, \"http\", \"permanent\", 0, 0);"
echo "Service: OK"
# Insert target
sqlite3 "$DB" "INSERT INTO targets (id, account_id, service_id, host, port, protocol, target_id, target_type, enabled, skip_tls_verify, request_timeout, session_idle_timeout, agent_network, disable_access_log) VALUES (\$TGT_ID, \"\$ACCOUNT_ID\", \"\$SVC_ID\", \"$BACKEND_IP\", $PORT, \"$PROTOCOL\", \"\$ZONE_ID\", \"subnet\", 1, 0, 0, 0, 0, 0);"
echo "Target: OK"
# Verify
sqlite3 -column "$DB" "SELECT s.name, t.host, t.port, t.protocol FROM services s JOIN targets t ON s.id=t.service_id WHERE s.name=\"$DOMAIN\";"
echo ""
echo "Restarting proxy stack..."
cd /root && docker compose restart netbird-server 2>/dev/null
sleep 15
docker compose restart proxy 2>/dev/null
echo "Done. Verify with: curl -sI https://$DOMAIN"
REMOTESCRIPT
+9 -6
View File
@@ -11,31 +11,33 @@ set -euo pipefail
# ── CT ID → PVE Node mapping (maintained HERE, not in prose contracts) ── # ── CT ID → PVE Node mapping (maintained HERE, not in prose contracts) ──
declare -A CT_NODES=( declare -A CT_NODES=(
# amdpve (192.168.68.15) # amdpve (192.168.68.15)
[100]=amdpve # abiba
[105]=amdpve # kagentz
[111]=amdpve # tdunna [111]=amdpve # tdunna
[112]=amdpve # tanko [112]=amdpve # tanko
[113]=amdpve # baggy [113]=amdpve # baggy
[115]=amdpve # scottdenya [115]=amdpve # scottdenya
# minipve (192.168.68.12) # minipve (192.168.68.12)
[102]=minipve # adguard (was acerpve)
[104]=minipve # authentik [104]=minipve # authentik
[110]=minipve # gitea [110]=minipve # gitea
[114]=minipve # mumuni
[116]=minipve # syslog-api [116]=minipve # syslog-api
[119]=minipve # infisical-vault
# storepve (192.168.68.6) # storepve (192.168.68.6)
[106]=storepve # ra-h-os [106]=storepve # ra-h-os
[107]=storepve # proxmox-backup [107]=storepve # proxmox-backup
[108]=storepve # media [108]=storepve # media
[117]=storepve # zulip [117]=storepve # zulip
# acerpve (192.168.68.9) [118]=storepve # jdownloader
[102]=acerpve # adguard # acerpve (192.168.68.9) — no CTs (bare metal GPU .8)
# hwepve (192.168.68.4)
[100]=hwepve # abiba (was amdpve)
[105]=hwepve # kagentz (was amdpve)
[114]=hwepve # mumuni (was minipve)
# ocupve (192.168.68.5) — no CTs (bare metal GPU .110) # ocupve (192.168.68.5) — no CTs (bare metal GPU .110)
# #
# REMOVED CTs (migrated to bare metal, decommissioned, or VMs): # REMOVED CTs (migrated to bare metal, decommissioned, or VMs):
# 101 llm-gpu → bare metal 192.168.68.8 (RTX 3090) # 101 llm-gpu → bare metal 192.168.68.8 (RTX 3090)
# 103 ocu-llm → bare metal 192.168.68.110 (RTX 5070) # 103 ocu-llm → bare metal 192.168.68.110 (RTX 5070)
# 109 docker-vm → KVM VM 192.168.68.7 (use direct SSH) # 109 docker-vm → KVM VM 192.168.68.7 (use direct SSH)
# 118 jitsi → stopped, not in service
) )
# Each node must be root-accessible via SSH hostname # Each node must be root-accessible via SSH hostname
@@ -48,6 +50,7 @@ declare -A NODE_IPS=(
[storepve]=192.168.68.6 [storepve]=192.168.68.6
[acerpve]=192.168.68.9 [acerpve]=192.168.68.9
[ocupve]=192.168.68.5 [ocupve]=192.168.68.5
[hwepve]=192.168.68.4
) )
resolve_node() { resolve_node() {
+8 -6
View File
@@ -46,17 +46,19 @@ You are a code reviewer for OpenProse infrastructure contracts in the Syslog Sol
The infrastructure-control.prose.md contract is the canonical reference for the cluster topology: The infrastructure-control.prose.md contract is the canonical reference for the cluster topology:
**Proxmox Cluster "Tabiri" (5 nodes):** **Proxmox Cluster "Tabiri" (6 nodes):**
- amdpve (192.168.68.15): abiba, kagentz, tanko, tdunna, baggy, scottdenya - amdpve (192.168.68.15): tanko, tdunna, baggy, scottdenya
- minipve (192.168.68.12): authentik, gitea, mumuni, syslog-api, jitsi - minipve (192.168.68.12): adguard, authentik, gitea, syslog-api, infisical-vault
- storepve (192.168.68.6): docker-vm, ra-h-os, PBS, media, zulip - storepve (192.168.68.6): docker-vm, ra-h-os, PBS, media, jdownloader, zulip
- acerpve (192.168.68.9): llm-gpu, adguard - acerpve (192.168.68.9): llm-gpu
- ocupve (192.168.68.5): ocu-llm - ocupve (192.168.68.5): ocu-llm
- hwepve (192.168.68.4): abiba, kagentz, mumuni
**CT IDs (verified 2026-07-04 against PVE API):** **CT IDs (verified 2026-07-24 against PVE API):**
100:abiba 102:adguard 104:authentik 105:kagentz 106:ra-h-os 100:abiba 102:adguard 104:authentik 105:kagentz 106:ra-h-os
107:pbs 108:media 110:gitea 111:tdunna 112:tanko 107:pbs 108:media 110:gitea 111:tdunna 112:tanko
113:baggy 114:mumuni 115:scottdenya 116:syslog-api 117:zulip 113:baggy 114:mumuni 115:scottdenya 116:syslog-api 117:zulip
118:jdownloader 119:infisical-vault
**NO CT 122, CT 123, or .19 exist in the cluster.** **NO CT 122, CT 123, or .19 exist in the cluster.**
+91
View File
@@ -0,0 +1,91 @@
#!/bin/bash
# swap-gpu-dense-model.sh — Swap RTX 3090 from qwen3.6-27B-code to SmartCode-Fable-5
# Run when download completes: ssh root@192.168.68.8 'bash -s' < this script
#
# Usage: bash swap-gpu-dense-model.sh
# Requires: new model at /home/llmuser/models/SmartCode-Fable-5-27B-UD-Q4_K_XL.gguf
set -e
MODEL_PATH="/home/llmuser/models/SmartCode-Fable-5-27B-UD-Q4_K_XL.gguf"
OLD_WRAPPER="/home/llmuser/llama-wrapper.sh"
echo "═══ Swapping gpu-dense to SmartCode-Fable-5 ═══"
# 1. Verify model file
if [ ! -f "$MODEL_PATH" ]; then
echo "❌ Model not found at $MODEL_PATH"
echo " Download: curl -L -o $MODEL_PATH <huggingface-url>"
exit 1
fi
MODEL_SIZE=$(ls -lh "$MODEL_PATH" | awk '{print $5}')
echo "✅ Model found: $MODEL_SIZE"
# 2. Create new wrapper script for SmartCode-Fable-5
cat > /home/llmuser/llama-fable-wrapper.sh << 'WRAPPER'
#!/bin/bash
# SmartCode-Fable-5 llama-server wrapper for RTX 3090
# Sampler settings from model card: temp 0.9, top-p 0.95, top-k 60, repeat-penalty off
PORT=8080
GHOST_PID=$(ss -tlnp 2>/dev/null | grep -Po ":${PORT}\s+.*pid=\K[0-9]+" | head -1)
if [ -n "$GHOST_PID" ] && [ "$GHOST_PID" != "$$" ]; then
echo "[wrapper] Port $PORT occupied by ghost pid $GHOST_PID — cleaning up" >&2
kill -9 "$GHOST_PID" 2>/dev/null
sleep 2
fi
exec /usr/local/bin/llama-server \
--model /home/llmuser/models/SmartCode-Fable-5-27B-UD-Q4_K_XL.gguf \
--ctx-size 131072 \
--cache-type-k q4_0 \
--cache-type-v q4_0 \
--flash-attn 1 \
--cont-batching \
--parallel 1 \
--batch-size 2048 \
--ubatch-size 1024 \
--n-gpu-layers 99 \
--temp 0.9 \
--top-p 0.95 \
--top-k 60 \
--min-p 0.0 \
--repeat-penalty 1.0 \
--api-key not-needed \
--port 8080 \
--host 0.0.0.0
WRAPPER
chmod 755 /home/llmuser/llama-fable-wrapper.sh
echo "✅ Created /home/llmuser/llama-fable-wrapper.sh"
# 3. Update systemd service to use new wrapper
echo "📝 Updating systemd service..."
sed -i 's|ExecStart=/home/llmuser/llama-wrapper.sh|ExecStart=/home/llmuser/llama-fable-wrapper.sh|' /etc/systemd/system/llama-server.service
systemctl daemon-reload
# 4. Stop old server, start new
echo "🔄 Restarting llama-server..."
systemctl stop llama-server
sleep 3
systemctl start llama-server
sleep 8
# 5. Verify
echo ""
echo "═══ Verification ═══"
systemctl is-active llama-server
echo ""
echo "Port 8080:"
ss -tlnp 2>/dev/null | grep ":8080" | head -1
echo ""
echo "GPU VRAM:"
nvidia-smi --query-gpu=memory.used,memory.total,utilization.gpu --format=csv,noheader 2>/dev/null
echo ""
echo "=== Health check ==="
curl -s --max-time 5 http://localhost:8080/health 2>/dev/null
echo ""
echo ""
echo "✅ Swap complete. Test via LiteLLM:"
echo " curl -s http://192.168.68.116/v1/chat/completions -H 'Authorization: Bearer <key>' -H 'Content-Type: application/json' -d '{\"model\":\"gpu-dense\",\"messages\":[{\"role\":\"user\",\"content\":\"write hello world in python\"}],\"max_tokens\":100}'"
+3 -3
View File
@@ -16,7 +16,7 @@ Runs every 15 minutes in the background. Also triggers on session start.
## Requires ## Requires
- **Zulip API key** for `abiba-bot@chat.sysloggh.net` in `$ZULIP_API_KEY` - **Zulip API key** for `abiba-bot@chat.sysloggh.net` in `$ZULIP_API_KEY`
- **SSH access** to Tanko (192.168.68.122), Mumuni (192.168.68.123, lxc/114 on hwepve since 2026-07-20), and Agent Zero Docker host (192.168.68.14) - **SSH access** to Tanko (192.168.68.122), Mumuni (192.168.68.24, inside Abiba CT100 on hwepve), and Agent Zero Docker host (192.168.68.14)
- **PM2** on localhost for pi process management - **PM2** on localhost for pi process management
- **Network access** to `chat.sysloggh.net`, `localhost:9200` - **Network access** to `chat.sysloggh.net`, `localhost:9200`
- **Write access** to `/root/zulip-health-monitor.log` and `/tmp/zulip-monitor-debounce` - **Write access** to `/root/zulip-health-monitor.log` and `/tmp/zulip-monitor-debounce`
@@ -182,13 +182,13 @@ grep -a "Finalized\|Failed to finalize" /root/.pm2/logs/abiba-zulip-out.log | ta
| `last_error` set | Log and monitor | | `last_error` set | Log and monitor |
| Crash loop >10/h | Alert user | | Crash loop >10/h | Alert user |
### Step 3: Platform B — Hermes (Tanko .122, Mumuni .123) ### Step 3: Platform B — Hermes (Tanko .122, Mumuni .24)
**B1: Gateway State** **B1: Gateway State**
```bash ```bash
ssh root@192.168.68.122 "cat ~/.hermes/gateway_state.json" ssh root@192.168.68.122 "cat ~/.hermes/gateway_state.json"
ssh root@192.168.68.123 "cat ~/.hermes/gateway_state.json" ssh root@192.168.68.24 "cat ~/.hermes/gateway_state.json" # Mumuni inside Abiba CT100
``` ```
Check `platforms.zulip.state`: `connected` ✅ | `disconnected` ❌ | `error` ❌ | missing → not installed. Check `platforms.zulip.state`: `connected` ✅ | `disconnected` ❌ | `error` ❌ | missing → not installed.
+1 -1
View File
@@ -65,7 +65,7 @@ triggers:
|------|----|------|---------| |------|----|------|---------|
| Zulip server | 192.168.68.19 | root | Docker: `zulip-zulip-1` | | Zulip server | 192.168.68.19 | root | Docker: `zulip-zulip-1` |
| Abiba (pi) | localhost | root | PM2: `abiba-zulip` | | Abiba (pi) | localhost | root | PM2: `abiba-zulip` |
| Mumuni | 192.168.68.123 | root | `hermes gateway restart` | | Mumuni | 192.168.68.24 (CT100 abiba) | root | `hermes gateway restart` |
| Tanko | 192.168.68.122 | jerome | `PATH=$PATH:/home/jerome/.hermes/hermes-agent hermes gateway restart` | | Tanko | 192.168.68.122 | jerome | `PATH=$PATH:/home/jerome/.hermes/hermes-agent hermes gateway restart` |
## Debounce ## Debounce