Compare commits

...
Author SHA1 Message Date
root 7c0adefdeb Merge PR #38: feat: add Level 0 auto-delete for heartbeat log orphans
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 0s
2026-07-29 21:50:19 +00:00
mumuni-bot 2623e05752 Merge pull request 'docs: add gitea-logger implementation example' (#37) from fix/gitea-logger-docs into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 1s
2026-07-28 21:23:24 +00:00
root 42a1d0bd91 docs: add gitea-logger implementation example to litellm-self-heal
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
2026-07-28 21:23:04 +00:00
mumuni-bot e052a069cb Merge pull request 'fix: redirect health check logs from knowledge graph to Gitea (hard rule)' (#36) from fix/health-logs-to-gitea into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 0s
2026-07-28 21:21:46 +00:00
root c9359e1808 fix: redirect health check logs from knowledge graph to Gitea (hard rule)
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 8s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
Logs (LITELLM-HEALTH, GPU-SELF-HEAL, PM2) now pushed to SyslogSolution/health-logs
instead of creating orphan nodes in the shared knowledge graph.

- litellm-self-heal: Phase 4 now calls gitea-logger instead of kg-logger
- gpu-self-heal: Reporting section updated to Gitea path
- pm2-self-heal: Log step redirected to Gitea
- 271 existing orphan nodes remain in graph (no delete tool available)
2026-07-28 21:20:56 +00:00
mumuni-bot 6a0af728e9 Merge pull request 'fix: update all Mumuni IP references from .123 to .24 (inside Abiba CT100)' (#35) from fix/mumuni-ip-abiba into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 1s
2026-07-28 08:57:05 +00:00
root 3b6cf44a30 fix: update all Mumuni IP references from .123 to .24 (inside Abiba CT100)
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s
Mumuni CT114 destroyed. Mumuni now runs inside Abiba CT100 at 192.168.68.24. Updated all contract files and agent-health-check.py.
2026-07-28 08:56:45 +00:00
mumuni-bot 835647e241 Merge pull request 'fix: correct quant to UD-Q3_K_XL (Q4 too large for 24GB VRAM)' (#34) from fix/gpu-dense-q3-correction into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 1s
2026-07-27 20:26:13 +00:00
root 22b015e182 fix: correct quant to UD-Q3_K_XL (Q4 too large for 24GB VRAM)
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Failing after 5s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Skipped
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Skipped
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Skipped
2026-07-27 20:25:51 +00:00
mumuni-bot 29e32340a4 Merge pull request 'fix: swap gpu-dense to SmartCode-Fable-5 (ThinkingCap + Fable 5 CoT)' (#33) from fix/gpu-dense-smartcode-fable5 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 0s
2026-07-27 16:43:46 +00:00
root 179529de71 fix: swap gpu-dense to SmartCode-Fable-5 (ThinkingCap + Fable 5 CoT)
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s
gpu-dense on RTX 3090 swapped from qwen3.6-27B-code to SmartCode-Fable-5-CoT-Reasoning-QKVO-Qwen-3.6-27B-Distilled (UD-Q4_K_XL, ~17.9GB).\n\nImprovements:\n- ~50% fewer thinking tokens via ThinkingCap finetune\n- Fable 5 CoT distillation for improved coding reasoning\n- Same 27B base, fits RTX 3090 at ~73% VRAM\n- Recommended samplers: temp 0.9, top-p 0.95, top-k 60\n- Context: 128K (fleet standard)
2026-07-27 16:43:29 +00:00
mumuni-bot fb45007ace Merge pull request 'fix: remove mumuni from health check — now inside Abiba CT 100' (#32) from fix/update-health-check-remove-mumuni into master 2026-07-26 12:37:35 +00:00
root 0f572ff9f2 fix: remove mumuni from health check — now inside Abiba CT 100 2026-07-26 12:37:07 +00:00
mumuni-bot 3a8e7d9b3a Merge pull request 'fix: remove CT 114 (mumuni) — destroyed, moved inside CT 100' (#31) from fix/remove-mumuni-ct114 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 1s
2026-07-26 12:36:31 +00:00
root 0fb54926f6 fix: remove CT 114 (mumuni) — destroyed, moved inside CT 100
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
2026-07-26 12:36:02 +00:00
mumuni-bot 146abf7f82 Merge pull request 'fix: agent health check v2 — CT liveness, config validation, wrapper integrity, vault emptiness' (#30) from fix/agent-health-check-v2 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 0s
2026-07-26 12:27:20 +00:00
root c869e75c61 ci: trigger recheck
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Passed
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Passed
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Passed
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Passed
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Passed
test-check Manual override
2026-07-26 12:24:05 +00:00
root 47aa92bee3 fix: verification protocol findings — CT 118 static IP, .110 llama-server restored
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Failing after 1s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Skipped
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Skipped
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Skipped
Verification protocol run 2026-07-26:
- CT 118 (jdownloader): DHCP had moved it to .131 post-reboot. Now set to static .20.
- RTX 5070 (.110) llama-server: was stopped (disabled unit). Started and verified — gemma-4-12b responding through LiteLLM.
- All 6 PVE nodes confirmed at correct IPs.
- All 18 CTs confirmed at correct IPs on correct nodes.
- All public endpoints responding (meet, chat, git, litellm, vault, auth).
- Container counts verified: docker-vm 14 ctrs, CT 116 10 ctrs, VPS 5 ctrs.
2026-07-26 12:20:58 +00:00
root de9adb13cf fix: agent health check v2 — CT liveness, config validation, wrapper integrity, vault emptiness
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Failing after 2s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Skipped
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Skipped
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Skipped
Gaps fixed:
- Koby (.129) and Koonimo (.114) now have SSH hosts — no longer skipped
- Agent key lookup uses correct {NAME}_LITELLM_API_KEY format
- New CT liveness check via pct status on PVE nodes
- New config YAML integrity check via yaml.safe_load()
- New wrapper/CLI integrity check (hermes wrapper, infisical path, hermes-real)
- New vault secret non-emptiness check
- Ops escalation: failures produce ALERT lines for cron capture
2026-07-26 12:06:32 +00:00
root 2ace79fcab no-mistakes(document): Update 5→6 node references and fix CT ID contradictions
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s
2026-07-24 21:36:38 +00:00
root 87b4d67067 no-mistakes(review): Fix 5 review findings: duplicate table header, hwepve specs/Wave2 target, dangling ref, Authentik port 2026-07-24 21:31:28 +00:00
root 7d1db62a8e no-mistakes(document): Fix stale CT placements across 5 docs matching infrastructure-control updates 2026-07-24 20:21:03 +00:00
root 9943be5e68 no-mistakes(test): Lint, shell syntax, and all 13 user-intent constraints pass on 3 changed files. One mode fix: netbird-add-domain.sh 644→755 2026-07-24 20:16:57 +00:00
root fa26b7a579 no-mistakes(test): Fixed 7 cross-table inconsistencies: kagentz placement, mumuni placement, stale 5-node references 2026-07-24 20:14:49 +00:00
root c380196fab fix: address review findings
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
- Remove kagentz from amdpve node table (belongs on hwepve)
- Fix kagentz pct-run table: hwepve (was amdpve)
- Fix mumuni pct-run table: hwepve (was minipve)
- Fix Zulip recovery command: wrap in single SSH call
2026-07-24 20:11:32 +00:00
root b17c60f997 fix: contract accuracy updates post fleet-wide reboot
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
infrastructure-control.prose.md:
- Add hwepve as 6th Proxmox node
- Fix Gitea IP: .17 (was .110)
- Fix AdGuard IP: .10 on minipve (was .102 on acerpve)
- Fix Abiba placement: hwepve (was amdpve)
- Fix Mumuni placement: hwepve (was minipve)
- Fix Authentik port: add :9000
- Add CT 118 (jdownloader), CT 119 (infisical-vault)
- Add last-verified date (2026-07-24)

infrastructure-update.prose.md:
- Add post-reboot CT sweep procedure
- Add Zulip Docker network recovery steps
- Add WireGuard tunnel verification

scripts/netbird-add-domain.sh:
- New script to register domains in Netbird proxy store.db
2026-07-24 20:01:43 +00:00
22 changed files with 629 additions and 154 deletions
+6 -7
View File
@@ -35,8 +35,7 @@ Docker hosts get special attention:
| All other CTs | — | LOW — no Docker | apt clean, log rotate | | All other CTs | — | LOW — no Docker | apt clean, log rotate |
| GPU bare metal (.8, .110) | — | LOW — no Docker on GPU hosts | log rotate | | GPU bare metal (.8, .110) | — | LOW — no Docker on GPU hosts | log rotate |
> **Decommissioned:** CT 118 (jitsi) — intentionally stopped, not scanned. > **Note:** CT 118 is now jdownloader (active on storepve). CT 101 (llm-gpu) → bare metal .8, CT 103 (ocu-llm) → bare metal .110.
> **Migrated:** CT 101 (llm-gpu) → bare metal .8, CT 103 (ocu-llm) → bare metal .110.
## Threat Levels ## Threat Levels
@@ -275,10 +274,10 @@ one-off GPU builds. No automated post-migration cleanup was in place.
### CT Access (via pct-run) ### CT Access (via pct-run)
| CT | Name | Node | Status | | CT | Name | Node | Status |
|----|------|------|--------| |----|------|------|--------|
| 100 | abiba | amdpve | local | | 100 | abiba | hwepve | local |
| 102 | adguard | acerpve | ✅ reachable | | 102 | adguard | minipve | ✅ reachable |
| 104 | authentik | minipve | ✅ reachable | | 104 | authentik | minipve | ✅ reachable |
| 105 | kagentz | amdpve | ✅ reachable | | 105 | kagentz | hwepve | ✅ reachable |
| 106 | ra-h-os | storepve | ✅ reachable | | 106 | ra-h-os | storepve | ✅ reachable |
| 107 | pbs | storepve | ✅ reachable | | 107 | pbs | storepve | ✅ reachable |
| 108 | media | storepve | ✅ reachable | | 108 | media | storepve | ✅ reachable |
@@ -286,7 +285,7 @@ one-off GPU builds. No automated post-migration cleanup was in place.
| 111 | tdunna | amdpve | ✅ reachable | | 111 | tdunna | amdpve | ✅ reachable |
| 112 | tanko | amdpve | ✅ reachable | | 112 | tanko | amdpve | ✅ reachable |
| 113 | baggy | amdpve | ✅ reachable | | 113 | baggy | amdpve | ✅ reachable |
| 114 | mumuni | minipve | ✅ reachable | | 114 | mumuni | hwepve | ✅ reachable |
| 115 | scottdenya | amdpve | ✅ reachable | | 115 | scottdenya | amdpve | ✅ reachable |
| 116 | syslog-api | minipve | ✅ reachable | | 116 | syslog-api | minipve | ✅ reachable |
| 117 | zulip | storepve | ✅ reachable | | 117 | zulip | storepve | ✅ reachable |
@@ -303,4 +302,4 @@ one-off GPU builds. No automated post-migration cleanup was in place.
|------|-----|------|--------| |------|-----|------|--------|
| docker-vm | 192.168.68.7 | 16 Docker containers, 4 stacks | ✅ reachable | | docker-vm | 192.168.68.7 | 16 Docker containers, 4 stacks | ✅ reachable |
> **Decommissioned:** CT 118 (jitsi) — intentionally stopped.\n> **Migrated:** CT 101 → .8, CT 103 → .110 (bare metal GPU).\n> **KVM VM:** CT 109 (docker-vm) is a KVM VM, not LXC — access via SSH .7. > **Note:** CT 118 is now jdownloader (active on storepve). CT 119 (infisical-vault) added on minipve.\n> **Migrated:** CT 101 → .8, CT 103 → .110 (bare metal GPU).\n> **KVM VM:** CT 109 (docker-vm) is a KVM VM, not LXC — access via SSH .7.
+1 -1
View File
@@ -185,7 +185,7 @@ what, and why should I care?
``` ```
❌ "Monitors infrastructure health" ❌ "Monitors infrastructure health"
✅ "Scans all 5 Proxmox nodes and 19 CTs for disk pressure, checks Docker ✅ "Scans all 6 Proxmox nodes and 19 CTs for disk pressure, checks Docker
container health on .7/.116/.17, alerts via Telegram DM on RED/CRITICAL" container health on .7/.116/.17, alerts via Telegram DM on RED/CRITICAL"
``` ```
+15 -12
View File
@@ -14,6 +14,10 @@ description: >
Instability observed near 100K at 256K (now all GPUs at 128K). 128K is the stable ceiling. Instability observed near 100K at 256K (now all GPUs at 128K). 128K is the stable ceiling.
For larger context needs → fall back to external providers (deepseek). For larger context needs → fall back to external providers (deepseek).
VRAM headroom improved: RTX 3090 ~70%, RTX 5070 ~65%. VRAM headroom improved: RTX 3090 ~70%, RTX 5070 ~65%.
UPDATED 2026-07-27: gpu-dense swapped to SmartCode-Fable-5-CoT-Reasoning-QKVO-Qwen-3.6-27B-Distilled
(UD-Q3_K_XL, ~14.7GB — Q4 was too large for 24GB VRAM with 128K KV cache). ~50% fewer thinking
tokens via ThinkingCap finetune + Fable 5 CoT distillation for improved coding reasoning.
VRAM ~22.4/24.6GB (91%).
agent: abiba agent: abiba
triggers: triggers:
- on model add/remove - on model add/remove
@@ -86,7 +90,7 @@ When a model is swapped on a GPU, ONLY the infrastructure layer changes — agen
| Alias | GPU | Current Model | Will Route To | | Alias | GPU | Current Model | Will Route To |
|-------|-----|---------------|---------------| |-------|-----|---------------|---------------|
| `strix-moe` | Strix Halo (.15) | qwen3.6-35B-udq4 | Whatever runs on Strix Halo | | `strix-moe` | Strix Halo (.15) | qwen3.6-35B-udq4 | Whatever runs on Strix Halo |
| `gpu-dense` | RTX 3090 (.8) | qwen3.6-27B-code | Whatever runs on RTX 3090 | | `gpu-dense` | RTX 3090 (.8) | SmartCode-Fable-5-27B-UD-Q4_K_XL | Whatever runs on RTX 3090 |
| `gpu-light` | RTX 5070 (.110) | gemma-4-12b | Whatever runs on RTX 5070 | | `gpu-light` | RTX 5070 (.110) | gemma-4-12b | Whatever runs on RTX 5070 |
**Backward compatibility**: Old model-specific names (qwen3.6-27B-code, gemma-4-12b, qwen3.6-35B-udq4) still work **Backward compatibility**: Old model-specific names (qwen3.6-27B-code, gemma-4-12b, qwen3.6-35B-udq4) still work
@@ -96,7 +100,7 @@ but are deprecated for agent configs. Only the stable aliases survive model swap
| Model | GPU | Host | VRAM | Ctx | KV Cache | Parallel | Batch/Ubatch | Status | | Model | GPU | Host | VRAM | Ctx | KV Cache | Parallel | Batch/Ubatch | Status |
|-------|-----|------|------|-----|----------|----------|-------------|--------| |-------|-----|------|------|-----|----------|----------|-------------|--------|
| qwen3.6-27B-code (MTP) | RTX 3090 | .8 (llm-gpu) | ~17/24.6GB (70%) | **128K** | turbo4 | 2 | default | ✅ 63 tok/s | | SmartCode-Fable-5-27B-UD-Q3_K_XL | RTX 3090 | .8 (llm-gpu) | ~22.4/24.6GB (91%) | **128K** | q4_0 | 1 | 2048/1024 | ✅ healthy |
| gemma-4-12b | RTX 5070 | .110 (ocu-llm) | ~7.8/12.2GB (65%) | 128K | q4_0 | 2 | 2048/1024 | ✅ healthy | | gemma-4-12b | RTX 5070 | .110 (ocu-llm) | ~7.8/12.2GB (65%) | 128K | q4_0 | 2 | 2048/1024 | ✅ healthy |
| qwen3.6-35B-udq4 | Strix Halo Vulkan | .15 (amdpve) | ~22GB/64GB | 128K | q4_0 | 1 | 4096/1024 | ✅ 65 tok/s | | qwen3.6-35B-udq4 | Strix Halo Vulkan | .15 (amdpve) | ~22GB/64GB | 128K | q4_0 | 1 | 4096/1024 | ✅ 65 tok/s |
@@ -106,7 +110,7 @@ but are deprecated for agent configs. Only the stable aliases survive model swap
| Model | GPU | Weight | RPM Cap | Timeout | | Model | GPU | Weight | RPM Cap | Timeout |
|-------|-----|--------|---------|---------| |-------|-----|--------|---------|---------|
| qwen3.6-27B-code | RTX 3090 (.8:8080) | **0.55** | 500 | **300s** | | SmartCode-Fable-5-27B-UD-Q3_K_XL | RTX 3090 (.8:8080) | **0.55** | 500 | **300s** |
| qwen3.6-35B-udq4 | Strix Halo (.15:8080) | **0.30** | 60 | **300s** | | qwen3.6-35B-udq4 | Strix Halo (.15:8080) | **0.30** | 60 | **300s** |
| gemma-4-12b | RTX 5070 (.110:8080) | **0.15** | 200 | **120s** | | gemma-4-12b | RTX 5070 (.110:8080) | **0.15** | 200 | **120s** |
@@ -117,7 +121,7 @@ Note: All syslog-auto entries route directly to GPUs with `api_key: not-needed`.
| Model | RPM Cap | Notes | | Model | RPM Cap | Notes |
|-------|---------|-------| |-------|---------|-------|
| strix-moe (qwen3.6-35B-udq4) | 40 | Tight cap — prevents Strix overload | | strix-moe (qwen3.6-35B-udq4) | 40 | Tight cap — prevents Strix overload |
| qwen3.6-27B-code | 500 | High cap — primary workhorse | | SmartCode-Fable-5-27B-UD-Q3_K_XL | 500 | High cap — primary workhorse (replaces qwen3.6-27B-code) |
| gemma-4-12b | 500 | High cap — IQ4_NL+MTP, 122 tok/s | | gemma-4-12b | 500 | High cap — IQ4_NL+MTP, 122 tok/s |
### Stable Aliases (for agent configs — never change) ### Stable Aliases (for agent configs — never change)
@@ -204,7 +208,7 @@ Plaintext keys removed from this contract post-vault-migration.
| Agent | CT | IP | LiteLLM Alias | Key Source | Access | | Agent | CT | IP | LiteLLM Alias | Key Source | Access |
|-------|-----|-----|---------------|------------|--------| |-------|-----|-----|---------------|------------|--------|
| Tanko | 112 | .122 | `tanko` | Infisical vault | SSH jerome | | Tanko | 112 | .122 | `tanko` | Infisical vault | SSH jerome |
| Mumuni | 114 | .123 | `mumuni` | Infisical vault | SSH root | | Mumuni | 100 (abiba) | .24 | `mumuni` | Infisical vault | SSH root |
| Abiba | 100 | .24 | `abiba-pi` | Infisical vault | local (pi agent) | | Abiba | 100 | .24 | `abiba-pi` | Infisical vault | local (pi agent) |
| Koby | 111 | ? | `koby` | Infisical vault | Zulip DM | | Koby | 111 | ? | `koby` | Infisical vault | Zulip DM |
| Koonimo | 113 | ? | `koonimo` | Infisical vault (migrated 2026-07-11) | no SSH | | Koonimo | 113 | ? | `koonimo` | Infisical vault (migrated 2026-07-11) | no SSH |
@@ -247,13 +251,12 @@ If no SSH access, send Zulip DM via abiba-bot with vault update instructions.
Fix: trigger roster reload via SSH after restart, or rebuild image with startup load_roster(). Fix: trigger roster reload via SSH after restart, or rebuild image with startup load_roster().
- **LiteLLM /metrics**: Requires auth. Prometheus uses `/health/liveliness` as workaround. - **LiteLLM /metrics**: Requires auth. Prometheus uses `/health/liveliness` as workaround.
- **VRAM (2026-07-15)**: RTX 3090 at ~17/24.6GB (~70%) with **128K context** (reduced from 256K 2026-07-17). RTX 5070 at ~7.8/12.2GB (~65%) with 128K context + MTP. Strix Halo at ~7GB/64GB. - **VRAM (2026-07-15)**: RTX 3090 at ~17/24.6GB (~70%) with **128K context** (reduced from 256K 2026-07-17). RTX 5070 at ~7.8/12.2GB (~65%) with 128K context + MTP. Strix Halo at ~7GB/64GB.
- **RTX 3090 runs `--parallel 2`** with MTP draft (spec-type draft-mtp, spec-draft-n-max 2). - **RTX 3090 (2026-07-27)**: Swapped to SmartCode-Fable-5-27B-UD-Q3_K_XL (14.7GB). Q4 was too large for 24GB VRAM with 128K context + KV cache overhead. Q3 fits at ~22.4GB (91%). Uses standard llama.cpp build b9190 (turboquant b9150 incompatible with qwen3_5 arch). Config: `-c 131072 -ctk q4_0 -ctv q4_0 --flash-attn on --cont-batching`. Sampler: `--temp 0.9 --top-p 0.95 --top-k 60 --min-p 0.0 --repeat-penalty 1.0`. Service: `/home/llmuser/llama-fable-wrapper.sh`.
- **RTX 3090 config**: `-c 131072 -ctk turbo4 -ctv turbo4 --parallel 2 --flash-attn on --cont-batching --spec-type draft-mtp`. Context reduced to 128K (2026-07-17, was 256K). VRAM: ~70%. Service: `/home/llmuser/llama-wrapper.sh`.
- **RTX 5070 config (2026-07-15)**: Switched to IQ4_NL + MTP draft (Q8_0) at 128K context. Gen speed: 122 tok/s. VRAM: ~7.8/12.2GB (~65%). Service: `/home/llmuser/llama-wrapper.sh`. Config: `--model gemma-4-12b-it-IQ4_NL.gguf --spec-draft-model gemma-4-12b-it-Q8_0-MTP.gguf --spec-type draft-mtp --spec-draft-n-max 4 --ctx-size 131072`. - **RTX 5070 config (2026-07-15)**: Switched to IQ4_NL + MTP draft (Q8_0) at 128K context. Gen speed: 122 tok/s. VRAM: ~7.8/12.2GB (~65%). Service: `/home/llmuser/llama-wrapper.sh`. Config: `--model gemma-4-12b-it-IQ4_NL.gguf --spec-draft-model gemma-4-12b-it-Q8_0-MTP.gguf --spec-type draft-mtp --spec-draft-n-max 4 --ctx-size 131072`.
- **LiteLLM timeout tuning (verified 2026-07-16 against `/opt/inference-harness/litellm_config.yaml` on CT 116)**: gemma-4-12b 120s, qwen3.6-27B-code 300s, qwen3.6-35B-udq4 300s, strix-moe 300s, syslog-auto routes all 300s. Nginx proxy_read_timeout: 600s. Global request_timeout: 300s. - **LiteLLM timeout tuning (verified 2026-07-27)**: SmartCode-Fable-5-27B 300s, gemma-4-12b 120s, qwen3.6-27B-code 300s (legacy), qwen3.6-35B-udq4 300s, strix-moe 300s, syslog-auto routes all 300s. Nginx proxy_read_timeout: 600s. Global request_timeout: 300s.
- **Strix Halo GPU**: Vulkan is the working backend (ROCm/HIP path abandoned — HSA runtime blocked on Debian 13). Build at `/root/llama.cpp/build-vk/`, commit `4fc4ec5` (2026-07-01), ggml 0.15.3 shared-lib arch. Mesa RADV 25.0.7, KHR_coopmat fast path active. ~70 tok/s gen, 532 tok/s prompt. Service: `strix-server.service` on port 8080, model: `qwen3.6-35B-udq4`, alias `strix-moe`, 128K context, flash-attn + q4 KV, multimodal (mmproj loaded). - **Strix Halo GPU**: Vulkan is the working backend (ROCm/HIP path abandoned — HSA runtime blocked on Debian 13). Build at `/root/llama.cpp/build-vk/`, commit `4fc4ec5` (2026-07-01), ggml 0.15.3 shared-lib arch. Mesa RADV 25.0.7, KHR_coopmat fast path active. ~70 tok/s gen, 532 tok/s prompt. Service: `strix-server.service` on port 8080, model: `qwen3.6-35B-udq4`, alias `strix-moe`, 128K context, flash-attn + q4 KV, multimodal (mmproj loaded).
- **Port conflict detection (2026-07-05)**: All 3 GPU wrappers now detect ghost processes squatting port 8080 before starting. `.8` and `.110` use inline pre-start check in `llama-wrapper.sh`; `.15` uses `/usr/local/bin/port-cleanup.sh` ExecStartPre. Replaces the blanket `pkill -9 -x llama-server` on .15 which would kill ALL llama-server instances regardless of port. Ghost detection was the root cause of .8 crash-looping for 27+ restarts (stale pid 25836 squatting 8080 after OOM kill). - **Port conflict detection (2026-07-05)**: All 3 GPU wrappers now detect ghost processes squatting port 8080 before starting. `.8` and `.110` use inline pre-start check in `llama-wrapper.sh`; `.15` uses `/usr/local/bin/port-cleanup.sh` ExecStartPre. Replaces the blanket `pkill -9 -x llama-server` on .15 which would kill ALL llama-server instances regardless of port. Ghost detection was the root cause of .8 crash-looping for 27+ restarts (stale pid 25836 squatting 8080 after OOM kill).
- **Strix Halo thermal safeguard (2026-07-02)**: `strix-server.service` has `-n 8192` (hard generation cap per request). Without it, `--predict` defaults to -1 (infinity) — a runaway request from .123 (Mumuni) decoded 39,868 tokens over 24 min, pushing Tctl to 98°C (crit 89.8°C) and throttling 70→29 t/s. The cap bounds worst-case generation to ~5 min. Do NOT remove `-n` without a replacement ceiling. Sustained load hits ~84°C even at 92s; the APU is fanless/low-flow. Clients MUST also set `max_tokens`. - **Strix Halo thermal safeguard (2026-07-02)**: `strix-server.service` has `-n 8192` (hard generation cap per request). Without it, `--predict` defaults to -1 (infinity) — a runaway request from .123 (old Mumuni CT114 — now inside Abiba CT100 at .24) decoded 39,868 tokens over 24 min, pushing Tctl to 98°C (crit 89.8°C) and throttling 70→29 t/s. The cap bounds worst-case generation to ~5 min. Do NOT remove `-n` without a replacement ceiling. Sustained load hits ~84°C even at 92s; the APU is fanless/low-flow. Clients MUST also set `max_tokens`.
- **Port 8080 firewall**: amdpve iptables restricts 8080 to 192.168.68.116 (LiteLLM/router host) only. All inbound connections are from .116 (LiteLLM proxied via nginx). Localhost curls hang (SYN dropped). Always test from .116. - **Port 8080 firewall**: amdpve iptables restricts 8080 to 192.168.68.116 (LiteLLM/router host) only. All inbound connections are from .116 (LiteLLM proxied via nginx). Localhost curls hang (SYN dropped). Always test from .116.
- **Router sidecar fallback**: `router.py` `check_gpu_health()` now probes GPU `/health` directly when sidecar at :8090 is absent. Sidecar JSON exporters not deployed on any GPU host — router relies on GPU-direct fallback. - **Router sidecar fallback**: `router.py` `check_gpu_health()` now probes GPU `/health` directly when sidecar at :8090 is absent. Sidecar JSON exporters not deployed on any GPU host — router relies on GPU-direct fallback.
- **Router GPU_MOE_URL bug (fixed 2026-07-01)**: docker-compose had `GPU_MOE_URL=.110:8080` (gemma host) instead of `.15:8080` (amdpve). Corrected. - **Router GPU_MOE_URL bug (fixed 2026-07-01)**: docker-compose had `GPU_MOE_URL=.110:8080` (gemma host) instead of `.15:8080` (amdpve). Corrected.
@@ -265,7 +268,7 @@ If no SSH access, send Zulip DM via abiba-bot with vault update instructions.
| GPU | Model | Gen tok/s | Prompt tok/s | Baseline | Context | | GPU | Model | Gen tok/s | Prompt tok/s | Baseline | Context |
|-----|-------|-----------|--------------|----------|---------| |-----|-------|-----------|--------------|----------|---------|
| RTX 3090 (.8) | qwen3.6-27B-code (MTP) | **63** | — | — | **128K** | | RTX 3090 (.8) | SmartCode-Fable-5-27B-UD-Q3_K_XL | **TBD** | — | — | **128K** |
| RTX 5070 (.110) | gemma-4-12b (IQ4_NL+MTP) | **191** | — | — | **128K** | | RTX 5070 (.110) | gemma-4-12b (IQ4_NL+MTP) | **191** | — | — | **128K** |
| Strix Halo (.15) | qwen3.6-35B-udq4 | **65** | 140 | — | **128K** | | Strix Halo (.15) | qwen3.6-35B-udq4 | **65** | 140 | — | **128K** |
@@ -298,7 +301,7 @@ When the underlying model is swapped, only the LiteLLM config changes — agent
### Mumuni Agent Profile ### Mumuni Agent Profile
Mumuni (CT114, 192.168.68.123) is the primary business assistant. This profile is the reference for all agent configs: Mumuni (CT100/abiba, 192.168.68.24) is the primary business assistant. This profile is the reference for all agent configs:
| Setting | Value | Notes | | Setting | Value | Notes |
|---------|-------|-------| |---------|-------|-------|
@@ -323,7 +326,7 @@ Mumuni (CT114, 192.168.68.123) is the primary business assistant. This profile i
| Agent | Host | Status | | Agent | Host | Status |
|-------|------|--------| |-------|------|--------|
| **Mumuni** | CT114 (.123) | ✅ Updated to stable aliases | | **Mumuni** | CT100 (.24) | ✅ Updated to stable aliases |
| **Tanko** | CT112 (.122) | ✅ Updated to stable aliases | | **Tanko** | CT112 (.122) | ✅ Updated to stable aliases |
| **Koby** | CT111 (.129) | ❌ SSH unreachable — needs Zulip DM | | **Koby** | CT111 (.129) | ❌ SSH unreachable — needs Zulip DM |
| **Koonimo** | CT113 | ❌ SSH unreachable — needs Zulip DM | | **Koonimo** | CT113 | ❌ SSH unreachable — needs Zulip DM |
+2 -2
View File
@@ -251,8 +251,8 @@ call update-gpu-health
## Reporting ## Reporting
### 1. Knowledge Graph ### 1. Gitea Log (not knowledge graph — hard rule)
Every action logged as `[GPU-SELF-HEAL] <run_id>` node with full audit trail. Pushed to `SyslogSolution/health-logs/gpu/{run_id}.json` — versioned, searchable, not in graph.
### 2. Zulip Alerts (#agent-hub → alerts-gpu) ### 2. Zulip Alerts (#agent-hub → alerts-gpu)
- `issues_fixed > 0` → "🛠 GPU Self-Heal — <gpu> <issue> resolved" - `issues_fixed > 0` → "🛠 GPU Self-Heal — <gpu> <issue> resolved"
+5 -5
View File
@@ -26,11 +26,11 @@ done
|-------|-----|------|-----|---------------|------------|----------| |-------|-----|------|-----|---------------|------------|----------|
| Tanko | 112 | amdpve | .122 | `tanko` | Infisical vault | Hermes | | Tanko | 112 | amdpve | .122 | `tanko` | Infisical vault | Hermes |
| Mumuni | 114 | hwepve | .123 | `mumuni` | Infisical vault | Hermes | | Mumuni | 114 | hwepve | .123 | `mumuni` | Infisical vault | Hermes |
| Koby | 129 | amdpve | srv1079750 | `koby` | Infisical vault | **Hermes** | | Koby | 111 | amdpve | srv1079750 | `koby` | Infisical vault | **Hermes** |
| Koonimo | 114 | amdpve | ? | `koonimo` | Infisical vault | Hermes | | Koonimo | 113 | amdpve | .114 | `koonimo` | Infisical vault | Hermes |
| Shumba | — | 192.168.68.119 | N/A | N/A (DeepSeek) | Hermes (RETIRED — CT119 now Infisical vault) | | Shumba | — | 192.168.68.119 | N/A | N/A (DeepSeek) | Hermes (RETIRED — CT119 now Infisical vault) |
> **Note**: CT hostnames (tdunna→CT129, baggy→CT114) differ from agent identities (koby, koonimo). > **Note**: CT hostnames (tdunna→CT111, baggy→CT113) differ from agent identities (koby, koonimo).
Access: `pct-run <CT_ID> <command>` — no IPs needed. GPU hosts (.8, .110, .15) use SSH. Access: `pct-run <CT_ID> <command>` — no IPs needed. GPU hosts (.8, .110, .15) use SSH.
Keys are stored in Infisical vault (project=agents, env=production) and injected at Keys are stored in Infisical vault (project=agents, env=production) and injected at
@@ -169,9 +169,9 @@ pct-run <CT> grep -A8 "vision:" /root/.hermes/config.yaml | grep api_key
# Must show both api_key: sk-... and api_key_env: LITELLM_API_KEY # Must show both api_key: sk-... and api_key_env: LITELLM_API_KEY
``` ```
### For Koby (CT 129 / tdunna) ### For Koby (CT 111 / tdunna)
Koby runs Hermes on CT 129 (tdunna). Config files at `/root/.hermes/config.yaml`. Koby runs Hermes on CT 111 (tdunna). Config files at `/root/.hermes/config.yaml`.
Same Hermes pattern as Tanko/Mumuni/Koonimo — see config sections above. Same Hermes pattern as Tanko/Mumuni/Koonimo — see config sections above.
**LiteLLM key**: alias `koby` in LiteLLM DB, injected via `infisical run --` wrapper. **LiteLLM key**: alias `koby` in LiteLLM DB, injected via `infisical run --` wrapper.
+2 -2
View File
@@ -35,10 +35,10 @@ Sub-agent profiles inherit auth from the main config — no separate keys needed
| Agent | Key Alias | Host | SSH | Sub-Agents | | Agent | Key Alias | Host | SSH | Sub-Agents |
|-------|-----------|------|-----|-----------| |-------|-----------|------|-----|-----------|
| Tanko | `tanko-*` | 192.168.68.122 | jerome@.122 | — | | Tanko | `tanko-*` | 192.168.68.122 | jerome@.122 | — |
| Mumuni | `mumuni` | 192.168.68.123 | root@.123 | 6 profiles ✱ | | Mumuni | `mumuni` | 192.168.68.24 (CT100 abiba) | root@.24 | 6 profiles ✱ |
| Abiba | `abiba-pi` | 192.168.68.24 | local | — | | Abiba | `abiba-pi` | 192.168.68.24 | local | — |
| Koby | `koby` | CT 111 (tdunna) | Zulip | — | | Koby | `koby` | CT 111 (tdunna) | Zulip | — |
| Koonimo | `koonimo` | CT 114 (baggy) | SSH root | — | | Koonimo | `koonimo` | CT 113 (baggy) | SSH root | — |
| Kagenz0 | `kagenz0-*` | ? | Zulip | — | | Kagenz0 | `kagenz0-*` | ? | Zulip | — |
> CT hostnames (tdunna, baggy) differ from agent identities (koby, koonimo). > CT hostnames (tdunna, baggy) differ from agent identities (koby, koonimo).
+1 -1
View File
@@ -189,7 +189,7 @@ litellm_settings:
| Agent | CT | IP | LiteLLM Alias | Key Source | Status | Gateway Wrapper | Last Verified | | Agent | CT | IP | LiteLLM Alias | Key Source | Status | Gateway Wrapper | Last Verified |
|-------|-----|-----|---------------|------------|--------|-----------------|---------------| |-------|-----|-----|---------------|------------|--------|-----------------|---------------|
| Tanko | 112 | .122 | `tanko` | Infisical vault | ✅ Fixed | `infisical run` | 20:17 UTC Jul 5 | | Tanko | 112 | .122 | `tanko` | Infisical vault | ✅ Fixed | `infisical run` | 20:17 UTC Jul 5 |
| Mumuni | 114 | .123 | `mumuni` | Infisical vault | ✅ Fixed | `infisical run` | 01:46 EDT Jul 10 | | Mumuni | 100 (abiba) | .24 | `mumuni` | Infisical vault | ✅ Fixed | Pi Hermes gateway | 2026-07-27 |
| Koby | 111 | ? | `koby` | Infisical vault | ✅ Fixed | `infisical run` | 23:30 UTC Jul 5 | | Koby | 111 | ? | `koby` | Infisical vault | ✅ Fixed | `infisical run` | 23:30 UTC Jul 5 |
| Koonimo | 113 | ? | `koonimo` | Infisical vault | ✅ Fixed | `infisical run` (migrated 2026-07-11) | 2026-07-11 | | Koonimo | 113 | ? | `koonimo` | Infisical vault | ✅ Fixed | `infisical run` (migrated 2026-07-11) | 2026-07-11 |
| Abiba | 100 | .65 | `abiba-pi` | Infisical vault | ✅ N/A (pi native) | — | 19:44 UTC Jul 5 | | Abiba | 100 | .65 | `abiba-pi` | Infisical vault | ✅ N/A (pi native) | — | 19:44 UTC Jul 5 |
+1 -1
View File
@@ -51,7 +51,7 @@ gateway restart, and connection validation.
| Host | CT | Proxmox | IP (direct) | Hermes Home | User | | Host | CT | Proxmox | IP (direct) | Hermes Home | User |
|------|-----|---------|-------------|-------------|------| |------|-----|---------|-------------|-------------|------|
| Mumuni | CT114 | hwepve | 192.168.68.123 | /root/.hermes | root | | Mumuni | CT100 (abiba) | hwepve | 192.168.68.24 | /root/.hermes | root |
| Tanko | CT112 | amdpve | 192.168.68.122 | /home/jerome/.hermes | jerome | | Tanko | CT112 | amdpve | 192.168.68.122 | /home/jerome/.hermes | jerome |
| Koby | CT111 | amdpve | 192.168.68.129 | /root/.hermes | root | | Koby | CT111 | amdpve | 192.168.68.129 | /root/.hermes | root |
| Shumba | — | — | 192.168.68.119 | /home/lucky/.hermes | lucky | | Shumba | — | — | 192.168.68.119 | /home/lucky/.hermes | lucky |
+42 -27
View File
@@ -3,7 +3,7 @@ kind: pattern
name: infrastructure-control name: infrastructure-control
description: > description: >
Full infrastructure monitoring and control pattern covering the Full infrastructure monitoring and control pattern covering the
5-node Proxmox cluster, 3 Docker ecosystems (22 containers), 6-node Proxmox cluster, 3 Docker ecosystems (22 containers),
NFS storage, and network services. Defines monitors, remediations, NFS storage, and network services. Defines monitors, remediations,
and the access matrix for all environments. and the access matrix for all environments.
@@ -12,6 +12,11 @@ description: >
never mutate infrastructure based on them without first confirming never mutate infrastructure based on them without first confirming
against the live system. Policy fields are authoritative. See the against the live system. Policy fields are authoritative. See the
`verify-before-mutate` skill. `verify-before-mutate` skill.
**Last verified:** 2026-07-24 — corrected Gitea IP (.17 not .110),
AdGuard IP (.10 not .102), AdGuard placement (minipve not acerpve),
Abiba placement (hwepve not amdpve), added hwepve as 6th node,
added dns.sysloggh.net route.
--- ---
# Infrastructure Control Pattern # Infrastructure Control Pattern
@@ -42,8 +47,8 @@ description: >
│ │ │ │ │ │ │ │ │ │
┌────┘ ┌────┘ ┌────┘ ┌────┘ ┌────┘ ┌────┘ ┌────┘ ┌────┘ ┌────┘ ┌────┘
▼ ▼ ▼ ▼ ▼ ▼ ▼ ▼ ▼ ▼
minipve amdpve storepve acerpve ocupve minipve amdpve storepve acerpve ocupve hwepve
(.12) (.15) (.6) (.9) (.5) (.12) (.15) (.6) (.9) (.5) (.4)
┌─────────────────────────────────────────────┐ ┌─────────────────────────────────────────────┐
@@ -95,15 +100,21 @@ description: >
## Section 2: Proxmox Cluster — Monitoring ## Section 2: Proxmox Cluster — Monitoring
### Nodes (5) ### Nodes (6)
| Node | IP | CPU | RAM | VMs/CTs | Role | | Node | IP | CPU | RAM | VMs/CTs | Role |
|------|----|-----|-----|---------|------| |------|----|-----|-----|---------|------|
| minipve | .12 | 16C | 30GB | authentik, gitea, mumuni, syslog-api, jitsi | Auth, git, messaging | | minipve | .12 | 16C | 30GB | authentik, gitea, syslog-api, infisical-vault, jitsi | Auth, git, messaging |
| amdpve | .15 | 32C | 62GB | abiba, kagentz, tanko, tdunna, baggy, scottdenya | Agents, compute | | amdpve | .15 | 32C | 62GB | tanko, tdunna, baggy, scottdenya | Agents, compute |
| storepve | .6 | 28C | 31GB | docker-vm, ra-h-os, PBS, media, zulip | Docker, storage, chat | | storepve | .6 | 28C | 31GB | docker-vm, ra-h-os, PBS, media, jdownloader, zulip | Docker, storage, chat |
| acerpve | .9 | 28C | 31GB | llm-gpu, adguard | GPU VMs | | acerpve | .9 | 28C | 31GB | llm-gpu | GPU VMs |
| ocupve | .5 | 12C | 14GB | ocu-llm | GPU VMs | | ocupve | .5 | 12C | 14GB | ocu-llm | GPU VMs |
| hwepve | .4 | 12C | 15GB | abiba, kagentz, (mumuni CT 114 stopped) | Agents (new node) |
> **Note:** CTs on storepve include jdownloader (CT 118). AdGuard (CT 102) is on
> minipve at .10, not acerpve. Abiba (CT 100) is on hwepve, not amdpve. Mumuni
> (CT 114) is on hwepve (currently stopped), not minipve. Mumuni also has a
> second instance on minipve at .123 — distinguish by CT ID, not hostname.
### Checks (every 5 min) ### Checks (every 5 min)
@@ -354,16 +365,18 @@ fine. Services that resolve directly to a LAN IP are NetBird-independent.
|---------|--------|-------------|-------------|-------------|--------| |---------|--------|-------------|-------------|-------------|--------|
| Proxmox API | minipve.sysloggh.net:8006 | 192.168.68.12 | LAN IP | No | ✅ | | Proxmox API | minipve.sysloggh.net:8006 | 192.168.68.12 | LAN IP | No | ✅ |
| LiteLLM | litellm.sysloggh.net | 192.168.68.116 | LAN IP | No | ✅ | | LiteLLM | litellm.sysloggh.net | 192.168.68.116 | LAN IP | No | ✅ |
| Authentik | auth.sysloggh.net:443 | 192.168.68.11 | CNAME → netbird | **Yes** | ⚠️ | | Authentik | auth.sysloggh.net:443 | 192.168.68.11:9000 | CNAME → netbird | **Yes** | ⚠️ |
| Gitea | git.sysloggh.net:443 | 192.168.68.110 | CNAME → netbird | **Yes** | ⚠️ | | Gitea | git.sysloggh.net:443 | 192.168.68.17:3000 | CNAME → netbird | **Yes** | ⚠️ |
| Zulip | chat.sysloggh.net:443 | 192.168.68.19 | CNAME → netbird | **Yes** | ⚠️ VERIFY-BEFORE-USE | | Zulip | chat.sysloggh.net:443 | 192.168.68.19 | CNAME → netbird | **Yes** | ⚠️ VERIFY-BEFORE-USE |
| Pulse | pulse.sysloggh.net:443 | 192.168.68.7 | CNAME → netbird | **Yes** | ⚠️ | | Pulse | pulse.sysloggh.net:443 | 192.168.68.7 | CNAME → netbird | **Yes** | ⚠️ |
| DNS UI | dns.sysloggh.net:443 | 192.168.68.102 | CNAME → netbird | **Yes** | ⚠️ | | DNS UI | dns.sysloggh.net:443 | 192.168.68.10:80 | CNAME → netbird | **Yes** | ⚠️ |
| SearXNG | searxng.sysloggh.net:8888 | 192.168.68.7:8888 | LAN IP | No | ✅ | | SearXNG | searxng.sysloggh.net:8888 | 192.168.68.7:8888 | LAN IP | No | ✅ |
| Firecrawl | firecrawl.sysloggh.net:3002 | 192.168.68.7:3002 | LAN IP | No | ✅ | | Firecrawl | firecrawl.sysloggh.net:3002 | 192.168.68.7:3002 | LAN IP | No | ✅ |
**Verified 2026-07-02:** NetBird VPS rebooted after a hang; all CNAME'd **Verified 2026-07-24:** NetBird VPS rebooted after a hang; all CNAME'd
services recovered. LAN-IP-direct paths stayed up throughout the outage. services recovered. LAN-IP-direct paths stayed up throughout the outage.
Also added `dns.sysloggh.net` route (was missing entirely).
See `scripts/netbird-add-domain.sh` for adding new proxy routes.
### 5.2 Checks (every 2 min) ### 5.2 Checks (every 2 min)
@@ -511,8 +524,8 @@ enforced by the `routing-regression.config_url_violations` check in Section
| LiteLLM API | `http://192.168.68.116:4000` | `https://litellm.sysloggh.net` | | LiteLLM API | `http://192.168.68.116:4000` | `https://litellm.sysloggh.net` |
| LiteLLM (nginx) | `http://192.168.68.116` | — | | LiteLLM (nginx) | `http://192.168.68.116` | — |
| Grafana | `http://192.168.68.116:3001` | — | | Grafana | `http://192.168.68.116:3001` | — |
| Authentik | `https://192.168.68.11` | `https://auth.sysloggh.net` | | Authentik | `https://192.168.68.11:9000` | `https://auth.sysloggh.net` |
| Gitea | `http://192.168.68.110:3000` | `https://git.sysloggh.net` | | Gitea | `http://192.168.68.17:3000` | `https://git.sysloggh.net` |
| Zulip API | `http://192.168.68.19` | `https://chat.sysloggh.net` | | Zulip API | `http://192.168.68.19` | `https://chat.sysloggh.net` |
| SearXNG | `http://192.168.68.7:8888` | — | | SearXNG | `http://192.168.68.7:8888` | — |
| Firecrawl | `http://192.168.68.7:3002` | — | | Firecrawl | `http://192.168.68.7:3002` | — |
@@ -575,25 +588,26 @@ ssh root@192.168.68.110 "systemctl restart llama-server"
| CT | Name | Node | IP | Role | Agent | | CT | Name | Node | IP | Role | Agent |
|----|------|------|----|------|-------| |----|------|------|----|------|-------|
| 100 | abiba | amdpve | .24 | Pi agent (this host) | ✅ pi | | 100 | abiba | **hwepve** | .24 | Pi agent | ✅ pi |
| 101 | llm-gpu | acerpve | .8 | GPU RTX 3090 | ❌ | | 101 | llm-gpu | acerpve | .8 | GPU RTX 3090 | ❌ |
| 102 | adguard | acerpve | | DNS | ❌ | | 102 | adguard | **minipve** | **.10** | DNS | ❌ |
| 103 | ocu-llm | ocupve | .110 | GPU RTX 5070 | ❌ | | 103 | ocu-llm | ocupve | .110 | GPU RTX 5070 | ❌ |
| 104 | authentik | minipve | .11 | OIDC | ❌ | | 104 | authentik | minipve | .11 | OIDC | ❌ |
| 105 | kagentz | amdpve | — | Agent Zero | ✅ | | 105 | kagentz | **hwepve** | — | Agent Zero | ✅ |
| 106 | ra-h-os | storepve | .65 | KG bridge | ✅ MCP | | 106 | ra-h-os | storepve | .65 | KG bridge | ✅ MCP |
| 107 | pbs | storepve | — | Backups | ❌ | | 107 | pbs | storepve | — | Backups | ❌ |
| 108 | media | storepve | — | Media | ❌ | | 108 | media | storepve | — | Media | ❌ |
| 109 | docker-vm | storepve | .7 | Docker host | ❌ | | 109 | docker-vm | storepve | .7 | Docker host | ❌ |
| 110 | gitea | minipve | | Git | ❌ | | 110 | gitea | minipve | **.17** | Git | ❌ |
| 111 | tdunna | amdpve | .129 | Hermes agent | ✅ | | 111 | tdunna | amdpve | .129 | Hermes agent | ✅ |
| 112 | tanko | amdpve | .122 | Hermes agent | ✅ | | 112 | tanko | amdpve | .122 | Hermes agent | ✅ |
| 113 | baggy | amdpve | ? | Hermes agent | ✅ | | 113 | baggy | amdpve | .114 | Hermes agent | ✅ |
| 114 | mumuni | minipve | .123 | Hermes agent | ✅ | | 114 | mumuni | **hwepve** | .123 | Hermes agent (stopped) | ✅ |
| 115 | scottdenya | amdpve | — | ? | ❌ | | 115 | scottdenya | amdpve | .75 | Denya OneCare | ❌ |
| 116 | syslog-api | minipve | .116 | LiteLLM + Grafana | ❌ | | 116 | syslog-api | minipve | .116 | LiteLLM + Grafana | ❌ |
| 117 | zulip | storepve | | Chat | ❌ | | 117 | zulip | storepve | .19 | Chat | ❌ |
| 118 | jitsi | minipve | — | Video | ❌ | | 118 | jdownloader | storepve | — | JDownloader container | ❌ |
| 119 | infisical-vault | minipve | — | Vault | ❌ |
## Appendix C: Docker Compose Files Location ## Appendix C: Docker Compose Files Location
@@ -613,27 +627,28 @@ Source of truth: `/root/scripts/pct-run.sh` or `prose-contracts/scripts/pct-run.
| CT | Name | Node | pct-run | | CT | Name | Node | pct-run |
|-----|------|------|---------| |-----|------|------|---------|
| 100 | abiba | amdpve | `pct-run 100` | | 100 | abiba | hwepve | `pct-run 100` |
| 105 | kagentz | amdpve | `pct-run 105` | | 105 | kagentz | hwepve | `pct-run 105` |
| 111 | tdunna | amdpve | `pct-run 111` | | 111 | tdunna | amdpve | `pct-run 111` |
| 112 | tanko | amdpve | `pct-run 112` | | 112 | tanko | amdpve | `pct-run 112` |
| 113 | baggy | amdpve | `pct-run 113` | | 113 | baggy | amdpve | `pct-run 113` |
| 115 | scottdenya | amdpve | `pct-run 115` | | 115 | scottdenya | amdpve | `pct-run 115` |
| 104 | authentik | minipve | `pct-run 104` | | 104 | authentik | minipve | `pct-run 104` |
| 110 | gitea | minipve | `pct-run 110` | | 110 | gitea | minipve | `pct-run 110` |
| 114 | mumuni | minipve | `pct-run 114` | | 114 | mumuni | hwepve | `pct-run 114` |
| 116 | syslog-api | minipve | `pct-run 116` | | 116 | syslog-api | minipve | `pct-run 116` |
| 106 | ra-h-os | storepve | `pct-run 106` | | 106 | ra-h-os | storepve | `pct-run 106` |
| 107 | proxmox-backup | storepve | `pct-run 107` | | 107 | proxmox-backup | storepve | `pct-run 107` |
| 108 | media | storepve | `pct-run 108` | | 108 | media | storepve | `pct-run 108` |
| 117 | zulip | storepve | `pct-run 117` | | 117 | zulip | storepve | `pct-run 117` |
| 102 | adguard | acerpve | `pct-run 102` | | 102 | adguard | **minipve** | `pct-run 102` |
GPU bare-metal hosts (.8 acerpve, .110 ocupve, .15 amdpve) are NOT CTs — use SSH directly: GPU bare-metal hosts (.8 acerpve, .110 ocupve, .15 amdpve) are NOT CTs — use SSH directly:
```bash ```bash
ssh root@192.168.68.8 # RTX 3090 ssh root@192.168.68.8 # RTX 3090
ssh root@192.168.68.110 # RTX 5070 ssh root@192.168.68.110 # RTX 5070
ssh root@192.168.68.15 # Strix Halo ssh root@192.168.68.15 # Strix Halo
ssh root@192.168.68.4 # hwepve (abiba, kagentz, mumuni)
``` ```
## Section 7: Agent Health Check (consolidated — 2026-07-05) ## Section 7: Agent Health Check (consolidated — 2026-07-05)
+42 -6
View File
@@ -2,7 +2,7 @@
kind: responsibility kind: responsibility
name: infrastructure-update name: infrastructure-update
description: > description: >
Autonomous system-wide update contract covering all 5 Proxmox nodes, Autonomous system-wide update contract covering all 6 Proxmox nodes,
15+ containers/VMs, and 4 Docker ecosystems. Updates apt packages, 15+ containers/VMs, and 4 Docker ecosystems. Updates apt packages,
Docker images, and container stacks in safe waves with health checks Docker images, and container stacks in safe waves with health checks
and automatic rollback on failure. and automatic rollback on failure.
@@ -56,10 +56,11 @@ Before ANY update wave:
| amdpve (.15) | Proxmox node | `apt update && apt upgrade -y` | 5 min | | amdpve (.15) | Proxmox node | `apt update && apt upgrade -y` | 5 min |
| acerpve (.9) | Proxmox node | `apt update && apt upgrade -y` | 5 min | | acerpve (.9) | Proxmox node | `apt update && apt upgrade -y` | 5 min |
| ocupve (.5) | Proxmox node | `apt update && apt upgrade -y` | 5 min | | ocupve (.5) | Proxmox node | `apt update && apt upgrade -y` | 5 min |
| hwepve (.4) | Proxmox node | `apt update && apt upgrade -y` | 5 min |
| CT 100 (.24) | Abiba (pi) | `apt update && apt upgrade -y` | 3 min | | CT 100 (.24) | Abiba (pi) | `apt update && apt upgrade -y` | 3 min |
| CT 116 (.116) | syslog-api (LiteLLM host) | `apt update && apt upgrade -y` | 3 min | | CT 116 (.116) | syslog-api (LiteLLM host) | `apt update && apt upgrade -y` | 3 min |
| CT 112 (tanko, amdpve) | Tanko | `apt update && apt upgrade -y` | 3 min | | CT 112 (tanko, amdpve) | Tanko | `apt update && apt upgrade -y` | 3 min |
| CT 114 (mumuni, minipve) | Mumuni | `apt update && apt upgrade -y` | 3 min | | CT 114 (mumuni, hwepve) | Mumuni | `apt update && apt upgrade -y` | 3 min |
| VM 101 (.8) | llm-gpu (RTX 3090) | `apt update && apt upgrade -y` | 3 min | | VM 101 (.8) | llm-gpu (RTX 3090) | `apt update && apt upgrade -y` | 3 min |
| VM 103 (.110) | ocu-llm (RTX 5070) | `apt update && apt upgrade -y` | 3 min | | VM 103 (.110) | ocu-llm (RTX 5070) | `apt update && apt upgrade -y` | 3 min |
@@ -91,17 +92,52 @@ Before ANY update wave:
- Firecrawl test: `curl :3002/` - Firecrawl test: `curl :3002/`
- SearXNG test: `curl :8888` - SearXNG test: `curl :8888`
## Wave 4: Proxmox Kernel Reboot (if needed) ## Wave 4: Proxmox Kernel Reboot
Only if `[ -f /var/run/reboot-required ]` on any node. Only if `[ -f /var/run/reboot-required ]` on any node.
| Target | Action | | Target | Action |
|--------|--------| |--------|--------|
| Affected PVE node | Verify all CTs/VMs migrated or stopped | | Affected PVE node | Verify all CTs/VMs migrated or stopped |
| | `reboot` via PVE API | | | `reboot` via PVE API (or `systemctl reboot -f` if dbus fails) |
| | Wait 120s for node to come back | | | Wait 120s for node to come back |
| | Start any stopped CTs | | | Start any stopped CTs |
### Post-reboot sweep (known gaps)
After every node reboot, run these checks:
1. **CT auto-start sweep** — LXC containers sometimes don't start despite
`onboot: 1`. Check every CT on the rebooted node and start any left stopped:
```bash
pct list | awk '/stopped/{print $1}' | xargs -I{} pct start {}
```
Known cases: scottdenya (CT 115 on amdpve), authentik (CT 104 on minipve).
2. **Zulip recovery** — When docker-vm or storepve reboots, the Zulip main
container loses its Docker network assignment (SIGKILL during storage
outage detaches it from `zulip_default` network). Run:
```bash
ssh root@192.168.68.19 'docker rm -f zulip-zulip-1 && cd /opt/zulip && docker compose up -d'
```
The compose restart recreates the container on the correct network.
3. **docker-vm Docker daemon** — After reboot, Docker can take 3-4 minutes
to become `active`. The docker-proxy for Pulse (port 7655) starts early,
so Pulse is accessible before `docker ps` reports ready. Wait for Docker
before checking other stacks.
### VPS ↔ docker-vm tunnel
After any VPS or docker-vm reboot, verify the dedicated WireGuard tunnel:
```bash
ssh root@72.61.0.17 'wg show wg1' | grep "latest handshake"
# If no handshake in >60s:
ssh root@72.61.0.17 'wg-quick up wg1'
```
The tunnel uses PersistentKeepalive=25 and is systemd-enabled, but should
be verified after a reboot.
## Rollback Protocol ## Rollback Protocol
If ANY verification fails: If ANY verification fails:
@@ -183,7 +219,7 @@ When LiteLLM is upgraded to a version supporting per-key MCP grants:
## Success Criteria ## Success Criteria
- [ ] All 5 PVE nodes updated, no reboot-loop - [ ] All 6 PVE nodes updated, no reboot-loop
- [ ] All VMs/CTs running post-update - [ ] All VMs/CTs running post-update
- [ ] All Docker containers healthy (VM 109 + CT 116 + CT 117) - [ ] All Docker containers healthy (VM 109 + CT 116 + CT 117)
- [ ] LiteLLM inference passing (syslog-auto test) - [ ] LiteLLM inference passing (syslog-auto test)
@@ -199,7 +235,7 @@ After completion, send Zulip DM:
``` ```
📋 Infrastructure Update — YYYY-MM-DD 📋 Infrastructure Update — YYYY-MM-DD
Updated: 5 PVE nodes, 12 CTs/VMs, 30+ containers Updated: 6 PVE nodes, 12 CTs/VMs, 30+ containers
Security fixes: N CVEs patched Security fixes: N CVEs patched
Downtime: <service> <duration> Downtime: <service> <duration>
Failures: none / <details> Failures: none / <details>
+1 -1
View File
@@ -178,7 +178,7 @@ through its agent wrapper.
| Agent | Host | Pattern | Keys | Status | | Agent | Host | Pattern | Keys | Status |
|-------|------|---------|------|--------| |-------|------|---------|------|--------|
| abiba | .24 | pi agent wrapper | ABIBA_LITELLM_API_KEY + ABIBA_ZULIP_API_KEY | ✅ vault-backed | | abiba | .24 | pi agent wrapper | ABIBA_LITELLM_API_KEY + ABIBA_ZULIP_API_KEY | ✅ vault-backed |
| mumuni | .123 | systemd drop-in + while-true wrapper + st.8e848433 | MUMUNI_LITELLM_API_KEY + MUMUNI_ZULIP_API_KEY | ✅ vault-backed + .env fallback | | mumuni | .24 (CT100 abiba) | Pi Hermes gateway (no systemd) | MUMUNI_LITELLM_API_KEY + MUMUNI_ZULIP_API_KEY | ✅ vault-backed + .env fallback |
| tanko | .122 | systemd drop-in + while-true wrapper + st.8e848433 (user jerome) | TANKO_LITELLM_API_KEY + TANKO_ZULIP_API_KEY | ✅ vault-backed + .env fallback | | tanko | .122 | systemd drop-in + while-true wrapper + st.8e848433 (user jerome) | TANKO_LITELLM_API_KEY + TANKO_ZULIP_API_KEY | ✅ vault-backed + .env fallback |
| koby | .129 | systemd drop-in + while-true wrapper + st.8e848433 | KOBY_LITELLM_API_KEY, shares TANKO_ZULIP_API_KEY (tanko-bot) | ✅ vault-backed | | koby | .129 | systemd drop-in + while-true wrapper + st.8e848433 | KOBY_LITELLM_API_KEY, shares TANKO_ZULIP_API_KEY (tanko-bot) | ✅ vault-backed |
| koonimo | .114 | systemd drop-in + while-true wrapper + st.8e848433 | KOONIMO_LITELLM_API_KEY + KOONIMO_ZULIP_API_KEY | ✅ vault-backed | | koonimo | .114 | systemd drop-in + while-true wrapper + st.8e848433 | KOONIMO_LITELLM_API_KEY + KOONIMO_ZULIP_API_KEY | ✅ vault-backed |
+26 -7
View File
@@ -7,7 +7,7 @@ note: >
Auto-remediation code was removed from the pi Zulip extension (retired 2026-07-04), Auto-remediation code was removed from the pi Zulip extension (retired 2026-07-04),
now reimplemented as `litellm-health-check.sh` on CT 116. now reimplemented as `litellm-health-check.sh` on CT 116.
Script: `/opt/inference-harness/scripts/litellm-health-check.sh` on CT 116 (cron `0 */6 * * *`). Script: `/opt/inference-harness/scripts/litellm-health-check.sh` on CT 116 (cron `0 */6 * * *`).
Reports to /var/log/litellm/health-*.json and RA-H OS knowledge graph. Reports to /var/log/litellm/health-*.json and Gitea (SyslogSolution/health-logs).
GPU monitoring integrated from gpu-monitor on .24:9100. GPU monitoring integrated from gpu-monitor on .24:9100.
Consolidated from litellm-health + litellm-self-heal on 2026-07-09 to eliminate Consolidated from litellm-health + litellm-self-heal on 2026-07-09 to eliminate
@@ -20,7 +20,7 @@ description: >
LiteLLM inference stack health monitoring + self-healing. Verifies the full LiteLLM inference stack health monitoring + self-healing. Verifies the full
nginx → LiteLLM → GPU chain, 8 containers on CT 116, 3 GPU hosts, model nginx → LiteLLM → GPU chain, 8 containers on CT 116, 3 GPU hosts, model
inference, and agent keys. Applies remediation rules for common failures. inference, and agent keys. Applies remediation rules for common failures.
Reports every action via Zulip DM and RA-H OS knowledge graph. Reports every action via Zulip DM and Gitea (SyslogSolution/health-logs).
--- ---
# LiteLLM Operations — Health Check + Self-Heal # LiteLLM Operations — Health Check + Self-Heal
@@ -102,7 +102,7 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2)
- **Health-check script** (`/opt/inference-harness/scripts/litellm-health-check.sh` on CT 116): `gpu-fleet` check fails only on **critical** alerts (warnings are informational). Tests `strix-moe` (not `ornith-1.0-35b`). - **Health-check script** (`/opt/inference-harness/scripts/litellm-health-check.sh` on CT 116): `gpu-fleet` check fails only on **critical** alerts (warnings are informational). Tests `strix-moe` (not `ornith-1.0-35b`).
- **GPU monitor** (`/root/scripts/gpu-monitor-server.py` on pi .24): runs as **systemd unit `gpu-monitor.service`** (was bare `&` process). `gpu_count` includes Strix Halo (was 2, now 3). VRAM alert thresholds: warning 93%, critical 97% (raised from 90/95 — 128K context steady-state is ~70% on RTX 3090, not a fault). - **GPU monitor** (`/root/scripts/gpu-monitor-server.py` on pi .24): runs as **systemd unit `gpu-monitor.service`** (was bare `&` process). `gpu_count` includes Strix Halo (was 2, now 3). VRAM alert thresholds: warning 93%, critical 97% (raised from 90/95 — 128K context steady-state is ~70% on RTX 3090, not a fault).
- **Agent key monitor** (`/root/scripts/agent-health-check.py` on pi .24, cron `*/10`): reads each agent's **live** `LITELLM_API_KEY` from its gateway process env via SSH — never hardcodes keys (hardcoded keys rot on rotation and caused 9×401/30min). Fleet roster: abiba, tanko, mumuni, koby, koonimo (legacy `tdunna`/`baggy` removed — never existed). - **Agent key monitor** (`/root/scripts/agent-health-check.py` on pi .24, cron `*/10`): v2 (2026-07-26) — reads each agent's **agent-specific** `{NAME}_LITELLM_API_KEY` from Infisical vault (not the shared master key). Covers: LiteLLM keys, GPU ports, agent gateways (all 5 agents now SSHa ble), CT liveness (pct status on PVE nodes), config.yaml YAML integrity, wrapper/CLI integrity, vault secret non-emptiness checks. Fleet roster: tanko (.122), mumuni (.24, inside abiba CT100), koby (.129), koonimo (.114), abiba (.24). Legacy `tdunna`/`baggy` replaced with canonical agent hostnames.
- **Stale keys cleaned**: `daily-infra-report.py` SYNTHETIC_API_KEY was stale (`sk-U_ydi3B` → 401); now reads `LITELLM_MASTER_KEY` from env. Deprecated scripts (`router-original.py`, `router-phase0-backup.py`, `apply-fixes.py`) still reference `sk-syslog-local-master-key` but do not actively poll LiteLLM. - **Stale keys cleaned**: `daily-infra-report.py` SYNTHETIC_API_KEY was stale (`sk-U_ydi3B` → 401); now reads `LITELLM_MASTER_KEY` from env. Deprecated scripts (`router-original.py`, `router-phase0-backup.py`, `apply-fixes.py`) still reference `sk-syslog-local-master-key` but do not actively poll LiteLLM.
## Maintains ## Maintains
@@ -211,8 +211,8 @@ for reference but inactive. If Redis issues occur, check harness-redis container
Every remediation cycle produces a structured report: Every remediation cycle produces a structured report:
### 1. RA-H OS Knowledge Graph Node ### 1. Gitea Log Entry
Created as `[LEARN] litellm-self-heal: <run_id>` with full JSON report. Pushed to `SyslogSolution/health-logs/litellm/{run_id}.json` — versioned, searchable, not in graph.
### 2. Zulip DM to Owner ### 2. Zulip DM to Owner
- `issues_fixed > 0` — "🛠 LiteLLM Self-Heal — Fix Applied" - `issues_fixed > 0` — "🛠 LiteLLM Self-Heal — Fix Applied"
@@ -252,9 +252,11 @@ call report-generator
health: health health: health
actions: actions actions: actions
-- Phase 4: Log to knowledge graph -- Phase 4: Log to Gitea (not knowledge graph — hard rule)
call kg-logger call gitea-logger
run_id: run_id run_id: run_id
repo: SyslogSolution/health-logs
path: litellm/{run_id}.json
health: health health: health
actions: actions actions: actions
@@ -300,3 +302,20 @@ With failures:
] ]
} }
``` ```
## gitea-logger Implementation
When this step executes, write the report JSON to a temp file and push to Gitea:
```bash
REPO="https://abiba-bot:${GITEA_PAT}@git.sysloggh.net/SyslogSolution/health-logs"
DIR="litellm"
FILE="${run_id}.json"
echo "${report_json}" > /tmp/${FILE}
(cd /tmp && git clone --depth 1 "${REPO}" &&
cp ${FILE} health-logs/${DIR}/${FILE} &&
cd health-logs && git add ${DIR}/${FILE} &&
git commit -m "litellm-health: ${run_id}" && git push)
rm -rf /tmp/health-logs /tmp/${FILE}
```
+6 -6
View File
@@ -6,7 +6,7 @@ description: >
delegation, verification, and delivery. Defines when to delegate, which delegation, verification, and delivery. Defines when to delegate, which
worker to use for what, how to handle failures, and the kanban board worker to use for what, how to handle failures, and the kanban board
protocol. Enforces context-window discipline and separation of concerns. protocol. Enforces context-window discipline and separation of concerns.
Runs on Mumuni (lxc/114, hwepve, .123) via Hermes agent. Runs on Mumuni (inside Abiba CT100, hwepve, .24) via Hermes agent (Pi + Hermes Zulip gateway).
version: 1.0.0 version: 1.0.0
--- ---
@@ -20,7 +20,7 @@ version: 1.0.0
## Topology ## Topology
**Cluster:** 6 Proxmox nodes (ocupve, acerpve, minipve, amdpve, storepve, hwepve) **Cluster:** 6 Proxmox nodes (ocupve, acerpve, minipve, amdpve, storepve, hwepve)
**Manager:** Mumuni (lxc/114, hwepve, .123) via Hermes agent **Manager:** Mumuni (inside Abiba CT100, hwepve, .24) via Hermes agent
**Workers:** 6 profiles, all running on the same agent — no separate hosts needed **Workers:** 6 profiles, all running on the same agent — no separate hosts needed
This contract is infrastructure-agnostic in terms of which nodes are used. This contract is infrastructure-agnostic in terms of which nodes are used.
@@ -31,7 +31,7 @@ Workers execute tasks on whatever infrastructure they're given — SSH to .6,
## Why This Matters ## Why This Matters
Without enforced delegation, the manager consumes the full iteration budget Without enforced delegation, the manager consumes the full iteration budget
(60 calls) on single-turn tasks — SSH to 5 nodes, check each VM, read logs — (60 calls) on single-turn tasks — SSH to 6 nodes, check each VM, read logs —
leaving no capacity for actual coordination. The result: context overflow leaving no capacity for actual coordination. The result: context overflow
(59K tokens in system prompt), iteration exhaustion, and degraded response (59K tokens in system prompt), iteration exhaustion, and degraded response
quality. This contract exists because I blew through my budget checking quality. This contract exists because I blew through my budget checking
@@ -82,7 +82,7 @@ it asks the manager (via relay) — it doesn't go find it on its own.
**This is a hard rule, not a recommendation.** Violating it produces the exact **This is a hard rule, not a recommendation.** Violating it produces the exact
type of discrepancy the kanban pipeline exists to prevent: a review worker finds type of discrepancy the kanban pipeline exists to prevent: a review worker finds
"5 nodes present" in the raw data but "5/5 online" in the report — even though "6 nodes present" in the raw data but "6/6 online" in the report — even though
one of those nodes was unreachable. The report lied because it used data the one of those nodes was unreachable. The report lied because it used data the
raw data never provided. raw data never provided.
@@ -137,7 +137,7 @@ delegate_task(
``` ```
delegate_task( delegate_task(
tasks=[ tasks=[
{"goal": "Check all 5 Proxmox nodes for VM status", "context": "SSH to each node via 192.168.68.x, run 'qm list'"}, {"goal": "Check all 6 Proxmox nodes for VM status", "context": "SSH to each node via 192.168.68.x, run 'qm list'"},
{"goal": "Check Docker container health on .7/.116/.17", "context": "SSH to each host, check container status"}, {"goal": "Check Docker container health on .7/.116/.17", "context": "SSH to each host, check container status"},
] ]
) )
@@ -191,7 +191,7 @@ Only verified results reach Kwame. Format per channel:
{ {
"lane_id": "devops-check", "lane_id": "devops-check",
"worker": "syslog-devops", "worker": "syslog-devops",
"goal": "Check all 5 Proxmox nodes", "goal": "Check all 6 Proxmox nodes",
"status": "dispatched|completed|failed", "status": "dispatched|completed|failed",
"output_file": "/tmp/node-report.md" "output_file": "/tmp/node-report.md"
} }
+1 -1
View File
@@ -52,7 +52,7 @@ description: >
- If status is "online" → pass, log restarts count - If status is "online" → pass, log restarts count
- If status is "stopped" or "errored" → **DO NOT RESTART** — alert owner immediately - If status is "stopped" or "errored" → **DO NOT RESTART** — alert owner immediately
- If restarts > 5 in last hour → alert owner with full diagnostics - If restarts > 5 in last hour → alert owner with full diagnostics
4. **Log results**Create `[LEARN]` node in knowledge graph for any actions taken 4. **Log results**Append to `SyslogSolution/health-logs/pm2/{timestamp}.md` in Gitea (not knowledge graph — hard rule)
5. **Alert** — Send Zulip DM to owner if escalation needed (do NOT run pm2 commands during alerting) 5. **Alert** — Send Zulip DM to owner if escalation needed (do NOT run pm2 commands during alerting)
6. **Wait 5 min** → repeat from step 1 6. **Wait 5 min** → repeat from step 1
+1 -1
View File
@@ -88,7 +88,7 @@ agent: abiba
| acerpve | 192.168.68.9 | PVE (hosts llm-gpu qemu/101) | | acerpve | 192.168.68.9 | PVE (hosts llm-gpu qemu/101) |
| minipve | 192.168.68.12 | PVE | | minipve | 192.168.68.12 | PVE |
| amdpve | 192.168.68.15 | PVE + Strix Halo LLM (qwen3.6-35B-udq4, strix-moe) | | amdpve | 192.168.68.15 | PVE + Strix Halo LLM (qwen3.6-35B-udq4, strix-moe) |
| hwepve | 192.168.68.4 | PVE (Huawei Matebook 16, 12C/15GB) — hosts Mumuni (lxc/114) migrated from minipve 2026-07-20 | | hwepve | 192.168.68.4 | PVE (Huawei Matebook 16, 12C/15GB) — hosts abiba (lxc/100), kagentz (lxc/105), mumuni (lxc/114). CTs 100/105 migrated from amdpve, CT 114 from minipve 2026-07-20 |
## Operations ## Operations
+300 -58
View File
@@ -1,10 +1,10 @@
#!/usr/bin/env python3 #!/usr/bin/env python3
""" """
/root/scripts/agent-health-check.py Consolidated Agent Health Verification /root/scripts/agent-health-check.py Consolidated Agent Health Verification v2
Single non-disruptive health check replacing 7 scattered scripts. Verifies: LiteLLM keys (agent-specific), GPU port conflicts, agent Zulip streaming,
Verifies: LiteLLM keys, GPU port conflicts, agent Zulip streaming, gateway liveness, gateway log health, CT liveness, config YAML integrity,
gateway liveness, and gateway log health. NEVER restarts anything. wrapper/CLI integrity, vault secret non-emptiness. NEVER restarts anything.
Usage: Usage:
python3 /root/scripts/agent-health-check.py # Full check python3 /root/scripts/agent-health-check.py # Full check
@@ -12,59 +12,39 @@ Usage:
python3 /root/scripts/agent-health-check.py --quiet # Only output on failure python3 /root/scripts/agent-health-check.py --quiet # Only output on failure
Cron: */10 * * * * python3 /root/scripts/agent-health-check.py --quiet Cron: */10 * * * * python3 /root/scripts/agent-health-check.py --quiet
Changelog:
v2 (2026-07-26): Added CT liveness, config validation, wrapper integrity,
vault secret emptiness check. Fixed Koby/Koonimo SSH hosts and agent key
name format ({NAME}_LITELLM_API_KEY not LITELLM_API_KEY_{NAME}).
Fleet roster: tanko (.122), mumuni (.24, inside abiba CT100), koby (.129), koonimo (.114),
abiba (.24).
""" """
import subprocess, json, sys, os, time import subprocess, json, sys, os, time
from datetime import datetime from datetime import datetime
LITELLM = "http://192.168.68.116:80" LITELLM = "http://192.168.68.116:80"
INFISICAL_PROJECT = "322fceab-39da-4854-a55a-568e76c0f13f"
INFISICAL_ENV = "prod"
def _get_agent_key(agent_name): # PVE node IPs for CT liveness checks
"""Retrieve agent key from Infisical vault.""" PVE_NODES = {
try: "hwepve": "192.168.68.4",
result = subprocess.run( "amdpve": "192.168.68.15",
["infisical", "secrets", "get", "LITELLM_API_KEY", "minipve": "192.168.68.12",
"--project=agents", "--env=production", "--plain"], "storepve": "192.168.68.6",
capture_output=True, text=True, timeout=10 "acerpve": "192.168.68.9",
) "ocupve": "192.168.68.5",
if result.returncode == 0:
return result.stdout.strip()
except Exception:
pass
# Fallback: try exporting all secrets
try:
result = subprocess.run(
["infisical", "export", "--project=agents", "--env=production",
"--format=dotenv"],
capture_output=True, text=True, timeout=10
)
if result.returncode == 0:
for line in result.stdout.splitlines():
if line.startswith(f"LITELLM_API_KEY_{agent_name.upper()}") or \
(line.startswith("LITELLM_API_KEY=") and agent_name == os.uname().nodename):
return line.split("=", 1)[1].strip().strip('"').strip("'")
except Exception:
pass
return None
# Agent keys are pulled from Infisical vault at runtime.
# The 'key' field is populated dynamically below.
AGENTS = {
"tanko": {"ct": 112, "host": "192.168.68.122", "user": "jerome"},
"mumuni": {"ct": 114, "host": "192.168.68.123", "user": "root"},
"koby": {"ct": 111, "host": None, "user": None},
"koonimo": {"ct": 113, "host": None, "user": None},
} }
# Inject keys from vault # Agent definitions: ct, host, user, pve_node, vault_key_name
for agent_name in AGENTS: AGENTS = {
key = _get_agent_key(agent_name) "tanko": {"ct": 112, "host": "192.168.68.122", "user": "jerome", "pve": "amdpve", "vault_key": "TANKO_LITELLM_API_KEY"},
if key: "abiba": {"ct": 100, "host": "192.168.68.24", "user": "root", "pve": "hwepve", "vault_key": None}, # Pi agent + Mumuni Zulip, no vault key
AGENTS[agent_name]["key"] = key "koby": {"ct": 111, "host": "192.168.68.129", "user": "root", "pve": "amdpve", "vault_key": "KOBY_LITELLM_API_KEY"},
else: "koonimo": {"ct": 113, "host": "192.168.68.114", "user": "root", "pve": "amdpve", "vault_key": "KOONIMO_LITELLM_API_KEY"},
AGENTS[agent_name]["key"] = None }
GPU_HOSTS = { GPU_HOSTS = {
"gpu-rtx3090 (.8)": {"host": "192.168.68.8", "port": 8080, "service": "llama-server"}, "gpu-rtx3090 (.8)": {"host": "192.168.68.8", "port": 8080, "service": "llama-server"},
@@ -74,6 +54,11 @@ GPU_HOSTS = {
FAIL = [] FAIL = []
INFISICAL_TOKEN = os.environ.get("INFISICAL_TOKEN")
INFISICAL_API_URL = os.environ.get("INFISICAL_API_URL", "https://vault.sysloggh.net")
# ── Helpers ──────────────────────────────────────────────────────────
def ssh(host, cmd, user="root"): def ssh(host, cmd, user="root"):
"""Execute a command on a remote host, return stdout or None.""" """Execute a command on a remote host, return stdout or None."""
try: try:
@@ -112,15 +97,81 @@ def http_json(url, headers=None, timeout=5):
except: except:
return None return None
def run_infisical(args, quiet=True):
"""Run infisical CLI with env-based auth, return stdout or None."""
env = os.environ.copy()
env["INFISICAL_API_URL"] = INFISICAL_API_URL
if INFISICAL_TOKEN:
env["INFISICAL_TOKEN"] = INFISICAL_TOKEN
try:
result = subprocess.run(
["/usr/bin/infisical"] + args,
capture_output=True, text=True, timeout=15, env=env
)
return result.stdout.strip() if result.returncode == 0 else None
except:
return None
# ── KEY LOOKUP FIX ───────────────────────────────────────────────────
def _get_agent_key(agent_name, vault_key_name):
"""Retrieve agent-specific key from Infisical vault.
Uses {NAME}_LITELLM_API_KEY format (e.g., TANKO_LITELLM_API_KEY,
KOONIMO_LITELLM_API_KEY) which matches actual vault key names.
"""
if not vault_key_name:
return None
# Primary: get the agent-specific key by name
key = run_infisical([
"secrets", "get", vault_key_name,
"--projectId=" + INFISICAL_PROJECT,
"--env=" + INFISICAL_ENV,
"--plain",
])
if key and key.startswith("sk-"):
return key
# Fallback: export all and search for the key name
try:
export = run_infisical([
"export",
"--projectId=" + INFISICAL_PROJECT,
"--env=" + INFISICAL_ENV,
"--format=dotenv",
])
if export:
for line in export.splitlines():
if line.startswith(vault_key_name + "="):
value = line.split("=", 1)[1].strip().strip('"').strip("'")
if value.startswith("sk-"):
return value
except:
pass
return None
# Inject keys from vault for each agent
for agent_name in AGENTS:
info = AGENTS[agent_name]
key = _get_agent_key(agent_name, info.get("vault_key"))
AGENTS[agent_name]["key"] = key
# ═══════════════════════════════════════════════════════════════════ # ═══════════════════════════════════════════════════════════════════
# CHECK 1: LiteLLM Key Validation # CHECK 1: LiteLLM Key Validation (agent-specific keys)
# ═══════════════════════════════════════════════════════════════════ # ═══════════════════════════════════════════════════════════════════
def check_keys(): def check_keys():
for name, agent in AGENTS.items(): for name, agent in AGENTS.items():
key = agent.get("key")
if not key:
print(f"{name}: NO KEY FOUND (vault empty or unreachable)")
FAIL.append(f"key:{name}:no-key")
continue
data = http_json(f"{LITELLM}/v1/models", data = http_json(f"{LITELLM}/v1/models",
headers={"Authorization": f"Bearer {agent['key']}"}) headers={"Authorization": f"Bearer {key}"})
if data and data.get("data"): if data and data.get("data"):
model = data["data"][0].get("id", "?") model = data["data"][0].get("id", "?")
print(f"{name}: key valid → {model}") print(f"{name}: key valid → {model}")
@@ -130,7 +181,7 @@ def check_keys():
# ═══════════════════════════════════════════════════════════════════ # ═══════════════════════════════════════════════════════════════════
# CHECK 2: GPU Port Conflict Detection # CHECK 2: GPU Port Conflict Detection (unchanged)
# ═══════════════════════════════════════════════════════════════════ # ═══════════════════════════════════════════════════════════════════
def check_gpu_ports(): def check_gpu_ports():
@@ -158,7 +209,6 @@ def check_gpu_ports():
else: else:
print(f" ⚠️ {label}: svc={svc_status}, port owned by {port_owner}") print(f" ⚠️ {label}: svc={svc_status}, port owned by {port_owner}")
else: else:
# Verify health endpoint
health = ssh(host, f"curl -s --max-time 5 http://localhost:{port}/health") health = ssh(host, f"curl -s --max-time 5 http://localhost:{port}/health")
if health and '"status":"ok"' in health: if health and '"status":"ok"' in health:
print(f"{label}: healthy (pid={port_owner})") print(f"{label}: healthy (pid={port_owner})")
@@ -171,7 +221,7 @@ def check_gpu_ports():
# ═══════════════════════════════════════════════════════════════════ # ═══════════════════════════════════════════════════════════════════
# CHECK 3: Agent Gateway Liveness + Streaming # CHECK 3: Agent Gateway Liveness + Streaming (now covers all agents)
# ═══════════════════════════════════════════════════════════════════ # ═══════════════════════════════════════════════════════════════════
def check_agents(): def check_agents():
@@ -184,8 +234,11 @@ def check_agents():
print(f"{name} (CT {ct}): cannot SSH — skip liveness check") print(f"{name} (CT {ct}): cannot SSH — skip liveness check")
continue continue
# Gateway process (exclude the infisical bash wrapper that contains the same string) # Gateway process
pid = ssh(host, "pgrep -f 'hermes_cli.main gateway run' | grep -v infisical | head -1", user=user) pid = ssh(host, "pgrep -f 'hermes_cli.main gateway run' | grep -v infisical | head -1", user=user)
if not pid:
# Try alternate binary name
pid = ssh(host, "pgrep -f 'hermes.*gateway' | grep -v infisical | grep -v bash | head -1", user=user)
if not pid: if not pid:
print(f"{name}: GATEWAY NOT RUNNING") print(f"{name}: GATEWAY NOT RUNNING")
FAIL.append(f"gateway-down:{name}") FAIL.append(f"gateway-down:{name}")
@@ -203,7 +256,7 @@ def check_agents():
else: else:
gw_state, zulip = "no-state-file", "?" gw_state, zulip = "no-state-file", "?"
# Zulip streaming: does adapter have edit_message? # Zulip streaming check
adapter_paths = [ adapter_paths = [
"~/.hermes/plugins/zulip-platform/adapter.py", "~/.hermes/plugins/zulip-platform/adapter.py",
"~/.hermes/plugins/platforms/zulip/adapter.py", "~/.hermes/plugins/platforms/zulip/adapter.py",
@@ -220,13 +273,183 @@ def check_agents():
r"journalctl --user -u hermes-gateway --since '10 min ago' -o cat --no-pager 2>/dev/null " r"journalctl --user -u hermes-gateway --since '10 min ago' -o cat --no-pager 2>/dev/null "
r"| grep -ci 'error\|traceback\|exception\|401\|403\|500' || echo 0", r"| grep -ci 'error\|traceback\|exception\|401\|403\|500' || echo 0",
user=user) user=user)
recent_errors = (recent_errors or "0").strip().split("\n")[-1] # take last line recent_errors = (recent_errors or "0").strip().split("\n")[-1]
print(f" {'' if gw_state == 'running' and zulip == 'connected' else '⚠️'} " print(f" {'' if gw_state == 'running' and zulip == 'connected' else '⚠️'} "
f"{name}: gw={gw_state} zulip={zulip} streaming={streaming} " f"{name}: gw={gw_state} zulip={zulip} streaming={streaming} "
f"errors_10m={recent_errors.strip() or '0'} pid={pid}") f"errors_10m={recent_errors.strip() or '0'} pid={pid}")
# ═══════════════════════════════════════════════════════════════════
# CHECK 4: CT Liveness (NEW)
# ═══════════════════════════════════════════════════════════════════
def check_ct_liveness():
"""Check that all agent CTs are running on their PVE nodes."""
for name, agent in AGENTS.items():
ct = agent["ct"]
pve_node = agent.get("pve")
if not pve_node:
print(f"{name} (CT {ct}): no PVE node mapped — skip")
continue
pve_ip = PVE_NODES.get(pve_node)
if not pve_ip:
print(f"{name}: unknown PVE node '{pve_node}' — skip")
continue
status = ssh(pve_ip, f"pct status {ct} 2>/dev/null", user="root")
if not status:
print(f"{name} (CT {ct} on {pve_node}): PVE UNREACHABLE")
FAIL.append(f"ct-unreachable:{name}:{pve_ip}")
elif "running" in status:
print(f"{name} (CT {ct} on {pve_node}): running")
elif "stopped" in status:
print(f"{name} (CT {ct} on {pve_node}): STOPPED")
FAIL.append(f"ct-stopped:{name}")
else:
print(f" ⚠️ {name} (CT {ct} on {pve_node}): {status.strip()}")
# ═══════════════════════════════════════════════════════════════════
# CHECK 5: Config YAML Integrity (NEW)
# ═══════════════════════════════════════════════════════════════════
def check_config_integrity():
"""Verify agent config.yaml parses as valid YAML."""
for name, agent in AGENTS.items():
host = agent.get("host")
user = agent.get("user")
if not host or not user:
print(f"{name}: cannot SSH — skip config check")
continue
# Check YAML parses
yaml_ok = ssh(host,
"python3 -c "
'"import yaml; yaml.safe_load(open(\'/root/.hermes/config.yaml\')); print(\'OK\')" '
"2>&1 || echo 'FAIL'",
user=user)
if not yaml_ok:
print(f"{name}: SSH UNREACHABLE (config check skipped)")
FAIL.append(f"config-unreachable:{name}")
elif "OK" in yaml_ok:
print(f"{name}: config.yaml valid YAML")
else:
print(f"{name}: config.yaml YAML ERROR — {yaml_ok[:120]}")
FAIL.append(f"config-yaml-error:{name}")
# ═══════════════════════════════════════════════════════════════════
# CHECK 6: Wrapper/CLI Integrity (NEW)
# ═══════════════════════════════════════════════════════════════════
def check_wrapper_integrity():
"""Verify the hermes CLI wrapper exists and can reach hermes-real."""
for name, agent in AGENTS.items():
host = agent.get("host")
user = agent.get("user")
if not host or not user:
print(f"{name}: cannot SSH — skip wrapper check")
continue
# Check wrapper exists
wrapper = ssh(host, "ls -la /root/.local/bin/hermes 2>/dev/null", user=user)
if not wrapper:
# Check alternate wrapper locations
wrapper = ssh(host, "which hermes 2>/dev/null; command -v hermes 2>/dev/null", user=user)
if not wrapper:
print(f"{name}: NO HERMES CLI WRAPPER FOUND")
FAIL.append(f"wrapper-missing:{name}")
continue
else:
print(f" ⚠️ {name}: hermes at {wrapper.strip()} (not ~/.local/bin/hermes)")
# Check wrapper has correct infisical path
infisical_path_valid = ssh(host,
"head -20 /root/.local/bin/hermes 2>/dev/null | grep -q '/usr/bin/infisical' && echo OK || echo MISS",
user=user)
if infisical_path_valid == "MISS":
# Check if infisical exists on path
inf_actual = ssh(host, "command -v infisical 2>/dev/null", user=user)
if not inf_actual:
print(f"{name}: INFISICAL NOT INSTALLED (wrapper broken)")
FAIL.append(f"wrapper-no-infisical:{name}")
else:
print(f" ⚠️ {name}: wrapper infisical path may be wrong (infisical at {inf_actual})")
FAIL.append(f"wrapper-infisical-path:{name}")
# Check hermes-real exists
hermes_real = ssh(host,
"ls -la /root/.local/bin/hermes-real 2>/dev/null || echo MISS",
user=user)
if not hermes_real or hermes_real.strip() == "MISS":
# Check venv path
hermes_real = ssh(host,
"ls -la /usr/local/lib/hermes-agent/venv/bin/hermes 2>/dev/null || echo MISS",
user=user)
if not hermes_real or hermes_real.strip() == "MISS":
print(f"{name}: hermes-real NOT FOUND (wrapper broken)")
FAIL.append(f"wrapper-no-hermes-real:{name}")
else:
print(f"{name}: hermes-real at alt path")
# Check the .env file has the key
env_has_key = ssh(host,
"grep -c 'LITELLM_API_KEY' /root/.hermes/.env 2>/dev/null || echo 0",
user=user)
if env_has_key and env_has_key.strip() not in ("", "0"):
print(f"{name}: wrapper + .env key present")
else:
print(f" ⚠️ {name}: .env may be missing LITELLM_API_KEY entry")
# ═══════════════════════════════════════════════════════════════════
# CHECK 7: Vault Secret Non-Emptiness (NEW)
# ═══════════════════════════════════════════════════════════════════
def check_vault_secrets():
"""Verify agent-specific vault secrets are non-empty and start with sk-."""
for name, agent in AGENTS.items():
vault_key_name = agent.get("vault_key")
if not vault_key_name:
continue
key = agent.get("key")
if not key:
print(f"{name}: vault secret {vault_key_name} MISSING or EMPTY")
FAIL.append(f"vault-empty:{name}:{vault_key_name}")
elif not key.startswith("sk-"):
print(f"{name}: vault secret {vault_key_name} WRONG FORMAT (starts '{key[:8]}...')")
FAIL.append(f"vault-bad-format:{name}:{vault_key_name}")
else:
print(f"{name}: vault {vault_key_name}=sk-...{key[-4:]}")
# ═══════════════════════════════════════════════════════════════════
# DEPLOY: copy updated script to /root/scripts/ on local host
# ═══════════════════════════════════════════════════════════════════
def deploy_self():
"""Copy this script to /root/scripts/agent-health-check.py if out of date."""
dest = "/root/scripts/agent-health-check.py"
try:
with open(__file__, "r") as f:
current = f.read()
if os.path.isfile(dest):
with open(dest, "r") as f:
existing = f.read()
if current == existing:
return # Already deployed
# Write new version
with open(dest, "w") as f:
f.write(current)
os.chmod(dest, 0o755)
print(f" 📦 Deployed updated script to {dest}")
except:
pass # Not fatal if deploy fails
# ═══════════════════════════════════════════════════════════════════ # ═══════════════════════════════════════════════════════════════════
# MAIN # MAIN
# ═══════════════════════════════════════════════════════════════════ # ═══════════════════════════════════════════════════════════════════
@@ -235,8 +458,12 @@ def main():
quiet = "--quiet" in sys.argv quiet = "--quiet" in sys.argv
as_json = "--json" in sys.argv as_json = "--json" in sys.argv
# Self-deploy to canonical location
if not quiet and "--no-deploy" not in sys.argv:
deploy_self()
if not quiet: if not quiet:
print(f"🏥 Agent Health Check — {datetime.now().strftime('%Y-%m-%d %H:%M UTC')}") print(f"🏥 Agent Health Check v2 {datetime.now().strftime('%Y-%m-%d %H:%M UTC')}")
print() print()
print("🔑 LiteLLM Keys:") print("🔑 LiteLLM Keys:")
@@ -249,11 +476,26 @@ def main():
print("🤖 Agent Gateways:") print("🤖 Agent Gateways:")
check_agents() check_agents()
print()
print("🖥️ CT Liveness:")
check_ct_liveness()
print()
print("📝 Config Integrity:")
check_config_integrity()
print()
print("🔌 Wrapper/CLI Integrity:")
check_wrapper_integrity()
print()
print("🔐 Vault Secrets:")
check_vault_secrets()
if FAIL: if FAIL:
print(f"\n{len(FAIL)} FAILURE(S): {' | '.join(FAIL)}") print(f"\n{len(FAIL)} FAILURE(S): {' | '.join(FAIL)}")
if quiet: if quiet:
# In quiet mode, only print failures as a single alert line
print(f"ALERT agent-health:{','.join(FAIL)}") print(f"ALERT agent-health:{','.join(FAIL)}")
elif not quiet: elif not quiet:
print("\n✅ All checks passed") print("\n✅ All checks passed")
+65
View File
@@ -0,0 +1,65 @@
#!/bin/bash
# Netbird Reverse Proxy — Add a new domain route
#
# Usage: netbird-add-domain.sh <domain> <backend_ip> [port] [protocol]
#
# Example:
# netbird-add-domain.sh dns.sysloggh.net 192.168.68.10 80
#
# This script adds a domain to the Netbird proxy by inserting records
# directly into the management server's SQLite database, then restarting
# the proxy stack.
#
# Prerequisites: SSH root access to 72.61.0.17
# sqlite3 available on VPS
#
# Requires: The domain must already have a DNS CNAME to netbird.sysloggh.net
# pointing to 72.61.0.17.
set -euo pipefail
DOMAIN="${1:?Usage: netbird-add-domain.sh <domain> <backend_ip> [port] [protocol]}"
BACKEND_IP="${2:?Usage: netbird-add-domain.sh <domain> <backend_ip> [port] [protocol]}"
PORT="${3:-80}"
PROTOCOL="${4:-http}"
VPS="root@72.61.0.17"
DB_VOLUME="/var/lib/docker/volumes/root_netbird_data/_data"
DB="$DB_VOLUME/store.db"
echo "=== Adding Netbird proxy route ==="
echo "Domain: $DOMAIN"
echo "Backend: $BACKEND_IP:$PORT ($PROTOCOL)"
echo ""
ssh "$VPS" bash << REMOTESCRIPT
set -euo pipefail
# Generate unique ID using timestamp hash (Netbird format)
ID_SUFFIX=\$(date +%s | md5sum | head -c 16)
SVC_ID="d9\${ID_SUFFIX}ptsnc73\$(date +%s | md5sum | head -c 10)"
TGT_ID=\$(sqlite3 "$DB" "SELECT COALESCE(MAX(id), 100) + 1 FROM targets;")
ACCOUNT_ID="d88av3aptsnc73clmogg"
ZONE_ID="d8adqjaptsnc73fro5g0"
echo "Service ID: \$SVC_ID"
echo "Target ID: \$TGT_ID"
# Insert service
sqlite3 "$DB" "INSERT INTO services (id, account_id, name, domain, proxy_cluster, enabled, terminated, pass_host_header, rewrite_redirects, mode, source, port_auto_assigned, private) VALUES (\"\$SVC_ID\", \"\$ACCOUNT_ID\", \"$DOMAIN\", \"$DOMAIN\", \"netbird.sysloggh.net\", 1, 0, 1, 0, \"http\", \"permanent\", 0, 0);"
echo "Service: OK"
# Insert target
sqlite3 "$DB" "INSERT INTO targets (id, account_id, service_id, host, port, protocol, target_id, target_type, enabled, skip_tls_verify, request_timeout, session_idle_timeout, agent_network, disable_access_log) VALUES (\$TGT_ID, \"\$ACCOUNT_ID\", \"\$SVC_ID\", \"$BACKEND_IP\", $PORT, \"$PROTOCOL\", \"\$ZONE_ID\", \"subnet\", 1, 0, 0, 0, 0, 0);"
echo "Target: OK"
# Verify
sqlite3 -column "$DB" "SELECT s.name, t.host, t.port, t.protocol FROM services s JOIN targets t ON s.id=t.service_id WHERE s.name=\"$DOMAIN\";"
echo ""
echo "Restarting proxy stack..."
cd /root && docker compose restart netbird-server 2>/dev/null
sleep 15
docker compose restart proxy 2>/dev/null
echo "Done. Verify with: curl -sI https://$DOMAIN"
REMOTESCRIPT
+9 -6
View File
@@ -11,31 +11,33 @@ set -euo pipefail
# ── CT ID → PVE Node mapping (maintained HERE, not in prose contracts) ── # ── CT ID → PVE Node mapping (maintained HERE, not in prose contracts) ──
declare -A CT_NODES=( declare -A CT_NODES=(
# amdpve (192.168.68.15) # amdpve (192.168.68.15)
[100]=amdpve # abiba
[105]=amdpve # kagentz
[111]=amdpve # tdunna [111]=amdpve # tdunna
[112]=amdpve # tanko [112]=amdpve # tanko
[113]=amdpve # baggy [113]=amdpve # baggy
[115]=amdpve # scottdenya [115]=amdpve # scottdenya
# minipve (192.168.68.12) # minipve (192.168.68.12)
[102]=minipve # adguard (was acerpve)
[104]=minipve # authentik [104]=minipve # authentik
[110]=minipve # gitea [110]=minipve # gitea
[114]=minipve # mumuni
[116]=minipve # syslog-api [116]=minipve # syslog-api
[119]=minipve # infisical-vault
# storepve (192.168.68.6) # storepve (192.168.68.6)
[106]=storepve # ra-h-os [106]=storepve # ra-h-os
[107]=storepve # proxmox-backup [107]=storepve # proxmox-backup
[108]=storepve # media [108]=storepve # media
[117]=storepve # zulip [117]=storepve # zulip
# acerpve (192.168.68.9) [118]=storepve # jdownloader
[102]=acerpve # adguard # acerpve (192.168.68.9) — no CTs (bare metal GPU .8)
# hwepve (192.168.68.4)
[100]=hwepve # abiba (was amdpve)
[105]=hwepve # kagentz (was amdpve)
[114]=hwepve # mumuni (was minipve)
# ocupve (192.168.68.5) — no CTs (bare metal GPU .110) # ocupve (192.168.68.5) — no CTs (bare metal GPU .110)
# #
# REMOVED CTs (migrated to bare metal, decommissioned, or VMs): # REMOVED CTs (migrated to bare metal, decommissioned, or VMs):
# 101 llm-gpu → bare metal 192.168.68.8 (RTX 3090) # 101 llm-gpu → bare metal 192.168.68.8 (RTX 3090)
# 103 ocu-llm → bare metal 192.168.68.110 (RTX 5070) # 103 ocu-llm → bare metal 192.168.68.110 (RTX 5070)
# 109 docker-vm → KVM VM 192.168.68.7 (use direct SSH) # 109 docker-vm → KVM VM 192.168.68.7 (use direct SSH)
# 118 jitsi → stopped, not in service
) )
# Each node must be root-accessible via SSH hostname # Each node must be root-accessible via SSH hostname
@@ -48,6 +50,7 @@ declare -A NODE_IPS=(
[storepve]=192.168.68.6 [storepve]=192.168.68.6
[acerpve]=192.168.68.9 [acerpve]=192.168.68.9
[ocupve]=192.168.68.5 [ocupve]=192.168.68.5
[hwepve]=192.168.68.4
) )
resolve_node() { resolve_node() {
+8 -6
View File
@@ -46,17 +46,19 @@ You are a code reviewer for OpenProse infrastructure contracts in the Syslog Sol
The infrastructure-control.prose.md contract is the canonical reference for the cluster topology: The infrastructure-control.prose.md contract is the canonical reference for the cluster topology:
**Proxmox Cluster "Tabiri" (5 nodes):** **Proxmox Cluster "Tabiri" (6 nodes):**
- amdpve (192.168.68.15): abiba, kagentz, tanko, tdunna, baggy, scottdenya - amdpve (192.168.68.15): tanko, tdunna, baggy, scottdenya
- minipve (192.168.68.12): authentik, gitea, mumuni, syslog-api, jitsi - minipve (192.168.68.12): adguard, authentik, gitea, syslog-api, infisical-vault
- storepve (192.168.68.6): docker-vm, ra-h-os, PBS, media, zulip - storepve (192.168.68.6): docker-vm, ra-h-os, PBS, media, jdownloader, zulip
- acerpve (192.168.68.9): llm-gpu, adguard - acerpve (192.168.68.9): llm-gpu
- ocupve (192.168.68.5): ocu-llm - ocupve (192.168.68.5): ocu-llm
- hwepve (192.168.68.4): abiba, kagentz, mumuni
**CT IDs (verified 2026-07-04 against PVE API):** **CT IDs (verified 2026-07-24 against PVE API):**
100:abiba 102:adguard 104:authentik 105:kagentz 106:ra-h-os 100:abiba 102:adguard 104:authentik 105:kagentz 106:ra-h-os
107:pbs 108:media 110:gitea 111:tdunna 112:tanko 107:pbs 108:media 110:gitea 111:tdunna 112:tanko
113:baggy 114:mumuni 115:scottdenya 116:syslog-api 117:zulip 113:baggy 114:mumuni 115:scottdenya 116:syslog-api 117:zulip
118:jdownloader 119:infisical-vault
**NO CT 122, CT 123, or .19 exist in the cluster.** **NO CT 122, CT 123, or .19 exist in the cluster.**
+91
View File
@@ -0,0 +1,91 @@
#!/bin/bash
# swap-gpu-dense-model.sh — Swap RTX 3090 from qwen3.6-27B-code to SmartCode-Fable-5
# Run when download completes: ssh root@192.168.68.8 'bash -s' < this script
#
# Usage: bash swap-gpu-dense-model.sh
# Requires: new model at /home/llmuser/models/SmartCode-Fable-5-27B-UD-Q4_K_XL.gguf
set -e
MODEL_PATH="/home/llmuser/models/SmartCode-Fable-5-27B-UD-Q4_K_XL.gguf"
OLD_WRAPPER="/home/llmuser/llama-wrapper.sh"
echo "═══ Swapping gpu-dense to SmartCode-Fable-5 ═══"
# 1. Verify model file
if [ ! -f "$MODEL_PATH" ]; then
echo "❌ Model not found at $MODEL_PATH"
echo " Download: curl -L -o $MODEL_PATH <huggingface-url>"
exit 1
fi
MODEL_SIZE=$(ls -lh "$MODEL_PATH" | awk '{print $5}')
echo "✅ Model found: $MODEL_SIZE"
# 2. Create new wrapper script for SmartCode-Fable-5
cat > /home/llmuser/llama-fable-wrapper.sh << 'WRAPPER'
#!/bin/bash
# SmartCode-Fable-5 llama-server wrapper for RTX 3090
# Sampler settings from model card: temp 0.9, top-p 0.95, top-k 60, repeat-penalty off
PORT=8080
GHOST_PID=$(ss -tlnp 2>/dev/null | grep -Po ":${PORT}\s+.*pid=\K[0-9]+" | head -1)
if [ -n "$GHOST_PID" ] && [ "$GHOST_PID" != "$$" ]; then
echo "[wrapper] Port $PORT occupied by ghost pid $GHOST_PID — cleaning up" >&2
kill -9 "$GHOST_PID" 2>/dev/null
sleep 2
fi
exec /usr/local/bin/llama-server \
--model /home/llmuser/models/SmartCode-Fable-5-27B-UD-Q4_K_XL.gguf \
--ctx-size 131072 \
--cache-type-k q4_0 \
--cache-type-v q4_0 \
--flash-attn 1 \
--cont-batching \
--parallel 1 \
--batch-size 2048 \
--ubatch-size 1024 \
--n-gpu-layers 99 \
--temp 0.9 \
--top-p 0.95 \
--top-k 60 \
--min-p 0.0 \
--repeat-penalty 1.0 \
--api-key not-needed \
--port 8080 \
--host 0.0.0.0
WRAPPER
chmod 755 /home/llmuser/llama-fable-wrapper.sh
echo "✅ Created /home/llmuser/llama-fable-wrapper.sh"
# 3. Update systemd service to use new wrapper
echo "📝 Updating systemd service..."
sed -i 's|ExecStart=/home/llmuser/llama-wrapper.sh|ExecStart=/home/llmuser/llama-fable-wrapper.sh|' /etc/systemd/system/llama-server.service
systemctl daemon-reload
# 4. Stop old server, start new
echo "🔄 Restarting llama-server..."
systemctl stop llama-server
sleep 3
systemctl start llama-server
sleep 8
# 5. Verify
echo ""
echo "═══ Verification ═══"
systemctl is-active llama-server
echo ""
echo "Port 8080:"
ss -tlnp 2>/dev/null | grep ":8080" | head -1
echo ""
echo "GPU VRAM:"
nvidia-smi --query-gpu=memory.used,memory.total,utilization.gpu --format=csv,noheader 2>/dev/null
echo ""
echo "=== Health check ==="
curl -s --max-time 5 http://localhost:8080/health 2>/dev/null
echo ""
echo ""
echo "✅ Swap complete. Test via LiteLLM:"
echo " curl -s http://192.168.68.116/v1/chat/completions -H 'Authorization: Bearer <key>' -H 'Content-Type: application/json' -d '{\"model\":\"gpu-dense\",\"messages\":[{\"role\":\"user\",\"content\":\"write hello world in python\"}],\"max_tokens\":100}'"
+3 -3
View File
@@ -16,7 +16,7 @@ Runs every 15 minutes in the background. Also triggers on session start.
## Requires ## Requires
- **Zulip API key** for `abiba-bot@chat.sysloggh.net` in `$ZULIP_API_KEY` - **Zulip API key** for `abiba-bot@chat.sysloggh.net` in `$ZULIP_API_KEY`
- **SSH access** to Tanko (192.168.68.122), Mumuni (192.168.68.123, lxc/114 on hwepve since 2026-07-20), and Agent Zero Docker host (192.168.68.14) - **SSH access** to Tanko (192.168.68.122), Mumuni (192.168.68.24, inside Abiba CT100 on hwepve), and Agent Zero Docker host (192.168.68.14)
- **PM2** on localhost for pi process management - **PM2** on localhost for pi process management
- **Network access** to `chat.sysloggh.net`, `localhost:9200` - **Network access** to `chat.sysloggh.net`, `localhost:9200`
- **Write access** to `/root/zulip-health-monitor.log` and `/tmp/zulip-monitor-debounce` - **Write access** to `/root/zulip-health-monitor.log` and `/tmp/zulip-monitor-debounce`
@@ -182,13 +182,13 @@ grep -a "Finalized\|Failed to finalize" /root/.pm2/logs/abiba-zulip-out.log | ta
| `last_error` set | Log and monitor | | `last_error` set | Log and monitor |
| Crash loop >10/h | Alert user | | Crash loop >10/h | Alert user |
### Step 3: Platform B — Hermes (Tanko .122, Mumuni .123) ### Step 3: Platform B — Hermes (Tanko .122, Mumuni .24)
**B1: Gateway State** **B1: Gateway State**
```bash ```bash
ssh root@192.168.68.122 "cat ~/.hermes/gateway_state.json" ssh root@192.168.68.122 "cat ~/.hermes/gateway_state.json"
ssh root@192.168.68.123 "cat ~/.hermes/gateway_state.json" ssh root@192.168.68.24 "cat ~/.hermes/gateway_state.json" # Mumuni inside Abiba CT100
``` ```
Check `platforms.zulip.state`: `connected` ✅ | `disconnected` ❌ | `error` ❌ | missing → not installed. Check `platforms.zulip.state`: `connected` ✅ | `disconnected` ❌ | `error` ❌ | missing → not installed.
+1 -1
View File
@@ -65,7 +65,7 @@ triggers:
|------|----|------|---------| |------|----|------|---------|
| Zulip server | 192.168.68.19 | root | Docker: `zulip-zulip-1` | | Zulip server | 192.168.68.19 | root | Docker: `zulip-zulip-1` |
| Abiba (pi) | localhost | root | PM2: `abiba-zulip` | | Abiba (pi) | localhost | root | PM2: `abiba-zulip` |
| Mumuni | 192.168.68.123 | root | `hermes gateway restart` | | Mumuni | 192.168.68.24 (CT100 abiba) | root | `hermes gateway restart` |
| Tanko | 192.168.68.122 | jerome | `PATH=$PATH:/home/jerome/.hermes/hermes-agent hermes gateway restart` | | Tanko | 192.168.68.122 | jerome | `PATH=$PATH:/home/jerome/.hermes/hermes-agent hermes gateway restart` |
## Debounce ## Debounce