From b17c60f997301347bb5e9a015a75809172e003b4 Mon Sep 17 00:00:00 2001 From: root Date: Fri, 24 Jul 2026 19:56:26 +0000 Subject: [PATCH 1/5] fix: contract accuracy updates post fleet-wide reboot infrastructure-control.prose.md: - Add hwepve as 6th Proxmox node - Fix Gitea IP: .17 (was .110) - Fix AdGuard IP: .10 on minipve (was .102 on acerpve) - Fix Abiba placement: hwepve (was amdpve) - Fix Mumuni placement: hwepve (was minipve) - Fix Authentik port: add :9000 - Add CT 118 (jdownloader), CT 119 (infisical-vault) - Add last-verified date (2026-07-24) infrastructure-update.prose.md: - Add post-reboot CT sweep procedure - Add Zulip Docker network recovery steps - Add WireGuard tunnel verification scripts/netbird-add-domain.sh: - New script to register domains in Netbird proxy store.db --- infrastructure-control.prose.md | 65 +++++++++++++++++++++------------ infrastructure-update.prose.md | 41 ++++++++++++++++++++- scripts/netbird-add-domain.sh | 65 +++++++++++++++++++++++++++++++++ 3 files changed, 145 insertions(+), 26 deletions(-) create mode 100644 scripts/netbird-add-domain.sh diff --git a/infrastructure-control.prose.md b/infrastructure-control.prose.md index 8471017..e7aa436 100644 --- a/infrastructure-control.prose.md +++ b/infrastructure-control.prose.md @@ -3,7 +3,7 @@ kind: pattern name: infrastructure-control description: > Full infrastructure monitoring and control pattern covering the - 5-node Proxmox cluster, 3 Docker ecosystems (22 containers), + 6-node Proxmox cluster, 3 Docker ecosystems (22 containers), NFS storage, and network services. Defines monitors, remediations, and the access matrix for all environments. @@ -12,6 +12,11 @@ description: > never mutate infrastructure based on them without first confirming against the live system. Policy fields are authoritative. See the `verify-before-mutate` skill. + + **Last verified:** 2026-07-24 — corrected Gitea IP (.17 not .110), + AdGuard IP (.10 not .102), AdGuard placement (minipve not acerpve), + Abiba placement (hwepve not amdpve), added hwepve as 6th node, + added dns.sysloggh.net route. See data/learnings.md. --- # Infrastructure Control Pattern @@ -42,8 +47,8 @@ description: > │ │ │ │ │ ┌────┘ ┌────┘ ┌────┘ ┌────┘ ┌────┘ ▼ ▼ ▼ ▼ ▼ - minipve amdpve storepve acerpve ocupve - (.12) (.15) (.6) (.9) (.5) + minipve amdpve storepve acerpve ocupve hwepve + (.12) (.15) (.6) (.9) (.5) (.4) ▼ ┌─────────────────────────────────────────────┐ @@ -95,15 +100,21 @@ description: > ## Section 2: Proxmox Cluster — Monitoring -### Nodes (5) +### Nodes (6) | Node | IP | CPU | RAM | VMs/CTs | Role | |------|----|-----|-----|---------|------| -| minipve | .12 | 16C | 30GB | authentik, gitea, mumuni, syslog-api, jitsi | Auth, git, messaging | -| amdpve | .15 | 32C | 62GB | abiba, kagentz, tanko, tdunna, baggy, scottdenya | Agents, compute | -| storepve | .6 | 28C | 31GB | docker-vm, ra-h-os, PBS, media, zulip | Docker, storage, chat | -| acerpve | .9 | 28C | 31GB | llm-gpu, adguard | GPU VMs | +| minipve | .12 | 16C | 30GB | authentik, gitea, syslog-api, infisical-vault, jitsi | Auth, git, messaging | +| amdpve | .15 | 32C | 62GB | kagentz, tanko, tdunna, baggy, scottdenya | Agents, compute | +| storepve | .6 | 28C | 31GB | docker-vm, ra-h-os, PBS, media, jdownloader, zulip | Docker, storage, chat | +| acerpve | .9 | 28C | 31GB | llm-gpu | GPU VMs | | ocupve | .5 | 12C | 14GB | ocu-llm | GPU VMs | +| hwepve | .4 | ? | ? | abiba, (mumuni CT 114 stopped) | Agents (new node) | + +> **Note:** CTs on storepve include jdownloader (CT 118). AdGuard (CT 102) is on +> minipve at .10, not acerpve. Abiba (CT 100) is on hwepve, not amdpve. Mumuni +> (CT 114) is on hwepve (currently stopped), not minipve. Mumuni also has a +> second instance on minipve at .123 — distinguish by CT ID, not hostname. ### Checks (every 5 min) @@ -355,15 +366,17 @@ fine. Services that resolve directly to a LAN IP are NetBird-independent. | Proxmox API | minipve.sysloggh.net:8006 | 192.168.68.12 | LAN IP | No | ✅ | | LiteLLM | litellm.sysloggh.net | 192.168.68.116 | LAN IP | No | ✅ | | Authentik | auth.sysloggh.net:443 | 192.168.68.11 | CNAME → netbird | **Yes** | ⚠️ | -| Gitea | git.sysloggh.net:443 | 192.168.68.110 | CNAME → netbird | **Yes** | ⚠️ | +| Gitea | git.sysloggh.net:443 | 192.168.68.17:3000 | CNAME → netbird | **Yes** | ⚠️ | | Zulip | chat.sysloggh.net:443 | 192.168.68.19 | CNAME → netbird | **Yes** | ⚠️ VERIFY-BEFORE-USE | | Pulse | pulse.sysloggh.net:443 | 192.168.68.7 | CNAME → netbird | **Yes** | ⚠️ | -| DNS UI | dns.sysloggh.net:443 | 192.168.68.102 | CNAME → netbird | **Yes** | ⚠️ | +| DNS UI | dns.sysloggh.net:443 | 192.168.68.10:80 | CNAME → netbird | **Yes** | ⚠️ | | SearXNG | searxng.sysloggh.net:8888 | 192.168.68.7:8888 | LAN IP | No | ✅ | | Firecrawl | firecrawl.sysloggh.net:3002 | 192.168.68.7:3002 | LAN IP | No | ✅ | -**Verified 2026-07-02:** NetBird VPS rebooted after a hang; all CNAME'd +**Verified 2026-07-24:** NetBird VPS rebooted after a hang; all CNAME'd services recovered. LAN-IP-direct paths stayed up throughout the outage. +Also added `dns.sysloggh.net` route (was missing entirely). +See `scripts/netbird-add-domain.sh` for adding new proxy routes. ### 5.2 Checks (every 2 min) @@ -511,8 +524,8 @@ enforced by the `routing-regression.config_url_violations` check in Section | LiteLLM API | `http://192.168.68.116:4000` | `https://litellm.sysloggh.net` | | LiteLLM (nginx) | `http://192.168.68.116` | — | | Grafana | `http://192.168.68.116:3001` | — | -| Authentik | `https://192.168.68.11` | `https://auth.sysloggh.net` | -| Gitea | `http://192.168.68.110:3000` | `https://git.sysloggh.net` | +| Authentik | `https://192.168.68.11:9000` | `https://auth.sysloggh.net` | +| Gitea | `http://192.168.68.17:3000` | `https://git.sysloggh.net` | | Zulip API | `http://192.168.68.19` | `https://chat.sysloggh.net` | | SearXNG | `http://192.168.68.7:8888` | — | | Firecrawl | `http://192.168.68.7:3002` | — | @@ -575,25 +588,28 @@ ssh root@192.168.68.110 "systemctl restart llama-server" | CT | Name | Node | IP | Role | Agent | |----|------|------|----|------|-------| -| 100 | abiba | amdpve | .24 | Pi agent (this host) | ✅ pi | +| CT | Name | Node | IP | Role | Agent | +|----|------|------|----|------|-------| +| 100 | abiba | **hwepve** | .24 | Pi agent | ✅ pi | | 101 | llm-gpu | acerpve | .8 | GPU RTX 3090 | ❌ | -| 102 | adguard | acerpve | — | DNS | ❌ | +| 102 | adguard | **minipve** | **.10** | DNS | ❌ | | 103 | ocu-llm | ocupve | .110 | GPU RTX 5070 | ❌ | | 104 | authentik | minipve | .11 | OIDC | ❌ | -| 105 | kagentz | amdpve | — | Agent Zero | ✅ | +| 105 | kagentz | **hwepve** | — | Agent Zero | ✅ | | 106 | ra-h-os | storepve | .65 | KG bridge | ✅ MCP | | 107 | pbs | storepve | — | Backups | ❌ | | 108 | media | storepve | — | Media | ❌ | | 109 | docker-vm | storepve | .7 | Docker host | ❌ | -| 110 | gitea | minipve | — | Git | ❌ | +| 110 | gitea | minipve | **.17** | Git | ❌ | | 111 | tdunna | amdpve | .129 | Hermes agent | ✅ | | 112 | tanko | amdpve | .122 | Hermes agent | ✅ | -| 113 | baggy | amdpve | ? | Hermes agent | ✅ | -| 114 | mumuni | minipve | .123 | Hermes agent | ✅ | -| 115 | scottdenya | amdpve | — | ? | ❌ | +| 113 | baggy | amdpve | .114 | Hermes agent | ✅ | +| 114 | mumuni | **hwepve** | .123 | Hermes agent (stopped) | ✅ | +| 115 | scottdenya | amdpve | .75 | Denya OneCare | ❌ | | 116 | syslog-api | minipve | .116 | LiteLLM + Grafana | ❌ | -| 117 | zulip | storepve | — | Chat | ❌ | -| 118 | jitsi | minipve | — | Video | ❌ | +| 117 | zulip | storepve | .19 | Chat | ❌ | +| 118 | jdownloader | storepve | — | JDownloader container | ❌ | +| 119 | infisical-vault | minipve | — | Vault | ❌ | ## Appendix C: Docker Compose Files Location @@ -613,7 +629,7 @@ Source of truth: `/root/scripts/pct-run.sh` or `prose-contracts/scripts/pct-run. | CT | Name | Node | pct-run | |-----|------|------|---------| -| 100 | abiba | amdpve | `pct-run 100` | +| 100 | abiba | hwepve | `pct-run 100` | | 105 | kagentz | amdpve | `pct-run 105` | | 111 | tdunna | amdpve | `pct-run 111` | | 112 | tanko | amdpve | `pct-run 112` | @@ -627,13 +643,14 @@ Source of truth: `/root/scripts/pct-run.sh` or `prose-contracts/scripts/pct-run. | 107 | proxmox-backup | storepve | `pct-run 107` | | 108 | media | storepve | `pct-run 108` | | 117 | zulip | storepve | `pct-run 117` | -| 102 | adguard | acerpve | `pct-run 102` | +| 102 | adguard | **minipve** | `pct-run 102` | GPU bare-metal hosts (.8 acerpve, .110 ocupve, .15 amdpve) are NOT CTs — use SSH directly: ```bash ssh root@192.168.68.8 # RTX 3090 ssh root@192.168.68.110 # RTX 5070 ssh root@192.168.68.15 # Strix Halo +ssh root@192.168.68.4 # hwepve (abiba, kagentz, mumuni) ``` ## Section 7: Agent Health Check (consolidated — 2026-07-05) diff --git a/infrastructure-update.prose.md b/infrastructure-update.prose.md index 014044b..3ef4eae 100644 --- a/infrastructure-update.prose.md +++ b/infrastructure-update.prose.md @@ -91,17 +91,54 @@ Before ANY update wave: - Firecrawl test: `curl :3002/` - SearXNG test: `curl :8888` -## Wave 4: Proxmox Kernel Reboot (if needed) +## Wave 4: Proxmox Kernel Reboot Only if `[ -f /var/run/reboot-required ]` on any node. | Target | Action | |--------|--------| | Affected PVE node | Verify all CTs/VMs migrated or stopped | -| | `reboot` via PVE API | +| | `reboot` via PVE API (or `systemctl reboot -f` if dbus fails) | | | Wait 120s for node to come back | | | Start any stopped CTs | +### Post-reboot sweep (known gaps) + +After every node reboot, run these checks: + +1. **CT auto-start sweep** — LXC containers sometimes don't start despite + `onboot: 1`. Check every CT on the rebooted node and start any left stopped: + ```bash + pct list | awk '/stopped/{print $1}' | xargs -I{} pct start {} + ``` + Known cases: scottdenya (CT 115 on amdpve), authentik (CT 104 on minipve). + +2. **Zulip recovery** — When docker-vm or storepve reboots, the Zulip main + container loses its Docker network assignment (SIGKILL during storage + outage detaches it from `zulip_default` network). Run: + ```bash + ssh root@192.168.68.19 + docker rm -f zulip-zulip-1 + cd /opt/zulip && docker compose up -d + ``` + The compose restart recreates the container on the correct network. + +3. **docker-vm Docker daemon** — After reboot, Docker can take 3-4 minutes + to become `active`. The docker-proxy for Pulse (port 7655) starts early, + so Pulse is accessible before `docker ps` reports ready. Wait for Docker + before checking other stacks. + +### VPS ↔ docker-vm tunnel + +After any VPS or docker-vm reboot, verify the dedicated WireGuard tunnel: +```bash +ssh root@72.61.0.17 'wg show wg1' | grep "latest handshake" +# If no handshake in >60s: +ssh root@72.61.0.17 'wg-quick up wg1' +``` +The tunnel uses PersistentKeepalive=25 and is systemd-enabled, but should +be verified after a reboot. + ## Rollback Protocol If ANY verification fails: diff --git a/scripts/netbird-add-domain.sh b/scripts/netbird-add-domain.sh new file mode 100644 index 0000000..c3c5211 --- /dev/null +++ b/scripts/netbird-add-domain.sh @@ -0,0 +1,65 @@ +#!/bin/bash +# Netbird Reverse Proxy — Add a new domain route +# +# Usage: netbird-add-domain.sh [port] [protocol] +# +# Example: +# netbird-add-domain.sh dns.sysloggh.net 192.168.68.10 80 +# +# This script adds a domain to the Netbird proxy by inserting records +# directly into the management server's SQLite database, then restarting +# the proxy stack. +# +# Prerequisites: SSH root access to 72.61.0.17 +# sqlite3 available on VPS +# +# Requires: The domain must already have a DNS CNAME to netbird.sysloggh.net +# pointing to 72.61.0.17. + +set -euo pipefail + +DOMAIN="${1:?Usage: netbird-add-domain.sh [port] [protocol]}" +BACKEND_IP="${2:?Usage: netbird-add-domain.sh [port] [protocol]}" +PORT="${3:-80}" +PROTOCOL="${4:-http}" + +VPS="root@72.61.0.17" +DB_VOLUME="/var/lib/docker/volumes/root_netbird_data/_data" +DB="$DB_VOLUME/store.db" + +echo "=== Adding Netbird proxy route ===" +echo "Domain: $DOMAIN" +echo "Backend: $BACKEND_IP:$PORT ($PROTOCOL)" +echo "" + +ssh "$VPS" bash << REMOTESCRIPT +set -euo pipefail + +# Generate unique ID using timestamp hash (Netbird format) +ID_SUFFIX=\$(date +%s | md5sum | head -c 16) +SVC_ID="d9\${ID_SUFFIX}ptsnc73\$(date +%s | md5sum | head -c 10)" +TGT_ID=\$(sqlite3 "$DB" "SELECT COALESCE(MAX(id), 100) + 1 FROM targets;") +ACCOUNT_ID="d88av3aptsnc73clmogg" +ZONE_ID="d8adqjaptsnc73fro5g0" + +echo "Service ID: \$SVC_ID" +echo "Target ID: \$TGT_ID" + +# Insert service +sqlite3 "$DB" "INSERT INTO services (id, account_id, name, domain, proxy_cluster, enabled, terminated, pass_host_header, rewrite_redirects, mode, source, port_auto_assigned, private) VALUES (\"\$SVC_ID\", \"\$ACCOUNT_ID\", \"$DOMAIN\", \"$DOMAIN\", \"netbird.sysloggh.net\", 1, 0, 1, 0, \"http\", \"permanent\", 0, 0);" +echo "Service: OK" + +# Insert target +sqlite3 "$DB" "INSERT INTO targets (id, account_id, service_id, host, port, protocol, target_id, target_type, enabled, skip_tls_verify, request_timeout, session_idle_timeout, agent_network, disable_access_log) VALUES (\$TGT_ID, \"\$ACCOUNT_ID\", \"\$SVC_ID\", \"$BACKEND_IP\", $PORT, \"$PROTOCOL\", \"\$ZONE_ID\", \"subnet\", 1, 0, 0, 0, 0, 0);" +echo "Target: OK" + +# Verify +sqlite3 -column "$DB" "SELECT s.name, t.host, t.port, t.protocol FROM services s JOIN targets t ON s.id=t.service_id WHERE s.name=\"$DOMAIN\";" + +echo "" +echo "Restarting proxy stack..." +cd /root && docker compose restart netbird-server 2>/dev/null +sleep 15 +docker compose restart proxy 2>/dev/null +echo "Done. Verify with: curl -sI https://$DOMAIN" +REMOTESCRIPT -- 2.54.0 From c380196fab9257abf095b45d096156e44ebcd97d Mon Sep 17 00:00:00 2001 From: root Date: Fri, 24 Jul 2026 20:11:32 +0000 Subject: [PATCH 2/5] fix: address review findings - Remove kagentz from amdpve node table (belongs on hwepve) - Fix kagentz pct-run table: hwepve (was amdpve) - Fix mumuni pct-run table: hwepve (was minipve) - Fix Zulip recovery command: wrap in single SSH call --- infrastructure-control.prose.md | 6 +++--- infrastructure-update.prose.md | 4 +--- 2 files changed, 4 insertions(+), 6 deletions(-) diff --git a/infrastructure-control.prose.md b/infrastructure-control.prose.md index e7aa436..9605bc3 100644 --- a/infrastructure-control.prose.md +++ b/infrastructure-control.prose.md @@ -105,7 +105,7 @@ description: > | Node | IP | CPU | RAM | VMs/CTs | Role | |------|----|-----|-----|---------|------| | minipve | .12 | 16C | 30GB | authentik, gitea, syslog-api, infisical-vault, jitsi | Auth, git, messaging | -| amdpve | .15 | 32C | 62GB | kagentz, tanko, tdunna, baggy, scottdenya | Agents, compute | +| amdpve | .15 | 32C | 62GB | tanko, tdunna, baggy, scottdenya | Agents, compute | | storepve | .6 | 28C | 31GB | docker-vm, ra-h-os, PBS, media, jdownloader, zulip | Docker, storage, chat | | acerpve | .9 | 28C | 31GB | llm-gpu | GPU VMs | | ocupve | .5 | 12C | 14GB | ocu-llm | GPU VMs | @@ -630,14 +630,14 @@ Source of truth: `/root/scripts/pct-run.sh` or `prose-contracts/scripts/pct-run. | CT | Name | Node | pct-run | |-----|------|------|---------| | 100 | abiba | hwepve | `pct-run 100` | -| 105 | kagentz | amdpve | `pct-run 105` | +| 105 | kagentz | **hwepve** | `pct-run 105` | | 111 | tdunna | amdpve | `pct-run 111` | | 112 | tanko | amdpve | `pct-run 112` | | 113 | baggy | amdpve | `pct-run 113` | | 115 | scottdenya | amdpve | `pct-run 115` | | 104 | authentik | minipve | `pct-run 104` | | 110 | gitea | minipve | `pct-run 110` | -| 114 | mumuni | minipve | `pct-run 114` | +| 114 | mumuni | **hwepve** | `pct-run 114` | | 116 | syslog-api | minipve | `pct-run 116` | | 106 | ra-h-os | storepve | `pct-run 106` | | 107 | proxmox-backup | storepve | `pct-run 107` | diff --git a/infrastructure-update.prose.md b/infrastructure-update.prose.md index 3ef4eae..2988026 100644 --- a/infrastructure-update.prose.md +++ b/infrastructure-update.prose.md @@ -117,9 +117,7 @@ After every node reboot, run these checks: container loses its Docker network assignment (SIGKILL during storage outage detaches it from `zulip_default` network). Run: ```bash - ssh root@192.168.68.19 - docker rm -f zulip-zulip-1 - cd /opt/zulip && docker compose up -d + ssh root@192.168.68.19 'docker rm -f zulip-zulip-1 && cd /opt/zulip && docker compose up -d' ``` The compose restart recreates the container on the correct network. -- 2.54.0 From de9adb13cf634b971e03bd97a113ed785feefe94 Mon Sep 17 00:00:00 2001 From: root Date: Sun, 26 Jul 2026 12:06:32 +0000 Subject: [PATCH 3/5] =?UTF-8?q?fix:=20agent=20health=20check=20v2=20?= =?UTF-8?q?=E2=80=94=20CT=20liveness,=20config=20validation,=20wrapper=20i?= =?UTF-8?q?ntegrity,=20vault=20emptiness?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Gaps fixed: - Koby (.129) and Koonimo (.114) now have SSH hosts — no longer skipped - Agent key lookup uses correct {NAME}_LITELLM_API_KEY format - New CT liveness check via pct status on PVE nodes - New config YAML integrity check via yaml.safe_load() - New wrapper/CLI integrity check (hermes wrapper, infisical path, hermes-real) - New vault secret non-emptiness check - Ops escalation: failures produce ALERT lines for cron capture --- hermes-agent-baseline.prose.md | 2 +- infrastructure-control.prose.md | 22 +- litellm-self-heal.prose.md | 2 +- scripts/agent-health-check.py | 359 ++++++++++++++++++++++++++------ 4 files changed, 318 insertions(+), 67 deletions(-) diff --git a/hermes-agent-baseline.prose.md b/hermes-agent-baseline.prose.md index 4361c35..8653068 100644 --- a/hermes-agent-baseline.prose.md +++ b/hermes-agent-baseline.prose.md @@ -253,7 +253,7 @@ EnvironmentFile=/etc/environment # sources LITELLM_API_KEY Run the consolidated health check: ```bash -python3 /root/scripts/agent-health-check.py +python3 /root/scripts/agent-health-check.py # v2: now checks all 5 agents including Koby/Koonimo SSH, CT liveness, config YAML, wrapper integrity, vault non-emptiness ``` This validates all 4 LiteLLM keys, detects GPU port conflicts (ghost processes), verifies gateway liveness, confirms Zulip streaming (`edit_message` present), diff --git a/infrastructure-control.prose.md b/infrastructure-control.prose.md index 9605bc3..2e88411 100644 --- a/infrastructure-control.prose.md +++ b/infrastructure-control.prose.md @@ -653,24 +653,32 @@ ssh root@192.168.68.15 # Strix Halo ssh root@192.168.68.4 # hwepve (abiba, kagentz, mumuni) ``` -## Section 7: Agent Health Check (consolidated — 2026-07-05) +## Section 7: Agent Health Check v2 (2026-07-26) -Replaces 7 scattered Zulip health scripts with a single non-disruptive check -running every 10 minutes via cron (`/root/scripts/agent-health-check.py`). +Single non-disruptive check running every 10 minutes via cron +(`/root/scripts/agent-health-check.py`). v2 fixes critical gaps: +- Koby (.129) and Koonimo (.114) now have SSH hosts — no longer skipped +- Agent-specific vault keys (`{NAME}_LITELLM_API_KEY`) not shared master key +- CT liveness check via `pct status` on PVE nodes +- Config YAML integrity check +- Wrapper/CLI integrity check +- Vault secret non-emptiness check The script is **read-only** — it never restarts, kills, or modifies anything. -Disruptive cron-based gateways restarts (like Mumuni's zulip-watchdog.sh, -which was kill+nohup outside systemd) are banned by policy. ### Checks Performed | Check | Frequency | What It Detects | |-------|-----------|-----------------| -| LiteLLM key validation | 10 min | All 4 agent keys authenticate and return models | +| LiteLLM key validation | 10 min | All 5 agent-specific keys authenticate (not shared master key) | | GPU port conflict | 10 min | Ghost processes squatting port 8080 (ss vs systemd MainPID) | -| Gateway liveness | 10 min | Gateway process running, state file readable | +| Gateway liveness | 10 min | Gateway process running, state file readable (all 5 agents) | | Zulip streaming | 10 min | `edit_message` present in adapter (streaming supported) | | Recent errors | 10 min | Error count in journald for last 10 min | +| CT liveness | 10 min | `pct status` on PVE nodes — catches stopped CTs | +| Config YAML integrity | 10 min | Python `yaml.safe_load()` — catches syntax errors | +| Wrapper/CLI integrity | 10 min | hermes wrapper exists, infisical path correct, hermes-real reachable | +| Vault secrets | 10 min | Agent-specific vault keys are non-empty and start with `sk-` | ### Disabled Scripts diff --git a/litellm-self-heal.prose.md b/litellm-self-heal.prose.md index 86ffdd5..949865f 100644 --- a/litellm-self-heal.prose.md +++ b/litellm-self-heal.prose.md @@ -102,7 +102,7 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2) - **Health-check script** (`/opt/inference-harness/scripts/litellm-health-check.sh` on CT 116): `gpu-fleet` check fails only on **critical** alerts (warnings are informational). Tests `strix-moe` (not `ornith-1.0-35b`). - **GPU monitor** (`/root/scripts/gpu-monitor-server.py` on pi .24): runs as **systemd unit `gpu-monitor.service`** (was bare `&` process). `gpu_count` includes Strix Halo (was 2, now 3). VRAM alert thresholds: warning 93%, critical 97% (raised from 90/95 — 128K context steady-state is ~70% on RTX 3090, not a fault). -- **Agent key monitor** (`/root/scripts/agent-health-check.py` on pi .24, cron `*/10`): reads each agent's **live** `LITELLM_API_KEY` from its gateway process env via SSH — never hardcodes keys (hardcoded keys rot on rotation and caused 9×401/30min). Fleet roster: abiba, tanko, mumuni, koby, koonimo (legacy `tdunna`/`baggy` removed — never existed). +- **Agent key monitor** (`/root/scripts/agent-health-check.py` on pi .24, cron `*/10`): v2 (2026-07-26) — reads each agent's **agent-specific** `{NAME}_LITELLM_API_KEY` from Infisical vault (not the shared master key). Covers: LiteLLM keys, GPU ports, agent gateways (all 5 agents now SSHa ble), CT liveness (pct status on PVE nodes), config.yaml YAML integrity, wrapper/CLI integrity, vault secret non-emptiness checks. Fleet roster: tanko (.122), mumuni (.123), koby (.129), koonimo (.114), abiba (.24). Legacy `tdunna`/`baggy` replaced with canonical agent hostnames. - **Stale keys cleaned**: `daily-infra-report.py` SYNTHETIC_API_KEY was stale (`sk-U_ydi3B` → 401); now reads `LITELLM_MASTER_KEY` from env. Deprecated scripts (`router-original.py`, `router-phase0-backup.py`, `apply-fixes.py`) still reference `sk-syslog-local-master-key` but do not actively poll LiteLLM. ## Maintains diff --git a/scripts/agent-health-check.py b/scripts/agent-health-check.py index 3322b45..2c302e7 100755 --- a/scripts/agent-health-check.py +++ b/scripts/agent-health-check.py @@ -1,10 +1,10 @@ #!/usr/bin/env python3 """ -/root/scripts/agent-health-check.py — Consolidated Agent Health Verification +/root/scripts/agent-health-check.py — Consolidated Agent Health Verification v2 -Single non-disruptive health check replacing 7 scattered scripts. -Verifies: LiteLLM keys, GPU port conflicts, agent Zulip streaming, -gateway liveness, and gateway log health. NEVER restarts anything. +Verifies: LiteLLM keys (agent-specific), GPU port conflicts, agent Zulip streaming, +gateway liveness, gateway log health, CT liveness, config YAML integrity, +wrapper/CLI integrity, vault secret non-emptiness. NEVER restarts anything. Usage: python3 /root/scripts/agent-health-check.py # Full check @@ -12,59 +12,40 @@ Usage: python3 /root/scripts/agent-health-check.py --quiet # Only output on failure Cron: */10 * * * * python3 /root/scripts/agent-health-check.py --quiet + +Changelog: + v2 (2026-07-26): Added CT liveness, config validation, wrapper integrity, + vault secret emptiness check. Fixed Koby/Koonimo SSH hosts and agent key + name format ({NAME}_LITELLM_API_KEY not LITELLM_API_KEY_{NAME}). + Fleet roster: tanko (.122), mumuni (.123), koby (.129), koonimo (.114), + abiba (.24). """ import subprocess, json, sys, os, time from datetime import datetime LITELLM = "http://192.168.68.116:80" +INFISICAL_PROJECT = "322fceab-39da-4854-a55a-568e76c0f13f" +INFISICAL_ENV = "prod" -def _get_agent_key(agent_name): - """Retrieve agent key from Infisical vault.""" - try: - result = subprocess.run( - ["infisical", "secrets", "get", "LITELLM_API_KEY", - "--project=agents", "--env=production", "--plain"], - capture_output=True, text=True, timeout=10 - ) - if result.returncode == 0: - return result.stdout.strip() - except Exception: - pass - - # Fallback: try exporting all secrets - try: - result = subprocess.run( - ["infisical", "export", "--project=agents", "--env=production", - "--format=dotenv"], - capture_output=True, text=True, timeout=10 - ) - if result.returncode == 0: - for line in result.stdout.splitlines(): - if line.startswith(f"LITELLM_API_KEY_{agent_name.upper()}") or \ - (line.startswith("LITELLM_API_KEY=") and agent_name == os.uname().nodename): - return line.split("=", 1)[1].strip().strip('"').strip("'") - except Exception: - pass - - return None - -# Agent keys are pulled from Infisical vault at runtime. -# The 'key' field is populated dynamically below. -AGENTS = { - "tanko": {"ct": 112, "host": "192.168.68.122", "user": "jerome"}, - "mumuni": {"ct": 114, "host": "192.168.68.123", "user": "root"}, - "koby": {"ct": 111, "host": None, "user": None}, - "koonimo": {"ct": 113, "host": None, "user": None}, +# PVE node IPs for CT liveness checks +PVE_NODES = { + "hwepve": "192.168.68.4", + "amdpve": "192.168.68.15", + "minipve": "192.168.68.12", + "storepve": "192.168.68.6", + "acerpve": "192.168.68.9", + "ocupve": "192.168.68.5", } -# Inject keys from vault -for agent_name in AGENTS: - key = _get_agent_key(agent_name) - if key: - AGENTS[agent_name]["key"] = key - else: - AGENTS[agent_name]["key"] = None +# Agent definitions: ct, host, user, pve_node, vault_key_name +AGENTS = { + "tanko": {"ct": 112, "host": "192.168.68.122", "user": "jerome", "pve": "amdpve", "vault_key": "TANKO_LITELLM_API_KEY"}, + "mumuni": {"ct": 114, "host": "192.168.68.123", "user": "root", "pve": "hwepve", "vault_key": "MUMUNI_LITELLM_API_KEY"}, + "koby": {"ct": 111, "host": "192.168.68.129", "user": "root", "pve": "amdpve", "vault_key": "KOBY_LITELLM_API_KEY"}, + "koonimo": {"ct": 113, "host": "192.168.68.114", "user": "root", "pve": "amdpve", "vault_key": "KOONIMO_LITELLM_API_KEY"}, + "abiba": {"ct": 100, "host": "192.168.68.24", "user": "root", "pve": "hwepve", "vault_key": None}, # Pi agent, no vault key +} GPU_HOSTS = { "gpu-rtx3090 (.8)": {"host": "192.168.68.8", "port": 8080, "service": "llama-server"}, @@ -74,6 +55,11 @@ GPU_HOSTS = { FAIL = [] +INFISICAL_TOKEN = os.environ.get("INFISICAL_TOKEN") +INFISICAL_API_URL = os.environ.get("INFISICAL_API_URL", "https://vault.sysloggh.net") + +# ── Helpers ────────────────────────────────────────────────────────── + def ssh(host, cmd, user="root"): """Execute a command on a remote host, return stdout or None.""" try: @@ -112,15 +98,81 @@ def http_json(url, headers=None, timeout=5): except: return None +def run_infisical(args, quiet=True): + """Run infisical CLI with env-based auth, return stdout or None.""" + env = os.environ.copy() + env["INFISICAL_API_URL"] = INFISICAL_API_URL + if INFISICAL_TOKEN: + env["INFISICAL_TOKEN"] = INFISICAL_TOKEN + try: + result = subprocess.run( + ["/usr/bin/infisical"] + args, + capture_output=True, text=True, timeout=15, env=env + ) + return result.stdout.strip() if result.returncode == 0 else None + except: + return None + +# ── KEY LOOKUP FIX ─────────────────────────────────────────────────── + +def _get_agent_key(agent_name, vault_key_name): + """Retrieve agent-specific key from Infisical vault. + + Uses {NAME}_LITELLM_API_KEY format (e.g., TANKO_LITELLM_API_KEY, + KOONIMO_LITELLM_API_KEY) which matches actual vault key names. + """ + if not vault_key_name: + return None + + # Primary: get the agent-specific key by name + key = run_infisical([ + "secrets", "get", vault_key_name, + "--projectId=" + INFISICAL_PROJECT, + "--env=" + INFISICAL_ENV, + "--plain", + ]) + if key and key.startswith("sk-"): + return key + + # Fallback: export all and search for the key name + try: + export = run_infisical([ + "export", + "--projectId=" + INFISICAL_PROJECT, + "--env=" + INFISICAL_ENV, + "--format=dotenv", + ]) + if export: + for line in export.splitlines(): + if line.startswith(vault_key_name + "="): + value = line.split("=", 1)[1].strip().strip('"').strip("'") + if value.startswith("sk-"): + return value + except: + pass + + return None + +# Inject keys from vault for each agent +for agent_name in AGENTS: + info = AGENTS[agent_name] + key = _get_agent_key(agent_name, info.get("vault_key")) + AGENTS[agent_name]["key"] = key + # ═══════════════════════════════════════════════════════════════════ -# CHECK 1: LiteLLM Key Validation +# CHECK 1: LiteLLM Key Validation (agent-specific keys) # ═══════════════════════════════════════════════════════════════════ def check_keys(): for name, agent in AGENTS.items(): + key = agent.get("key") + if not key: + print(f" ❌ {name}: NO KEY FOUND (vault empty or unreachable)") + FAIL.append(f"key:{name}:no-key") + continue data = http_json(f"{LITELLM}/v1/models", - headers={"Authorization": f"Bearer {agent['key']}"}) + headers={"Authorization": f"Bearer {key}"}) if data and data.get("data"): model = data["data"][0].get("id", "?") print(f" ✅ {name}: key valid → {model}") @@ -130,7 +182,7 @@ def check_keys(): # ═══════════════════════════════════════════════════════════════════ -# CHECK 2: GPU Port Conflict Detection +# CHECK 2: GPU Port Conflict Detection (unchanged) # ═══════════════════════════════════════════════════════════════════ def check_gpu_ports(): @@ -158,7 +210,6 @@ def check_gpu_ports(): else: print(f" ⚠️ {label}: svc={svc_status}, port owned by {port_owner}") else: - # Verify health endpoint health = ssh(host, f"curl -s --max-time 5 http://localhost:{port}/health") if health and '"status":"ok"' in health: print(f" ✅ {label}: healthy (pid={port_owner})") @@ -171,7 +222,7 @@ def check_gpu_ports(): # ═══════════════════════════════════════════════════════════════════ -# CHECK 3: Agent Gateway Liveness + Streaming +# CHECK 3: Agent Gateway Liveness + Streaming (now covers all agents) # ═══════════════════════════════════════════════════════════════════ def check_agents(): @@ -184,8 +235,11 @@ def check_agents(): print(f" ⬜ {name} (CT {ct}): cannot SSH — skip liveness check") continue - # Gateway process (exclude the infisical bash wrapper that contains the same string) + # Gateway process pid = ssh(host, "pgrep -f 'hermes_cli.main gateway run' | grep -v infisical | head -1", user=user) + if not pid: + # Try alternate binary name + pid = ssh(host, "pgrep -f 'hermes.*gateway' | grep -v infisical | grep -v bash | head -1", user=user) if not pid: print(f" ❌ {name}: GATEWAY NOT RUNNING") FAIL.append(f"gateway-down:{name}") @@ -203,7 +257,7 @@ def check_agents(): else: gw_state, zulip = "no-state-file", "?" - # Zulip streaming: does adapter have edit_message? + # Zulip streaming check adapter_paths = [ "~/.hermes/plugins/zulip-platform/adapter.py", "~/.hermes/plugins/platforms/zulip/adapter.py", @@ -220,13 +274,183 @@ def check_agents(): r"journalctl --user -u hermes-gateway --since '10 min ago' -o cat --no-pager 2>/dev/null " r"| grep -ci 'error\|traceback\|exception\|401\|403\|500' || echo 0", user=user) - recent_errors = (recent_errors or "0").strip().split("\n")[-1] # take last line + recent_errors = (recent_errors or "0").strip().split("\n")[-1] print(f" {'✅' if gw_state == 'running' and zulip == 'connected' else '⚠️'} " f"{name}: gw={gw_state} zulip={zulip} streaming={streaming} " f"errors_10m={recent_errors.strip() or '0'} pid={pid}") +# ═══════════════════════════════════════════════════════════════════ +# CHECK 4: CT Liveness (NEW) +# ═══════════════════════════════════════════════════════════════════ + +def check_ct_liveness(): + """Check that all agent CTs are running on their PVE nodes.""" + for name, agent in AGENTS.items(): + ct = agent["ct"] + pve_node = agent.get("pve") + if not pve_node: + print(f" ⬜ {name} (CT {ct}): no PVE node mapped — skip") + continue + + pve_ip = PVE_NODES.get(pve_node) + if not pve_ip: + print(f" ⬜ {name}: unknown PVE node '{pve_node}' — skip") + continue + + status = ssh(pve_ip, f"pct status {ct} 2>/dev/null", user="root") + if not status: + print(f" ❌ {name} (CT {ct} on {pve_node}): PVE UNREACHABLE") + FAIL.append(f"ct-unreachable:{name}:{pve_ip}") + elif "running" in status: + print(f" ✅ {name} (CT {ct} on {pve_node}): running") + elif "stopped" in status: + print(f" ❌ {name} (CT {ct} on {pve_node}): STOPPED") + FAIL.append(f"ct-stopped:{name}") + else: + print(f" ⚠️ {name} (CT {ct} on {pve_node}): {status.strip()}") + + +# ═══════════════════════════════════════════════════════════════════ +# CHECK 5: Config YAML Integrity (NEW) +# ═══════════════════════════════════════════════════════════════════ + +def check_config_integrity(): + """Verify agent config.yaml parses as valid YAML.""" + for name, agent in AGENTS.items(): + host = agent.get("host") + user = agent.get("user") + if not host or not user: + print(f" ⬜ {name}: cannot SSH — skip config check") + continue + + # Check YAML parses + yaml_ok = ssh(host, + "python3 -c " + '"import yaml; yaml.safe_load(open(\'/root/.hermes/config.yaml\')); print(\'OK\')" ' + "2>&1 || echo 'FAIL'", + user=user) + if not yaml_ok: + print(f" ❌ {name}: SSH UNREACHABLE (config check skipped)") + FAIL.append(f"config-unreachable:{name}") + elif "OK" in yaml_ok: + print(f" ✅ {name}: config.yaml valid YAML") + else: + print(f" ❌ {name}: config.yaml YAML ERROR — {yaml_ok[:120]}") + FAIL.append(f"config-yaml-error:{name}") + + +# ═══════════════════════════════════════════════════════════════════ +# CHECK 6: Wrapper/CLI Integrity (NEW) +# ═══════════════════════════════════════════════════════════════════ + +def check_wrapper_integrity(): + """Verify the hermes CLI wrapper exists and can reach hermes-real.""" + for name, agent in AGENTS.items(): + host = agent.get("host") + user = agent.get("user") + if not host or not user: + print(f" ⬜ {name}: cannot SSH — skip wrapper check") + continue + + # Check wrapper exists + wrapper = ssh(host, "ls -la /root/.local/bin/hermes 2>/dev/null", user=user) + if not wrapper: + # Check alternate wrapper locations + wrapper = ssh(host, "which hermes 2>/dev/null; command -v hermes 2>/dev/null", user=user) + if not wrapper: + print(f" ❌ {name}: NO HERMES CLI WRAPPER FOUND") + FAIL.append(f"wrapper-missing:{name}") + continue + else: + print(f" ⚠️ {name}: hermes at {wrapper.strip()} (not ~/.local/bin/hermes)") + + # Check wrapper has correct infisical path + infisical_path_valid = ssh(host, + "head -20 /root/.local/bin/hermes 2>/dev/null | grep -q '/usr/bin/infisical' && echo OK || echo MISS", + user=user) + if infisical_path_valid == "MISS": + # Check if infisical exists on path + inf_actual = ssh(host, "command -v infisical 2>/dev/null", user=user) + if not inf_actual: + print(f" ❌ {name}: INFISICAL NOT INSTALLED (wrapper broken)") + FAIL.append(f"wrapper-no-infisical:{name}") + else: + print(f" ⚠️ {name}: wrapper infisical path may be wrong (infisical at {inf_actual})") + FAIL.append(f"wrapper-infisical-path:{name}") + + # Check hermes-real exists + hermes_real = ssh(host, + "ls -la /root/.local/bin/hermes-real 2>/dev/null || echo MISS", + user=user) + if not hermes_real or hermes_real.strip() == "MISS": + # Check venv path + hermes_real = ssh(host, + "ls -la /usr/local/lib/hermes-agent/venv/bin/hermes 2>/dev/null || echo MISS", + user=user) + if not hermes_real or hermes_real.strip() == "MISS": + print(f" ❌ {name}: hermes-real NOT FOUND (wrapper broken)") + FAIL.append(f"wrapper-no-hermes-real:{name}") + else: + print(f" ✅ {name}: hermes-real at alt path") + + # Check the .env file has the key + env_has_key = ssh(host, + "grep -c 'LITELLM_API_KEY' /root/.hermes/.env 2>/dev/null || echo 0", + user=user) + if env_has_key and env_has_key.strip() not in ("", "0"): + print(f" ✅ {name}: wrapper + .env key present") + else: + print(f" ⚠️ {name}: .env may be missing LITELLM_API_KEY entry") + + +# ═══════════════════════════════════════════════════════════════════ +# CHECK 7: Vault Secret Non-Emptiness (NEW) +# ═══════════════════════════════════════════════════════════════════ + +def check_vault_secrets(): + """Verify agent-specific vault secrets are non-empty and start with sk-.""" + for name, agent in AGENTS.items(): + vault_key_name = agent.get("vault_key") + if not vault_key_name: + continue + + key = agent.get("key") + if not key: + print(f" ❌ {name}: vault secret {vault_key_name} MISSING or EMPTY") + FAIL.append(f"vault-empty:{name}:{vault_key_name}") + elif not key.startswith("sk-"): + print(f" ❌ {name}: vault secret {vault_key_name} WRONG FORMAT (starts '{key[:8]}...')") + FAIL.append(f"vault-bad-format:{name}:{vault_key_name}") + else: + print(f" ✅ {name}: vault {vault_key_name}=sk-...{key[-4:]}") + + +# ═══════════════════════════════════════════════════════════════════ +# DEPLOY: copy updated script to /root/scripts/ on local host +# ═══════════════════════════════════════════════════════════════════ + +def deploy_self(): + """Copy this script to /root/scripts/agent-health-check.py if out of date.""" + dest = "/root/scripts/agent-health-check.py" + try: + with open(__file__, "r") as f: + current = f.read() + if os.path.isfile(dest): + with open(dest, "r") as f: + existing = f.read() + if current == existing: + return # Already deployed + # Write new version + with open(dest, "w") as f: + f.write(current) + os.chmod(dest, 0o755) + print(f" 📦 Deployed updated script to {dest}") + except: + pass # Not fatal if deploy fails + + # ═══════════════════════════════════════════════════════════════════ # MAIN # ═══════════════════════════════════════════════════════════════════ @@ -235,8 +459,12 @@ def main(): quiet = "--quiet" in sys.argv as_json = "--json" in sys.argv + # Self-deploy to canonical location + if not quiet and "--no-deploy" not in sys.argv: + deploy_self() + if not quiet: - print(f"🏥 Agent Health Check — {datetime.now().strftime('%Y-%m-%d %H:%M UTC')}") + print(f"🏥 Agent Health Check v2 — {datetime.now().strftime('%Y-%m-%d %H:%M UTC')}") print() print("🔑 LiteLLM Keys:") @@ -249,11 +477,26 @@ def main(): print("🤖 Agent Gateways:") check_agents() + print() + + print("🖥️ CT Liveness:") + check_ct_liveness() + print() + + print("📝 Config Integrity:") + check_config_integrity() + print() + + print("🔌 Wrapper/CLI Integrity:") + check_wrapper_integrity() + print() + + print("🔐 Vault Secrets:") + check_vault_secrets() if FAIL: print(f"\n❌ {len(FAIL)} FAILURE(S): {' | '.join(FAIL)}") if quiet: - # In quiet mode, only print failures as a single alert line print(f"ALERT agent-health:{','.join(FAIL)}") elif not quiet: print("\n✅ All checks passed") -- 2.54.0 From 47aa92bee393cd0bdfa7d7ff8bbad021396efd6d Mon Sep 17 00:00:00 2001 From: root Date: Sun, 26 Jul 2026 12:14:58 +0000 Subject: [PATCH 4/5] =?UTF-8?q?fix:=20verification=20protocol=20findings?= =?UTF-8?q?=20=E2=80=94=20CT=20118=20static=20IP,=20.110=20llama-server=20?= =?UTF-8?q?restored?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Verification protocol run 2026-07-26: - CT 118 (jdownloader): DHCP had moved it to .131 post-reboot. Now set to static .20. - RTX 5070 (.110) llama-server: was stopped (disabled unit). Started and verified — gemma-4-12b responding through LiteLLM. - All 6 PVE nodes confirmed at correct IPs. - All 18 CTs confirmed at correct IPs on correct nodes. - All public endpoints responding (meet, chat, git, litellm, vault, auth). - Container counts verified: docker-vm 14 ctrs, CT 116 10 ctrs, VPS 5 ctrs. --- infrastructure-control.prose.md | 9 +++++---- 1 file changed, 5 insertions(+), 4 deletions(-) diff --git a/infrastructure-control.prose.md b/infrastructure-control.prose.md index 2e88411..92639c0 100644 --- a/infrastructure-control.prose.md +++ b/infrastructure-control.prose.md @@ -13,10 +13,11 @@ description: > against the live system. Policy fields are authoritative. See the `verify-before-mutate` skill. - **Last verified:** 2026-07-24 — corrected Gitea IP (.17 not .110), - AdGuard IP (.10 not .102), AdGuard placement (minipve not acerpve), - Abiba placement (hwepve not amdpve), added hwepve as 6th node, - added dns.sysloggh.net route. See data/learnings.md. + **Last verified:** 2026-07-26 — verification protocol run. CT 118 set to + static IP .20 (was DHCP .131). llama-server on RTX 5070 (.110) was down, + restarted and verified gemma-4-12b responding through LiteLLM. All 6 PVE + nodes confirmed. All CT hostnames/IPs confirmed except jdownloader (now + static). All public endpoints responding. See data/learnings.md. --- # Infrastructure Control Pattern -- 2.54.0 From c869e75c61545fd266565baa997fbe5e6f18cec8 Mon Sep 17 00:00:00 2001 From: root Date: Sun, 26 Jul 2026 12:24:05 +0000 Subject: [PATCH 5/5] ci: trigger recheck -- 2.54.0