diff --git a/disk-gc-threat-response.prose.md b/disk-gc-threat-response.prose.md index 14bcf33..fa0cf34 100644 --- a/disk-gc-threat-response.prose.md +++ b/disk-gc-threat-response.prose.md @@ -274,10 +274,10 @@ one-off GPU builds. No automated post-migration cleanup was in place. ### CT Access (via pct-run) | CT | Name | Node | Status | |----|------|------|--------| -| 100 | abiba | hwepve | local | +| 100 | abiba | minipve | local | | 102 | adguard | minipve | ✅ reachable | | 104 | authentik | minipve | ✅ reachable | -| 105 | kagentz | hwepve | ✅ reachable | +| 105 | kagentz | minipve | ✅ reachable | | 106 | ra-h-os | storepve | ✅ reachable | | 107 | pbs | storepve | ✅ reachable | | 108 | media | storepve | ✅ reachable | @@ -285,7 +285,6 @@ one-off GPU builds. No automated post-migration cleanup was in place. | 111 | tdunna | amdpve | ✅ reachable | | 112 | tanko | amdpve | ✅ reachable | | 113 | baggy | amdpve | ✅ reachable | -| 114 | mumuni | hwepve | ✅ reachable | | 115 | scottdenya | amdpve | ✅ reachable | | 116 | syslog-api | minipve | ✅ reachable | | 117 | zulip | storepve | ✅ reachable | diff --git a/docs/AUTHORING-GUIDE.md b/docs/AUTHORING-GUIDE.md index bd68ee9..bfe4c8b 100644 --- a/docs/AUTHORING-GUIDE.md +++ b/docs/AUTHORING-GUIDE.md @@ -185,7 +185,7 @@ what, and why should I care? ``` ❌ "Monitors infrastructure health" -✅ "Scans all 6 Proxmox nodes and 19 CTs for disk pressure, checks Docker +✅ "Scans all 5 Proxmox nodes and 19 CTs for disk pressure, checks Docker container health on .7/.116/.17, alerts via Telegram DM on RED/CRITICAL" ``` diff --git a/hermes-agent-baseline.prose.md b/hermes-agent-baseline.prose.md index 7e13e0f..6cdbb78 100644 --- a/hermes-agent-baseline.prose.md +++ b/hermes-agent-baseline.prose.md @@ -25,7 +25,7 @@ done | Agent | CT | Node | IP | LiteLLM Alias | Key Source | Platform | |-------|-----|------|-----|---------------|------------|----------| | Tanko | 112 | amdpve | .122 | `tanko` | Infisical vault | Hermes | -| Mumuni | 100 | hwepve | .24 | `mumuni` | Infisical vault | Hermes | +| Mumuni | 100 | minipve | .24 | `mumuni` | Infisical vault | Hermes | | Koby | 111 | amdpve | .129 | `koby` | Infisical vault | **Hermes** | | Koonimo | 113 | amdpve | .114 | `koonimo` | Infisical vault | Hermes | | Shumba | — | 192.168.68.119 | N/A | N/A (DeepSeek) | Hermes (RETIRED — CT119 now Infisical vault) | diff --git a/hermes-zulip-restore.prose.md b/hermes-zulip-restore.prose.md index 4a55150..c2d5b6d 100644 --- a/hermes-zulip-restore.prose.md +++ b/hermes-zulip-restore.prose.md @@ -51,7 +51,7 @@ gateway restart, and connection validation. | Host | CT | Proxmox | IP (direct) | Hermes Home | User | |------|-----|---------|-------------|-------------|------| -| Mumuni | CT100 (abiba) | hwepve | 192.168.68.24 | /root/.hermes | root | +| Mumuni | CT100 (abiba) | minipve | 192.168.68.24 | /root/.hermes | root | | Tanko | CT112 | amdpve | 192.168.68.122 | /home/jerome/.hermes | jerome | | Koby | CT111 | amdpve | 192.168.68.129 | /root/.hermes | root | | Shumba | — | — | 192.168.68.119 | /home/lucky/.hermes | lucky | diff --git a/infrastructure-control.prose.md b/infrastructure-control.prose.md index c1f02de..dbebbce 100644 --- a/infrastructure-control.prose.md +++ b/infrastructure-control.prose.md @@ -3,7 +3,7 @@ kind: pattern name: infrastructure-control description: > Full infrastructure monitoring and control pattern covering the - 6-node Proxmox cluster, 3 Docker ecosystems (22 containers), + 5-node Proxmox cluster, 3 Docker ecosystems (22 containers), NFS storage, and network services. Defines monitors, remediations, and the access matrix for all environments. @@ -13,10 +13,11 @@ description: > against the live system. Policy fields are authoritative. See the `verify-before-mutate` skill. - **Last verified:** 2026-07-24 — corrected Gitea IP (.17 not .110), - AdGuard IP (.10 not .102), AdGuard placement (minipve not acerpve), - Abiba placement (hwepve not amdpve), added hwepve as 6th node, - added dns.sysloggh.net route. + **Last verified:** 2026-08-15 — hwepve removed from Tabiri cluster + (now 5 nodes: minipve, amdpve, storepve, acerpve, ocupve). hwepve + (192.168.68.4) is a standalone PVE node + NetBird routing peer; + London relocation pending. CTs 100 (abiba) and 105 (kagentz) moved + to minipve. --- # Infrastructure Control Pattern @@ -34,21 +35,21 @@ description: > ┌─────────────┐ ┌──────────────┐ ┌──────────────┐ │ Abiba │ │ Tanko │ │ Mumuni │ │ (pi) │ │ (Hermes) │ │ (Hermes) │ - │ CT 100 │ │ CT 112 │ │ CT 114 │ + │ CT 100 │ │ CT 112 │ │ CT 100 │ └──────┬──────┘ └──────┬───────┘ └──────┬───────┘ │ │ │ └──────────────────┼────────────────────┘ ▼ - ┌──────────────────────────────────────┐ - │ Proxmox Cluster API │ - │ minipve.sysloggh.net:443 │ - │ (monitoring@pve!mumuni token) │ - └────┬──────┬──────┬──────┬──────┬─────┘ - │ │ │ │ │ - ┌────┘ ┌────┘ ┌────┘ ┌────┘ ┌────┘ - ▼ ▼ ▼ ▼ ▼ - minipve amdpve storepve acerpve ocupve hwepve - (.12) (.15) (.6) (.9) (.5) (.4) + ┌────────────────────────────────────┐ + │ Proxmox Cluster API │ + │ minipve.sysloggh.net:443 │ + │ (monitoring@pve!mumuni token) │ + └────┬──────┬──────┬──────┬──────────┘ + │ │ │ │ + ┌────┘ ┌────┘ ┌────┘ ┌────┘ + ▼ ▼ ▼ ▼ + minipve amdpve storepve acerpve ocupve + (.12) (.15) (.6) (.9) (.5) ▼ ┌─────────────────────────────────────────────┐ @@ -100,21 +101,29 @@ description: > ## Section 2: Proxmox Cluster — Monitoring -### Nodes (6) +### Nodes (5) | Node | IP | CPU | RAM | VMs/CTs | Role | |------|----|-----|-----|---------|------| -| minipve | .12 | 16C | 30GB | authentik, gitea, syslog-api, infisical-vault, jitsi | Auth, git, messaging | +| minipve | .12 | 16C | 30GB | abiba, kagentz, authentik, gitea, syslog-api, infisical-vault, jitsi | Auth, git, messaging | | amdpve | .15 | 32C | 62GB | tanko, tdunna, baggy, scottdenya | Agents, compute | | storepve | .6 | 28C | 31GB | docker-vm, ra-h-os, PBS, media, jdownloader, zulip | Docker, storage, chat | | acerpve | .9 | 28C | 31GB | llm-gpu | GPU VMs | | ocupve | .5 | 12C | 14GB | ocu-llm | GPU VMs | -| hwepve | .4 | 12C | 15GB | abiba, kagentz, (mumuni CT 114 stopped) | Agents (new node) | > **Note:** CTs on storepve include jdownloader (CT 118). AdGuard (CT 102) is on -> minipve at .10, not acerpve. Abiba (CT 100) is on hwepve, not amdpve. Mumuni -> (CT 114) is on hwepve (currently stopped), not minipve. Mumuni also has a -> second instance on minipve at .123 — distinguish by CT ID, not hostname. +> minipve at .10, not acerpve. Abiba (CT 100) and kagentz (CT 105) are on +> minipve (moved from hwepve 2026-08-15). Mumuni runs inside Abiba CT100 +> (.24); CT 114 (mumuni) no longer exists in the cluster. +> +> **hwepve (192.168.68.4) — STANDALONE (removed from Tabiri 2026-08-15):** +> Huawei MateBook 16 (KLVL-WXX9), pve-manager/9.2.10, kernel 7.0.14-8-pve. +> Zero VMs/CTs. Being relocated to London as a standalone PVE node + NetBird +> routing peer (relocation pending). Localizations applied: timezone +> Europe/London, lid-switch ignore, sleep/suspend/hibernate targets masked, +> cluster-shared storage removed (remaining: local, local-lvm, storage, +> mediastore). prometheus-node-exporter active on :9100; net.ipv4.ip_forward=1; +> NetBird client not yet installed (enrollment pending setup key). ### Checks (every 5 min) @@ -592,12 +601,12 @@ ssh root@192.168.68.110 "systemctl restart llama-server" | CT | Name | Node | IP | Role | Agent | |----|------|------|----|------|-------| -| 100 | abiba | **hwepve** | .24 | Pi agent | ✅ pi | +| 100 | abiba | minipve | .24 | Pi agent | ✅ pi | | 101 | llm-gpu | acerpve | .8 | GPU RTX 3090 | ❌ | | 102 | adguard | **minipve** | **.10** | DNS | ❌ | | 103 | ocu-llm | ocupve | .110 | GPU RTX 5070 | ❌ | | 104 | authentik | minipve | .11 | OIDC | ❌ | -| 105 | kagentz | **hwepve** | — | Agent Zero | ✅ | +| 105 | kagentz | minipve | — | Agent Zero | ✅ | | 106 | ra-h-os | storepve | .65 | KG bridge | ✅ MCP | | 107 | pbs | storepve | — | Backups | ❌ | | 108 | media | storepve | — | Media | ❌ | @@ -606,7 +615,6 @@ ssh root@192.168.68.110 "systemctl restart llama-server" | 111 | tdunna | amdpve | .129 | Hermes agent | ✅ | | 112 | tanko | amdpve | .122 | Hermes agent | ✅ | | 113 | baggy | amdpve | .114 | Hermes agent | ✅ | -| 114 | mumuni | **hwepve** | .123 | Hermes agent (stopped) | ✅ | | 115 | scottdenya | amdpve | .75 | Denya OneCare | ❌ | | 116 | syslog-api | minipve | .116 | LiteLLM + Grafana | ❌ | | 117 | zulip | storepve | .19 | Chat | ❌ | @@ -631,15 +639,14 @@ Source of truth: `/root/scripts/pct-run.sh` or `prose-contracts/scripts/pct-run. | CT | Name | Node | pct-run | |-----|------|------|---------| -| 100 | abiba | hwepve | `pct-run 100` | -| 105 | kagentz | hwepve | `pct-run 105` | +| 100 | abiba | minipve | `pct-run 100` | +| 105 | kagentz | minipve | `pct-run 105` | | 111 | tdunna | amdpve | `pct-run 111` | | 112 | tanko | amdpve | `pct-run 112` | | 113 | baggy | amdpve | `pct-run 113` | | 115 | scottdenya | amdpve | `pct-run 115` | | 104 | authentik | minipve | `pct-run 104` | | 110 | gitea | minipve | `pct-run 110` | -| 114 | mumuni | hwepve | `pct-run 114` | | 116 | syslog-api | minipve | `pct-run 116` | | 106 | ra-h-os | storepve | `pct-run 106` | | 107 | proxmox-backup | storepve | `pct-run 107` | @@ -652,7 +659,7 @@ GPU bare-metal hosts (.8 acerpve, .110 ocupve, .15 amdpve) are NOT CTs — use S ssh root@192.168.68.8 # RTX 3090 ssh root@192.168.68.110 # RTX 5070 ssh root@192.168.68.15 # Strix Halo -ssh root@192.168.68.4 # hwepve (abiba, kagentz, mumuni) +ssh root@192.168.68.4 # hwepve — standalone London node + NetBird routing peer (relocation pending) ``` ## Section 7: Agent Health Check (consolidated — 2026-07-05) diff --git a/infrastructure-update.prose.md b/infrastructure-update.prose.md index d94ed50..1c5fe41 100644 --- a/infrastructure-update.prose.md +++ b/infrastructure-update.prose.md @@ -2,7 +2,7 @@ kind: responsibility name: infrastructure-update description: > - Autonomous system-wide update contract covering all 6 Proxmox nodes, + Autonomous system-wide update contract covering all 5 Proxmox nodes, 15+ containers/VMs, and 4 Docker ecosystems. Updates apt packages, Docker images, and container stacks in safe waves with health checks and automatic rollback on failure. @@ -56,11 +56,10 @@ Before ANY update wave: | amdpve (.15) | Proxmox node | `apt update && apt upgrade -y` | 5 min | | acerpve (.9) | Proxmox node | `apt update && apt upgrade -y` | 5 min | | ocupve (.5) | Proxmox node | `apt update && apt upgrade -y` | 5 min | -| hwepve (.4) | Proxmox node | `apt update && apt upgrade -y` | 5 min | | CT 100 (.24) | Abiba (pi) | `apt update && apt upgrade -y` | 3 min | | CT 116 (.116) | syslog-api (LiteLLM host) | `apt update && apt upgrade -y` | 3 min | | CT 112 (tanko, amdpve) | Tanko | `apt update && apt upgrade -y` | 3 min | -| CT 100 (mumuni/abiba, hwepve) | Mumuni | `apt update && apt upgrade -y` | 3 min | +| CT 100 (mumuni/abiba, minipve) | Mumuni | `apt update && apt upgrade -y` | 3 min | | VM 101 (.8) | llm-gpu (RTX 3090) | `apt update && apt upgrade -y` | 3 min | | VM 103 (.110) | ocu-llm (RTX 5070) | `apt update && apt upgrade -y` | 3 min | @@ -163,9 +162,9 @@ Before Wave 1, snapshot these files: /etc/systemd/system/strix-server.service (amdpve .15 — strix-moe) /etc/systemd/system/llama-server.service (VM 101 .8, VM 103 .110) # Hermes agent configs (key enforcement — 2026-07-10) -/root/.hermes/config.yaml (Mumuni CT 114, Tanko CT 112, etc.) -/root/.config/systemd/user/hermes-gateway.service (Mumuni CT 114 — EnvironmentFile fixed) -/etc/environment (Mumuni CT 114 — LITELLM_API_KEY) +/root/.hermes/config.yaml (Mumuni inside CT 100, Tanko CT 112, etc.) +/root/.config/systemd/user/hermes-gateway.service (Mumuni inside CT 100 — EnvironmentFile fixed) +/etc/environment (Mumuni inside CT 100 — LITELLM_API_KEY) ``` ## MCP Gateway (2026-07-10) @@ -219,7 +218,7 @@ When LiteLLM is upgraded to a version supporting per-key MCP grants: ## Success Criteria -- [ ] All 6 PVE nodes updated, no reboot-loop +- [ ] All 5 PVE nodes updated, no reboot-loop - [ ] All VMs/CTs running post-update - [ ] All Docker containers healthy (VM 109 + CT 116 + CT 117) - [ ] LiteLLM inference passing (syslog-auto test) @@ -235,7 +234,7 @@ After completion, send Zulip DM: ``` 📋 Infrastructure Update — YYYY-MM-DD -Updated: 6 PVE nodes, 12 CTs/VMs, 30+ containers +Updated: 5 PVE nodes, 12 CTs/VMs, 30+ containers Security fixes: N CVEs patched Downtime: Failures: none /
diff --git a/litellm-health.prose.md b/litellm-health.prose.md index 9b8c904..2cd8792 100644 --- a/litellm-health.prose.md +++ b/litellm-health.prose.md @@ -26,7 +26,7 @@ description: > Scraped by Prometheus with Bearer master key; endpoint returns 307 → /metrics/. - Alertmanager (harness-alertmanager :9093) + zulip-bridge (:9102) deliver firing alerts to #agent-hub > alerts-infra via abiba-bot. Added 2026-08-09. - - Prometheus node job covers ALL 6 PVE nodes (.4/.5/.6/.9/.12/.15:9100). + - Prometheus node job covers ALL 5 PVE nodes (.5/.6/.9/.12/.15:9100). --- ## Architecture (v4.0.0 — Direct: nginx → LiteLLM → GPU) diff --git a/mumuni-delegation-prose-contract.prose.md b/mumuni-delegation-prose-contract.prose.md index 5921059..eec0b4b 100644 --- a/mumuni-delegation-prose-contract.prose.md +++ b/mumuni-delegation-prose-contract.prose.md @@ -6,7 +6,7 @@ description: > delegation, verification, and delivery. Defines when to delegate, which worker to use for what, how to handle failures, and the kanban board protocol. Enforces context-window discipline and separation of concerns. - Runs on Mumuni (inside Abiba CT100, hwepve, .24) via Hermes agent (Pi + Hermes Zulip gateway). + Runs on Mumuni (inside Abiba CT100, minipve, .24) via Hermes agent (Pi + Hermes Zulip gateway). version: 1.0.0 --- @@ -19,8 +19,8 @@ version: 1.0.0 ## Topology -**Cluster:** 6 Proxmox nodes (ocupve, acerpve, minipve, amdpve, storepve, hwepve) -**Manager:** Mumuni (inside Abiba CT100, hwepve, .24) via Hermes agent +**Cluster:** 5 Proxmox nodes (ocupve, acerpve, minipve, amdpve, storepve) +**Manager:** Mumuni (inside Abiba CT100, minipve, .24) via Hermes agent **Workers:** 6 profiles, all running on the same agent — no separate hosts needed This contract is infrastructure-agnostic in terms of which nodes are used. @@ -31,7 +31,7 @@ Workers execute tasks on whatever infrastructure they're given — SSH to .6, ## Why This Matters Without enforced delegation, the manager consumes the full iteration budget -(60 calls) on single-turn tasks — SSH to 6 nodes, check each VM, read logs — +(60 calls) on single-turn tasks — SSH to 5 nodes, check each VM, read logs — leaving no capacity for actual coordination. The result: context overflow (59K tokens in system prompt), iteration exhaustion, and degraded response quality. This contract exists because I blew through my budget checking @@ -82,7 +82,7 @@ it asks the manager (via relay) — it doesn't go find it on its own. **This is a hard rule, not a recommendation.** Violating it produces the exact type of discrepancy the kanban pipeline exists to prevent: a review worker finds -"6 nodes present" in the raw data but "6/6 online" in the report — even though +"5 nodes present" in the raw data but "5/5 online" in the report — even though one of those nodes was unreachable. The report lied because it used data the raw data never provided. @@ -137,7 +137,7 @@ delegate_task( ``` delegate_task( tasks=[ - {"goal": "Check all 6 Proxmox nodes for VM status", "context": "SSH to each node via 192.168.68.x, run 'qm list'"}, + {"goal": "Check all 5 Proxmox nodes for VM status", "context": "SSH to each node via 192.168.68.x, run 'qm list'"}, {"goal": "Check Docker container health on .7/.116/.17", "context": "SSH to each host, check container status"}, ] ) @@ -191,7 +191,7 @@ Only verified results reach Kwame. Format per channel: { "lane_id": "devops-check", "worker": "syslog-devops", - "goal": "Check all 6 Proxmox nodes", + "goal": "Check all 5 Proxmox nodes", "status": "dispatched|completed|failed", "output_file": "/tmp/node-report.md" } diff --git a/proxmox-monitor.prose.md b/proxmox-monitor.prose.md index bc8819c..90fadaa 100644 --- a/proxmox-monitor.prose.md +++ b/proxmox-monitor.prose.md @@ -5,7 +5,7 @@ description: > Proxmox cluster + Docker monitoring via the existing Grafana/Prometheus stack on CT 116. Replaces Pulse with file-provisioned Grafana dashboards. Three exporters feed Prometheus: prometheus-pve-exporter (cluster-aware, single - instance), node_exporter (all 6 PVE nodes), and a custom docker-stats-exporter + instance), node_exporter (all 5 PVE nodes), and a custom docker-stats-exporter (Docker 29 / containerd image-store compatible, since cAdvisor cannot resolve the layerdb). Dashboards exposed at http://192.168.68.116:3001/ (direct LAN, not behind nginx). agent: abiba @@ -33,8 +33,8 @@ agent: abiba | Exporter | Host:Port | Scope | Notes | |----------|-----------|-------|-------| -| prometheus-pve-exporter | .116:9221 (container) | All 6 nodes + guests + 36 storage pools | Single instance, cluster-aware via amdpve API. Config `/opt/monitoring/pve.yml` (token `monitoring@pve!prometheus`, PVEAuditor role). Metric schema is label-based (`id=node/amdpve`, `id=lxc/100`). | -| node_exporter | .5/.6/.9/.12/.15/.4:9100 (systemd) | Per-node CPU/mem/disk/net/temp | Installed via apt on all 6 PVE nodes, enabled (reboot-persistent). Collectors: textfile, systemd, tcpstat, ethtool. hwepve (.4) added 2026-07-19. | +| prometheus-pve-exporter | .116:9221 (container) | All 5 nodes + 14 guests + 36 storage pools | Single instance, cluster-aware via amdpve API. Config `/opt/monitoring/pve.yml` (token `monitoring@pve!prometheus`, PVEAuditor role). Metric schema is label-based (`id=node/amdpve`, `id=lxc/100`). | +| node_exporter | .5/.6/.9/.12/.15:9100 (systemd) | Per-node CPU/mem/disk/net/temp | Installed via apt on all 5 PVE nodes, enabled (reboot-persistent). Collectors: textfile, systemd, tcpstat, ethtool. | | docker-stats-exporter | .116:9324 (container) | 10 Docker containers on .116 | **Custom** (cAdvisor v0.51 incompatible with Docker 29 containerd image store — layerdb gone). Uses Docker Engine API over unix socket. Script `/opt/monitoring/docker-stats-exporter.py`. | ## PVE API Token @@ -48,7 +48,7 @@ agent: abiba | UID | Title | Panels | Source | |-----|-------|--------|--------| -| proxmox-cluster | Proxmox Cluster Overview | 16 | cluster status, 6-node CPU/mem/disk/load gauges, guests table, storage pools, guest CPU/mem timeseries | +| proxmox-cluster | Proxmox Cluster Overview | 16 | cluster status, 5-node CPU/mem/disk/load gauges, guests table, storage pools, guest CPU/mem timeseries | | proxmox-node | Proxmox Node Detail | 13 | per-node CPU per-core, memory, network, disk IO/IOPS/latency, temperature, disk space (variable: $node) | | docker-containers | Docker Containers | 10 | per-container CPU/mem/network, restarts, memory limit ratio (variable: $container) | | gpu-fleet | GPU Fleet | 7 | (existing, preserved in DB, not provisioned) | @@ -77,9 +77,9 @@ agent: abiba | `/opt/monitoring/grafana/dashboards/build-dashboards.py` | .116 | dashboard JSON generator | | `/opt/monitoring/grafana/dashboards/json/*.json` | .116 | provisioned dashboard definitions | | `/opt/monitoring/grafana/datasources/prometheus.yml` | .116 | datasource provisioning | -| `/etc/default/prometheus-node-exporter` | .5/.6/.9/.12/.15/.4 | node_exporter collector config | +| `/etc/default/prometheus-node-exporter` | .5/.6/.9/.12/.15 | node_exporter collector config | -## Cluster "Tabiri" — 6 Nodes +## Cluster "Tabiri" — 5 Nodes | Node | IP | Role | |------|----|----| @@ -88,7 +88,6 @@ agent: abiba | acerpve | 192.168.68.9 | PVE (hosts llm-gpu qemu/101) | | minipve | 192.168.68.12 | PVE | | amdpve | 192.168.68.15 | PVE + Strix Halo LLM (qwen3.6-35B-udq4, strix-moe) | -| hwepve | 192.168.68.4 | PVE (Huawei Matebook 16, 12C/15GB) — hosts abiba (lxc/100), kagentz (lxc/105), mumuni (lxc/114). CTs 100/105 migrated from amdpve, CT 114 from minipve 2026-07-20 | ## Operations diff --git a/scripts/agent-health-check.py b/scripts/agent-health-check.py index f9c85d6..2929669 100755 --- a/scripts/agent-health-check.py +++ b/scripts/agent-health-check.py @@ -30,7 +30,6 @@ INFISICAL_ENV = "prod" # PVE node IPs for CT liveness checks PVE_NODES = { - "hwepve": "192.168.68.4", "amdpve": "192.168.68.15", "minipve": "192.168.68.12", "storepve": "192.168.68.6", @@ -41,7 +40,7 @@ PVE_NODES = { # Agent definitions: ct, host, user, pve_node, vault_key_name AGENTS = { "tanko": {"ct": 112, "host": "192.168.68.122", "user": "jerome", "pve": "amdpve", "vault_key": "TANKO_LITELLM_API_KEY"}, - "abiba": {"ct": 100, "host": "192.168.68.24", "user": "root", "pve": "hwepve", "vault_key": None}, # Pi agent + Mumuni Zulip, no vault key + "abiba": {"ct": 100, "host": "192.168.68.24", "user": "root", "pve": "minipve", "vault_key": None}, # Pi agent + Mumuni Zulip, no vault key "koby": {"ct": 111, "host": "192.168.68.129", "user": "root", "pve": "amdpve", "vault_key": "KOBY_LITELLM_API_KEY"}, "koonimo": {"ct": 113, "host": "192.168.68.114", "user": "root", "pve": "amdpve", "vault_key": "KOONIMO_LITELLM_API_KEY"}, } diff --git a/scripts/daily-infra-report.py b/scripts/daily-infra-report.py index ada9523..23bf7d2 100755 --- a/scripts/daily-infra-report.py +++ b/scripts/daily-infra-report.py @@ -322,7 +322,7 @@ def collect(): mumuni_data = {} mumuni_platforms = mumuni_data.get("platforms", {}) report["agents"]["mumuni"] = { - "platform": "hermes", "ct": 114, "ip": "192.168.68.24", + "platform": "hermes", "ct": 100, "ip": "192.168.68.24", "gateway_state": mumuni_data.get("gateway_state", "unknown"), "telegram_state": mumuni_platforms.get("telegram", {}).get("state", "unknown"), "zulip_state": mumuni_platforms.get("zulip", {}).get("state", "not_installed"), diff --git a/scripts/pct-run.sh b/scripts/pct-run.sh index ebf6b4c..afbf56a 100755 --- a/scripts/pct-run.sh +++ b/scripts/pct-run.sh @@ -28,10 +28,8 @@ declare -A CT_NODES=( [117]=storepve # zulip [118]=storepve # jdownloader # acerpve (192.168.68.9) — no CTs (bare metal GPU .8) - # hwepve (192.168.68.4) - [100]=hwepve # abiba (was amdpve) - [105]=hwepve # kagentz (was amdpve) - [114]=hwepve # mumuni (was minipve) + [100]=minipve # abiba (was hwepve) + [105]=minipve # kagentz (was hwepve) # ocupve (192.168.68.5) — no CTs (bare metal GPU .110) # # REMOVED CTs (migrated to bare metal, decommissioned, or VMs): @@ -50,7 +48,6 @@ declare -A NODE_IPS=( [storepve]=192.168.68.6 [acerpve]=192.168.68.9 [ocupve]=192.168.68.5 - [hwepve]=192.168.68.4 ) resolve_node() { @@ -77,7 +74,7 @@ main() { if [[ $# -lt 1 ]]; then echo "Usage: pct-run [command...]" >&2 echo " pct-run 112 cat /etc/hostname" >&2 - echo " pct-run 114 systemctl status hermes-gateway" >&2 + echo " pct-run 100 systemctl status hermes-gateway" >&2 echo "" echo "Known CTs:" >&2 for ct in $(echo "${!CT_NODES[@]}" | tr ' ' '\n' | sort -n); do diff --git a/scripts/prose-ai-review.sh b/scripts/prose-ai-review.sh index ab5228b..eaf8a4c 100755 --- a/scripts/prose-ai-review.sh +++ b/scripts/prose-ai-review.sh @@ -46,18 +46,17 @@ You are a code reviewer for OpenProse infrastructure contracts in the Syslog Sol The infrastructure-control.prose.md contract is the canonical reference for the cluster topology: -**Proxmox Cluster "Tabiri" (6 nodes):** +**Proxmox Cluster "Tabiri" (5 nodes):** - amdpve (192.168.68.15): tanko, tdunna, baggy, scottdenya -- minipve (192.168.68.12): adguard, authentik, gitea, syslog-api, infisical-vault +- minipve (192.168.68.12): abiba, kagentz, adguard, authentik, gitea, syslog-api, infisical-vault - storepve (192.168.68.6): docker-vm, ra-h-os, PBS, media, jdownloader, zulip - acerpve (192.168.68.9): llm-gpu - ocupve (192.168.68.5): ocu-llm -- hwepve (192.168.68.4): abiba, kagentz, mumuni **CT IDs (verified 2026-07-24 against PVE API):** 100:abiba 102:adguard 104:authentik 105:kagentz 106:ra-h-os 107:pbs 108:media 110:gitea 111:tdunna 112:tanko -113:baggy 114:mumuni 115:scottdenya 116:syslog-api 117:zulip +113:baggy 115:scottdenya 116:syslog-api 117:zulip 118:jdownloader 119:infisical-vault **NO CT 122, CT 123, or .19 exist in the cluster.** @@ -65,7 +64,7 @@ The infrastructure-control.prose.md contract is the canonical reference for the **CRITICAL RULES (never regress):** 1. NO /grafana/ nginx route — it was tried and reverted on 2026-07-02. Grafana is direct LAN at :3001. 2. NO .19 IP — Zulip is CT 117 on storepve. -3. NO CT 122/123 — Tanko=CT 112, Mumuni=CT 114. +3. NO CT 122/123 — Tanko=CT 112, Mumuni=CT 100 (inside Abiba). No CT 114 anywhere. 4. Strix Halo :8080 is FIREWALLED to .116 only — cannot be probed from abiba (.24). 5. abiba-zulip PM2 process is DECOMMISSIONED (2026-07-04) — abiba uses Telegram only. diff --git a/zulip-health.prose.md b/zulip-health.prose.md index 6173222..2f32e42 100644 --- a/zulip-health.prose.md +++ b/zulip-health.prose.md @@ -16,7 +16,7 @@ Runs every 15 minutes in the background. Also triggers on session start. ## Requires - **Zulip API key** for `abiba-bot@chat.sysloggh.net` in `$ZULIP_API_KEY` -- **SSH access** to Tanko (192.168.68.122), Mumuni (192.168.68.24, inside Abiba CT100 on hwepve), and Agent Zero Docker host (192.168.68.14) +- **SSH access** to Tanko (192.168.68.122), Mumuni (192.168.68.24, inside Abiba CT100 on minipve), and Agent Zero Docker host (192.168.68.14) - **PM2** on localhost for pi process management - **Network access** to `chat.sysloggh.net`, `localhost:9200` - **Write access** to `/root/zulip-health-monitor.log` and `/tmp/zulip-monitor-debounce` @@ -87,7 +87,7 @@ streamed in real-time via Zulip's `PATCH /api/v1/messages/{id}` API: - Adapter implements `edit_message()` using `_api_patch()` helper - Gateway stream consumer progressively edits the Zulip message - User sees real-time agent thinking instead of waiting for full response -- Verified: Tanko (CT 112) and Mumuni (CT 114) both have streaming active +- Verified: Tanko (CT 112) and Mumuni (inside Abiba CT 100) both have streaming active ### Verification ```bash diff --git a/zulip-resilience-v3.prose.md b/zulip-resilience-v3.prose.md index 4825441..d2c0b19 100644 --- a/zulip-resilience-v3.prose.md +++ b/zulip-resilience-v3.prose.md @@ -420,7 +420,7 @@ Backup v2 before starting: `cp index.js index.js.v2-backup-$(date +%Y%m%d-%H%M%S |-------|----------|-------------|--------------|-------------| | **Abiba** | pi (CT 100) | ✅ Connected | API key missing from Infisical injection; poll timeout noise | Added .env fallback; AbortError treated as empty poll (no retry); poll timeout 65s→90s | | **Tanko** | Hermes (CT 112) | ✅ Connected | Gateway disconnected since Jul 11; watchdog restart didn't re-establish Zulip | Full gateway restart (kill wrapper, let infisical-gateway.sh respawn) | -| **Mumuni** | Hermes (CT 114) | ✅ Connected | No issues found | None needed | +| **Mumuni** | Hermes (inside Abiba CT 100) | ✅ Connected | No issues found | None needed | ### Key Fixes Applied