PR #117 follow-up (verify PASS-WITH-FINDINGS): 1. infrastructure-update.prose.md: - Update access table: agent keys now have per-key MCP grants (2026-09-18) - Strike-through old limitation: per-key grants now work - Mark Migration Path as COMPLETED 2026-09-18 2. hermes-config-template.prose.md: - Remove hedge ('may have been upgraded') - State fact: per-key MCP grants verified 2026-09-18 This resolves the contradiction where one file asserted per-key MCP access and the other denied it.
14 KiB
kind, name, description, agent, triggers, version
| kind | name | description | agent | triggers | version | |||
|---|---|---|---|---|---|---|---|---|
| responsibility | infrastructure-update | Autonomous system-wide update contract covering all 5 Proxmox nodes, 15+ containers/VMs, and 5 Docker ecosystems (docker-vm .7, CT 116 .116, CT 117, hwpve .11, NetBird VPS 72.61.0.17). Updates apt packages, Docker images, and container stacks in safe waves with health checks and automatic rollback on failure. | abiba |
|
1.4.0 |
Maintains
- update-status: { phase, node, action, result, timestamp }
- update-history: array of past update runs with results
- security-state: { cve_count, last_patched, pending_updates }
Pre-Flight Checklist
Before ANY update wave:
- ✅ All Proxmox nodes online (
GET /api2/json/nodes) - ✅ All critical VMs/CTs running (VM 109 docker-vm, CT 116 syslog-api, CT 117 zulip, CT 106 ra-h-os)
- ✅ GPU bare-metal hosts reachable: .8 (RTX 3090), .110 (RTX 5070), .15 (Strix Halo)
- ✅ Docker healthy on VM 109 (.7), CT 116 (.116)
- ✅ LiteLLM health check passing (port 4000, /mcp-rest/tools/list with master key)
- ✅ LiteLLM MCP gateway serving RA-H OS tools (90 tools)
- ✅ Zulip server reachable
- ✅ GPU fleet healthy (all 3 GPUs: RTX 3090, RTX 5070, RX 7600)
- ✅ Disk >20% free on all nodes
- 📋 Snapshot critical configs (LiteLLM, nginx, docker-compose files)
Wave 1: Storage & Infra Nodes (lowest impact)
| Target | Type | Command | Timeout |
|---|---|---|---|
| storepve (.6) | Proxmox node | apt update && apt upgrade -y |
5 min |
| minipve (.12) | Proxmox node | apt update && apt upgrade -y |
5 min |
| VM 109 (.7) | Docker host VM | apt update && apt upgrade -y |
5 min |
| CT 106 (.65) | RA-H OS | apt update && apt upgrade -y |
3 min |
| CT 117 (zulip, storepve) | Zulip | apt update && apt upgrade -y |
3 min |
Verify after Wave 1:
- Docker healthy on .7:
docker ps - RA-H OS MCP responding:
curl 192.168.68.65:3100/mcp - Zulip responding:
curl https://chat.sysloggh.net/api/v1/server_settings
Wave 2: Compute & Agent Nodes
| Target | Type | Command | Timeout |
|---|---|---|---|
| amdpve (.15) | Proxmox node | apt update && apt upgrade -y |
5 min |
| acerpve (.9) | Proxmox node | apt update && apt upgrade -y |
5 min |
| ocupve (.5) | Proxmox node | apt update && apt upgrade -y |
5 min |
| CT 100 (.24) | Abiba (pi) | apt update && apt upgrade -y |
3 min |
| CT 116 (.116) | syslog-api (LiteLLM host) | apt update && apt upgrade -y |
3 min |
| CT 112 (tanko, amdpve) | Tanko | apt update && apt upgrade -y |
3 min |
| CT 105 (kagentz, minipve) | Mumuni | apt update && apt upgrade -y |
3 min |
| VM 101 (.8) | llm-gpu (RTX 3090) | apt update && apt upgrade -y |
3 min |
| VM 103 (.110) | ocu-llm (RTX 5070) | apt update && apt upgrade -y |
3 min |
Verify after Wave 2:
- All VMs/CTs running: check via Proxmox API
- LiteLLM healthy:
curl localhost:4000/health/liveliness(via CT 116) - LiteLLM MCP tools:
curl localhost:4000/mcp-rest/tools/list -H "Authorization: Bearer $MASTER_KEY"→ 90 tools - GPU servers responding: check :8080 on VM 101, VM 103; check strix-moe via router (http://192.168.68.116/health/unified — .15:8080 is firewalled to .116 only)
- Zulip agents connected: check Mumuni/Tanko gateway state
- Abiba PM2 processes online:
pm2 status
Wave 3: Docker Image Updates
| Host | Stack | Command |
|---|---|---|
| VM 109 (.7) | Firecrawl | cd /opt/search-stack/firecrawl-source && docker compose pull && docker compose up -d |
| VM 109 (.7) | SearXNG | cd /opt/search-stack/searxng && docker compose pull && docker compose up -d |
| VM 109 (.7) | Home stack (Pulse, Stirling PDF) — JDownloader moved to CT 118 LXC 2026-08-01 | cd /opt/home_stack && docker compose pull && docker compose up -d |
| VM 109 (.7) | Audiobookshelf | cd /opt/audiobookshelf && docker compose pull && docker compose up -d |
| CT 116 (.116) | Inference Harness (LiteLLM, Prometheus, Grafana) | cd /opt/inference-harness && docker compose pull && docker compose up -d |
| CT 116 (.116, via minipve) | Trove docker agent (trove-agent-docker) | pct exec 116 -- bash -c 'cd /opt/trove-agent && docker compose pull && docker compose up -d' |
| CT 117 (storepve) | Zulip | pct exec 117 -- bash -c 'cd /opt/zulip && docker compose pull && docker compose up -d' (from storepve; compose recreates on zulip_default network) |
| CT 117 (storepve) | Jitsi | pct exec 117 -- bash -c 'cd /opt/jitsi && docker compose pull && docker compose up -d' (from storepve) |
| hwpve (.11) | Authentik (server, worker, postgres) | ssh root@192.168.68.11 'cd /root && docker compose pull && docker compose up -d' |
| NetBird VPS (72.61.0.17) | NetBird (server, dashboard, proxy, traefik, crowdsec) | ssh root@72.61.0.17 'cd /root && docker compose pull && docker compose up -d' |
| VM 109 (.7) | Trove test | cd /opt/trove-test && docker compose pull && docker compose up -d |
| VM 109 (.7) | docker-stats | cd /opt/docker-stats && docker compose pull && docker compose up -d |
| CT 116 (.116) | Monitoring (Grafana, Prometheus, Alertmanager, PVE exporter) | cd /opt/monitoring && docker compose pull && docker compose up -d |
Verify after Wave 3:
- All containers healthy:
docker pson each host - End-to-end inference test:
curl localhost:4000/v1/chat/completions(via CT 116) with syslog-auto - MCP integration test:
curl localhost:4000/mcp-rest/tools/list -H "Authorization: Bearer $MASTER_KEY"→ 90 tools (23 RA-H OS + 67 GitHub) - Zulip test: send test message to #agent-hub
- Dashboard loading:
curl localhost:3001/(via CT 116) - Firecrawl test:
curl -X POST http://192.168.68.7:3002/v1/search -H 'Content-Type: application/json' -d '{"query":"health","limit":1}'→"success":true(GET/returns 200) - Authentik test:
curl http://192.168.68.11:9000/→ 302 redirect to login - NetBird test:
curl -s -o /dev/null -w '%{http_code}' https://netbird.sysloggh.net/→ 200 - harness-litellm cold start: allow 3-5 min after recreate — reports unhealthy and :4000 refuses connections while loading config/DB, then recovers to 200 on its own (verified 2026-09-08)
- SearXNG test:
curl :8888 - Digest-pin sweep:
grep -rn '@sha256:' /opt/*/docker-compose.y*on every host — digest-pinned images are INVISIBLE todocker compose pull(the pin re-pulls the same digest forever, so new releases never appear). Flag every pin in the run report and propose un-pinning to a floating tag with user approval before editing. Found 2026-09-10: audiobookshelf was digest-pinned at 2.34.0 (container created 2026-07-18) and silently missed by every sweep; dockhand stack was also pinned (stack removed 2026-09-10, unused). After un-pinning audiobookshelf to :latest it updated to 2.36.0 and verified HTTP 200. - Version-pin awareness: a fixed version tag (e.g.
image: ...litellm:1.99.1) is a no-op fordocker compose pulljust like a digest pin, so the stack silently stops advancing. CT 116harness-litellmis INTENTIONALLY pinned to1.99.1(registrymain-stable/latestcurrently resolve to1.100.1, sha256:a3715fa7 — a bleeding-edge jump explicitly declined 2026-09-11). Every run must look up the newest STABLE release tag for any version-pinned image, bump the pin deliberately with user approval, recreate, and re-verify. Never silently revert a pin to a floating tag.
Wave 4: Proxmox Kernel Reboot
Only if [ -f /var/run/reboot-required ] on any node.
| Target | Action |
|---|---|
| Affected PVE node | Verify all CTs/VMs migrated or stopped |
reboot via PVE API (or systemctl reboot -f if dbus fails) |
|
| Wait 120s for node to come back | |
| Start any stopped CTs |
Post-reboot sweep (known gaps)
After every node reboot, run these checks:
-
CT auto-start sweep — LXC containers sometimes don't start despite
onboot: 1. Check every CT on the rebooted node and start any left stopped:pct list | awk '/stopped/{print $1}' | xargs -I{} pct start {}Known cases: scottdenya (CT 115 on amdpve), authentik (CT 104 on minipve).
-
Zulip recovery — When docker-vm or storepve reboots, the Zulip main container loses its Docker network assignment (SIGKILL during storage outage detaches it from
zulip_defaultnetwork). Run:ssh root@192.168.68.19 'docker rm -f zulip-zulip-1 && cd /opt/zulip && docker compose up -d'The compose restart recreates the container on the correct network.
-
docker-vm Docker daemon — After reboot, Docker can take 3-4 minutes to become
active. The docker-proxy for Pulse (port 7655) starts early, so Pulse is accessible beforedocker psreports ready. Wait for Docker before checking other stacks.
VPS ↔ docker-vm tunnel
After any VPS or docker-vm reboot, verify the dedicated WireGuard tunnel:
ssh root@72.61.0.17 'wg show wg1' | grep "latest handshake"
# If no handshake in >60s:
ssh root@72.61.0.17 'wg-quick up wg1'
The tunnel uses PersistentKeepalive=25 and is systemd-enabled, but should be verified after a reboot.
Rollback Protocol
If ANY verification fails:
- Apt rollback: Restore from Proxmox snapshot if taken, or
apt install <pkg>=<old_version> - Docker rollback:
docker compose down && docker compose up -d(uses cached images) - Config rollback: Restore from
/tmp/infra-update-backup-<date>/snapshots - Escalate: Send Zulip DM with failure details if auto-rollback fails
Config Backup
Before Wave 1, snapshot these files:
/opt/inference-harness/docker-compose.yml (CT 116 .116) ⚡ contains MCP_SERVER env vars
/opt/inference-harness/litellm_config.yaml (CT 116 .116) ⚡ contains mcp_servers.ra_h_os
/opt/monitoring/prometheus.yml (CT 116 .116)
/etc/nginx/nginx.conf (harness-nginx on CT 116)
/opt/search-stack/firecrawl-source/docker-compose.yaml (VM 109 .7)
/opt/search-stack/searxng/docker-compose.yml (VM 109 .7)
/opt/home_stack/docker-compose.yml (VM 109 .7)
/opt/audiobookshelf/docker-compose.yml (VM 109 .7)
/root/compose.yml (hwpve .11 — Authentik server/worker/postgres)
/root/docker-compose.yml (NetBird VPS — netbird server/dashboard/proxy, traefik, crowdsec)
/root/.pi/agent/extensions/config.yaml (CT 100 .24)
/etc/systemd/system/strix-server.service (amdpve .15 — strix-moe)
/etc/systemd/system/llama-server.service (VM 101 .8, VM 103 .110)
# Hermes agent configs (key enforcement — 2026-07-10)
/home/hermes/.hermes/config.yaml (Mumuni kagentz CT105; Tanko CT112 uses /home/jerome/.hermes)
/etc/systemd/system/hermes-gateway.service (Mumuni kagentz CT105 — system unit, User=hermes)
/etc/environment (LITELLM_API_KEY — legacy path, Mumuni now keys via Infisical)
MCP Gateway (2026-07-10)
LiteLLM CT 116 now serves as an authenticated MCP gateway for RA-H OS tools.
Configuration
litellm_config.yaml (/opt/inference-harness/litellm_config.yaml):
mcp_servers:
ra_h_os:
url: "http://192.168.68.65:3100/mcp"
transport: "http"
auth_type: "none"
docker-compose.yml env vars:
- MCP_SERVER_RAHOS_URL=http://192.168.68.65:3100/mcp
- MCP_SERVER_RAHOS_TRANSPORT=http
Access
| Key | MCP Access |
|---|---|
| Master key | ✅ Full — 90 tools (vault-injected) |
| Agent keys (mumuni, tanko, etc.) | ✅ Per-key grants supported (as of 2026-09-18 verification) |
Known Limitations
Per-key MCP server grants not functional — only master key has access(resolved 2026-09-18: per-key grants now work)- Responses API (
/v1/responses) with MCP tools broken on llama.cpp backends - HTTP 307 redirect on
/mcp→ use/mcp/(trailing slash) or/mcp-rest/endpoints api_mode: responsesin Hermes appends/v1/responsesto base_url → base_url must end at/v1, never/responses(double-path bug)
Migration Path (COMPLETED 2026-09-18)
Per-key MCP grants are now supported:
- ✅ Agent keys granted MCP access via
allowed_mcpfield - ✅ Hermes
mcp_servers.litellm.urlset tohttps://litellm.sysloggh.net/mcp - ✅
headers: {x-litellm-api-key: "Bearer <literal_key>"}added to MCP config
Security-Specific Updates
| Check | Command | Action |
|---|---|---|
| CVE count | apt list --upgradable 2>/dev/null | grep -i security | wc -l |
Report in update summary |
| Kernel vulns | uname -r vs latest available |
Flag if >2 versions behind |
| Docker CVEs | docker scout quickview or trivy image |
Flag critical CVEs |
| SSL certs | openssl s_client -connect chat.sysloggh.net:443 </dev/null 2>/dev/null | openssl x509 -noout -dates |
Alert if <30 days |
Success Criteria
- All 5 PVE nodes updated, no reboot-loop
- All VMs/CTs running post-update
- All Docker containers healthy (VM 109 + CT 116 + CT 117 + hwpve .11 + NetBird VPS)
- LiteLLM inference passing (syslog-auto test)
- Zulip server + all 3 agents connected
- GPU fleet at full capacity (3/3)
- LiteLLM MCP gateway healthy (90 tools via master key)
- Zero security CVEs remaining
- <10 min total downtime per service
Report
After completion, send Zulip DM:
📋 Infrastructure Update — YYYY-MM-DD
Updated: 5 PVE nodes, 12 CTs/VMs, 30+ containers
Security fixes: N CVEs patched
Downtime: <service> <duration>
Failures: none / <details>
Configs backed up: /tmp/infra-update-backup-YYYYMMDD/