Files
prose-contracts/infrastructure-update.prose.md
root f77d6ca1d1
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 8s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 2s
fix: correct field name to allowed_mcp_servers
PR #118 finding F1 (low): The deployed LiteLLM on CT 116 uses
allowed_mcp_servers (193 occurrences in installed package), not
bare allowed_mcp. One-word doc fix.
2026-09-18 18:49:20 +00:00

14 KiB

kind, name, description, agent, triggers, version
kind name description agent triggers version
responsibility infrastructure-update Autonomous system-wide update contract covering all 5 Proxmox nodes, 15+ containers/VMs, and 5 Docker ecosystems (docker-vm .7, CT 116 .116, CT 117, hwpve .11, NetBird VPS 72.61.0.17). Updates apt packages, Docker images, and container stacks in safe waves with health checks and automatic rollback on failure. abiba
on "infra update" command
weekly (Sunday 03:00 America/New_York) via Agent Zero scheduler task "weekly-fleet-docker-update" (qSOOVzsU) — implemented 2026-09-08
on security advisory relay from Mumuni
1.4.0

Maintains

  • update-status: { phase, node, action, result, timestamp }
  • update-history: array of past update runs with results
  • security-state: { cve_count, last_patched, pending_updates }

Pre-Flight Checklist

Before ANY update wave:

  1. ✅ All Proxmox nodes online (GET /api2/json/nodes)
  2. ✅ All critical VMs/CTs running (VM 109 docker-vm, CT 116 syslog-api, CT 117 zulip, CT 106 ra-h-os)
  3. ✅ GPU bare-metal hosts reachable: .8 (RTX 3090), .110 (RTX 5070), .15 (Strix Halo)
  4. ✅ Docker healthy on VM 109 (.7), CT 116 (.116)
  5. ✅ LiteLLM health check passing (port 4000, /mcp-rest/tools/list with master key)
  6. ✅ LiteLLM MCP gateway serving RA-H OS tools (90 tools)
  7. ✅ Zulip server reachable
  8. ✅ GPU fleet healthy (all 3 GPUs: RTX 3090, RTX 5070, RX 7600)
  9. ✅ Disk >20% free on all nodes
  10. 📋 Snapshot critical configs (LiteLLM, nginx, docker-compose files)

Wave 1: Storage & Infra Nodes (lowest impact)

Target Type Command Timeout
storepve (.6) Proxmox node apt update && apt upgrade -y 5 min
minipve (.12) Proxmox node apt update && apt upgrade -y 5 min
VM 109 (.7) Docker host VM apt update && apt upgrade -y 5 min
CT 106 (.65) RA-H OS apt update && apt upgrade -y 3 min
CT 117 (zulip, storepve) Zulip apt update && apt upgrade -y 3 min

Verify after Wave 1:

  • Docker healthy on .7: docker ps
  • RA-H OS MCP responding: curl 192.168.68.65:3100/mcp
  • Zulip responding: curl https://chat.sysloggh.net/api/v1/server_settings

Wave 2: Compute & Agent Nodes

Target Type Command Timeout
amdpve (.15) Proxmox node apt update && apt upgrade -y 5 min
acerpve (.9) Proxmox node apt update && apt upgrade -y 5 min
ocupve (.5) Proxmox node apt update && apt upgrade -y 5 min
CT 100 (.24) Abiba (pi) apt update && apt upgrade -y 3 min
CT 116 (.116) syslog-api (LiteLLM host) apt update && apt upgrade -y 3 min
CT 112 (tanko, amdpve) Tanko apt update && apt upgrade -y 3 min
CT 105 (kagentz, minipve) Mumuni apt update && apt upgrade -y 3 min
VM 101 (.8) llm-gpu (RTX 3090) apt update && apt upgrade -y 3 min
VM 103 (.110) ocu-llm (RTX 5070) apt update && apt upgrade -y 3 min

Verify after Wave 2:

  • All VMs/CTs running: check via Proxmox API
  • LiteLLM healthy: curl localhost:4000/health/liveliness (via CT 116)
  • LiteLLM MCP tools: curl localhost:4000/mcp-rest/tools/list -H "Authorization: Bearer $MASTER_KEY" → 90 tools
  • GPU servers responding: check :8080 on VM 101, VM 103; check strix-moe via router (http://192.168.68.116/health/unified — .15:8080 is firewalled to .116 only)
  • Zulip agents connected: check Mumuni/Tanko gateway state
  • Abiba PM2 processes online: pm2 status

Wave 3: Docker Image Updates

Host Stack Command
VM 109 (.7) Firecrawl cd /opt/search-stack/firecrawl-source && docker compose pull && docker compose up -d
VM 109 (.7) SearXNG cd /opt/search-stack/searxng && docker compose pull && docker compose up -d
VM 109 (.7) Home stack (Pulse, Stirling PDF) — JDownloader moved to CT 118 LXC 2026-08-01 cd /opt/home_stack && docker compose pull && docker compose up -d
VM 109 (.7) Audiobookshelf cd /opt/audiobookshelf && docker compose pull && docker compose up -d
CT 116 (.116) Inference Harness (LiteLLM, Prometheus, Grafana) cd /opt/inference-harness && docker compose pull && docker compose up -d
CT 116 (.116, via minipve) Trove docker agent (trove-agent-docker) pct exec 116 -- bash -c 'cd /opt/trove-agent && docker compose pull && docker compose up -d'
CT 117 (storepve) Zulip pct exec 117 -- bash -c 'cd /opt/zulip && docker compose pull && docker compose up -d' (from storepve; compose recreates on zulip_default network)
CT 117 (storepve) Jitsi pct exec 117 -- bash -c 'cd /opt/jitsi && docker compose pull && docker compose up -d' (from storepve)
hwpve (.11) Authentik (server, worker, postgres) ssh root@192.168.68.11 'cd /root && docker compose pull && docker compose up -d'
NetBird VPS (72.61.0.17) NetBird (server, dashboard, proxy, traefik, crowdsec) ssh root@72.61.0.17 'cd /root && docker compose pull && docker compose up -d'
VM 109 (.7) Trove test cd /opt/trove-test && docker compose pull && docker compose up -d
VM 109 (.7) docker-stats cd /opt/docker-stats && docker compose pull && docker compose up -d
CT 116 (.116) Monitoring (Grafana, Prometheus, Alertmanager, PVE exporter) cd /opt/monitoring && docker compose pull && docker compose up -d

Verify after Wave 3:

  • All containers healthy: docker ps on each host
  • End-to-end inference test: curl localhost:4000/v1/chat/completions (via CT 116) with syslog-auto
  • MCP integration test: curl localhost:4000/mcp-rest/tools/list -H "Authorization: Bearer $MASTER_KEY" → 90 tools (23 RA-H OS + 67 GitHub)
  • Zulip test: send test message to #agent-hub
  • Dashboard loading: curl localhost:3001/ (via CT 116)
  • Firecrawl test: curl -X POST http://192.168.68.7:3002/v1/search -H 'Content-Type: application/json' -d '{"query":"health","limit":1}' → "success":true (GET / returns 200)
  • Authentik test: curl http://192.168.68.11:9000/ → 302 redirect to login
  • NetBird test: curl -s -o /dev/null -w '%{http_code}' https://netbird.sysloggh.net/ → 200
  • harness-litellm cold start: allow 3-5 min after recreate — reports unhealthy and :4000 refuses connections while loading config/DB, then recovers to 200 on its own (verified 2026-09-08)
  • SearXNG test: curl :8888
  • Digest-pin sweep: grep -rn '@sha256:' /opt/*/docker-compose.y* on every host — digest-pinned images are INVISIBLE to docker compose pull (the pin re-pulls the same digest forever, so new releases never appear). Flag every pin in the run report and propose un-pinning to a floating tag with user approval before editing. Found 2026-09-10: audiobookshelf was digest-pinned at 2.34.0 (container created 2026-07-18) and silently missed by every sweep; dockhand stack was also pinned (stack removed 2026-09-10, unused). After un-pinning audiobookshelf to :latest it updated to 2.36.0 and verified HTTP 200.
  • Version-pin awareness: a fixed version tag (e.g. image: ...litellm:1.99.1) is a no-op for docker compose pull just like a digest pin, so the stack silently stops advancing. CT 116 harness-litellm is INTENTIONALLY pinned to 1.99.1 (registry main-stable/latest currently resolve to 1.100.1, sha256:a3715fa7 — a bleeding-edge jump explicitly declined 2026-09-11). Every run must look up the newest STABLE release tag for any version-pinned image, bump the pin deliberately with user approval, recreate, and re-verify. Never silently revert a pin to a floating tag.

Wave 4: Proxmox Kernel Reboot

Only if [ -f /var/run/reboot-required ] on any node.

Target Action
Affected PVE node Verify all CTs/VMs migrated or stopped
reboot via PVE API (or systemctl reboot -f if dbus fails)
Wait 120s for node to come back
Start any stopped CTs

Post-reboot sweep (known gaps)

After every node reboot, run these checks:

  1. CT auto-start sweep — LXC containers sometimes don't start despite onboot: 1. Check every CT on the rebooted node and start any left stopped:

    pct list | awk '/stopped/{print $1}' | xargs -I{} pct start {}
    

    Known cases: scottdenya (CT 115 on amdpve), authentik (CT 104 on minipve).

  2. Zulip recovery — When docker-vm or storepve reboots, the Zulip main container loses its Docker network assignment (SIGKILL during storage outage detaches it from zulip_default network). Run:

    ssh root@192.168.68.19 'docker rm -f zulip-zulip-1 && cd /opt/zulip && docker compose up -d'
    

    The compose restart recreates the container on the correct network.

  3. docker-vm Docker daemon — After reboot, Docker can take 3-4 minutes to become active. The docker-proxy for Pulse (port 7655) starts early, so Pulse is accessible before docker ps reports ready. Wait for Docker before checking other stacks.

VPS ↔ docker-vm tunnel

After any VPS or docker-vm reboot, verify the dedicated WireGuard tunnel:

ssh root@72.61.0.17 'wg show wg1' | grep "latest handshake"
# If no handshake in >60s:
ssh root@72.61.0.17 'wg-quick up wg1'

The tunnel uses PersistentKeepalive=25 and is systemd-enabled, but should be verified after a reboot.

Rollback Protocol

If ANY verification fails:

  1. Apt rollback: Restore from Proxmox snapshot if taken, or apt install <pkg>=<old_version>
  2. Docker rollback: docker compose down && docker compose up -d (uses cached images)
  3. Config rollback: Restore from /tmp/infra-update-backup-<date>/ snapshots
  4. Escalate: Send Zulip DM with failure details if auto-rollback fails

Config Backup

Before Wave 1, snapshot these files:

/opt/inference-harness/docker-compose.yml               (CT 116 .116) ⚡ contains MCP_SERVER env vars
/opt/inference-harness/litellm_config.yaml              (CT 116 .116) ⚡ contains mcp_servers.ra_h_os
/opt/monitoring/prometheus.yml                          (CT 116 .116)
/etc/nginx/nginx.conf                                   (harness-nginx on CT 116)
/opt/search-stack/firecrawl-source/docker-compose.yaml  (VM 109 .7)
/opt/search-stack/searxng/docker-compose.yml            (VM 109 .7)
/opt/home_stack/docker-compose.yml                      (VM 109 .7)
/opt/audiobookshelf/docker-compose.yml                  (VM 109 .7)
/root/compose.yml                                       (hwpve .11 — Authentik server/worker/postgres)
/root/docker-compose.yml                                (NetBird VPS — netbird server/dashboard/proxy, traefik, crowdsec)
/root/.pi/agent/extensions/config.yaml                  (CT 100 .24)
/etc/systemd/system/strix-server.service                (amdpve .15 — strix-moe)
/etc/systemd/system/llama-server.service                (VM 101 .8, VM 103 .110)
# Hermes agent configs (key enforcement — 2026-07-10)
/home/hermes/.hermes/config.yaml                        (Mumuni kagentz CT105; Tanko CT112 uses /home/jerome/.hermes)
/etc/systemd/system/hermes-gateway.service              (Mumuni kagentz CT105 — system unit, User=hermes)
/etc/environment                                        (LITELLM_API_KEY — legacy path, Mumuni now keys via Infisical)

MCP Gateway (2026-07-10)

LiteLLM CT 116 now serves as an authenticated MCP gateway for RA-H OS tools.

Configuration

litellm_config.yaml (/opt/inference-harness/litellm_config.yaml):

mcp_servers:
  ra_h_os:
    url: "http://192.168.68.65:3100/mcp"
    transport: "http"
    auth_type: "none"

docker-compose.yml env vars:

- MCP_SERVER_RAHOS_URL=http://192.168.68.65:3100/mcp
- MCP_SERVER_RAHOS_TRANSPORT=http

Access

Key MCP Access
Master key ✅ Full — 90 tools (vault-injected)
Agent keys (mumuni, tanko, etc.) ✅ Per-key grants supported (as of 2026-09-18 verification)

Known Limitations

  • Per-key MCP server grants not functional — only master key has access (resolved 2026-09-18: per-key grants now work)
  • Responses API (/v1/responses) with MCP tools broken on llama.cpp backends
  • HTTP 307 redirect on /mcp → use /mcp/ (trailing slash) or /mcp-rest/ endpoints
  • api_mode: responses in Hermes appends /v1/responses to base_url → base_url must end at /v1, never /responses (double-path bug)

Migration Path (COMPLETED 2026-09-18)

Per-key MCP grants are now supported:

  1. ✅ Agent keys granted MCP access via allowed_mcp_servers field
  2. ✅ Hermes mcp_servers.litellm.url set to https://litellm.sysloggh.net/mcp
  3. ✅ headers: {x-litellm-api-key: "Bearer <literal_key>"} added to MCP config

Security-Specific Updates

Check Command Action
CVE count apt list --upgradable 2>/dev/null | grep -i security | wc -l Report in update summary
Kernel vulns uname -r vs latest available Flag if >2 versions behind
Docker CVEs docker scout quickview or trivy image Flag critical CVEs
SSL certs openssl s_client -connect chat.sysloggh.net:443 </dev/null 2>/dev/null | openssl x509 -noout -dates Alert if <30 days

Success Criteria

  • All 5 PVE nodes updated, no reboot-loop
  • All VMs/CTs running post-update
  • All Docker containers healthy (VM 109 + CT 116 + CT 117 + hwpve .11 + NetBird VPS)
  • LiteLLM inference passing (syslog-auto test)
  • Zulip server + all 3 agents connected
  • GPU fleet at full capacity (3/3)
  • LiteLLM MCP gateway healthy (90 tools via master key)
  • Zero security CVEs remaining
  • <10 min total downtime per service

Report

After completion, send Zulip DM:

📋 Infrastructure Update — YYYY-MM-DD

Updated: 5 PVE nodes, 12 CTs/VMs, 30+ containers
Security fixes: N CVEs patched
Downtime: <service> <duration>
Failures: none / <details>
Configs backed up: /tmp/infra-update-backup-YYYYMMDD/