- All GPUs: --parallel 1 → 2 (6 concurrent slots, was 3) - .8 RTX 3090: ctx 256K→128K, VRAM 96%→83%, turbo4 KV cache - .110 RTX 5070: ctx 256K→128K, ubatch 4096→512 (was inverted), VRAM 90%→77%, q4_0 KV - .15 Strix Halo: parallel 1→2, 256K ctx (41GB free), q8_0 KV, AMD metrics via /sys/class/drm - LiteLLM: gemma timeout 25→120s, qwen timeout 40→90s, syslog-auto (qwen) 40→90s - Agent configs: context_length 262144 for syslog-auto, 131072 for direct qwen/gemma - Updated health-check operation, agent config implications, benchmark table (fixed model↔GPU mapping)
6.9 KiB
6.9 KiB
kind, name, description, agent, triggers, version
| kind | name | description | agent | triggers | version | |||
|---|---|---|---|---|---|---|---|---|
| responsibility | infrastructure-update | Autonomous system-wide update contract covering all 5 Proxmox nodes, 15+ containers/VMs, and 4 Docker ecosystems. Updates apt packages, Docker images, and container stacks in safe waves with health checks and automatic rollback on failure. | abiba |
|
1.1.0 |
Maintains
- update-status: { phase, node, action, result, timestamp }
- update-history: array of past update runs with results
- security-state: { cve_count, last_patched, pending_updates }
Pre-Flight Checklist
Before ANY update wave:
- ✅ All Proxmox nodes online (
GET /api2/json/nodes) - ✅ All critical VMs/CTs running (VM 101 llm-gpu, VM 103 ocu-llm, VM 109 docker-vm, CT 116 syslog-api, CT 117 zulip, CT 106 ra-h-os)
- ✅ Docker healthy on VM 109 (.7), CT 116 (.116)
- ✅ LiteLLM health check passing
- ✅ Zulip server reachable
- ✅ GPU fleet healthy (all 3 GPUs: RTX 3090, RTX 5070, RX 7600)
- ✅ Disk >20% free on all nodes
- 📋 Snapshot critical configs (LiteLLM, nginx, docker-compose files)
Wave 1: Storage & Infra Nodes (lowest impact)
| Target | Type | Command | Timeout |
|---|---|---|---|
| storepve (.6) | Proxmox node | apt update && apt upgrade -y |
5 min |
| minipve (.12) | Proxmox node | apt update && apt upgrade -y |
5 min |
| VM 109 (.7) | Docker host VM | apt update && apt upgrade -y |
5 min |
| CT 106 (.65) | RA-H OS | apt update && apt upgrade -y |
3 min |
| CT 117 (zulip, storepve) | Zulip | apt update && apt upgrade -y |
3 min |
Verify after Wave 1:
- Docker healthy on .7:
docker ps - RA-H OS MCP responding:
curl 192.168.68.65:3100/mcp - Zulip responding:
curl https://chat.sysloggh.net/api/v1/server_settings
Wave 2: Compute & Agent Nodes
| Target | Type | Command | Timeout |
|---|---|---|---|
| amdpve (.15) | Proxmox node | apt update && apt upgrade -y |
5 min |
| acerpve (.9) | Proxmox node | apt update && apt upgrade -y |
5 min |
| ocupve (.5) | Proxmox node | apt update && apt upgrade -y |
5 min |
| CT 100 (.24) | Abiba (pi) | apt update && apt upgrade -y |
3 min |
| CT 116 (.116) | syslog-api (LiteLLM host) | apt update && apt upgrade -y |
3 min |
| CT 112 (tanko, amdpve) | Tanko | apt update && apt upgrade -y |
3 min |
| CT 114 (mumuni, minipve) | Mumuni | apt update && apt upgrade -y |
3 min |
| VM 101 (.8) | llm-gpu (RTX 3090) | apt update && apt upgrade -y |
3 min |
| VM 103 (.110) | ocu-llm (RTX 5070) | apt update && apt upgrade -y |
3 min |
Verify after Wave 2:
- All VMs/CTs running: check via Proxmox API
- LiteLLM healthy:
curl localhost:4000/health/liveliness(via CT 116) - GPU servers responding: check :8080 on VM 101, VM 103; check ornith via router (http://192.168.68.116/health/unified — .15:8080 is firewalled to .116 only)
- Zulip agents connected: check Mumuni/Tanko gateway state
- Abiba PM2 processes online:
pm2 status
Wave 3: Docker Image Updates
| Host | Stack | Command |
|---|---|---|
| VM 109 (.7) | Firecrawl | cd /opt/search-stack/firecrawl-source && docker compose pull && docker compose up -d |
| VM 109 (.7) | SearXNG | cd /opt/search-stack/searxng && docker compose pull && docker compose up -d |
| VM 109 (.7) | Home stack (Pulse, Stirling PDF, JDownloader 2) | cd /opt/home_stack && docker compose pull && docker compose up -d |
| VM 109 (.7) | Audiobookshelf | cd /opt/audiobookshelf && docker compose pull && docker compose up -d |
| CT 116 (.116) | Inference Harness (LiteLLM, Prometheus, Grafana) | cd /opt/inference-harness && docker compose pull && docker compose up -d |
| CT 117 (zulip, storepve) | Zulip | docker pull zulip/docker-zulip:latest && docker restart zulip-zulip-1 |
Verify after Wave 3:
- All containers healthy:
docker pson each host - End-to-end inference test:
curl localhost:4000/v1/chat/completions(via CT 116) with syslog-auto - Zulip test: send test message to #agent-hub
- Dashboard loading:
curl localhost:3001/(via CT 116) - Firecrawl test:
curl :3002/ - SearXNG test:
curl :8888
Wave 4: Proxmox Kernel Reboot (if needed)
Only if [ -f /var/run/reboot-required ] on any node.
| Target | Action |
|---|---|
| Affected PVE node | Verify all CTs/VMs migrated or stopped |
reboot via PVE API |
|
| Wait 120s for node to come back | |
| Start any stopped CTs |
Rollback Protocol
If ANY verification fails:
- Apt rollback: Restore from Proxmox snapshot if taken, or
apt install <pkg>=<old_version> - Docker rollback:
docker compose down && docker compose up -d(uses cached images) - Config rollback: Restore from
/tmp/infra-update-backup-<date>/snapshots - Escalate: Send Zulip DM with failure details if auto-rollback fails
Config Backup
Before Wave 1, snapshot these files:
/opt/inference-harness/docker-compose.yml (CT 116 .116)
/opt/inference-harness/litellm_config.yaml (CT 116 .116)
/opt/monitoring/prometheus.yml (CT 116 .116)
/etc/nginx/nginx.conf (harness-nginx on CT 116)
/opt/search-stack/firecrawl-source/docker-compose.yaml (VM 109 .7)
/opt/search-stack/searxng/docker-compose.yml (VM 109 .7)
/opt/home_stack/docker-compose.yml (VM 109 .7)
/opt/audiobookshelf/docker-compose.yml (VM 109 .7)
/root/.pi/agent/extensions/config.yaml (CT 100 .24)
/etc/systemd/system/ornith-server.service (amdpve .15)
/etc/systemd/system/llama-server.service (VM 101 .8, VM 103 .110)
Run: mkdir -p /tmp/infra-update-backup-$(date +%Y%m%d) && rsync -av ...
Security-Specific Updates
| Check | Command | Action |
|---|---|---|
| CVE count | apt list --upgradable 2>/dev/null | grep -i security | wc -l |
Report in update summary |
| Kernel vulns | uname -r vs latest available |
Flag if >2 versions behind |
| Docker CVEs | docker scout quickview or trivy image |
Flag critical CVEs |
| SSL certs | openssl s_client -connect chat.sysloggh.net:443 </dev/null 2>/dev/null | openssl x509 -noout -dates |
Alert if <30 days |
Success Criteria
- All 5 PVE nodes updated, no reboot-loop
- All VMs/CTs running post-update
- All Docker containers healthy (VM 109 + CT 116 + CT 117)
- LiteLLM inference passing (syslog-auto test)
- Zulip server + all 3 agents connected
- GPU fleet at full capacity (3/3)
- Zero security CVEs remaining
- <10 min total downtime per service
Report
After completion, send Zulip DM:
📋 Infrastructure Update — YYYY-MM-DD
Updated: 5 PVE nodes, 12 CTs/VMs, 30+ containers
Security fixes: N CVEs patched
Downtime: <service> <duration>
Failures: none / <details>
Configs backed up: /tmp/infra-update-backup-YYYYMMDD/