diff --git a/infrastructure-control.prose.md b/infrastructure-control.prose.md index 5a698bd..e8319b3 100644 --- a/infrastructure-control.prose.md +++ b/infrastructure-control.prose.md @@ -2,355 +2,332 @@ kind: pattern name: infrastructure-control description: > - Full infrastructure monitoring and control pattern covering the 5-node - Proxmox cluster, 3 Docker ecosystems (22 containers), storage fabric, - and network services. Declares what agents can observe, what they can - enforce, and where human escalation is required. + Full infrastructure monitoring and control pattern covering the + 5-node Proxmox cluster, 3 Docker ecosystems (22 containers), + NFS storage, and network services. Defines monitors, remediations, + and the access matrix for all environments. --- -## Architecture Overview +# Infrastructure Control Pattern + +## Topology ``` - ┌──────────────────────────────────────────┐ - │ OpenProse Contracts │ - │ (monitor → diagnose → remediate → log) │ - └─────┬──────────┬──────────┬──────────────┘ - │ │ │ - ┌────────────┼──────────┼──────────┼────────────┐ - ▼ ▼ ▼ ▼ ▼ - ┌──────────┐ ┌──────────┐ ┌──────┐ ┌──────┐ ┌──────────┐ - │ Proxmox │ │ Docker │ │Storage│ │Network│ │ Agents │ - │ Cluster │ │Ecosystems│ │Fabric │ │Services│ │(health) │ - │ 5 nodes │ │ 3 hosts │ │NFS+ZFS│ │DNS+Auth│ │6 agents │ - └──────────┘ └──────────┘ └──────┘ └──────┘ └──────────┘ + ┌─────────────────────────┐ + │ OpenProse Contract │ + │ (declares what's true) │ + └──────────┬──────────────┘ + │ + ┌────────────────────┼────────────────────┐ + ▼ ▼ ▼ + ┌─────────────┐ ┌──────────────┐ ┌──────────────┐ + │ Abiba │ │ Tanko │ │ Mumuni │ + │ (pi) │ │ (Hermes) │ │ (Hermes) │ + │ CT 100 │ │ CT 122 │ │ CT 114 │ + └──────┬──────┘ └──────┬───────┘ └──────┬───────┘ + │ │ │ + └──────────────────┼────────────────────┘ + ▼ + ┌──────────────────────────────────────┐ + │ Proxmox Cluster API │ + │ minipve.sysloggh.net:443 │ + │ (monitoring@pve!mumuni token) │ + └────┬──────┬──────┬──────┬──────┬─────┘ + │ │ │ │ │ + ┌────┘ ┌────┘ ┌────┘ ┌────┘ ┌────┘ + ▼ ▼ ▼ ▼ ▼ + minipve amdpve storepve acerpve ocupve + (.12) (.15) (.6) (.9) (.5) + + ▼ + ┌─────────────────────────────────────────────┐ + │ Docker Ecosystems │ + │ ┌────────────┬────────────┬──────────────┐ │ + │ │ docker-vm │ CT 116 │ Netbird │ │ + │ │ (.7) │ (.116) │ (72.61.0.17) │ │ + │ │ 11 ctrs │ 6 ctrs │ 5 ctrs │ │ + │ └────────────┴────────────┴──────────────┘ │ + └─────────────────────────────────────────────┘ ``` -## Layer 1: Proxmox Cluster +## Section 1: Access Matrix -### Cluster Topology +### API Tokens & Credentials + +| Resource | Auth Method | Credential Source | Status | +|----------|------------|-------------------|--------| +| Proxmox Cluster | PVE API Token | `monitoring@pve!mumuni=...` | ✅ | +| Proxmox Root | Password via API ticket | `root@pam:kakashi19` | ✅ | +| docker-vm (.7) | SSH root | SSH key | ✅ | +| CT 116 (syslog-api) | SSH root | SSH key | ✅ | +| Tanko CT (.122) | SSH jerome | SSH key | ✅ | +| Netbird (.17) | SSH root | SSH key | ✅ | +| Gitea | API token | abiba-bot token | ✅ | +| Zulip | Bot API key | per-bot tokens | ✅ | +| RA-H OS | MCP bridge | port 3100 | ✅ | + +### Reachability Matrix + +| From / To | PVE API | docker-vm (.7) | CT 116 | Tanko (.122) | Netbird (.17) | +|-----------|---------|----------------|--------|-------------|---------------| +| **Abiba** (CT 100) | ✅ :443 | ✅ SSH | ✅ SSH | ✅ SSH | ❌ no Netbird | +| **Tanko** (CT 122) | ❌ | ❌ | ❌ | ✅ | ❌ | +| **docker-vm** (.7) | ❌ | ✅ | ❌ | ❌ | ❌ | + +**Conclusion:** Only Abiba has cross-infrastructure access. All monitoring contracts run from Abiba. + +## Section 2: Proxmox Cluster — Monitoring + +### Nodes (5) + +| Node | IP | CPU | RAM | VMs/CTs | Role | +|------|----|-----|-----|---------|------| +| minipve | .12 | 16C | 30GB | authentik, gitea, mumuni, syslog-api, jitsi | Auth, git, messaging | +| amdpve | .15 | 32C | 62GB | abiba, kagentz, tanko, tdunna, baggy, scottdenya | Agents, compute | +| storepve | .6 | 28C | 31GB | docker-vm, ra-h-os, PBS, media, zulip | Docker, storage, chat | +| acerpve | .9 | 28C | 31GB | llm-gpu, adguard | GPU VMs | +| ocupve | .5 | 12C | 14GB | ocu-llm | GPU VMs | + +### Checks (every 5 min) ``` -minipve (.12) ─── 16C/30GB ─── authentik CT 104, gitea CT 110, -│ mumuni CT 114, syslog-api CT 116, -│ jitsi CT 118 -│ -amdpve (.15) ─── 32C/62GB ─── abiba CT 100, kagentz CT 105, -│ tanko CT 112, tdunna CT 111, -│ baggy CT 113, scottdenya CT 115 -│ -storepve (.6) ──── 28C/31GB ─── docker-vm VM 109, ra-h-os CT 106, -│ PBS CT 107, media CT 108, -│ zulip CT 117 -│ -acerpve (.9) ──── 28C/31GB ─── llm-gpu VM 101, adguard CT 102 -│ -ocupve (.5) ──── 12C/14GB ─── ocu-llm VM 103 -``` - -### Maintains (Monitoring) - -```yaml ## Maintains -- cluster.nodes: each node's status (online/offline), CPU%, RAM%, uptime -- cluster.vms: each VM/CT's status (running/stopped), CPU, RAM, disk usage -- cluster.storage: each storage target's usage %, active/inactive -- cluster.ha: high-availability state and fencing status + +- cluster-status: { online_nodes: int, offline_nodes: int, timestamp } +- node-cpu-usage: { node: pct, warnings: [] } +- node-memory-usage: { node: free_pct, warnings: [] } +- node-uptime: { node: seconds, just_rebooted: bool } + +## Checks + +- For each node in [minipve, amdpve, storepve, acerpve, ocupve]: + - GET /api2/json/nodes/{node}/status → check status == "online" + - GET /api2/json/nodes/{node}/status → cpu < 0.80 + - GET /api2/json/nodes/{node}/status → free_mem > 10% + - GET /api2/json/nodes/{node}/status → uptime > 300 (warn if just rebooted) + +- For each CT in inventory (19 total): + - GET /api2/json/cluster/resources → filter by type=lxc + - Warn if status != "running" + +- Storage pools: + - GET /api2/json/nodes/storepve/storage/mediastore → used < 80% + - GET /api2/json/nodes/storepve/storage/pbs-backup → last backup < 48h + +## Remediations + +- Node offline → alert via Zulip DM, escalate after 3x +- CT stopped → `pct start ` via PVE API, verify after 30s +- VM stopped → `qm start ` via PVE API, verify after 60s +- Storage > 80% → warn via Zulip DM +- Storage > 95% → crit via Zulip + relay to maintainer ``` -### Rules (Monitoring) +## Section 3: Docker Ecosystems — Monitoring + +### Ecosystem A: docker-vm (192.168.68.7) + +11 containers across 4 compose stacks: + +| Stack | Path | Containers | +|-------|------|-----------| +| **Firecrawl** | `/opt/search-stack/firecrawl-source/` | api, rabbitmq, postgres, playwright, redis | +| **SearXNG** | `/opt/search-stack/searxng/` | searxng, valkey | +| **Home stack** | `/opt/home_stack/` | jdownloader, bentopdf, pulse | +| **Audiobookshelf** | `/opt/audiobookshelf/` | audiobookshelf | + +### Ecosystem B: CT 116 syslog-api (192.168.68.116) + +| Container | Image | Port | +|-----------|-------|------| +| harness-litellm | berriai/litellm:1.90.0-rc.1 | :4001 | +| harness-nginx | nginx:alpine | :80 | +| harness-router | inference-harness-router | :9000 | +| harness-postgres | postgres:16-alpine | :5432 | +| harness-redis | redis:7-alpine | :6379 | +| harness-dashboard | inference-harness-dashboard | :3000 | + +### Ecosystem C: Netbird (72.61.0.17) + +| Container | Image | Role | +|-----------|-------|------| +| netbird-server | netbird | VPN controller | +| netbird-dashboard | netbird UI | Web management | +| netbird-proxy | nginx | TLS termination | +| netbird-crowdsec | crowdsecurity/crowdsec:v1.7.7 | WAF | +| netbird-traefik | traefik | Reverse proxy | + +### Checks (every 60s) ``` -Rule PM-1: Node Offline - Detect → PVE API shows node status != "online" - Action → Alert via Zulip DM with node name and last known uptime - Escalate → If >2 nodes offline, page owner - Requires: PVE API token - -Rule PM-2: Node Resource Pressure - Detect → CPU > 80% or RAM > 85% for 5 consecutive checks - Action → Log to KG, alert via Zulip DM with resource snapshot - Escalate → If sustained >90% for 1h, page owner - Requires: PVE API token - -Rule PM-3: VM/CT Unexpectedly Stopped - Detect → PVE API shows running VM now "stopped" - Action → `pct start ` or `qm start ` via PVE API - Verify → Re-check status after 30s - Escalate → If still stopped after 2 retries, alert owner - Requires: PVE API token with appropriate permissions - -Rule PM-4: Storage Usage Warning - Detect → Storage pool usage > 80% - Action → Log to KG, alert via Zulip DM - Escalate → If >90%, immediate alert (risk of VM freeze) - Requires: PVE API -``` - -### Control (Remediation) - -```yaml -## Enforce -- PM-3 auto-start: start critical VMs/CTs when unexpectedly stopped - - Critical: abiba (100), tanko (112), mumuni (114), syslog-api (116), - docker-vm (109), gitea (110), authentik (104), zulip (117) -- PM-4 storage: if mediastore > 85%, alert with biggest consumers - -## Prohibited -- Never migrate VMs/CTs between nodes automatically -- Never delete storage volumes -- Never modify cluster HA configuration -``` - -## Layer 2: Docker Ecosystems - -### Docker Hosts - -```yaml -Host: docker-vm (192.168.68.7, storepve VM 109) - Resources: 10GB RAM, 4 vCPU, 158GB disk (33% used) - Stacks: - firecrawl: api, rabbitmq, postgres, playwright, redis - searxng: searxng (:8888), valkey - home_stack: jdownloader (:5800), bentopdf (:8989), - pulse (:7655), audiobookshelf (:13378) - Mounts: /media/storage (3.6TB NFS), /media/mediastore (7.3TB NFS) - -Host: syslog-api (192.168.68.116, minipve CT 116) - Resources: 6GB RAM, 2 vCPU, 39GB disk - Stacks: - harness: litellm (:4001), nginx (:80), router (:9000), - postgres, redis, dashboard (:3000) - -Host: netbird (72.61.0.17, external VM) - Stacks: - netbird: server, dashboard, proxy, crowdsec, traefik -``` - -### Maintains (Monitoring) - -```yaml ## Maintains -- docker.hosts: each host's status, container count, resource usage -- docker.containers: each container's status (running/restarting/exited), - uptime, restart count, image version -- docker.stacks: each compose stack's service count and health -- docker.resources: per-host disk usage, Docker system df + +- docker-health: { ecosystem: string, healthy: int, unhealthy: int, total: int } +- container-status: { name: string, status: string, restarts: int, image: string } + +## Checks + +For each Docker host: + - docker ps --format "{{.Names}}" → every required container is running + - docker inspect --format "{{.State.Health.Status}}" → "healthy" + - docker inspect --format "{{.RestartCount}}" → < 3/hour + - df -h / | awk '{print $5}' → usage < 80% + +For docker-vm specifically: + - mountpoint -q /media/storage → NFS mounted + - mountpoint -q /media/mediastore → NFS mounted + - docker compose ls → expected stacks present + +## Remediations + +- Container unhealthy → `docker compose up -d `, wait 10s, re-check +- Container missing → `docker compose up -d` in its stack directory +- Docker daemon down → `systemctl restart docker` +- NFS mount lost → `mount -a`, if fails → alert (storepve issue) +- Disk > 80% → alert via Zulip DM +- Disk > 95% → crit + relay to maintainer ``` -### Rules (Monitoring) +## Section 4: Storage — Monitoring + +### NFS Mounts (docker-vm → storepve) + +| Mount | Export | Capacity | Used | Alert | +|-------|--------|----------|------|-------| +| /media/storage | storepve:/media/storage | 3.6TB | 264GB (8%) | None | +| /media/mediastore | storepve:/media/mediastore | 7.3TB | 5.2TB (75%) | ⚠️ warn at 80% | + +### Proxmox Storage Backends + +| Storage | Type | Content | Active | +|---------|------|---------|--------| +| local | dir | ISO, vztmpl, backup | ✅ | +| local-lvm | lvmthin | images, rootdir | ✅ | +| zfs-vm | nfs | rootdir, images | ✅ | +| mediastore | dir | images, backup, rootdir | ✅ | +| zfs-iso | nfs | vztmpl, iso | ✅ | +| pbs-backup | pbs | backup | ✅ | + +### Checks (every 10 min) ``` -Rule DC-1: Container Not Running - Detect → `docker ps` shows "exited" or "restarting" for critical container - Action → `docker compose up -d ` on the affected stack - Verify → Re-check after 10s - Escalate → If still failed after 3 retries, alert owner - Requires: SSH to docker-vm or CT 116 - -Rule DC-2: Container Restart Looping - Detect → Container restart count > 5 in last hour - Action → `docker logs --tail 50 ` → diagnose → KG log - Escalate → Always alert owner with diagnostic excerpt - Requires: SSH to docker-vm or CT 116 - -Rule DC-3: Docker Host Disk Pressure - Detect → `docker system df` shows < 10% reclaimable or disk > 85% - Action → `docker system prune -af` (safe: removes unused images/containers) - Verify → Re-check disk usage - Escalate → If still >85% after prune, alert owner - Requires: SSH to docker-vm - -Rule DC-4: Stack Health Check Failed - Detect → Container HEALTHCHECK status shows "unhealthy" - Action → `docker compose restart ` - Verify → Re-check health after 15s - Escalate → After 2 retries - Requires: SSH to affected host - -Rule DC-5: Firecrawl API Down - Detect → HTTP GET :3002/health returns non-200 - Action → `docker compose -f /opt/search-stack/firecrawl-source/docker-compose.yaml restart firecrawl-api-1` - Escalate → If still down after 2 retries - Requires: SSH to docker-vm - -Rule DC-6: SearXNG Down - Detect → HTTP GET :8888 returns non-200 - Action → `docker compose -f /opt/search-stack/searxng/docker-compose.yml restart searxng` - Escalate → If still down after 2 retries - Requires: SSH to docker-vm -``` - -### Control (Remediation) - -```yaml -## Enforce -- DC-1 auto-restart: restart critical containers when exited -- DC-3 auto-prune: prune unused Docker artifacts weekly -- DC-4/5/6: restart unhealthy stacks with verification - -## Prohibited -- Never remove volumes (data loss risk) -- Never modify docker-compose files remotely -- Never update container images without owner approval -``` - -## Layer 3: Storage Fabric - -### Storage Topology - -``` -storepve (.6) ── NFS ──► docker-vm: /media/storage (3.6TB, 8% used) - └── audiobookshelf, JDownloader downloads -storepve (.6) ── NFS ──► docker-vm: /media/mediastore (7.3TB, 75% used) - └── media library, configs -storepve ── PBS CT 107 ──► backup target for all VMs/CTs -``` - -### Maintains - -```yaml ## Maintains -- storage.nfs_mounts: each mount's availability, usage %, inode count -- storage.pbs: last successful backup timestamp, backup schedule compliance -- storage.mediastore_growth: weekly growth rate (high-use volume) + +- storage-usage: { mount: string, used_pct: float, growth_rate: float } +- backup-status: { last_success: timestamp, age_hours: int } +- snapshot-status: { pool: string, newest_age_hours: int } + +## Specific Alerts + +- mediastore growth rate > 10GB/day → warn (check what's writing) +- mediastore > 90% → crit (only 730GB remaining) +- No PBS backup in 48h → fail ``` -### Rules +## Section 5: Network Services — Monitoring + +| Service | Domain | Status | +|---------|--------|--------| +| Proxmox API | minipve.sysloggh.net:443 | ✅ | +| Authentik | auth.sysloggh.net:443 | ✅ | +| Gitea | git.sysloggh.net:443 | ✅ | +| Zulip | chat.sysloggh.net:443 | ✅ | +| LiteLLM | litellm.sysloggh.net:443 | ✅ | +| Pulse | pulse.sysloggh.net:443 | ✅ | +| SearXNG | searxng.sysloggh.net:8888 | ✅ | +| Firecrawl | firecrawl.sysloggh.net:3002 | ✅ | + +### Checks (every 2 min) ``` -Rule ST-1: NFS Mount Unreachable - Detect → `df` or `mountpoint -q` fails for NFS path - Action → Check storepve NFS server status via PVE API - Escalate → Immediate alert — containers may be read-only or crashing - Requires: PVE API + SSH to docker-vm - -Rule ST-2: Mediadisk Growth Alert - Detect → mediastore (7.3TB, 75%) growing > 50GB/week - Action → Identify top consumers: `du -sh /media/mediastore/* | sort -rh | head 10` - Log → KG node with growth trend - Escalate → If projected to exceed 90% within 30 days - -Rule ST-3: PBS Backup Compliance - Detect → Any CT/VM without a backup in the last 48h - Action → Log to KG, alert via Zulip DM - Requires: PVE API to check PBS tasks -``` - -## Layer 4: Network Services - -### Service Map - -| Service | Host | IP | Port | Purpose | -|---------|------|----|------|---------| -| adguard | CT 102 | .102 | 53, 80, 443 | DNS + DHCP | -| authentik | CT 104 | .104 | 443 | OIDC auth provider | -| gitea | CT 110 | .110 | 3000, 22 | Git hosting | -| zulip | CT 117 | .117 | 443 | Team chat | -| netbird | 72.61.0.17 | .17 | 443, 8006 | VPN controller | -| minipve | .12 | .12 | 443, 8006 | PVE API endpoint | -| PBS | CT 107 | .107 | 8007 | Proxmox Backup Server | - -### Maintains - -```yaml ## Maintains -- network.dns_resolution: critical hostnames resolve correctly -- network.auth_endpoints: authentik OIDC login flow works -- network.external_access: public endpoints respond (litellm, git, chat, auth, pulse) + +- dns-resolution: { services: [{name, resolves}] } +- ssl-expiry: { services: [{name, days_remaining}] } +- endpoint-reachability: { services: [{name, http_code}] } ``` -### Rules +## Section 6: Alert Routing & Escalation + +### Severity Levels + +| Severity | Channel | Format | Rate Limit | +|----------|---------|--------|-----------| +| **warn** | Zulip DM to owner | "⚠️ {check}: {detail}" | 1x/check/hour | +| **crit** | Zulip DM + #agent-hub | "🚨 {check}: {detail}" | Immediate | +| **escalated** | Zulip DM + relay to maintainer | "🔥 {check}: {detail}" | Immediate | + +### Escalation Chain ``` -Rule NW-1: DNS Failure - Detect → `host` or `nslookup` fails for known hostnames - Action → Check adguard CT 102 status via PVE API - Escalate → If authentik/gitea unreachable due to DNS → page owner - Requires: PVE API + dig/nslookup +Container unhealthy (3 consecutive failures) + → Alert owner + → 15 min no response → Relay to maintainer -Rule NW-2: Auth Endpoint Down - Detect → `GET https://auth.sysloggh.net` returns non-200 - Action → Check authentik CT 104 via PVE API → start if stopped - Escalate → Can't fix: relay to Mumuni (has authentik SSH access) - Requires: PVE API +Node offline + → Alert owner + → 5 min → Attempt restart via PVE API + → 2 failures → Escalate -Rule NW-3: Public Endpoint Unreachable - Detect → Any public URL fails (litellm, git, chat, pulse) - Action → Check container + nginx health on relevant host - Requires: SSH to affected host + curl +Storage > 95% + → Alert owner + maintainer relay + → Immediate action required + +SSL cert < 7 days + → Alert owner daily + → 3 days out → Escalate + +ZFS pool degraded + → Cannot auto-fix → Escalate immediately ``` -## Layer 5: Agent Health +## Appendix A: Quick Health Commands -### Agent Registry +```bash +# Full cluster status +PVE="https://minipve.sysloggh.net" +AUTH="Authorization: PVEAPIToken=monitoring@pve!mumuni=eafd56c5-93d4-4d40-a41d-e688be0987f3" +curl -sfk "$PVE/api2/json/cluster/resources" -H "$AUTH" -| Agent | Platform | CT | Status Endpoint | Last Seen | -|-------|----------|----|----------------|-----------| -| abiba | pi | 100 | :9200/health | live | -| tanko | Hermes | 122 | (via Gateway) | live | -| mumuni | Hermes | 114 | (via Gateway) | live | -| kagentz | Agent Zero | 105 | (via PVE) | running | -| koby | Hermes | 111 | (via PVE) | stopped | -| koonimo | Hermes | 113 | (via PVE) | stopped | +# Docker health from Abiba +ssh root@192.168.68.7 "docker ps --format '{{.Names}} {{.Status}}'" +ssh root@192.168.68.116 "docker ps --format '{{.Names}} {{.Status}}'" -### Rules - -``` -Rule AG-1: Agent VM/CT Stopped - Detect → PVE API shows agent CT as "stopped" - Action → `pct start ` via PVE API - Verify → Re-check after 15s - Escalate → If persistent, alert owner - Requires: PVE API - -Rule AG-2: Agent Not Responding - Detect → Health endpoint non-200 or Gateway disconnected - Action → Check CT is running, then check service status - Requires: PVE API + SSH +# Storage check +ssh root@192.168.68.7 "df -h /media/storage /media/mediastore" ``` -## Access Matrix +## Appendix B: CT Inventory -| Resource | Abiba (CT 100) | Tanko (CT 122) | Mumuni (CT 114) | -|----------|---------------|----------------|-----------------| -| **PVE API** (minipve:443) | ✅ API token | ❌ | ❌ | -| **docker-vm** (SSH .7) | ✅ root SSH | ❌ | ❌ | -| **CT 116** (SSH) | ✅ root SSH | ❌ | ❌ | -| **Gitea** | ✅ API token | Need to add | Need to add | -| **Zulip API** | ✅ bot token | ✅ bot token | ✅ bot token | -| **RA-H OS** | ✅ MCP bridge | ✅ MCP bridge | ✅ MCP bridge | -| **Netbird** (SSH .17) | ✅ root SSH | ❌ | ❌ | -| **Authentik** | ❌ No SSH | ❌ No SSH | ❌ No SSH | +| CT | Name | Node | IP | Role | Agent | +|----|------|------|----|------|-------| +| 100 | abiba | amdpve | .24 | Pi agent (this host) | ✅ pi | +| 101 | llm-gpu | acerpve | — | GPU VM | ❌ | +| 102 | adguard | acerpve | — | DNS | ❌ | +| 103 | ocu-llm | ocupve | — | GPU VLM | ❌ | +| 104 | authentik | minipve | .11 | OIDC | ❌ | +| 105 | kagentz | amdpve | — | Agent Zero | ✅ | +| 106 | ra-h-os | storepve | .65 | KG bridge | ✅ MCP | +| 107 | pbs | storepve | — | Backups | ❌ | +| 108 | media | storepve | — | Media | ❌ | +| 109 | docker-vm | storepve | .7 | Docker host | ❌ | +| 110 | gitea | minipve | — | Git | ❌ | +| 111 | tdunna | amdpve | — | ? | ❌ | +| 112 | tanko | amdpve | .122 | Hermes agent | ✅ | +| 113 | baggy | amdpve | — | ? | ❌ | +| 114 | mumuni | minipve | — | Hermes agent | ✅ | +| 115 | scottdenya | amdpve | — | ? | ❌ | +| 116 | syslog-api | minipve | .116 | LiteLLM stack | ❌ | +| 117 | zulip | storepve | — | Chat | ❌ | +| 118 | jitsi | minipve | — | Video | ❌ | -## Priority Matrix for Auto-Remediation +## Appendix C: Docker Compose Files Location -| Priority | Condition | Action | SLA | -|----------|-----------|--------|-----| -| **P0** | Critical CT/VM stopped | Auto-start via PVE API | 2 min | -| **P0** | LiteLLM/router down | Auto-restart container | 3 min | -| **P1** | Firecrawl/SearXNG down | Auto-restart container | 5 min | -| **P1** | Docker disk > 85% | Auto-prune | 5 min | -| **P2** | Storage usage > 80% | Alert + log | 15 min | -| **P2** | CT resource pressure | Alert + log | 15 min | -| **P3** | Non-critical container down | Alert + log | 30 min | -| **P3** | PBS backup missed | Alert + log | 1h | - -## Escalation Path - -``` -1. Auto-remediation ──► 2. Zulip DM to owner ──► 3. RA-H OS relay to agent with access - │ │ - ▼ ▼ - Kwame (Zulip DM) Mumuni (for authentik) - Tanko (for hermes agents) -``` - -## Known Gaps - -| Gap | What's Needed | -|-----|--------------| -| Authentik SSH | No agent has SSH — Mumuni needs key added | -| Koby/Koonimo stopped | CTs 111/113 are offline — need human intervention | -| Mumuni PVE access | No PVE token — needs `monitoring@pve!mumuni` | -| docker-vm backups | VM 109 isn't running PBS agent — backup status unknown | -| Netbird redundancy | Single point of failure for VPN | +| Host | Stack | Compose File | +|------|-------|-------------| +| docker-vm (.7) | Firecrawl | `/opt/search-stack/firecrawl-source/docker-compose.yaml` | +| docker-vm (.7) | SearXNG | `/opt/search-stack/searxng/docker-compose.yml` | +| docker-vm (.7) | Home stack | `/opt/home_stack/docker-compose.yml` | +| docker-vm (.7) | Audiobookshelf | `/opt/audiobookshelf/docker-compose.yml` | +| CT 116 (.116) | LiteLLM | `/root/docker-compose-litellm.yml` | +| Netbird (.17) | Netbird | Docker run (not compose) |