Full discovery and expansion based on PVE API exploration: Proxmox (5 nodes): - minipve (.12): authentik, gitea, mumuni, syslog-api, jitsi - amdpve (.15): abiba, kagentz, tanko, tdunna, baggy, scottdenya - storepve (.6): docker-vm, ra-h-os, PBS, media, zulip - acerpve (.9): llm-gpu, adguard - ocupve (.5): ocu-llm Docker (3 ecosystems, 22 containers): - docker-vm (.7): Firecrawl, SearXNG, Home stack, Audiobookshelf - CT 116: LiteLLM stack (6 containers) - Netbird (.17): VPN controller + dashboard + crowdsec + traefik Pattern now includes: access matrix, per-section checks/remediations, alert routing, escalation chain, quick-start commands, full CT inventory.
13 KiB
13 KiB
kind, name, description
| kind | name | description |
|---|---|---|
| pattern | infrastructure-control | Full infrastructure monitoring and control pattern covering the 5-node Proxmox cluster, 3 Docker ecosystems (22 containers), NFS storage, and network services. Defines monitors, remediations, and the access matrix for all environments. |
Infrastructure Control Pattern
Topology
┌─────────────────────────┐
│ OpenProse Contract │
│ (declares what's true) │
└──────────┬──────────────┘
│
┌────────────────────┼────────────────────┐
▼ ▼ ▼
┌─────────────┐ ┌──────────────┐ ┌──────────────┐
│ Abiba │ │ Tanko │ │ Mumuni │
│ (pi) │ │ (Hermes) │ │ (Hermes) │
│ CT 100 │ │ CT 122 │ │ CT 114 │
└──────┬──────┘ └──────┬───────┘ └──────┬───────┘
│ │ │
└──────────────────┼────────────────────┘
▼
┌──────────────────────────────────────┐
│ Proxmox Cluster API │
│ minipve.sysloggh.net:443 │
│ (monitoring@pve!mumuni token) │
└────┬──────┬──────┬──────┬──────┬─────┘
│ │ │ │ │
┌────┘ ┌────┘ ┌────┘ ┌────┘ ┌────┘
▼ ▼ ▼ ▼ ▼
minipve amdpve storepve acerpve ocupve
(.12) (.15) (.6) (.9) (.5)
▼
┌─────────────────────────────────────────────┐
│ Docker Ecosystems │
│ ┌────────────┬────────────┬──────────────┐ │
│ │ docker-vm │ CT 116 │ Netbird │ │
│ │ (.7) │ (.116) │ (72.61.0.17) │ │
│ │ 11 ctrs │ 6 ctrs │ 5 ctrs │ │
│ └────────────┴────────────┴──────────────┘ │
└─────────────────────────────────────────────┘
Section 1: Access Matrix
API Tokens & Credentials
| Resource | Auth Method | Credential Source | Status |
|---|---|---|---|
| Proxmox Cluster | PVE API Token | monitoring@pve!mumuni=... |
✅ |
| Proxmox Root | Password via API ticket | root@pam:kakashi19 |
✅ |
| docker-vm (.7) | SSH root | SSH key | ✅ |
| CT 116 (syslog-api) | SSH root | SSH key | ✅ |
| Tanko CT (.122) | SSH jerome | SSH key | ✅ |
| Netbird (.17) | SSH root | SSH key | ✅ |
| Gitea | API token | abiba-bot token | ✅ |
| Zulip | Bot API key | per-bot tokens | ✅ |
| RA-H OS | MCP bridge | port 3100 | ✅ |
Reachability Matrix
| From / To | PVE API | docker-vm (.7) | CT 116 | Tanko (.122) | Netbird (.17) |
|---|---|---|---|---|---|
| Abiba (CT 100) | ✅ :443 | ✅ SSH | ✅ SSH | ✅ SSH | ❌ no Netbird |
| Tanko (CT 122) | ❌ | ❌ | ❌ | ✅ | ❌ |
| docker-vm (.7) | ❌ | ✅ | ❌ | ❌ | ❌ |
Conclusion: Only Abiba has cross-infrastructure access. All monitoring contracts run from Abiba.
Section 2: Proxmox Cluster — Monitoring
Nodes (5)
| Node | IP | CPU | RAM | VMs/CTs | Role |
|---|---|---|---|---|---|
| minipve | .12 | 16C | 30GB | authentik, gitea, mumuni, syslog-api, jitsi | Auth, git, messaging |
| amdpve | .15 | 32C | 62GB | abiba, kagentz, tanko, tdunna, baggy, scottdenya | Agents, compute |
| storepve | .6 | 28C | 31GB | docker-vm, ra-h-os, PBS, media, zulip | Docker, storage, chat |
| acerpve | .9 | 28C | 31GB | llm-gpu, adguard | GPU VMs |
| ocupve | .5 | 12C | 14GB | ocu-llm | GPU VMs |
Checks (every 5 min)
## Maintains
- cluster-status: { online_nodes: int, offline_nodes: int, timestamp }
- node-cpu-usage: { node: pct, warnings: [] }
- node-memory-usage: { node: free_pct, warnings: [] }
- node-uptime: { node: seconds, just_rebooted: bool }
## Checks
- For each node in [minipve, amdpve, storepve, acerpve, ocupve]:
- GET /api2/json/nodes/{node}/status → check status == "online"
- GET /api2/json/nodes/{node}/status → cpu < 0.80
- GET /api2/json/nodes/{node}/status → free_mem > 10%
- GET /api2/json/nodes/{node}/status → uptime > 300 (warn if just rebooted)
- For each CT in inventory (19 total):
- GET /api2/json/cluster/resources → filter by type=lxc
- Warn if status != "running"
- Storage pools:
- GET /api2/json/nodes/storepve/storage/mediastore → used < 80%
- GET /api2/json/nodes/storepve/storage/pbs-backup → last backup < 48h
## Remediations
- Node offline → alert via Zulip DM, escalate after 3x
- CT stopped → `pct start <vmid>` via PVE API, verify after 30s
- VM stopped → `qm start <vmid>` via PVE API, verify after 60s
- Storage > 80% → warn via Zulip DM
- Storage > 95% → crit via Zulip + relay to maintainer
Section 3: Docker Ecosystems — Monitoring
Ecosystem A: docker-vm (192.168.68.7)
11 containers across 4 compose stacks:
| Stack | Path | Containers |
|---|---|---|
| Firecrawl | /opt/search-stack/firecrawl-source/ |
api, rabbitmq, postgres, playwright, redis |
| SearXNG | /opt/search-stack/searxng/ |
searxng, valkey |
| Home stack | /opt/home_stack/ |
jdownloader, bentopdf, pulse |
| Audiobookshelf | /opt/audiobookshelf/ |
audiobookshelf |
Ecosystem B: CT 116 syslog-api (192.168.68.116)
| Container | Image | Port |
|---|---|---|
| harness-litellm | berriai/litellm:1.90.0-rc.1 | :4001 |
| harness-nginx | nginx:alpine | :80 |
| harness-router | inference-harness-router | :9000 |
| harness-postgres | postgres:16-alpine | :5432 |
| harness-redis | redis:7-alpine | :6379 |
| harness-dashboard | inference-harness-dashboard | :3000 |
Ecosystem C: Netbird (72.61.0.17)
| Container | Image | Role |
|---|---|---|
| netbird-server | netbird | VPN controller |
| netbird-dashboard | netbird UI | Web management |
| netbird-proxy | nginx | TLS termination |
| netbird-crowdsec | crowdsecurity/crowdsec:v1.7.7 | WAF |
| netbird-traefik | traefik | Reverse proxy |
Checks (every 60s)
## Maintains
- docker-health: { ecosystem: string, healthy: int, unhealthy: int, total: int }
- container-status: { name: string, status: string, restarts: int, image: string }
## Checks
For each Docker host:
- docker ps --format "{{.Names}}" → every required container is running
- docker inspect <name> --format "{{.State.Health.Status}}" → "healthy"
- docker inspect <name> --format "{{.RestartCount}}" → < 3/hour
- df -h / | awk '{print $5}' → usage < 80%
For docker-vm specifically:
- mountpoint -q /media/storage → NFS mounted
- mountpoint -q /media/mediastore → NFS mounted
- docker compose ls → expected stacks present
## Remediations
- Container unhealthy → `docker compose up -d <name>`, wait 10s, re-check
- Container missing → `docker compose up -d` in its stack directory
- Docker daemon down → `systemctl restart docker`
- NFS mount lost → `mount -a`, if fails → alert (storepve issue)
- Disk > 80% → alert via Zulip DM
- Disk > 95% → crit + relay to maintainer
Section 4: Storage — Monitoring
NFS Mounts (docker-vm → storepve)
| Mount | Export | Capacity | Used | Alert |
|---|---|---|---|---|
| /media/storage | storepve:/media/storage | 3.6TB | 264GB (8%) | None |
| /media/mediastore | storepve:/media/mediastore | 7.3TB | 5.2TB (75%) | ⚠️ warn at 80% |
Proxmox Storage Backends
| Storage | Type | Content | Active |
|---|---|---|---|
| local | dir | ISO, vztmpl, backup | ✅ |
| local-lvm | lvmthin | images, rootdir | ✅ |
| zfs-vm | nfs | rootdir, images | ✅ |
| mediastore | dir | images, backup, rootdir | ✅ |
| zfs-iso | nfs | vztmpl, iso | ✅ |
| pbs-backup | pbs | backup | ✅ |
Checks (every 10 min)
## Maintains
- storage-usage: { mount: string, used_pct: float, growth_rate: float }
- backup-status: { last_success: timestamp, age_hours: int }
- snapshot-status: { pool: string, newest_age_hours: int }
## Specific Alerts
- mediastore growth rate > 10GB/day → warn (check what's writing)
- mediastore > 90% → crit (only 730GB remaining)
- No PBS backup in 48h → fail
Section 5: Network Services — Monitoring
| Service | Domain | Status |
|---|---|---|
| Proxmox API | minipve.sysloggh.net:443 | ✅ |
| Authentik | auth.sysloggh.net:443 | ✅ |
| Gitea | git.sysloggh.net:443 | ✅ |
| Zulip | chat.sysloggh.net:443 | ✅ |
| LiteLLM | litellm.sysloggh.net:443 | ✅ |
| Pulse | pulse.sysloggh.net:443 | ✅ |
| SearXNG | searxng.sysloggh.net:8888 | ✅ |
| Firecrawl | firecrawl.sysloggh.net:3002 | ✅ |
Checks (every 2 min)
## Maintains
- dns-resolution: { services: [{name, resolves}] }
- ssl-expiry: { services: [{name, days_remaining}] }
- endpoint-reachability: { services: [{name, http_code}] }
Section 6: Alert Routing & Escalation
Severity Levels
| Severity | Channel | Format | Rate Limit |
|---|---|---|---|
| warn | Zulip DM to owner | "⚠️ {check}: {detail}" | 1x/check/hour |
| crit | Zulip DM + #agent-hub | "🚨 {check}: {detail}" | Immediate |
| escalated | Zulip DM + relay to maintainer | "🔥 {check}: {detail}" | Immediate |
Escalation Chain
Container unhealthy (3 consecutive failures)
→ Alert owner
→ 15 min no response → Relay to maintainer
Node offline
→ Alert owner
→ 5 min → Attempt restart via PVE API
→ 2 failures → Escalate
Storage > 95%
→ Alert owner + maintainer relay
→ Immediate action required
SSL cert < 7 days
→ Alert owner daily
→ 3 days out → Escalate
ZFS pool degraded
→ Cannot auto-fix → Escalate immediately
Appendix A: Quick Health Commands
# Full cluster status
PVE="https://minipve.sysloggh.net"
AUTH="Authorization: PVEAPIToken=monitoring@pve!mumuni=eafd56c5-93d4-4d40-a41d-e688be0987f3"
curl -sfk "$PVE/api2/json/cluster/resources" -H "$AUTH"
# Docker health from Abiba
ssh root@192.168.68.7 "docker ps --format '{{.Names}} {{.Status}}'"
ssh root@192.168.68.116 "docker ps --format '{{.Names}} {{.Status}}'"
# Storage check
ssh root@192.168.68.7 "df -h /media/storage /media/mediastore"
Appendix B: CT Inventory
| CT | Name | Node | IP | Role | Agent |
|---|---|---|---|---|---|
| 100 | abiba | amdpve | .24 | Pi agent (this host) | ✅ pi |
| 101 | llm-gpu | acerpve | — | GPU VM | ❌ |
| 102 | adguard | acerpve | — | DNS | ❌ |
| 103 | ocu-llm | ocupve | — | GPU VLM | ❌ |
| 104 | authentik | minipve | .11 | OIDC | ❌ |
| 105 | kagentz | amdpve | — | Agent Zero | ✅ |
| 106 | ra-h-os | storepve | .65 | KG bridge | ✅ MCP |
| 107 | pbs | storepve | — | Backups | ❌ |
| 108 | media | storepve | — | Media | ❌ |
| 109 | docker-vm | storepve | .7 | Docker host | ❌ |
| 110 | gitea | minipve | — | Git | ❌ |
| 111 | tdunna | amdpve | — | ? | ❌ |
| 112 | tanko | amdpve | .122 | Hermes agent | ✅ |
| 113 | baggy | amdpve | — | ? | ❌ |
| 114 | mumuni | minipve | — | Hermes agent | ✅ |
| 115 | scottdenya | amdpve | — | ? | ❌ |
| 116 | syslog-api | minipve | .116 | LiteLLM stack | ❌ |
| 117 | zulip | storepve | — | Chat | ❌ |
| 118 | jitsi | minipve | — | Video | ❌ |
Appendix C: Docker Compose Files Location
| Host | Stack | Compose File |
|---|---|---|
| docker-vm (.7) | Firecrawl | /opt/search-stack/firecrawl-source/docker-compose.yaml |
| docker-vm (.7) | SearXNG | /opt/search-stack/searxng/docker-compose.yml |
| docker-vm (.7) | Home stack | /opt/home_stack/docker-compose.yml |
| docker-vm (.7) | Audiobookshelf | /opt/audiobookshelf/docker-compose.yml |
| CT 116 (.116) | LiteLLM | /root/docker-compose-litellm.yml |
| Netbird (.17) | Netbird | Docker run (not compose) |