PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
- Remove kagentz from amdpve node table (belongs on hwepve) - Fix kagentz pct-run table: hwepve (was amdpve) - Fix mumuni pct-run table: hwepve (was minipve) - Fix Zulip recovery command: wrap in single SSH call
681 lines
30 KiB
Markdown
681 lines
30 KiB
Markdown
---
|
|
kind: pattern
|
|
name: infrastructure-control
|
|
description: >
|
|
Full infrastructure monitoring and control pattern covering the
|
|
6-node Proxmox cluster, 3 Docker ecosystems (22 containers),
|
|
NFS storage, and network services. Defines monitors, remediations,
|
|
and the access matrix for all environments.
|
|
|
|
FIELD TRUST: Live-state fields (IPs, ports, hostnames, credentials,
|
|
container names, PIDs) are marked VERIFY-BEFORE-USE. They drift —
|
|
never mutate infrastructure based on them without first confirming
|
|
against the live system. Policy fields are authoritative. See the
|
|
`verify-before-mutate` skill.
|
|
|
|
**Last verified:** 2026-07-24 — corrected Gitea IP (.17 not .110),
|
|
AdGuard IP (.10 not .102), AdGuard placement (minipve not acerpve),
|
|
Abiba placement (hwepve not amdpve), added hwepve as 6th node,
|
|
added dns.sysloggh.net route. See data/learnings.md.
|
|
---
|
|
|
|
# Infrastructure Control Pattern
|
|
|
|
## Topology
|
|
|
|
```
|
|
┌─────────────────────────┐
|
|
│ OpenProse Contract │
|
|
│ (declares what's true) │
|
|
└──────────┬──────────────┘
|
|
│
|
|
┌────────────────────┼────────────────────┐
|
|
▼ ▼ ▼
|
|
┌─────────────┐ ┌──────────────┐ ┌──────────────┐
|
|
│ Abiba │ │ Tanko │ │ Mumuni │
|
|
│ (pi) │ │ (Hermes) │ │ (Hermes) │
|
|
│ CT 100 │ │ CT 112 │ │ CT 114 │
|
|
└──────┬──────┘ └──────┬───────┘ └──────┬───────┘
|
|
│ │ │
|
|
└──────────────────┼────────────────────┘
|
|
▼
|
|
┌──────────────────────────────────────┐
|
|
│ Proxmox Cluster API │
|
|
│ minipve.sysloggh.net:443 │
|
|
│ (monitoring@pve!mumuni token) │
|
|
└────┬──────┬──────┬──────┬──────┬─────┘
|
|
│ │ │ │ │
|
|
┌────┘ ┌────┘ ┌────┘ ┌────┘ ┌────┘
|
|
▼ ▼ ▼ ▼ ▼
|
|
minipve amdpve storepve acerpve ocupve hwepve
|
|
(.12) (.15) (.6) (.9) (.5) (.4)
|
|
|
|
▼
|
|
┌─────────────────────────────────────────────┐
|
|
│ Docker Ecosystems │
|
|
│ ┌────────────┬────────────┬──────────────┐ │
|
|
│ │ docker-vm │ CT 116 │ Netbird │ │
|
|
│ │ (.7) │ (.116) │ (72.61.0.17) │ │
|
|
│ │ 11 ctrs │ 6 ctrs │ 5 ctrs │ │
|
|
│ └────────────┴────────────┴──────────────┘ │
|
|
└─────────────────────────────────────────────┘
|
|
```
|
|
|
|
## Section 1: Access Matrix
|
|
|
|
### API Tokens & Credentials
|
|
|
|
| Resource | Auth Method | Credential Source | Status |
|
|
|----------|------------|-------------------|--------|
|
|
| Proxmox Cluster | PVE API Token | Infisical vault (`PROXMOX_API_TOKEN`) | ✅ |
|
|
| Proxmox Root | Password via API ticket | Infisical vault (`PROXMOX_ROOT_PASSWORD`) | ✅ |
|
|
| docker-vm (.7) | SSH root | SSH key | ✅ |
|
|
| CT 116 (syslog-api) | SSH root | SSH key | ✅ |
|
|
| Tanko CT (.122) | SSH jerome | id_ed25519 | ✅ |
|
|
| Mumuni CT (.123) | SSH root | id_ed25519 | ✅ |
|
|
| Baggy CT (113) | SSH jerome | ❌ no key access |
|
|
| Netbird (.17) | SSH root | SSH key | ✅ |
|
|
| Gitea | API token | Infisical vault (`GITEA_BOT_TOKEN`) | ✅ |
|
|
| Zulip | Bot API key | Infisical vault (`ZULIP_BOT_KEY`) | ✅ |
|
|
| RA-H OS | MCP bridge | port 3100 | ✅ |
|
|
| LiteLLM Admin | master key | Infisical vault (`LITELLM_MASTER_KEY`) | ✅ VERIFY-BEFORE-USE |
|
|
| Grafana | admin password | Infisical vault (`GRAFANA_ADMIN_PASSWORD`) | ✅ VERIFY-BEFORE-USE |
|
|
|
|
> **VERIFY-BEFORE-USE**: Credentials, IPs, ports, and hostnames in this
|
|
> contract are live-state fields. Test them against the live system before
|
|
> relying on them (e.g., `curl -H "Authorization: Bearer $KEY" .../v1/models`
|
|
> for a LiteLLM key; `pct config <ct>` for a CT IP). Drift is expected —
|
|
> see `verify-before-mutate` skill. Policy fields (rules, doctrines) are
|
|
> authoritative and do not need verification.
|
|
|
|
### Reachability Matrix
|
|
|
|
| From / To | PVE API | docker-vm (.7) | CT 116 | Tanko (.122) | Mumuni (.123) | Baggy (.114) |
|
|
|-----------|---------|----------------|--------|-------------|---------------|----------------|
|
|
| **Abiba** (.24) | ✅ :443 | ✅ SSH | ✅ SSH | ✅ SSH jerome | ✅ SSH root | ❌ SSH |
|
|
| **Tanko** (.122) | ❌ | ❌ | ❌ via NetBird | ✅ | ❌ | ❌ |
|
|
| **Mumuni** (.123) | ❌ | ❌ | ❌ | ❌ | ✅ | ❌ |
|
|
|
|
**Conclusion:** Only Abiba has cross-infrastructure access. All monitoring contracts run from Abiba.
|
|
|
|
## Section 2: Proxmox Cluster — Monitoring
|
|
|
|
### Nodes (6)
|
|
|
|
| Node | IP | CPU | RAM | VMs/CTs | Role |
|
|
|------|----|-----|-----|---------|------|
|
|
| minipve | .12 | 16C | 30GB | authentik, gitea, syslog-api, infisical-vault, jitsi | Auth, git, messaging |
|
|
| amdpve | .15 | 32C | 62GB | tanko, tdunna, baggy, scottdenya | Agents, compute |
|
|
| storepve | .6 | 28C | 31GB | docker-vm, ra-h-os, PBS, media, jdownloader, zulip | Docker, storage, chat |
|
|
| acerpve | .9 | 28C | 31GB | llm-gpu | GPU VMs |
|
|
| ocupve | .5 | 12C | 14GB | ocu-llm | GPU VMs |
|
|
| hwepve | .4 | ? | ? | abiba, (mumuni CT 114 stopped) | Agents (new node) |
|
|
|
|
> **Note:** CTs on storepve include jdownloader (CT 118). AdGuard (CT 102) is on
|
|
> minipve at .10, not acerpve. Abiba (CT 100) is on hwepve, not amdpve. Mumuni
|
|
> (CT 114) is on hwepve (currently stopped), not minipve. Mumuni also has a
|
|
> second instance on minipve at .123 — distinguish by CT ID, not hostname.
|
|
|
|
### Checks (every 5 min)
|
|
|
|
```
|
|
## Maintains
|
|
|
|
- cluster-status: { online_nodes: int, offline_nodes: int, timestamp }
|
|
- node-cpu-usage: { node: pct, warnings: [] }
|
|
- node-memory-usage: { node: free_pct, warnings: [] }
|
|
- node-uptime: { node: seconds, just_rebooted: bool }
|
|
|
|
## Checks
|
|
|
|
- For each node in [minipve, amdpve, storepve, acerpve, ocupve]:
|
|
- GET /api2/json/nodes/{node}/status → check status == "online"
|
|
- GET /api2/json/nodes/{node}/status → cpu < 0.80
|
|
- GET /api2/json/nodes/{node}/status → free_mem > 10%
|
|
- GET /api2/json/nodes/{node}/status → uptime > 300 (warn if just rebooted)
|
|
|
|
- For each CT in inventory (19 total):
|
|
- GET /api2/json/cluster/resources → filter by type=lxc
|
|
- Warn if status != "running"
|
|
|
|
- Storage pools:
|
|
- GET /api2/json/nodes/storepve/storage/mediastore → used < 80%
|
|
- GET /api2/json/nodes/storepve/storage/pbs-backup → last backup < 48h
|
|
|
|
## Remediations
|
|
|
|
- Node offline → alert via Zulip DM, escalate after 3x
|
|
- CT stopped → `pct start <vmid>` via PVE API, verify after 30s
|
|
- VM stopped → `qm start <vmid>` via PVE API, verify after 60s
|
|
- Storage > 80% → warn via Zulip DM
|
|
- Storage > 95% → crit via Zulip + relay to maintainer
|
|
```
|
|
|
|
## Section 3: Docker Ecosystems — Monitoring
|
|
|
|
### Ecosystem A: docker-vm (192.168.68.7)
|
|
|
|
16 containers across 4 compose stacks + trove agents:
|
|
|
|
| Stack | Path | Containers |
|
|
|-------|------|-----------|
|
|
| **Firecrawl** | `/opt/search-stack/firecrawl-source/` | api, rabbitmq, postgres, playwright, redis |
|
|
| **SearXNG** | `/opt/search-stack/searxng/` | searxng, valkey |
|
|
| **Home stack** | `/opt/home_stack/` | jdownloader, stirling-pdf, pulse |
|
|
| **Audiobookshelf** | `/opt/audiobookshelf/` | audiobookshelf |
|
|
| **Trove agents** | docker run (standalone) | trove-agent-proxmox, trove-test-agent-1, trove-test-server-1, docker-stats |
|
|
|
|
**Last verified:** 2026-07-09 — 16/16 containers running healthy.
|
|
|
|
### Home Stack Services
|
|
|
|
**Stirling-PDF** (deployed 2026-07-03, Authentik SSO 2026-07-03):
|
|
- URL: `https://pdf.sysloggh.net` (public) / `http://192.168.68.7:8989` (direct)
|
|
- Swagger: `http://192.168.68.7:8989/swagger-ui.html`
|
|
- Admin credentials: `admin` / `kakashi20stirling`
|
|
- API key: `adefaef837314afc37803747049c2e73f456da97699ff1b02b391347c4a3cb88`
|
|
- Authentik OAuth2: configured but disabled (requires paid Server license). Ready to enable: set `SECURITY_OAUTH2_ENABLED=true` + `SECURITY_LOGINMETHOD=all`
|
|
- Compose: `/opt/home_stack/docker-compose.yml`
|
|
- Control script: `/opt/home_stack/infra-control.sh`
|
|
|
|
**JDownloader**:
|
|
- URL: `http://192.168.68.7:5800` (web UI via VNC)
|
|
|
|
**Pulse** (Uptime Kuma):
|
|
- URL: `http://192.168.68.7:3001`
|
|
|
|
### Ecosystem B: CT 116 syslog-api (192.168.68.116)
|
|
|
|
8 containers in inference-harness stack:
|
|
|
|
| Container | Image | Port | Role |
|
|
|-----------|-------|------|------|
|
|
| harness-litellm | berriai/litellm:1.90.0-rc.1 | :4000→:4001 | API proxy, key mgmt, fallbacks |
|
|
| harness-router | inference-harness-router | :9000 (127.0.0.1) | GPU routing, slot booking, CB |
|
|
| harness-nginx | nginx:alpine | :80 | Entrypoint, /v1→LiteLLM, /dashboard/ |
|
|
| harness-postgres | postgres:16-alpine | :5432 | LiteLLM DB (keys, spend, config) |
|
|
| harness-redis | redis:7-alpine | :6379 | Router slots, circuit breakers |
|
|
| harness-dashboard | inference-harness-dashboard | :3000 | SyslogAI Harness UI |
|
|
| harness-grafana | grafana/grafana | :3000→:3001 (direct LAN, not behind nginx) | GPU + Proxmox + Docker dashboards |
|
|
| harness-prometheus | prom/prometheus | :9090 | Metrics scraper, 6 jobs |
|
|
|
|
**Nginx routing**:
|
|
- `/v1/*` → harness-litellm:4000 (API)
|
|
- `/admin/*` → harness-litellm:4000
|
|
- `/dashboard/` → harness-dashboard:3000
|
|
- `/litellm/*` → harness-litellm:4000
|
|
- `/health/*` → harness-litellm:4000/health/liveliness
|
|
- `/gpu/` → 192.168.68.24:9100 (fleet monitor)
|
|
|
|
**Prometheus targets**:
|
|
- 192.168.68.8:9400 (RTX 3090 — qwen)
|
|
- 192.168.68.110:9400 (RTX 5070 — gemma)
|
|
- 192.168.68.15:9400 (Strix Halo — qwen3.6-35B-udq4)
|
|
- 192.168.68.24:9401 (Router metrics exporter)
|
|
- harness-litellm:4000 (LiteLLM health)
|
|
|
|
### Ecosystem C: Netbird (72.61.0.17 — Hostinger srv1079750.hstgr.cloud)
|
|
|
|
5 containers, single compose stack at `/root/docker-compose.yml`:
|
|
|
|
| Container | Image | Role |
|
|
|-----------|-------|------|
|
|
| netbird-server | netbirdio/netbird-server:latest | VPN controller, management API |
|
|
| netbird-dashboard | netbirdio/dashboard:latest | Web management UI |
|
|
| netbird-proxy | netbirdio/reverse-proxy:latest | SNI router + TLS passthrough for *.sysloggh.net |
|
|
| netbird-crowdsec | crowdsecurity/crowdsec:v1.7.7 | WAF / threat intelligence |
|
|
| netbird-traefik | traefik:v3.6 | Edge reverse proxy, Let's Encrypt TLS |
|
|
|
|
**Architecture:** Internet → Traefik (TLS + HTTP→HTTPS redirect) → Netbird proxy
|
|
(TLS passthrough via HostSNI, PROXY protocol) → Netbird mesh → LAN backends.
|
|
Traefik handles the Netbird dashboard directly; all other `*.sysloggh.net`
|
|
domains are TLS-passthrough to the proxy on port 8443.
|
|
|
|
**Proxy config:** `/root/proxy.env` — token, ACME certs in `/certs/`, geolocation
|
|
DB auto-downloaded. Certs are per-domain Let's Encrypt via TLS-ALPN-01 challenge.
|
|
|
|
> **Dependency note:** NetBird is an **access layer only** — it exists so a
|
|
> human on a PC or away from the LAN can reach services by URL. It is **not**
|
|
> a dependency of any service backend. Per Section 7, every config and agent
|
|
> must reach backends by LAN IP; the `*.sysloggh.net` URLs are reserved for
|
|
> human browsers.
|
|
|
|
#### Known Issues
|
|
|
|
**A. Access log bloat (CRITICAL)** — 2026-07-03 discovery. The management server
|
|
writes an `access_log_entries` row for every proxied request (~45K/day). SQLite
|
|
has no built-in TTL. By 2026-07-03 the table reached 10.7M rows / 5.4GB, causing:
|
|
- DB writes slow → gRPC deadlines exceeded → proxy loses management connection
|
|
- On proxy restart, management server too slow to push route mappings
|
|
- Proxy starts HTTPS listener before routes arrive → all domains "unknown"
|
|
- **Fix:** batch-delete old entries, VACUUM. Prevention: weekly systemd timer (`netbird-cleanup.timer`) deletes entries >7 days old when DB exceeds 500MB.
|
|
- **Recurrence signal:** proxy logs show `unknown domain "git.sysloggh.net"` or `match: false` for known domains → management route sync failed.
|
|
|
|
**B. Stale management state after proxy restart.** Management server sometimes
|
|
stops pushing route mappings to a reconnecting proxy. Symptom: proxy starts,
|
|
mapping stream established, but no routes arrive → all domains `unknown`.
|
|
**Fix:** restart management server FIRST, THEN proxy. Order matters.
|
|
|
|
**C. Peer `sha-OWLpE8nS4Rp4IWzep+y2ayFeZ6C3lJ4eNRB29L5yQTE=` flapping.**
|
|
One peer (connecting through Traefik relay at 172.30.0.10) disconnects every
|
|
~2-3 minutes. This may be a symptom of DB bloat (healthcheck timeout due to
|
|
slow DB writes) rather than a network issue. Monitor after bloat fix.
|
|
|
|
#### Recovery Procedures
|
|
|
|
1. **Proxy losing all routes ("unknown domain" for all services):**
|
|
```bash
|
|
ssh root@72.61.0.17
|
|
cd /root
|
|
docker compose restart netbird-server
|
|
sleep 12
|
|
docker compose restart proxy
|
|
```
|
|
2. **Access log bloat detected (DB > 500MB):**
|
|
```bash
|
|
ssh root@72.61.0.17
|
|
/root/cleanup-access-logs.sh
|
|
```
|
|
3. **VPS hung (TCP accepts, SSH banner timeout):** hard reboot from Hostinger console.
|
|
|
|
### Checks (every 60s)
|
|
|
|
```
|
|
## Maintains
|
|
|
|
- docker-health: { ecosystem: string, healthy: int, unhealthy: int, total: int }
|
|
- container-status: { name: string, status: string, restarts: int, image: string }
|
|
|
|
## Checks
|
|
|
|
For each Docker host:
|
|
- docker ps --format "{{.Names}}" → every required container is running
|
|
- docker inspect <name> --format "{{.State.Health.Status}}" → "healthy"
|
|
- docker inspect <name> --format "{{.RestartCount}}" → < 3/hour
|
|
- df -h / | awk '{print $5}' → usage < 80%
|
|
|
|
For Netbird VPS specifically:
|
|
- netbird-store-db-size: du -m /var/lib/docker/volumes/root_netbird_data/_data/store.db → < 500MB
|
|
- netbird-access-log-count: sqlite3 store.db 'SELECT COUNT(*) FROM access_log_entries' → < 500K
|
|
- If DB > 500MB: alert + run /root/cleanup-access-logs.sh
|
|
- proxy-route-check: curl https://git.sysloggh.net/ → 2xx (proves routes loaded)
|
|
- proxy-sync-check: docker logs netbird-proxy --since 60s | grep 'Initial mapping sync complete' → present after restart
|
|
|
|
For docker-vm specifically:
|
|
- mountpoint -q /media/storage → NFS mounted
|
|
- mountpoint -q /media/mediastore → NFS mounted
|
|
- docker compose ls → expected stacks present
|
|
|
|
## Remediations
|
|
|
|
- Container unhealthy → `docker compose up -d <name>`, wait 10s, re-check
|
|
- Container missing → `docker compose up -d` in its stack directory
|
|
- Docker daemon down → `systemctl restart docker`
|
|
- NFS mount lost → `mount -a`, if fails → alert (storepve issue)
|
|
- Disk > 80% → alert via Zulip DM
|
|
- Disk > 95% → crit + relay to maintainer
|
|
```
|
|
|
|
## Section 4: Storage — Monitoring
|
|
|
|
### NFS Mounts (docker-vm → storepve)
|
|
|
|
| Mount | Export | Capacity | Used | Alert |
|
|
|-------|--------|----------|------|-------|
|
|
| /media/storage | storepve:/media/storage | 3.6TB | 264GB (8%) | None |
|
|
| /media/mediastore | storepve:/media/mediastore | 7.3TB | 5.2TB (75%) | ⚠️ warn at 80% |
|
|
|
|
### Proxmox Storage Backends
|
|
|
|
| Storage | Type | Content | Active |
|
|
|---------|------|---------|--------|
|
|
| local | dir | ISO, vztmpl, backup | ✅ |
|
|
| local-lvm | lvmthin | images, rootdir | ✅ |
|
|
| zfs-vm | nfs | rootdir, images | ✅ |
|
|
| mediastore | dir | images, backup, rootdir | ✅ |
|
|
| zfs-iso | nfs | vztmpl, iso | ✅ |
|
|
| pbs-backup | pbs | backup | ✅ |
|
|
|
|
### Checks (every 10 min)
|
|
|
|
```
|
|
## Maintains
|
|
|
|
- storage-usage: { mount: string, used_pct: float, growth_rate: float }
|
|
- backup-status: { last_success: timestamp, age_hours: int }
|
|
- snapshot-status: { pool: string, newest_age_hours: int }
|
|
|
|
## Specific Alerts
|
|
|
|
- mediastore growth rate > 10GB/day → warn (check what's writing)
|
|
- mediastore > 90% → crit (only 730GB remaining)
|
|
- No PBS backup in 48h → fail
|
|
```
|
|
|
|
## Section 5: Network Services — Monitoring
|
|
|
|
### 5.1 Service Inventory
|
|
|
|
The `Resolves To` column is the load-bearing one. Services that CNAME to
|
|
`netbird.sysloggh.net` (72.61.0.17) are routed through the NetBird reverse
|
|
proxy and **fail whenever NetBird blips**, even though their LAN backend is
|
|
fine. Services that resolve directly to a LAN IP are NetBird-independent.
|
|
|
|
| Service | Domain | LAN Backend | Resolves To | NetBird-dep | Status |
|
|
|---------|--------|-------------|-------------|-------------|--------|
|
|
| Proxmox API | minipve.sysloggh.net:8006 | 192.168.68.12 | LAN IP | No | ✅ |
|
|
| LiteLLM | litellm.sysloggh.net | 192.168.68.116 | LAN IP | No | ✅ |
|
|
| Authentik | auth.sysloggh.net:443 | 192.168.68.11 | CNAME → netbird | **Yes** | ⚠️ |
|
|
| Gitea | git.sysloggh.net:443 | 192.168.68.17:3000 | CNAME → netbird | **Yes** | ⚠️ |
|
|
| Zulip | chat.sysloggh.net:443 | 192.168.68.19 | CNAME → netbird | **Yes** | ⚠️ VERIFY-BEFORE-USE |
|
|
| Pulse | pulse.sysloggh.net:443 | 192.168.68.7 | CNAME → netbird | **Yes** | ⚠️ |
|
|
| DNS UI | dns.sysloggh.net:443 | 192.168.68.10:80 | CNAME → netbird | **Yes** | ⚠️ |
|
|
| SearXNG | searxng.sysloggh.net:8888 | 192.168.68.7:8888 | LAN IP | No | ✅ |
|
|
| Firecrawl | firecrawl.sysloggh.net:3002 | 192.168.68.7:3002 | LAN IP | No | ✅ |
|
|
|
|
**Verified 2026-07-24:** NetBird VPS rebooted after a hang; all CNAME'd
|
|
services recovered. LAN-IP-direct paths stayed up throughout the outage.
|
|
Also added `dns.sysloggh.net` route (was missing entirely).
|
|
See `scripts/netbird-add-domain.sh` for adding new proxy routes.
|
|
|
|
### 5.2 Checks (every 2 min)
|
|
|
|
```
|
|
## Maintains
|
|
|
|
- dns-resolution: { services: [{name, resolves_to, is_lan_ip}] }
|
|
- ssl-expiry: { services: [{name, days_remaining}] }
|
|
- endpoint-reachability: { services: [{name, http_code}] }
|
|
- netbird-dependency: { services: [{name, depends_on_netbird: bool}] }
|
|
```
|
|
|
|
### 5.3 Network Verification & Routing (every 5 min)
|
|
|
|
This block is what catches a NetBird blip before it becomes an outage. It
|
|
proves two independent paths for every service: the **URL path** (what a
|
|
browser uses, may cross NetBird) and the **LAN-IP path** (what configs and
|
|
agents must use, never crosses NetBird).
|
|
|
|
```
|
|
## Maintains
|
|
|
|
- dual-path-reachability: {
|
|
services: [{
|
|
name: string,
|
|
url_path: { http_code, tls_ok }, # the *.sysloggh.net path
|
|
lan_path: { ip, port, http_code }, # the IP:port path
|
|
lan_independent: bool # lan_path works when url_path fails
|
|
}]
|
|
}
|
|
- netbird-ingress-health: {
|
|
vps_up: bool, # TCP 22/80/443 accept on 72.61.0.17
|
|
ssh_banner_ok: bool, # banner exchange completes (catches hung box)
|
|
containers_up: int, # all 5 netbird-* containers running
|
|
proxy_routes_ok: int # traefik routers resolving (chat/git/auth return 2xx/3xx)
|
|
}
|
|
- routing-regression: {
|
|
new_netbird_cnames: [domain], # domains that newly CNAME to netbird.sysloggh.net
|
|
# → alert: a service just became NetBird-dependent
|
|
config_url_violations: [file] # configs/agents found referencing a *.sysloggh.net
|
|
# URL where a LAN IP is required (see Section 7)
|
|
}
|
|
|
|
## Checks
|
|
|
|
- For each service in 5.1:
|
|
- curl https://<domain>/ → record url_path (code, tls_ok)
|
|
- curl http://<lan_ip>:<port>/ → record lan_path
|
|
- If url_path fails AND lan_path succeeds → netbird-degraded, NOT service-down
|
|
- If BOTH fail → service-down (real outage)
|
|
|
|
- NetBird ingress (72.61.0.17):
|
|
- TCP probe 22, 80, 443 → all accept
|
|
- SSH banner exchange completes within 10s (catches the hung-box symptom
|
|
where TCP accepts but sshd never sends its banner)
|
|
- ssh root@72.61.0.17 'docker ps' → 5/5 netbird-* containers Up
|
|
- curl https://chat.sysloggh.net/ → 2xx/3xx (proxy route alive)
|
|
|
|
- DNS dependency drift:
|
|
- For each *.sysloggh.net, dig +short → flag any new CNAME → netbird.sysloggh.net
|
|
- A service moving FROM LAN-IP TO netbird CNAME is a regression (alert)
|
|
- A service moving FROM netbird CNAME TO LAN-IP is an improvement (log)
|
|
|
|
## Remediations
|
|
|
|
- NetBird VPS hung (TCP up, SSH banner timeout, TLS hang):
|
|
→ This is a host-level hang, not a service fault. Cannot self-remediate via SSH.
|
|
→ Alert crit immediately: "NetBird VPS hung — needs hard reboot from provider console"
|
|
→ Do NOT declare backend services down; their LAN-IP paths are still up.
|
|
- Single proxy route missing (e.g. git returns 000 but chat ok):
|
|
→ ssh root@72.61.0.17 'docker restart netbird-traefik'
|
|
→ re-check after 30s
|
|
- New NetBird CNAME detected:
|
|
→ Alert: "<domain> became NetBird-dependent — violates IP-first doctrine (Section 7)"
|
|
- Config URL violation detected:
|
|
→ Alert with file + line; do not auto-edit configs
|
|
```
|
|
|
|
## Section 6: Alert Routing & Escalation
|
|
|
|
### Severity Levels
|
|
|
|
| Severity | Channel | Format | Rate Limit |
|
|
|----------|---------|--------|-----------|
|
|
| **warn** | Zulip DM to owner | "⚠️ {check}: {detail}" | 1x/check/hour |
|
|
| **crit** | Zulip DM + #agent-hub | "🚨 {check}: {detail}" | Immediate |
|
|
| **escalated** | Zulip DM + relay to maintainer | "🔥 {check}: {detail}" | Immediate |
|
|
|
|
### Escalation Chain
|
|
|
|
```
|
|
Container unhealthy (3 consecutive failures)
|
|
→ Alert owner
|
|
→ 15 min no response → Relay to maintainer
|
|
|
|
Node offline
|
|
→ Alert owner
|
|
→ 5 min → Attempt restart via PVE API
|
|
→ 2 failures → Escalate
|
|
|
|
Storage > 95%
|
|
→ Alert owner + maintainer relay
|
|
→ Immediate action required
|
|
|
|
SSL cert < 7 days
|
|
→ Alert owner daily
|
|
→ 3 days out → Escalate
|
|
|
|
ZFS pool degraded
|
|
→ Cannot auto-fix → Escalate immediately
|
|
```
|
|
|
|
## Section 7: Configuration Doctrine — IP-First, URLs for Browsers
|
|
|
|
This doctrine is the structural fix for the NetBird-dependency outage. It is
|
|
enforced by the `routing-regression.config_url_violations` check in Section
|
|
5.3.
|
|
|
|
### Rule
|
|
|
|
1. **Service configs use LAN IP addresses, always.** Any file that wires a
|
|
service, agent, monitor, or integration to another internal service must
|
|
reference the backend by its `192.168.68.x` LAN IP and port — never by a
|
|
`*.sysloggh.net` URL. NetBird may blip at any time; LAN IPs do not.
|
|
2. **`*.sysloggh.net` URLs are reserved for human browser access.** They are
|
|
a convenience layer (split-horizon DNS on the LAN, NetBird reverse proxy
|
|
off the LAN) for a person at a keyboard. They must never be a load-bearing
|
|
dependency in code or config.
|
|
3. **Agents and monitors reach backends by LAN IP.** Abiba, Tanko, and Mumuni
|
|
all run on the LAN; there is no reason for them to hairpin through NetBird.
|
|
The daily infra report and all `ssh`/`curl` checks already follow this.
|
|
4. **DNS records should prefer LAN IPs over NetBird CNAMEs.** A service that
|
|
can resolve directly to its LAN IP (like `litellm` → `192.168.68.116` and
|
|
`minipve` → `192.168.68.12`) is NetBird-independent. Moving a record from
|
|
a NetBird CNAME to a LAN IP is an improvement; the reverse is a regression.
|
|
5. **NetBird is for remote/PC access only.** When you are out or on a PC,
|
|
NetBird carries you in. When you are on the LAN, NetBird is not in the
|
|
path. Configs must reflect this — they always assume LAN.
|
|
|
|
### Canonical IP map (configs MUST use these)
|
|
|
|
| Service | Use in configs | URL (browsers only) |
|
|
|---------|----------------|--------------------|
|
|
| Proxmox API | `https://192.168.68.12:8006` | `https://minipve.sysloggh.net` |
|
|
| LiteLLM API | `http://192.168.68.116:4000` | `https://litellm.sysloggh.net` |
|
|
| LiteLLM (nginx) | `http://192.168.68.116` | — |
|
|
| Grafana | `http://192.168.68.116:3001` | — |
|
|
| Authentik | `https://192.168.68.11:9000` | `https://auth.sysloggh.net` |
|
|
| Gitea | `http://192.168.68.17:3000` | `https://git.sysloggh.net` |
|
|
| Zulip API | `http://192.168.68.19` | `https://chat.sysloggh.net` |
|
|
| SearXNG | `http://192.168.68.7:8888` | — |
|
|
| Firecrawl | `http://192.168.68.7:3002` | — |
|
|
| RA-H OS bridge | `http://192.168.68.65:3100` | — |
|
|
|
|
### Violation examples (to be flagged by Section 5.3)
|
|
|
|
- ❌ An agent config setting `LITELLM_BASE_URL=https://litellm.sysloggh.net`
|
|
- ✅ Same config setting `LITELLM_BASE_URL=http://192.168.68.116:4000`
|
|
- ❌ A monitor curling `https://chat.sysloggh.net/api/v1/server_settings`
|
|
- ✅ A monitor curling `http://192.168.68.19/api/v1/server_settings` (and
|
|
optionally the URL as a separate browser-path check)
|
|
|
|
### Goal state
|
|
|
|
Every `*.sysloggh.net` record resolves to a LAN IP (split-horizon DNS on LAN,
|
|
NetBird reverse proxy off LAN). NetBird then becomes purely the off-LAN
|
|
ingress — if it blips, only remote browser users notice, and no agent,
|
|
monitor, or integration breaks.
|
|
|
|
## Appendix A: Quick Health Commands
|
|
|
|
```bash
|
|
# Full cluster status
|
|
PVE="https://minipve.sysloggh.net"
|
|
AUTH="Authorization: PVEAPIToken=monitoring@pve!mumuni=eafd56c5-93d4-4d40-a41d-e688be0987f3"
|
|
curl -sfk "$PVE/api2/json/cluster/resources" -H "$AUTH"
|
|
|
|
# Docker health from Abiba
|
|
ssh root@192.168.68.7 "docker ps --format '{{.Names}} {{.Status}}'"
|
|
ssh root@192.168.68.116 "docker ps --format '{{.Names}} {{.Status}}'"
|
|
|
|
# GPU fleet quick check
|
|
curl -s http://192.168.68.116/health/unified | jq .status
|
|
curl -s http://192.168.68.24:9100/gpu-data | jq .summary
|
|
|
|
# Grafana status
|
|
curl -s http://admin:$(infisical secrets get GRAFANA_ADMIN_PASSWORD --project=infrastructure --env=production --plain)@192.168.68.116:3001/api/health
|
|
|
|
# Prometheus targets
|
|
curl -s http://192.168.68.116:9090/api/v1/targets | jq '.data.activeTargets[] | {job: .labels.job, health: .health}'
|
|
|
|
# LiteLLM key check
|
|
curl -s -H "Authorization: Bearer $(infisical secrets get LITELLM_MASTER_KEY --project=infrastructure --env=production --plain)" \
|
|
http://192.168.68.116/litellm/key/list | jq '.keys[] | {alias: .key_alias, models: .models}'
|
|
|
|
# Storage check
|
|
ssh root@192.168.68.7 "df -h /media/storage /media/mediastore"
|
|
|
|
# Router roster reload (if needed)
|
|
curl -s -X POST http://192.168.68.116:9000/admin/roster/reload \
|
|
-H "Authorization: Bearer sk-admin-ee09fffd04978b61a1569ac670c68814"
|
|
|
|
# Restart stuck GPU (saturation watchdog alternative)
|
|
ssh root@192.168.68.8 "systemctl restart llama-server"
|
|
ssh root@192.168.68.110 "systemctl restart llama-server"
|
|
```
|
|
|
|
## Appendix B: CT Inventory
|
|
|
|
| CT | Name | Node | IP | Role | Agent |
|
|
|----|------|------|----|------|-------|
|
|
| CT | Name | Node | IP | Role | Agent |
|
|
|----|------|------|----|------|-------|
|
|
| 100 | abiba | **hwepve** | .24 | Pi agent | ✅ pi |
|
|
| 101 | llm-gpu | acerpve | .8 | GPU RTX 3090 | ❌ |
|
|
| 102 | adguard | **minipve** | **.10** | DNS | ❌ |
|
|
| 103 | ocu-llm | ocupve | .110 | GPU RTX 5070 | ❌ |
|
|
| 104 | authentik | minipve | .11 | OIDC | ❌ |
|
|
| 105 | kagentz | **hwepve** | — | Agent Zero | ✅ |
|
|
| 106 | ra-h-os | storepve | .65 | KG bridge | ✅ MCP |
|
|
| 107 | pbs | storepve | — | Backups | ❌ |
|
|
| 108 | media | storepve | — | Media | ❌ |
|
|
| 109 | docker-vm | storepve | .7 | Docker host | ❌ |
|
|
| 110 | gitea | minipve | **.17** | Git | ❌ |
|
|
| 111 | tdunna | amdpve | .129 | Hermes agent | ✅ |
|
|
| 112 | tanko | amdpve | .122 | Hermes agent | ✅ |
|
|
| 113 | baggy | amdpve | .114 | Hermes agent | ✅ |
|
|
| 114 | mumuni | **hwepve** | .123 | Hermes agent (stopped) | ✅ |
|
|
| 115 | scottdenya | amdpve | .75 | Denya OneCare | ❌ |
|
|
| 116 | syslog-api | minipve | .116 | LiteLLM + Grafana | ❌ |
|
|
| 117 | zulip | storepve | .19 | Chat | ❌ |
|
|
| 118 | jdownloader | storepve | — | JDownloader container | ❌ |
|
|
| 119 | infisical-vault | minipve | — | Vault | ❌ |
|
|
|
|
## Appendix C: Docker Compose Files Location
|
|
|
|
| Host | Stack | Compose File |
|
|
|------|-------|-------------|
|
|
| docker-vm (.7) | Firecrawl | `/opt/search-stack/firecrawl-source/docker-compose.yaml` |
|
|
| docker-vm (.7) | SearXNG | `/opt/search-stack/searxng/docker-compose.yml` |
|
|
| docker-vm (.7) | Home stack | `/opt/home_stack/docker-compose.yml` |
|
|
| docker-vm (.7) | Audiobookshelf | `/opt/audiobookshelf/docker-compose.yml` |
|
|
| CT 116 (.116) | Inference Harness | `/opt/inference-harness/docker-compose.yml` |
|
|
| Netbird (.17) | Netbird | Docker run (not compose) |
|
|
|
|
## CT Access (pct-run — no IPs needed)
|
|
|
|
All CTs are accessible via `pct-run <CT_ID> <command>`. No IP addresses required.
|
|
Source of truth: `/root/scripts/pct-run.sh` or `prose-contracts/scripts/pct-run.sh`.
|
|
|
|
| CT | Name | Node | pct-run |
|
|
|-----|------|------|---------|
|
|
| 100 | abiba | hwepve | `pct-run 100` |
|
|
| 105 | kagentz | **hwepve** | `pct-run 105` |
|
|
| 111 | tdunna | amdpve | `pct-run 111` |
|
|
| 112 | tanko | amdpve | `pct-run 112` |
|
|
| 113 | baggy | amdpve | `pct-run 113` |
|
|
| 115 | scottdenya | amdpve | `pct-run 115` |
|
|
| 104 | authentik | minipve | `pct-run 104` |
|
|
| 110 | gitea | minipve | `pct-run 110` |
|
|
| 114 | mumuni | **hwepve** | `pct-run 114` |
|
|
| 116 | syslog-api | minipve | `pct-run 116` |
|
|
| 106 | ra-h-os | storepve | `pct-run 106` |
|
|
| 107 | proxmox-backup | storepve | `pct-run 107` |
|
|
| 108 | media | storepve | `pct-run 108` |
|
|
| 117 | zulip | storepve | `pct-run 117` |
|
|
| 102 | adguard | **minipve** | `pct-run 102` |
|
|
|
|
GPU bare-metal hosts (.8 acerpve, .110 ocupve, .15 amdpve) are NOT CTs — use SSH directly:
|
|
```bash
|
|
ssh root@192.168.68.8 # RTX 3090
|
|
ssh root@192.168.68.110 # RTX 5070
|
|
ssh root@192.168.68.15 # Strix Halo
|
|
ssh root@192.168.68.4 # hwepve (abiba, kagentz, mumuni)
|
|
```
|
|
|
|
## Section 7: Agent Health Check (consolidated — 2026-07-05)
|
|
|
|
Replaces 7 scattered Zulip health scripts with a single non-disruptive check
|
|
running every 10 minutes via cron (`/root/scripts/agent-health-check.py`).
|
|
|
|
The script is **read-only** — it never restarts, kills, or modifies anything.
|
|
Disruptive cron-based gateways restarts (like Mumuni's zulip-watchdog.sh,
|
|
which was kill+nohup outside systemd) are banned by policy.
|
|
|
|
### Checks Performed
|
|
|
|
| Check | Frequency | What It Detects |
|
|
|-------|-----------|-----------------|
|
|
| LiteLLM key validation | 10 min | All 4 agent keys authenticate and return models |
|
|
| GPU port conflict | 10 min | Ghost processes squatting port 8080 (ss vs systemd MainPID) |
|
|
| Gateway liveness | 10 min | Gateway process running, state file readable |
|
|
| Zulip streaming | 10 min | `edit_message` present in adapter (streaming supported) |
|
|
| Recent errors | 10 min | Error count in journald for last 10 min |
|
|
|
|
### Disabled Scripts
|
|
|
|
| Script | Why Disabled |
|
|
|--------|-------------|
|
|
| `zulip-watchdog.sh` (Mumuni) | kill+nohup bypassed systemd, 27 restarts, pattern mismatch |
|
|
| `zulip-monitor.sh` (Abiba) | Replaced by agent-health-check.py + PM2 auto-restart |
|