diff --git a/infrastructure-monitoring.prose.md b/infrastructure-monitoring.prose.md index c55f98e..6ceae98 100644 --- a/infrastructure-monitoring.prose.md +++ b/infrastructure-monitoring.prose.md @@ -139,6 +139,79 @@ GPU .8 (RTX 3090) GPU .110 (RTX 5070) GPU .15 (Strix Halo) 4. Verify LiteLLM metrics flowing to Prometheus 5. ~~Update nginx to proxy `/monitoring/` → Grafana~~ (NOT recommended — nginx sub-path was tried for /grafana/ and reverted per proxmox-monitor; direct :3001 access is the standard) +## Execution + +### check-health + +**RUN LIVE, NEVER ECHO — every dispatch must execute the probes below with real tool calls; never repeat a prior report unless a live probe fails.** + +```bash +# Zulip API health (POST ping) +curl -s -o /dev/null -w '%{http_code}' -X POST https://chat.sysloggh.net/api/v1/messages -u 'abiba-bot@chat.sysloggh.net:KEY' +# Expected: 200 (HTTP 000 = unreachable/cache) + +# PM2 process health +pm2 jlist +# Expected: 5/5 online (abiba-telegram, abiba-zulip, zulip-watchdog, gitea-runner, spoton-service) + +# GPU exporters (may be down per DEPLOYMENT STATUS) +curl -s http://192.168.68.8:9400/metrics && echo " - OK" || echo " - FAIL" +curl -s http://192.168.68.110:9400/metrics && echo " - OK" || echo " - FAIL" +curl -s http://192.168.68.15:9400/metrics && echo " - OK" || echo " - FAIL" + +# Prometheus targets +curl -s http://192.168.68.116:9090/api/v1/targets | jq '.data.activeTargets' +# Expected: All targets UP (may show some down if exporters not deployed) + +# Grafana health +curl -s http://192.168.68.116:3001/api/health | jq '{status, version}' +# Expected: {"status":"ok","version":"..."} + +# LiteLLM metrics +curl -s http://192.168.68.116:4001/metrics | head -20 +# Expected: Prometheus-formatted metrics output +``` + +**Report format**: Summarize actual results from each probe. If any probe returns non-200 or empty output, flag as alert. + +### Phase 1: GPU Exporters + +**NVIDIA (.8 and .110)**: +1. Download `nvidia_gpu_exporter` binary +2. Create systemd service `nvidia-gpu-exporter.service` +3. Start and enable + +**AMD (.15)**: +1. Create Python exporter script at `/opt/amdgpu-exporter/exporter.py` +2. Parses `amdgpu_top --json -d 1000` output +3. Exposes key metrics at `:9400/metrics` via Python http.server +4. Create systemd service +5. Start and enable + +### Phase 2: Prometheus + +1. Create `/opt/monitoring/` directory on CT 116 +2. Write `prometheus.yml` with scrape configs for all targets +3. Add to docker-compose (or separate compose file) +4. Start container + +### Phase 3: Grafana + +1. Create `/opt/monitoring/grafana/` directories +2. Provision Prometheus datasource +3. Provision GPU fleet dashboard JSON +4. Provision LiteLLM dashboard JSON +5. Add to docker-compose +6. Start container + +### Phase 4: Verification + +1. Verify all 3 GPU exporters return 200 at :9400/metrics +2. Verify Prometheus targets all UP at :9090/targets +3. Verify Grafana accessible at :3001 with dashboards +4. Verify LiteLLM metrics flowing to Prometheus +5. ~~Update nginx to proxy `/monitoring/` → Grafana~~ (NOT recommended — nginx sub-path was tried for /grafana/ and reverted per proxmox-monitor; direct :3001 access is the standard) + ## Verification Commands ```bash