From 403fbcdd9f692667d359e6e89d49f5d9bb5a3669 Mon Sep 17 00:00:00 2001 From: mumuni-bot Date: Mon, 7 Sep 2026 07:08:49 +0000 Subject: [PATCH] fix: remove duplicated ## Execution block in infrastructure-monitoring PR #50's content was squash-merged out-of-band (2dfc3e1, 2026-08-30) while the PR itself stayed open, double-applying the check-health/Execution section. This drops the second copy; keeps one canonical Execution -> check-health -> Phases 1-4 -> Verification Commands flow. --- infrastructure-monitoring.prose.md | 73 ------------------------------ 1 file changed, 73 deletions(-) diff --git a/infrastructure-monitoring.prose.md b/infrastructure-monitoring.prose.md index e1f8172..62d8d4b 100644 --- a/infrastructure-monitoring.prose.md +++ b/infrastructure-monitoring.prose.md @@ -136,79 +136,6 @@ curl -s http://192.168.68.116:4001/metrics | head -20 **Report format**: Summarize actual results from each probe. If any probe returns non-200 or empty output, flag as alert. -### Phase 1: GPU Exporters - -**NVIDIA (.8 and .110)**: -1. Download `nvidia_gpu_exporter` binary -2. Create systemd service `nvidia-gpu-exporter.service` -3. Start and enable - -**AMD (.15)**: -1. Create Python exporter script at `/opt/amdgpu-exporter/exporter.py` -2. Parses `amdgpu_top --json -d 1000` output -3. Exposes key metrics at `:9400/metrics` via Python http.server -4. Create systemd service -5. Start and enable - -### Phase 2: Prometheus - -1. Create `/opt/monitoring/` directory on CT 116 -2. Write `prometheus.yml` with scrape configs for all targets -3. Add to docker-compose (or separate compose file) -4. Start container - -### Phase 3: Grafana - -1. Create `/opt/monitoring/grafana/` directories -2. Provision Prometheus datasource -3. Provision GPU fleet dashboard JSON -4. Provision LiteLLM dashboard JSON -5. Add to docker-compose -6. Start container - -### Phase 4: Verification - -1. Verify all 3 GPU exporters return 200 at :9400/metrics -2. Verify Prometheus targets all UP at :9090/targets -3. Verify Grafana accessible at :3001 with dashboards -4. Verify LiteLLM metrics flowing to Prometheus -5. ~~Update nginx to proxy `/monitoring/` → Grafana~~ (NOT recommended — nginx sub-path was tried for /grafana/ and reverted per proxmox-monitor; direct :3001 access is the standard) - -## Execution - -### check-health - -**RUN LIVE, NEVER ECHO — every dispatch must execute the probes below with real tool calls; never repeat a prior report unless a live probe fails.** - -```bash -# Zulip API health (POST ping) -curl -s -o /dev/null -w '%{http_code}' -X POST https://chat.sysloggh.net/api/v1/messages -u 'abiba-bot@chat.sysloggh.net:KEY' -# Expected: 200 (HTTP 000 = unreachable/cache) - -# PM2 process health -pm2 jlist -# Expected: 5/5 online (abiba-telegram, abiba-zulip, zulip-watchdog, gitea-runner, spoton-service) - -# GPU exporters (may be down per DEPLOYMENT STATUS) -curl -s http://192.168.68.8:9400/metrics && echo " - OK" || echo " - FAIL" -curl -s http://192.168.68.110:9400/metrics && echo " - OK" || echo " - FAIL" -curl -s http://192.168.68.15:9400/metrics && echo " - OK" || echo " - FAIL" - -# Prometheus targets -curl -s http://192.168.68.116:9090/api/v1/targets | jq '.data.activeTargets' -# Expected: All targets UP (may show some down if exporters not deployed) - -# Grafana health -curl -s http://192.168.68.116:3001/api/health | jq '{status, version}' -# Expected: {"status":"ok","version":"..."} - -# LiteLLM metrics -curl -s http://192.168.68.116:4001/metrics | head -20 -# Expected: Prometheus-formatted metrics output -``` - -**Report format**: Summarize actual results from each probe. If any probe returns non-200 or empty output, flag as alert. - ### Phase 1: GPU Exporters **NVIDIA (.8 and .110)**: