diff --git a/gpu-monitor.prose.md b/gpu-monitor.prose.md index ee98873..82bd0ee 100644 --- a/gpu-monitor.prose.md +++ b/gpu-monitor.prose.md @@ -97,7 +97,28 @@ This replaces the previous DM-only delivery. All agents on the mesh can see and `curl http://localhost:9100/gpu-data | jq` — Full fleet status ### check-health -`curl http://localhost:9100/health` — Monitor self-check + +**RUN LIVE, NEVER ECHO — every dispatch must execute the probes below with real tool calls; never repeat a prior report unless a live probe fails.** + +```bash +# GPU Monitor health + curl http://localhost:9100/health | jq +# Expected: 200 with {"status": "healthy", "cache_age_seconds": } + +# Router health (via nginx on port 80) +curl -s -o /dev/null -w '%{http_code}' http://192.168.68.116/health +# Expected: 200 (Router is up and responding) + +# LiteLLM health (via nginx on port 80) +curl -s -o /dev/null -w '%{http_code}' http://192.168.68.116/litellm/health +# Expected: 200 (LiteLLM is up and responding) + +# Dashboard + curl -s -o /dev/null -w '%{http_code}' http://192.168.68.116/dashboard/ +# Expected: 200 (Dashboard is up and responding) +``` + +**Report format**: Summarize actual results from each probe. If any probe returns non-200, flag as alert. ### view-dashboard Open `http://localhost:9100/` in browser — Live HTML dashboard diff --git a/infrastructure-monitoring.prose.md b/infrastructure-monitoring.prose.md index 62d8d4b..205aa03 100644 --- a/infrastructure-monitoring.prose.md +++ b/infrastructure-monitoring.prose.md @@ -108,7 +108,9 @@ GPU .8 (RTX 3090) GPU .110 (RTX 5070) GPU .15 (Strix Halo) ```bash # Zulip API health (POST ping) -curl -s -o /dev/null -w '%{http_code}' -X POST https://chat.sysloggh.net/api/v1/messages -u 'abiba-bot@chat.sysloggh.net:KEY' +source /etc/litellm-monitor.env +ZULIP_USER="abiba-bot@chat.sysloggh.net" +curl -s -o /dev/null -w '%{http_code}' -X POST https://chat.sysloggh.net/api/v1/messages -u "${ZULIP_USER}:${ZULIP_BOT_KEY}" # Expected: 200 (HTTP 000 = unreachable/cache) # PM2 process health @@ -120,6 +122,18 @@ curl -s http://192.168.68.8:9400/metrics && echo " - OK" || echo " - FAIL" curl -s http://192.168.68.110:9400/metrics && echo " - OK" || echo " - FAIL" curl -s http://192.168.68.15:9400/metrics && echo " - OK" || echo " - FAIL" +# Router health (via nginx on port 80) +curl -s -o /dev/null -w '%{http_code}' http://192.168.68.116/health +# Expected: 200 (Router is up and responding) + +# LiteLLM health (via nginx on port 80) +curl -s -o /dev/null -w '%{http_code}' http://192.168.68.116/litellm/health +# Expected: 200 (LiteLLM is up and responding) + +# PVE API (401 expected for unauthenticated probe — API is up over https) +curl -s -o /dev/null -w '%{http_code}' https://192.168.68.116:8006/api2/json +# Expected: 401 (unauthorized — API is up; 000 = unreachable, 500 = API down) + # Prometheus targets curl -s http://192.168.68.116:9090/api/v1/targets | jq '.data.activeTargets' # Expected: All targets UP (may show some down if exporters not deployed) @@ -128,7 +142,7 @@ curl -s http://192.168.68.116:9090/api/v1/targets | jq '.data.activeTargets' curl -s http://192.168.68.116:3001/api/health | jq '{status, version}' # Expected: {"status":"ok","version":"..."} -# LiteLLM metrics +# LiteLLM metrics (Prometheus endpoint) curl -s http://192.168.68.116:4001/metrics | head -20 # Expected: Prometheus-formatted metrics output ``` diff --git a/proxmox-monitor.prose.md b/proxmox-monitor.prose.md index 90fadaa..80ecd31 100644 --- a/proxmox-monitor.prose.md +++ b/proxmox-monitor.prose.md @@ -100,6 +100,32 @@ Edit `build-dashboards.py`, run it, `docker restart harness-grafana` ### check-targets `curl http://192.168.68.116:9090/api/v1/targets | jq '.data.activeTargets[] | {job:.labels.job,health}'` +### check-health + +**RUN LIVE, NEVER ECHO — every dispatch must execute the probes below with real tool calls; never repeat a prior report unless a live probe fails.** + +```bash +# Prometheus health (bound to 0.0.0.0:9090 on .116) +curl -s -o /dev/null -w '%{http_code}' http://192.168.68.116:9090/-/healthy +# Expected: 200 (Prometheus is up and healthy) + +# Grafana health (bound to 0.0.0.0:3001 on .116) +curl -s -o /dev/null -w '%{http_code}' http://192.168.68.116:3001/api/health +# Expected: 200 (Grafana is up and healthy) + +# Docker Stats exporter (bound to 127.0.0.1:9324 on .116 — must probe from .116 localhost) +ssh root@192.168.68.116 "curl -s -o /dev/null -w '%{http_code}' http://127.0.0.1:9324/metrics" +# Expected: 200 (docker-stats-exporter is up and responding) + +# PVE exporter (bound to 127.0.0.1:9221 on .116 — must probe from .116 localhost) +ssh root@192.168.68.116 "curl -s -o /dev/null -w '%{http_code}' http://127.0.0.1:9221/metrics" +# Expected: 200 (pve-exporter is up and responding) +``` + +**Report format**: Summarize actual results from each probe. If any probe returns non-200, flag as alert. + +**Note**: Docker Stats and PVE Exporter are bound to 127.0.0.1 (localhost-only) so they must be probed from .116 via SSH. Prometheus and Grafana are bound to 0.0.0.0 so they can be probed from the LAN. + ### restart-exporter `cd /opt/monitoring && docker compose restart pve-exporter docker-stats` diff --git a/zulip-health.prose.md b/zulip-health.prose.md index 64adba8..1e28d17 100644 --- a/zulip-health.prose.md +++ b/zulip-health.prose.md @@ -238,10 +238,11 @@ ssh root@192.168.68.24 "grep -E 'Finalized|Failed to finalize|Replied to' ~/.her **C1: A2A Server Health** ```bash -ssh root@192.168.68.14 "docker exec agent-zero curl -s --connect-timeout 5 http://127.0.0.1:8001/.well-known/agent.json" +# A2A server is on :50080 (not :8001) and is auth-gated (401 expected for unauthenticated) +ssh root@192.168.68.14 "curl -s --connect-timeout 5 -o /dev/null -w '%{http_code}' http://127.0.0.1:50080/a2a/" ``` -Expected: `{"name":"kagentz",...}`. Connection refused → A2A server down. +Expected: `401` (auth-gated, A2A server is up and responding) or `200` (if no auth required). Connection refused (000) → A2A server down. **C2: Adapter Process** @@ -262,12 +263,14 @@ Check: `processed=N` incrementing, `silence < 600s`, `reconnects` ≈ 0. **C4: A2A Response Verification** ```bash -ssh root@192.168.68.14 "docker exec agent-zero curl -s -X POST http://127.0.0.1:8001/a2a \ +# A2A server is on :50080 (not :8001) and is auth-gated (401 expected for unauthenticated) +ssh root@192.168.68.14 "curl -s -X POST http://127.0.0.1:50080/a2a \ -H 'Content-Type: application/json' \ + -H 'Authorization: Bearer $LITELLM_KEY' \ -d '{\"jsonrpc\":\"2.0\",\"method\":\"tasks/send\",\"params\":{\"message\":{\"role\":\"user\",\"parts\":[{\"text\":\"ping\"}]}},\"id\":1}'" ``` -Expected: task ID with "working" status. Poll for completion with `tasks/get`. +Expected: task ID with "working" status. Poll for completion with `tasks/get`. If 401, check LITELLM_KEY is set. **Platform C Actions**