fix: add Router+LiteLLM+PVE API probes to infrastructure-monitoring check-health

- Added Router health probe (http://192.168.68.116/health) — expected 200
- Added LiteLLM health probe (http://192.168.68.116/litellm/health) — expected 200
- Added PVE API probe (https://192.168.68.116:8006/api2/json) — 401 expected for unauthenticated
- Clarified that 401 means API is up, 000 means unreachable, 500 means API down

Closes: infra-monitoring-probe-alignment
This commit is contained in:
root
2026-09-08 07:20:41 +00:00
parent cc5fe0991c
commit a3e97ce72b
+13 -1
View File
@@ -122,6 +122,18 @@ curl -s http://192.168.68.8:9400/metrics && echo " - OK" || echo " - FAIL"
curl -s http://192.168.68.110:9400/metrics && echo " - OK" || echo " - FAIL"
curl -s http://192.168.68.15:9400/metrics && echo " - OK" || echo " - FAIL"
# Router health (via nginx on port 80)
curl -s -o /dev/null -w '%{http_code}' http://192.168.68.116/health
# Expected: 200 (Router is up and responding)
# LiteLLM health (via nginx on port 80)
curl -s -o /dev/null -w '%{http_code}' http://192.168.68.116/litellm/health
# Expected: 200 (LiteLLM is up and responding)
# PVE API (401 expected for unauthenticated probe — API is up over https)
curl -s -o /dev/null -w '%{http_code}' https://192.168.68.116:8006/api2/json
# Expected: 401 (unauthorized — API is up; 000 = unreachable, 500 = API down)
# Prometheus targets
curl -s http://192.168.68.116:9090/api/v1/targets | jq '.data.activeTargets'
# Expected: All targets UP (may show some down if exporters not deployed)
@@ -130,7 +142,7 @@ curl -s http://192.168.68.116:9090/api/v1/targets | jq '.data.activeTargets'
curl -s http://192.168.68.116:3001/api/health | jq '{status, version}'
# Expected: {"status":"ok","version":"..."}
# LiteLLM metrics
# LiteLLM metrics (Prometheus endpoint)
curl -s http://192.168.68.116:4001/metrics | head -20
# Expected: Prometheus-formatted metrics output
```