Merge pull request 'Add check-health section with live probes to infrastructure-monitoring contract' (#51) from fix/infra-check-health into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 1s
Merged by captain approval 2026-08-22: add check-health section with live probes to infrastructure-monitoring contract
This commit was merged in pull request #51.
This commit is contained in:
@@ -102,6 +102,39 @@ GPU .8 (RTX 3090) GPU .110 (RTX 5070) GPU .15 (Strix Halo)
|
||||
- Stack persists across reboots (systemd for exporters, Docker restart policy)
|
||||
|
||||
## Execution
|
||||
### check-health
|
||||
|
||||
**RUN LIVE, NEVER ECHO — every dispatch must execute the probes below with real tool calls; never repeat a prior report unless a live probe fails.**
|
||||
|
||||
```bash
|
||||
# Zulip API health (POST ping)
|
||||
curl -s -o /dev/null -w '%{http_code}' -X POST https://chat.sysloggh.net/api/v1/messages -u 'abiba-bot@chat.sysloggh.net:KEY'
|
||||
# Expected: 200 (HTTP 000 = unreachable/cache)
|
||||
|
||||
# PM2 process health
|
||||
pm2 jlist
|
||||
# Expected: 5/5 online (abiba-telegram, abiba-zulip, zulip-watchdog, gitea-runner, spoton-service)
|
||||
|
||||
# GPU exporters (may be down per DEPLOYMENT STATUS)
|
||||
curl -s http://192.168.68.8:9400/metrics && echo " - OK" || echo " - FAIL"
|
||||
curl -s http://192.168.68.110:9400/metrics && echo " - OK" || echo " - FAIL"
|
||||
curl -s http://192.168.68.15:9400/metrics && echo " - OK" || echo " - FAIL"
|
||||
|
||||
# Prometheus targets
|
||||
curl -s http://192.168.68.116:9090/api/v1/targets | jq '.data.activeTargets'
|
||||
# Expected: All targets UP (may show some down if exporters not deployed)
|
||||
|
||||
# Grafana health
|
||||
curl -s http://192.168.68.116:3001/api/health | jq '{status, version}'
|
||||
# Expected: {"status":"ok","version":"..."}
|
||||
|
||||
# LiteLLM metrics
|
||||
curl -s http://192.168.68.116:4001/metrics | head -20
|
||||
# Expected: Prometheus-formatted metrics output
|
||||
```
|
||||
|
||||
**Report format**: Summarize actual results from each probe. If any probe returns non-200 or empty output, flag as alert.
|
||||
|
||||
|
||||
### Phase 1: GPU Exporters
|
||||
|
||||
|
||||
Reference in New Issue
Block a user