PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 2s
- Implement four-state PBS GC logic in proxmox-monitor.sh: - probe-failed: unparseable JSON, store not found, or empty body → FAIL - running: collection in progress (last-run-endtime absent, upid present) → DO NOT FAIL - stale: no completed run within 48h → FAIL, naming last completed run age - healthy: completed within 48h → PASS, naming endtime and pending bytes - Add tests/test_pbs_gc_states.sh covering all four states - Proves the test bites on the pre-fix version (5/6 tests fail) - All 6 tests pass against the fixed version - Update contract-run.sh to map disk-gc-threat-response -> scripts/disk-gc-scan.py - Update Execution sections of host-scheduled contracts: infrastructure-monitoring, zulip-health, litellm-health, agent-health-check, disk-gc-threat-response, pm2-self-heal Adding note that execution is host-scheduled via cron, not agent session ack.
166 lines
6.7 KiB
Bash
Executable File
166 lines
6.7 KiB
Bash
Executable File
#!/bin/bash
|
|
# proxmox-monitor.sh — Proxmox Cluster + Docker Monitoring Health Check
|
|
# Implements proxmox-monitor.prose.md (check-health section)
|
|
#
|
|
# Legs: Prometheus, Grafana, Docker Stats Exporter, PVE Exporter, PBS GC
|
|
# All legs must return 200 for healthy status.
|
|
#
|
|
# Run: bash scripts/proxmox-monitor.sh
|
|
# Exits 0 if all probes pass, 1 if any fails.
|
|
|
|
set -uo pipefail
|
|
|
|
CT116_HOST="192.168.68.116"
|
|
FAILED=()
|
|
TIMESTAMP=$(date -u '+%Y-%m-%d %H:%M UTC')
|
|
|
|
echo "=== Proxmox Monitor — $TIMESTAMP ==="
|
|
echo "Executed from: $(pwd -P)"
|
|
echo ""
|
|
|
|
# 1. Prometheus health (bound to 0.0.0.0:9090 on .116)
|
|
PROM_CODE=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 http://${CT116_HOST}:9090/-/healthy 2>/dev/null)
|
|
PROM_CODE=$(printf '%s' "$PROM_CODE" | tr -d '[:space:]')
|
|
[ -n "$PROM_CODE" ] || PROM_CODE="000"
|
|
|
|
if [ "$PROM_CODE" = "200" ]; then
|
|
echo " ✅ Prometheus: alive"
|
|
else
|
|
echo " 🔴 Prometheus: probe-failed: ${CT116_HOST}:9090 (expected 200, got ${PROM_CODE})"
|
|
FAILED+=("prometheus")
|
|
fi
|
|
|
|
# 2. Grafana health (bound to 0.0.0.0:3001 on .116)
|
|
GRAF_CODE=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 http://${CT116_HOST}:3001/api/health 2>/dev/null)
|
|
GRAF_CODE=$(printf '%s' "$GRAF_CODE" | tr -d '[:space:]')
|
|
[ -n "$GRAF_CODE" ] || GRAF_CODE="000"
|
|
|
|
if [ "$GRAF_CODE" = "200" ]; then
|
|
echo " ✅ Grafana: alive"
|
|
else
|
|
echo " 🔴 Grafana: probe-failed: ${CT116_HOST}:3001 (expected 200, got ${GRAF_CODE})"
|
|
FAILED+=("grafana")
|
|
fi
|
|
|
|
# 3. Docker Stats exporter (bound to 127.0.0.1:9324 on .116 — probe from .116 localhost)
|
|
DOCKER_CODE=$(ssh -o ConnectTimeout=5 -o BatchMode=yes root@${CT116_HOST} \
|
|
"curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 http://127.0.0.1:9324/metrics" 2>/dev/null)
|
|
DOCKER_CODE=$(printf '%s' "$DOCKER_CODE" | tr -d '[:space:]')
|
|
[ -n "$DOCKER_CODE" ] || DOCKER_CODE="000"
|
|
|
|
if [ "$DOCKER_CODE" = "200" ]; then
|
|
echo " ✅ Docker Stats: alive"
|
|
else
|
|
echo " 🔴 Docker Stats: probe-failed: CT116:127.0.0.1:9324 (expected 200, got ${DOCKER_CODE})"
|
|
FAILED+=("docker-stats")
|
|
fi
|
|
|
|
# 4. PVE Exporter (bound to 127.0.0.1:9221 on .116 — probe from .116 localhost)
|
|
PVE_CODE=$(ssh -o ConnectTimeout=5 -o BatchMode=yes root@${CT116_HOST} \
|
|
"curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 http://127.0.0.1:9221/metrics" 2>/dev/null)
|
|
PVE_CODE=$(printf '%s' "$PVE_CODE" | tr -d '[:space:]')
|
|
[ -n "$PVE_CODE" ] || PVE_CODE="000"
|
|
|
|
if [ "$PVE_CODE" = "200" ]; then
|
|
echo " ✅ PVE Exporter: alive"
|
|
else
|
|
echo " 🔴 PVE Exporter: probe-failed: CT116:127.0.0.1:9221 (expected 200, got ${PVE_CODE})"
|
|
FAILED+=("pve-exporter")
|
|
fi
|
|
|
|
# 5. PBS GC liveness (storepve-datastore GC must have run within 48h)
|
|
# Use absolute path for pct to avoid PATH issues in non-interactive ssh
|
|
PBS_GC_OUTPUT=$(ssh -o ConnectTimeout=5 -o BatchMode=yes root@192.168.68.6 \
|
|
"/sbin/pct exec 107 -- /sbin/proxmox-backup-manager garbage-collection list --output-format json" 2>/dev/null)
|
|
PBS_GC_OUTPUT=$(printf '%s' "$PBS_GC_OUTPUT" | tr -d '[:space:]')
|
|
[ -n "$PBS_GC_OUTPUT" ] || PBS_GC_OUTPUT="000"
|
|
|
|
if [ "$PBS_GC_OUTPUT" = "000" ]; then
|
|
echo " 🔴 PBS GC: probe-failed: storepve:192.168.68.6 (expected JSON, got 000)"
|
|
FAILED+=("pbs-gc")
|
|
else
|
|
# Parse the JSON to get storepve-datastore's state with four distinct outcomes:
|
|
# 1. probe-failed: non-zero ssh status / empty / unparseable JSON
|
|
# 2. running: collection in progress (last-run-endtime absent or 0, but upid present)
|
|
# 3. stale: no completed run within 48h
|
|
# 4. healthy: completed within 48h
|
|
PBS_GC_RESULT=$(echo "$PBS_GC_OUTPUT" | python3 -c "
|
|
import sys, json
|
|
try:
|
|
data = json.load(sys.stdin)
|
|
for store in data:
|
|
if store['store'] == 'storepve-datastore':
|
|
endtime = store.get('last-run-endtime')
|
|
upid = store.get('upid')
|
|
pending = store.get('pending-bytes', 0)
|
|
|
|
# State 2: Running (collection in progress) — last-run-endtime absent while run is in progress
|
|
if (endtime is None or endtime == 0) and upid is not None:
|
|
print(f'running|{pending}')
|
|
break
|
|
|
|
# State 3: No completed run (never-run or stale)
|
|
if endtime is None or endtime == 0:
|
|
print(f'no-completed-run|{pending}')
|
|
break
|
|
|
|
# States 3 & 4: Completed (has endtime)
|
|
print(f'completed|{endtime}|{pending}')
|
|
break
|
|
else:
|
|
print(f'absent|0')
|
|
except json.JSONDecodeError:
|
|
print(f'unparseable|0')
|
|
" 2>/dev/null)
|
|
|
|
# Parse the state|endtime|pending format
|
|
PBS_GC_STATE=$(echo "$PBS_GC_RESULT" | cut -d'|' -f1)
|
|
|
|
if [ "$PBS_GC_STATE" = "unparseable" ]; then
|
|
# State 1: probe-failed (unparseable JSON)
|
|
echo " 🔴 PBS GC: probe-failed: storepve:192.168.68.6 (unparseable JSON)"
|
|
FAILED+=("pbs-gc")
|
|
elif [ "$PBS_GC_STATE" = "absent" ]; then
|
|
# State 1: probe-failed (store not found)
|
|
echo " 🔴 PBS GC: probe-failed: storepve:192.168.68.6 (storepve-datastore not found)"
|
|
FAILED+=("pbs-gc")
|
|
elif [ "$PBS_GC_STATE" = "running" ]; then
|
|
# State 2: collection in progress — do NOT fail
|
|
PENDING_BYTES=$(echo "$PBS_GC_RESULT" | cut -d'|' -f2)
|
|
echo " ⏳ PBS GC: running (started: in-progress, pending-bytes: ${PENDING_BYTES} B)"
|
|
elif [ "$PBS_GC_STATE" = "no-completed-run" ]; then
|
|
# State 3: no completed run within 48h
|
|
PENDING_BYTES=$(echo "$PBS_GC_RESULT" | cut -d'|' -f2)
|
|
echo " 🔴 PBS GC: no completed run within 48h (pending-bytes: ${PENDING_BYTES} B)"
|
|
FAILED+=("pbs-gc")
|
|
else
|
|
# States 3 & 4: completed (has endtime)
|
|
LAST_RUN_ENDTIME=$(echo "$PBS_GC_RESULT" | cut -d'|' -f2)
|
|
PENDING_BYTES=$(echo "$PBS_GC_RESULT" | cut -d'|' -f3)
|
|
|
|
# Convert epoch to age in hours
|
|
NOW_EPOCH=$(date -u +%s)
|
|
AGE_HOURS=$(( (NOW_EPOCH - LAST_RUN_ENDTIME) / 3600 ))
|
|
|
|
if [ $AGE_HOURS -gt 48 ]; then
|
|
# State 3: stale (no completed run within 48h)
|
|
echo " 🔴 PBS GC: stale — last completed run was ${AGE_HOURS}h ago (pending-bytes: ${PENDING_BYTES} B)"
|
|
FAILED+=("pbs-gc")
|
|
else
|
|
# State 4: healthy (completed within 48h)
|
|
echo " ✅ PBS GC: healthy — last completed run ${AGE_HOURS}h ago (pending-bytes: ${PENDING_BYTES} B)"
|
|
fi
|
|
fi
|
|
fi
|
|
|
|
# ── Summary ─────────────────────────────────────────────────────────────────
|
|
echo ""
|
|
if [ ${#FAILED[@]} -eq 0 ]; then
|
|
echo " ✅ All legs OK"
|
|
exit 0
|
|
else
|
|
for f in "${FAILED[@]}"; do
|
|
echo " 🔴 FAILED: $f"
|
|
done
|
|
exit 1
|
|
fi |