diff --git a/agent-health-check.prose.md b/agent-health-check.prose.md index e7a438d..ea36f55 100644 --- a/agent-health-check.prose.md +++ b/agent-health-check.prose.md @@ -38,7 +38,15 @@ Runs every 4 hours (2, 6, 10, 14, 18, 22 UTC at :35) via cron (`35 2,6,10,14,18, - ct_liveness: map of CT → active status - config_integrity: map of config file → valid/invalid -## Execution +(## Execution +) +**Execution model**: This contract is executed by a host-scheduled cron job (see +`scripts/contract-run.sh`). The cron job runs the monitoring script directly on the +target host and appends the result to `/var/log/contract-runs/.log`. An +agent-session acknowledgement (a `done:` line in the ops status log) is NOT +execution — it only proves the agent read the result and reported it. The actual +monitoring work happens in the host cron job. + ### check-health diff --git a/disk-gc-threat-response.prose.md b/disk-gc-threat-response.prose.md index a18a081..49023e6 100644 --- a/disk-gc-threat-response.prose.md +++ b/disk-gc-threat-response.prose.md @@ -159,7 +159,15 @@ from the `report_only_guests` YAML block above. - `timeout`: 300 seconds per CT (GC may take time on large docker hosts) - `retry`: 2 attempts for SSH failures before marking a CT unreachable -## Execution +(## Execution +) +**Execution model**: This contract is executed by a host-scheduled cron job (see +`scripts/contract-run.sh`). The cron job runs the monitoring script directly on the +target host and appends the result to `/var/log/contract-runs/.log`. An +agent-session acknowledgement (a `done:` line in the ops status log) is NOT +execution — it only proves the agent read the result and reported it. The actual +monitoring work happens in the host cron job. + ### Host filesystems: report-only, NEVER auto-delete diff --git a/infrastructure-monitoring.prose.md b/infrastructure-monitoring.prose.md index 52b1ac6..bb5db97 100644 --- a/infrastructure-monitoring.prose.md +++ b/infrastructure-monitoring.prose.md @@ -101,7 +101,15 @@ GPU .8 (RTX 3090) GPU .110 (RTX 5070) GPU .15 (Strix Halo) - LiteLLM metrics (requests, tokens, latency, errors) visible alongside GPU metrics - Stack persists across reboots (systemd for exporters, Docker restart policy) -## Execution +(## Execution +) +**Execution model**: This contract is executed by a host-scheduled cron job (see +`scripts/contract-run.sh`). The cron job runs the monitoring script directly on the +target host and appends the result to `/var/log/contract-runs/.log`. An +agent-session acknowledgement (a `done:` line in the ops status log) is NOT +execution — it only proves the agent read the result and reported it. The actual +monitoring work happens in the host cron job. + ### Liveness rule (scoped) diff --git a/litellm-health.prose.md b/litellm-health.prose.md index 0c95c1b..af7c5aa 100644 --- a/litellm-health.prose.md +++ b/litellm-health.prose.md @@ -110,7 +110,15 @@ contracts — read them there. Do not re-add retired names (`gemma-4-12b`, `gpu- | harness-grafana | grafana/grafana | :3000→:3001 | /api/health | | harness-prometheus | prom/prometheus | :9090 | /-/healthy | -## Execution +(## Execution +) +**Execution model**: This contract is executed by a host-scheduled cron job (see +`scripts/contract-run.sh`). The cron job runs the monitoring script directly on the +target host and appends the result to `/var/log/contract-runs/.log`. An +agent-session acknowledgement (a `done:` line in the ops status log) is NOT +execution — it only proves the agent read the result and reported it. The actual +monitoring work happens in the host cron job. + 1. **Read parameters** — Use provided values or defaults diff --git a/pm2-self-heal.prose.md b/pm2-self-heal.prose.md index 01a7d00..c386a44 100644 --- a/pm2-self-heal.prose.md +++ b/pm2-self-heal.prose.md @@ -50,7 +50,15 @@ description: > This was the root cause of the 18-restart accumulation. Threshold raised and PM2 counter reset on 2026-06-28. -## Execution +(## Execution +) +**Execution model**: This contract is executed by a host-scheduled cron job (see +`scripts/contract-run.sh`). The cron job runs the monitoring script directly on the +target host and appends the result to `/var/log/contract-runs/.log`. An +agent-session acknowledgement (a `done:` line in the ops status log) is NOT +execution — it only proves the agent read the result and reported it. The actual +monitoring work happens in the host cron job. + 1. **Check PM2 status** — Run `pm2 status --no-color` and parse the table (5th data column = PID, 8th = restarts, 9th = status) 2. **Check abiba-telegram** (safe to auto-restart): diff --git a/scripts/contract-run.sh b/scripts/contract-run.sh index f58461b..0203f52 100755 --- a/scripts/contract-run.sh +++ b/scripts/contract-run.sh @@ -12,7 +12,7 @@ # zulip-health -> scripts/zulip-monitor.sh # agent-health-check -> scripts/agent-health-check.py # litellm-health -> scripts/litellm-health-check.py -# disk-gc-threat-response -> (Python script, needs investigation) +# disk-gc-threat-response -> scripts/disk-gc-scan.py # pm2-self-heal -> scripts/pm2-self-heal.sh # # Exit codes: diff --git a/scripts/proxmox-monitor.sh b/scripts/proxmox-monitor.sh index 68b9010..ff18a7c 100755 --- a/scripts/proxmox-monitor.sh +++ b/scripts/proxmox-monitor.sh @@ -79,8 +79,11 @@ if [ "$PBS_GC_OUTPUT" = "000" ]; then echo " 🔴 PBS GC: probe-failed: storepve:192.168.68.6 (expected JSON, got 000)" FAILED+=("pbs-gc") else - # Parse the JSON to get storepve-datastore's state - # last-run-endtime: epoch timestamp of last completion (0 when not run or in progress) + # Parse the JSON to get storepve-datastore's state with four distinct outcomes: + # 1. probe-failed: non-zero ssh status / empty / unparseable JSON + # 2. running: collection in progress (last-run-endtime absent or 0, but upid present) + # 3. stale: no completed run within 48h + # 4. healthy: completed within 48h PBS_GC_RESULT=$(echo "$PBS_GC_OUTPUT" | python3 -c " import sys, json try: @@ -88,41 +91,64 @@ try: for store in data: if store['store'] == 'storepve-datastore': endtime = store.get('last-run-endtime') + upid = store.get('upid') pending = store.get('pending-bytes', 0) + + # State 2: Running (collection in progress) — last-run-endtime absent while run is in progress + if (endtime is None or endtime == 0) and upid is not None: + print(f'running|{pending}') + break + + # State 3: No completed run (never-run or stale) if endtime is None or endtime == 0: - print('never-run') - else: - print(f'{endtime}|{pending}') + print(f'no-completed-run|{pending}') + break + + # States 3 & 4: Completed (has endtime) + print(f'completed|{endtime}|{pending}') break else: - print('absent') + print(f'absent|0') except json.JSONDecodeError: - print('unparseable') + print(f'unparseable|0') " 2>/dev/null) - if [ -z "$PBS_GC_RESULT" ] || [ "$PBS_GC_RESULT" = "unparseable" ]; then + # Parse the state|endtime|pending format + PBS_GC_STATE=$(echo "$PBS_GC_RESULT" | cut -d'|' -f1) + + if [ "$PBS_GC_STATE" = "unparseable" ]; then + # State 1: probe-failed (unparseable JSON) echo " 🔴 PBS GC: probe-failed: storepve:192.168.68.6 (unparseable JSON)" FAILED+=("pbs-gc") - elif [ "$PBS_GC_RESULT" = "absent" ]; then - echo " 🔴 PBS GC: never-run (storepve-datastore not found in GC list)" + elif [ "$PBS_GC_STATE" = "absent" ]; then + # State 1: probe-failed (store not found) + echo " 🔴 PBS GC: probe-failed: storepve:192.168.68.6 (storepve-datastore not found)" FAILED+=("pbs-gc") - elif [ "$PBS_GC_RESULT" = "never-run" ]; then - echo " 🔴 PBS GC: never-run (storepve-datastore has no last-run-endtime)" + elif [ "$PBS_GC_STATE" = "running" ]; then + # State 2: collection in progress — do NOT fail + PENDING_BYTES=$(echo "$PBS_GC_RESULT" | cut -d'|' -f2) + echo " ⏳ PBS GC: running (started: in-progress, pending-bytes: ${PENDING_BYTES} B)" + elif [ "$PBS_GC_STATE" = "no-completed-run" ]; then + # State 3: no completed run within 48h + PENDING_BYTES=$(echo "$PBS_GC_RESULT" | cut -d'|' -f2) + echo " 🔴 PBS GC: no completed run within 48h (pending-bytes: ${PENDING_BYTES} B)" FAILED+=("pbs-gc") else - # Parse the endtime|pending format - LAST_RUN_ENDTIME=$(echo "$PBS_GC_RESULT" | cut -d'|' -f1) - PENDING_BYTES=$(echo "$PBS_GC_RESULT" | cut -d'|' -f2) + # States 3 & 4: completed (has endtime) + LAST_RUN_ENDTIME=$(echo "$PBS_GC_RESULT" | cut -d'|' -f2) + PENDING_BYTES=$(echo "$PBS_GC_RESULT" | cut -d'|' -f3) # Convert epoch to age in hours NOW_EPOCH=$(date -u +%s) AGE_HOURS=$(( (NOW_EPOCH - LAST_RUN_ENDTIME) / 3600 )) if [ $AGE_HOURS -gt 48 ]; then - echo " 🔴 PBS GC: stale (last run ${AGE_HOURS}h ago, pending-bytes: ${PENDING_BYTES} B)" + # State 3: stale (no completed run within 48h) + echo " 🔴 PBS GC: stale — last completed run was ${AGE_HOURS}h ago (pending-bytes: ${PENDING_BYTES} B)" FAILED+=("pbs-gc") else - echo " ✅ PBS GC: healthy (last run ${AGE_HOURS}h ago, pending-bytes: ${PENDING_BYTES} B)" + # State 4: healthy (completed within 48h) + echo " ✅ PBS GC: healthy — last completed run ${AGE_HOURS}h ago (pending-bytes: ${PENDING_BYTES} B)" fi fi fi diff --git a/tests/test_pbs_gc_states.sh b/tests/test_pbs_gc_states.sh new file mode 100755 index 0000000..f116471 --- /dev/null +++ b/tests/test_pbs_gc_states.sh @@ -0,0 +1,81 @@ +#!/bin/bash +# test_pbs_gc_states.sh — Tests for PBS GC four-state logic + +set -uo pipefail + +TEST_DIR="$(cd "$(dirname "$0")" && pwd)" +SCRIPTS_DIR="$(dirname "$TEST_DIR")/scripts" +PROXMOX_MONITOR="${1:-${SCRIPTS_DIR}/proxmox-monitor.sh}" + +PASS=0 +FAIL=0 + +run_test() { + local name="$1" + local json="$2" + local pattern="$3" + local behavior="$4" + + local wrapper monitor + wrapper=$(mktemp /tmp/pbs-test-wrapper.XXXXXX) + monitor=$(mktemp /tmp/pbs-test-monitor.XXXXXX) + + printf '%s\n' "$json" > "$wrapper" + cp "$PROXMOX_MONITOR" "$monitor" + + # Replace the SSH call with cat "$wrapper" + python3 /tmp/replace_ssh.py "$wrapper" "$monitor" + + local output exit_code + output=$(bash "$monitor" 2>&1) + exit_code=$? + + local ok=true + if [ "$behavior" = "fail" ]; then + if ! echo "$output" | grep -q "🔴 PBS GC"; then ok=false; fi + if [ $exit_code -eq 0 ]; then ok=false; fi + else + if ! echo "$output" | grep -q "$pattern"; then ok=false; fi + if echo "$output" | grep -q "🔴 PBS GC"; then ok=false; fi + fi + + if $ok; then + echo " ✅ $name" + PASS=$((PASS + 1)) + else + echo " 🔴 $name FAILED (exit=$exit_code)" + echo "$output" | grep "PBS GC" | sed 's/^/ /' + FAIL=$((FAIL + 1)) + fi + + rm -f "$wrapper" "$monitor" +} + +echo "=== PBS GC Four-State Tests ===" +echo "Script: $PROXMOX_MONITOR" +echo "" + +echo "1. probe-failed (unparseable JSON)" +run_test "unparseable-json" "NOT JSON {{{" "probe-failed" "fail" + +echo "2. probe-failed (store not found)" +run_test "store-missing" '[{"store": "other", "last-run-endtime": 1000}]' "probe-failed" "fail" + +echo "3. probe-failed (empty body)" +run_test "empty-body" "" "probe-failed" "fail" + +echo "4. running (in progress - upid set, no last-run-endtime)" +run_test "running" '[{"store": "storepve-datastore", "upid": "UPID:123:1:456:gc:root@pam", "pending-bytes": 100}]' "⏳ PBS GC: running" "pass" + +echo "5. stale (last run >48h)" +STALE=$(date -u -d "50 hours ago" +%s) +run_test "stale" "[{\"store\": \"storepve-datastore\", \"last-run-endtime\": $STALE, \"pending-bytes\": 200}]" "stale" "fail" + +echo "6. healthy (completed <48h)" +HEALTHY=$(date -u -d "1 hour ago" +%s) +run_test "healthy" "[{\"store\": \"storepve-datastore\", \"last-run-endtime\": $HEALTHY, \"pending-bytes\": 0}]" "✅ PBS GC: healthy" "pass" + +echo "" +echo "=== Results: $PASS passed, $FAIL failed ===" +[ $FAIL -eq 0 ] && echo "✅ All passed" || echo "🔴 Some failed" +[ $FAIL -eq 0 ] && exit 0 || exit 1 diff --git a/zulip-health.prose.md b/zulip-health.prose.md index c28fbfb..0e0202d 100644 --- a/zulip-health.prose.md +++ b/zulip-health.prose.md @@ -125,7 +125,15 @@ grep -c "async def edit_message" ~/.hermes/plugins/*/zulip*/adapter.py - **On `zulip-status` command**: Run on-demand and report to user - **On critical alert**: Escalate to relay message immediately, don't wait for schedule -## Execution +(## Execution +) +**Execution model**: This contract is executed by a host-scheduled cron job (see +`scripts/contract-run.sh`). The cron job runs the monitoring script directly on the +target host and appends the result to `/var/log/contract-runs/.log`. An +agent-session acknowledgement (a `done:` line in the ops status log) is NOT +execution — it only proves the agent read the result and reported it. The actual +monitoring work happens in the host cron job. + ### Liveness rule (scoped)