PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 9s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 10s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s
relay-785 / self-heal-contract-stale-script-pointers-20260924. gpu/ had been silent since 2026-09-14T18:02:03Z (12 days) and pm2/ had never posted a run. FINDING - the gpu-self-heal executor was never lost, only its schedule was. The brief concluded the mechanism was gone. It is not: on CT 116 /opt/inference-harness/scripts/gpu-self-heal.py exists (mtime 2026-09-21 11:43), is self-posting via gitea-logger.sh, and runs cleanly - I ran it and it posted health-logs/gpu/gpu-self-heal-20260926-144618.json. What was missing was the cron entry, dropped during a CT 116 /etc/cron.d rework on 2026-09-21. So this is a schedule restoration, not a resurrection. DECISIONS * gpu-self-heal -> RESTORE. Schedule re-added as /etc/cron.d/gpu-self-heal on CT 116, matching the observed historical cadence (:02 past 0/6/12/18). Proven with a REAL unattended cron fire (temporary */1 entry, removed after): gpu-self-heal-20260926-144802.json committed 14:48:03Z. The contract now names the executor, schedule, log and posting, which it previously did not. * pm2 health-logs posting -> RETIRE the claim. pm2-self-heal.prose.md required appending to health-logs/pm2/ as a 'hard rule', but scripts/pm2-self-heal.sh contains no Gitea or push code and never did; the directory has held only its init commit since 2026-07-28. Corrected to point at the contract-runner's durable per-run logs and failure note instead of adding a second, redundant posting path. DEAD-MAN'S-SWITCH - scripts/health-log-freshness.py + contract. Absence of logs must raise an alarm, and the producer cannot raise it, so this runs on CT 100, a different host from the producers, and fails when the newest health-logs/gpu entry is older than 12h (litellm 18h). Verified it would have caught the real gap: evaluated at 2026-09-20 the newest gpu/ entry was 126h old against a 12h limit -> STALE. Bugs found and fixed while testing, each caught by a test that bit: * the documented HEALTH_LOG_MAX_AGE_* override was never implemented; * a Gitea password held under GITEA_TOKEN was sent as an API token -> HTTP 401; now every candidate auth is tried and the first that works is used; * the ~/.git-credentials fallback filtered on GITEA_URL's host, which missed whenever GITEA_URL pointed at the internal IP -> zero candidates. Exit codes: 0 fresh, 1 stale-or-unreadable, 2 cannot run. A directory that cannot be read is a failure, never a skip. Evidence: live PASS; stale override names the directory and producer; no credential -> exit 2; whole thing runs green through contract-run.sh. prose-lint: PASSED.
226 lines
9.2 KiB
Bash
Executable File
226 lines
9.2 KiB
Bash
Executable File
#!/bin/bash
|
|
# contract-run.sh — Deterministic contract execution from machine scheduler
|
|
#
|
|
# Takes a contract name, resolves its script, runs it with timeout,
|
|
# logs output to $CONTRACT_RUN_LOG_DIR (default: /var/log/contract-runs/),
|
|
# and alerts on failure.
|
|
#
|
|
# Environment:
|
|
# CONTRACT_RUN_LOG_DIR Override the log directory (default: /var/log/contract-runs)
|
|
#
|
|
# Usage: bash scripts/contract-run.sh <contract-name>
|
|
#
|
|
# Contract names map to scripts as follows:
|
|
# infrastructure-monitoring -> scripts/infra-monitoring.sh
|
|
# proxmox-monitor -> scripts/proxmox-monitor.sh
|
|
# zulip-health -> scripts/zulip-monitor.sh
|
|
# agent-health-check -> scripts/agent-health-check.py
|
|
# litellm-health -> scripts/litellm-health-check.py
|
|
# disk-gc-threat-response -> scripts/disk-gc-scan.py
|
|
# pm2-self-heal -> scripts/pm2-self-heal.sh
|
|
# search-stack-visibility -> scripts/search-stack-check.py
|
|
# health-log-freshness -> scripts/health-log-freshness.py
|
|
#
|
|
# Execution copy: every contract pins the clone this script lives in (see
|
|
# docs/contract-execution-pinning.md). Before a contract runs, this wrapper
|
|
# proves the script it is about to execute byte-matches origin/master:
|
|
# CONTRACT_REVISION_PREFLIGHT=warn (default) log a refusal, still report
|
|
# CONTRACT_REVISION_PREFLIGHT=enforce refuse to report on a mismatch
|
|
# CONTRACT_REVISION_PREFLIGHT=off skip the check entirely
|
|
# A refusal names its class: cannot-verify:fetch-failed|ref-unresolvable,
|
|
# or mismatch:content|path-absent|detached-head|clone-ahead.
|
|
#
|
|
# Exit codes:
|
|
# 0 = contract passed
|
|
# 1 = contract failed (alert sent)
|
|
# 2 = probe failed (script missing, timeout, etc.)
|
|
|
|
set -uo pipefail
|
|
|
|
CONTRACT_NAME="$1"
|
|
SCRIPTS_DIR="$(cd "$(dirname "$0")" && pwd)"
|
|
LOG_DIR="${CONTRACT_RUN_LOG_DIR:-/var/log/contract-runs}"
|
|
TIMESTAMP=$(date -u '+%Y%m%d-%H%M%S')
|
|
LOG_FILE="${LOG_DIR}/${CONTRACT_NAME}-${TIMESTAMP}.log"
|
|
|
|
# Ensure log directory exists
|
|
mkdir -p "$LOG_DIR"
|
|
|
|
# Map contract name to script path
|
|
case "$CONTRACT_NAME" in
|
|
infrastructure-monitoring)
|
|
SCRIPT_PATH="${SCRIPTS_DIR}/infra-monitoring.sh"
|
|
INTERPRETER="bash"
|
|
;;
|
|
proxmox-monitor)
|
|
SCRIPT_PATH="${SCRIPTS_DIR}/proxmox-monitor.sh"
|
|
INTERPRETER="bash"
|
|
;;
|
|
zulip-health)
|
|
SCRIPT_PATH="${SCRIPTS_DIR}/zulip-monitor.sh"
|
|
INTERPRETER="bash"
|
|
;;
|
|
agent-health-check)
|
|
SCRIPT_PATH="${SCRIPTS_DIR}/agent-health-check.py"
|
|
INTERPRETER="python3"
|
|
;;
|
|
litellm-health)
|
|
SCRIPT_PATH="${SCRIPTS_DIR}/litellm-health-check.py"
|
|
INTERPRETER="python3"
|
|
;;
|
|
pm2-self-heal)
|
|
SCRIPT_PATH="${SCRIPTS_DIR}/pm2-self-heal.sh"
|
|
INTERPRETER="bash"
|
|
;;
|
|
disk-gc-threat-response)
|
|
SCRIPT_PATH="${SCRIPTS_DIR}/disk-gc-scan.py"
|
|
INTERPRETER="python3"
|
|
;;
|
|
search-stack-visibility)
|
|
SCRIPT_PATH="${SCRIPTS_DIR}/search-stack-check.py"
|
|
INTERPRETER="python3"
|
|
;;
|
|
health-log-freshness)
|
|
SCRIPT_PATH="${SCRIPTS_DIR}/health-log-freshness.py"
|
|
INTERPRETER="python3"
|
|
;;
|
|
*)
|
|
echo "Unknown contract: $CONTRACT_NAME" | tee -a "$LOG_FILE"
|
|
# Send alert for unknown contract
|
|
ALERT_MSG="🔴 Contract $CONTRACT_NAME: unknown contract name. Log: $LOG_FILE"
|
|
ZULIP_API_URL="${ZULIP_API_URL:-https://chat.sysloggh.net/api/v1}"
|
|
ZULIP_API_KEY="${ZULIP_API_KEY:-}"
|
|
ZULIP_USER="${ZULIP_USER:-abiba-bot@chat.sysloggh.net}"
|
|
if [ -n "$ZULIP_API_KEY" ] && command -v curl &> /dev/null; then
|
|
curl -sf -X POST "${ZULIP_API_URL}/messages" \
|
|
-u "${ZULIP_USER}:${ZULIP_API_KEY}" \
|
|
-d "type=private" \
|
|
-d "to=9" \
|
|
-d "content=${ALERT_MSG}" > /dev/null 2>&1 || true
|
|
fi
|
|
exit 2
|
|
;;
|
|
esac
|
|
|
|
# Check if script exists
|
|
if [ ! -f "$SCRIPT_PATH" ]; then
|
|
echo "Script not found: $SCRIPT_PATH" | tee -a "$LOG_FILE"
|
|
# Send alert for missing script
|
|
ALERT_MSG="🔴 Contract $CONTRACT_NAME: script not found at $SCRIPT_PATH. Log: $LOG_FILE"
|
|
ZULIP_API_URL="${ZULIP_API_URL:-https://chat.sysloggh.net/api/v1}"
|
|
ZULIP_API_KEY="${ZULIP_API_KEY:-}"
|
|
ZULIP_USER="${ZULIP_USER:-abiba-bot@chat.sysloggh.net}"
|
|
if [ -n "$ZULIP_API_KEY" ] && command -v curl &> /dev/null; then
|
|
curl -sf -X POST "${ZULIP_API_URL}/messages" \
|
|
-u "${ZULIP_USER}:${ZULIP_API_KEY}" \
|
|
-d "type=private" \
|
|
-d "to=9" \
|
|
-d "content=${ALERT_MSG}" > /dev/null 2>&1 || true
|
|
fi
|
|
exit 2
|
|
fi
|
|
|
|
# Run the script with timeout and capture output
|
|
echo "=== Contract: $CONTRACT_NAME ===" | tee "$LOG_FILE"
|
|
echo "Started: $(date -u '+%Y-%m-%d %H:%M:%S UTC')" | tee -a "$LOG_FILE"
|
|
echo "Script: $SCRIPT_PATH" | tee -a "$LOG_FILE"
|
|
echo "" | tee -a "$LOG_FILE"
|
|
|
|
# ── Revision preflight ───────────────────────────────────────────────────────
|
|
# A verdict is only meaningful if it came from the merged copy. This reports
|
|
# whether the executing copy matches, and distinguishes "could not check" from
|
|
# "this copy is wrong" so an operator can tell them apart.
|
|
# See docs/contract-execution-pinning.md.
|
|
#
|
|
# Default is WARN, not enforce: the guard gates every scheduled contract, and a
|
|
# legitimate state (branch mid-review, detached HEAD, briefly offline) would
|
|
# otherwise turn the whole fleet's monitoring into withheld verdicts. The
|
|
# criteria for flipping the default to enforce are written down in the doc.
|
|
REVISION_PREFLIGHT_MODE="${CONTRACT_REVISION_PREFLIGHT:-warn}"
|
|
REPO_ROOT="$(cd "${SCRIPTS_DIR}/.." && pwd)"
|
|
if [ "$REVISION_PREFLIGHT_MODE" != "off" ] && [ -x "${SCRIPTS_DIR}/revision-preflight.sh" ]; then
|
|
PREFLIGHT_OUT="$(mktemp)"
|
|
if "${SCRIPTS_DIR}/revision-preflight.sh" "$SCRIPT_PATH" "$REPO_ROOT" >"$PREFLIGHT_OUT" 2>&1; then
|
|
cat "$PREFLIGHT_OUT" | tee -a "$LOG_FILE"
|
|
else
|
|
cat "$PREFLIGHT_OUT" | tee -a "$LOG_FILE"
|
|
PREFLIGHT_REASON="$(grep -m1 '^REASON=' "$PREFLIGHT_OUT" | cut -d= -f2-)"
|
|
[ -n "$PREFLIGHT_REASON" ] || PREFLIGHT_REASON="unclassified"
|
|
if [ "$REVISION_PREFLIGHT_MODE" = "enforce" ]; then
|
|
echo "🚫 VERDICT WITHHELD: $PREFLIGHT_REASON" | tee -a "$LOG_FILE"
|
|
ALERT_MSG="🔴 Contract $CONTRACT_NAME: revision preflight REFUSED ($PREFLIGHT_REASON) — verdict withheld. Log: $LOG_FILE"
|
|
ZULIP_API_URL="${ZULIP_API_URL:-https://chat.sysloggh.net/api/v1}"
|
|
ZULIP_API_KEY="${ZULIP_API_KEY:-}"
|
|
ZULIP_USER="${ZULIP_USER:-abiba-bot@chat.sysloggh.net}"
|
|
if [ -n "$ZULIP_API_KEY" ] && command -v curl &> /dev/null; then
|
|
curl -sf -X POST "${ZULIP_API_URL}/messages" \
|
|
-u "${ZULIP_USER}:${ZULIP_API_KEY}" \
|
|
-d "type=private" \
|
|
-d "to=9" \
|
|
-d "content=${ALERT_MSG}" > /dev/null 2>&1 || true
|
|
fi
|
|
rm -f "$PREFLIGHT_OUT"
|
|
exit 2
|
|
fi
|
|
echo "⚠️ revision preflight: $PREFLIGHT_REASON — continuing because CONTRACT_REVISION_PREFLIGHT=$REVISION_PREFLIGHT_MODE" | tee -a "$LOG_FILE"
|
|
fi
|
|
rm -f "$PREFLIGHT_OUT"
|
|
fi
|
|
|
|
# Use timeout to prevent hangs (10 minutes default)
|
|
TIMEOUT=600
|
|
timeout "$TIMEOUT" $INTERPRETER "$SCRIPT_PATH" 2>&1 | tee -a "$LOG_FILE"
|
|
EXIT_CODE=${PIPESTATUS[0]}
|
|
|
|
# If timeout killed the process, EXIT_CODE will be 124
|
|
if [ $EXIT_CODE -eq 124 ]; then
|
|
echo "⏰ TIMEOUT: script exceeded ${TIMEOUT}s limit" | tee -a "$LOG_FILE"
|
|
fi
|
|
|
|
echo "" | tee -a "$LOG_FILE"
|
|
if [ $EXIT_CODE -eq 0 ]; then
|
|
echo "✅ VERDICT: PASS" | tee -a "$LOG_FILE"
|
|
exit 0
|
|
else
|
|
echo "🔴 VERDICT: FAIL (exit code $EXIT_CODE)" | tee -a "$LOG_FILE"
|
|
|
|
# Send alert (Zulip DM to user 9 + stream agent-hub topic alerts-infra)
|
|
# Using the same alert path as other monitors
|
|
ALERT_MSG="🔴 Contract $CONTRACT_NAME failed (exit $EXIT_CODE). Log: $LOG_FILE"
|
|
ALERT_SENT=false
|
|
|
|
# Take credentials from environment (ZULIP_API_KEY required)
|
|
ZULIP_API_URL="${ZULIP_API_URL:-https://chat.sysloggh.net/api/v1}"
|
|
ZULIP_API_KEY="${ZULIP_API_KEY:-}"
|
|
ZULIP_USER="${ZULIP_USER:-abiba-bot@chat.sysloggh.net}"
|
|
|
|
if [ -n "$ZULIP_API_KEY" ] && command -v curl &> /dev/null; then
|
|
# DM to user 9
|
|
DM_EXIT=0
|
|
curl -sf -X POST "${ZULIP_API_URL}/messages" \
|
|
-u "${ZULIP_USER}:${ZULIP_API_KEY}" \
|
|
-d "type=private" \
|
|
-d "to=9" \
|
|
-d "content=${ALERT_MSG}" > /dev/null 2>&1 || DM_EXIT=$?
|
|
|
|
# Stream agent-hub topic alerts-infra
|
|
STREAM_EXIT=0
|
|
curl -sf -X POST "${ZULIP_API_URL}/messages" \
|
|
-u "${ZULIP_USER}:${ZULIP_API_KEY}" \
|
|
-d "type=stream" \
|
|
-d "to=agent-hub" \
|
|
-d "topic=alerts-infra" \
|
|
-d "content=${ALERT_MSG}" > /dev/null 2>&1 || STREAM_EXIT=$?
|
|
|
|
if [ $DM_EXIT -eq 0 ] || [ $STREAM_EXIT -eq 0 ]; then
|
|
ALERT_SENT=true
|
|
else
|
|
echo "$(date -u '+%Y-%m-%dT%H:%M:%SZ') ALERT FAILURE: DM exit=$DM_EXIT, stream exit=$STREAM_EXIT" >> "$LOG_FILE"
|
|
fi
|
|
else
|
|
echo "$(date -u '+%Y-%m-%dT%H:%M:%SZ') ALERT SKIPPED: no ZULIP_API_KEY or curl" >> "$LOG_FILE"
|
|
fi
|
|
|
|
exit 1
|
|
fi
|