Files
root 226f2ad55e
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 9s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 10s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s
fix: reconcile the health-logs self-heal contracts, and add a dead-man's-switch
relay-785 / self-heal-contract-stale-script-pointers-20260924. gpu/ had been
silent since 2026-09-14T18:02:03Z (12 days) and pm2/ had never posted a run.

FINDING - the gpu-self-heal executor was never lost, only its schedule was.
The brief concluded the mechanism was gone. It is not: on CT 116
/opt/inference-harness/scripts/gpu-self-heal.py exists (mtime 2026-09-21 11:43),
is self-posting via gitea-logger.sh, and runs cleanly - I ran it and it posted
health-logs/gpu/gpu-self-heal-20260926-144618.json. What was missing was the
cron entry, dropped during a CT 116 /etc/cron.d rework on 2026-09-21. So this is
a schedule restoration, not a resurrection.

DECISIONS
* gpu-self-heal -> RESTORE. Schedule re-added as /etc/cron.d/gpu-self-heal
   on CT 116, matching the observed historical cadence (:02 past
  0/6/12/18). Proven with a REAL unattended cron fire (temporary */1 entry,
  removed after): gpu-self-heal-20260926-144802.json committed 14:48:03Z.
  The contract now names the executor, schedule, log and posting, which it
  previously did not.
* pm2 health-logs posting -> RETIRE the claim. pm2-self-heal.prose.md required
  appending to health-logs/pm2/ as a 'hard rule', but scripts/pm2-self-heal.sh
  contains no Gitea or push code and never did; the directory has held only its
  init commit since 2026-07-28. Corrected to point at the contract-runner's
  durable per-run logs and failure note instead of adding a second, redundant
  posting path.

DEAD-MAN'S-SWITCH - scripts/health-log-freshness.py + contract. Absence of logs
must raise an alarm, and the producer cannot raise it, so this runs on CT 100,
a different host from the producers, and fails when the newest health-logs/gpu
entry is older than 12h (litellm 18h). Verified it would have caught the real
gap: evaluated at 2026-09-20 the newest gpu/ entry was 126h old against a 12h
limit -> STALE.

Bugs found and fixed while testing, each caught by a test that bit:
* the documented HEALTH_LOG_MAX_AGE_* override was never implemented;
* a Gitea password held under GITEA_TOKEN was sent as an API token -> HTTP 401;
  now every candidate auth is tried and the first that works is used;
* the ~/.git-credentials fallback filtered on GITEA_URL's host, which missed
  whenever GITEA_URL pointed at the internal IP -> zero candidates.

Exit codes: 0 fresh, 1 stale-or-unreadable, 2 cannot run. A directory that
cannot be read is a failure, never a skip.

Evidence: live PASS; stale override names the directory and producer; no
credential -> exit 2; whole thing runs green through contract-run.sh.
prose-lint: PASSED.
2026-09-26 14:50:52 +00:00

226 lines
9.2 KiB
Bash
Executable File

#!/bin/bash
# contract-run.sh — Deterministic contract execution from machine scheduler
#
# Takes a contract name, resolves its script, runs it with timeout,
# logs output to $CONTRACT_RUN_LOG_DIR (default: /var/log/contract-runs/),
# and alerts on failure.
#
# Environment:
# CONTRACT_RUN_LOG_DIR Override the log directory (default: /var/log/contract-runs)
#
# Usage: bash scripts/contract-run.sh <contract-name>
#
# Contract names map to scripts as follows:
# infrastructure-monitoring -> scripts/infra-monitoring.sh
# proxmox-monitor -> scripts/proxmox-monitor.sh
# zulip-health -> scripts/zulip-monitor.sh
# agent-health-check -> scripts/agent-health-check.py
# litellm-health -> scripts/litellm-health-check.py
# disk-gc-threat-response -> scripts/disk-gc-scan.py
# pm2-self-heal -> scripts/pm2-self-heal.sh
# search-stack-visibility -> scripts/search-stack-check.py
# health-log-freshness -> scripts/health-log-freshness.py
#
# Execution copy: every contract pins the clone this script lives in (see
# docs/contract-execution-pinning.md). Before a contract runs, this wrapper
# proves the script it is about to execute byte-matches origin/master:
# CONTRACT_REVISION_PREFLIGHT=warn (default) log a refusal, still report
# CONTRACT_REVISION_PREFLIGHT=enforce refuse to report on a mismatch
# CONTRACT_REVISION_PREFLIGHT=off skip the check entirely
# A refusal names its class: cannot-verify:fetch-failed|ref-unresolvable,
# or mismatch:content|path-absent|detached-head|clone-ahead.
#
# Exit codes:
# 0 = contract passed
# 1 = contract failed (alert sent)
# 2 = probe failed (script missing, timeout, etc.)
set -uo pipefail
CONTRACT_NAME="$1"
SCRIPTS_DIR="$(cd "$(dirname "$0")" && pwd)"
LOG_DIR="${CONTRACT_RUN_LOG_DIR:-/var/log/contract-runs}"
TIMESTAMP=$(date -u '+%Y%m%d-%H%M%S')
LOG_FILE="${LOG_DIR}/${CONTRACT_NAME}-${TIMESTAMP}.log"
# Ensure log directory exists
mkdir -p "$LOG_DIR"
# Map contract name to script path
case "$CONTRACT_NAME" in
infrastructure-monitoring)
SCRIPT_PATH="${SCRIPTS_DIR}/infra-monitoring.sh"
INTERPRETER="bash"
;;
proxmox-monitor)
SCRIPT_PATH="${SCRIPTS_DIR}/proxmox-monitor.sh"
INTERPRETER="bash"
;;
zulip-health)
SCRIPT_PATH="${SCRIPTS_DIR}/zulip-monitor.sh"
INTERPRETER="bash"
;;
agent-health-check)
SCRIPT_PATH="${SCRIPTS_DIR}/agent-health-check.py"
INTERPRETER="python3"
;;
litellm-health)
SCRIPT_PATH="${SCRIPTS_DIR}/litellm-health-check.py"
INTERPRETER="python3"
;;
pm2-self-heal)
SCRIPT_PATH="${SCRIPTS_DIR}/pm2-self-heal.sh"
INTERPRETER="bash"
;;
disk-gc-threat-response)
SCRIPT_PATH="${SCRIPTS_DIR}/disk-gc-scan.py"
INTERPRETER="python3"
;;
search-stack-visibility)
SCRIPT_PATH="${SCRIPTS_DIR}/search-stack-check.py"
INTERPRETER="python3"
;;
health-log-freshness)
SCRIPT_PATH="${SCRIPTS_DIR}/health-log-freshness.py"
INTERPRETER="python3"
;;
*)
echo "Unknown contract: $CONTRACT_NAME" | tee -a "$LOG_FILE"
# Send alert for unknown contract
ALERT_MSG="🔴 Contract $CONTRACT_NAME: unknown contract name. Log: $LOG_FILE"
ZULIP_API_URL="${ZULIP_API_URL:-https://chat.sysloggh.net/api/v1}"
ZULIP_API_KEY="${ZULIP_API_KEY:-}"
ZULIP_USER="${ZULIP_USER:-abiba-bot@chat.sysloggh.net}"
if [ -n "$ZULIP_API_KEY" ] && command -v curl &> /dev/null; then
curl -sf -X POST "${ZULIP_API_URL}/messages" \
-u "${ZULIP_USER}:${ZULIP_API_KEY}" \
-d "type=private" \
-d "to=9" \
-d "content=${ALERT_MSG}" > /dev/null 2>&1 || true
fi
exit 2
;;
esac
# Check if script exists
if [ ! -f "$SCRIPT_PATH" ]; then
echo "Script not found: $SCRIPT_PATH" | tee -a "$LOG_FILE"
# Send alert for missing script
ALERT_MSG="🔴 Contract $CONTRACT_NAME: script not found at $SCRIPT_PATH. Log: $LOG_FILE"
ZULIP_API_URL="${ZULIP_API_URL:-https://chat.sysloggh.net/api/v1}"
ZULIP_API_KEY="${ZULIP_API_KEY:-}"
ZULIP_USER="${ZULIP_USER:-abiba-bot@chat.sysloggh.net}"
if [ -n "$ZULIP_API_KEY" ] && command -v curl &> /dev/null; then
curl -sf -X POST "${ZULIP_API_URL}/messages" \
-u "${ZULIP_USER}:${ZULIP_API_KEY}" \
-d "type=private" \
-d "to=9" \
-d "content=${ALERT_MSG}" > /dev/null 2>&1 || true
fi
exit 2
fi
# Run the script with timeout and capture output
echo "=== Contract: $CONTRACT_NAME ===" | tee "$LOG_FILE"
echo "Started: $(date -u '+%Y-%m-%d %H:%M:%S UTC')" | tee -a "$LOG_FILE"
echo "Script: $SCRIPT_PATH" | tee -a "$LOG_FILE"
echo "" | tee -a "$LOG_FILE"
# ── Revision preflight ───────────────────────────────────────────────────────
# A verdict is only meaningful if it came from the merged copy. This reports
# whether the executing copy matches, and distinguishes "could not check" from
# "this copy is wrong" so an operator can tell them apart.
# See docs/contract-execution-pinning.md.
#
# Default is WARN, not enforce: the guard gates every scheduled contract, and a
# legitimate state (branch mid-review, detached HEAD, briefly offline) would
# otherwise turn the whole fleet's monitoring into withheld verdicts. The
# criteria for flipping the default to enforce are written down in the doc.
REVISION_PREFLIGHT_MODE="${CONTRACT_REVISION_PREFLIGHT:-warn}"
REPO_ROOT="$(cd "${SCRIPTS_DIR}/.." && pwd)"
if [ "$REVISION_PREFLIGHT_MODE" != "off" ] && [ -x "${SCRIPTS_DIR}/revision-preflight.sh" ]; then
PREFLIGHT_OUT="$(mktemp)"
if "${SCRIPTS_DIR}/revision-preflight.sh" "$SCRIPT_PATH" "$REPO_ROOT" >"$PREFLIGHT_OUT" 2>&1; then
cat "$PREFLIGHT_OUT" | tee -a "$LOG_FILE"
else
cat "$PREFLIGHT_OUT" | tee -a "$LOG_FILE"
PREFLIGHT_REASON="$(grep -m1 '^REASON=' "$PREFLIGHT_OUT" | cut -d= -f2-)"
[ -n "$PREFLIGHT_REASON" ] || PREFLIGHT_REASON="unclassified"
if [ "$REVISION_PREFLIGHT_MODE" = "enforce" ]; then
echo "🚫 VERDICT WITHHELD: $PREFLIGHT_REASON" | tee -a "$LOG_FILE"
ALERT_MSG="🔴 Contract $CONTRACT_NAME: revision preflight REFUSED ($PREFLIGHT_REASON) — verdict withheld. Log: $LOG_FILE"
ZULIP_API_URL="${ZULIP_API_URL:-https://chat.sysloggh.net/api/v1}"
ZULIP_API_KEY="${ZULIP_API_KEY:-}"
ZULIP_USER="${ZULIP_USER:-abiba-bot@chat.sysloggh.net}"
if [ -n "$ZULIP_API_KEY" ] && command -v curl &> /dev/null; then
curl -sf -X POST "${ZULIP_API_URL}/messages" \
-u "${ZULIP_USER}:${ZULIP_API_KEY}" \
-d "type=private" \
-d "to=9" \
-d "content=${ALERT_MSG}" > /dev/null 2>&1 || true
fi
rm -f "$PREFLIGHT_OUT"
exit 2
fi
echo "⚠️ revision preflight: $PREFLIGHT_REASON — continuing because CONTRACT_REVISION_PREFLIGHT=$REVISION_PREFLIGHT_MODE" | tee -a "$LOG_FILE"
fi
rm -f "$PREFLIGHT_OUT"
fi
# Use timeout to prevent hangs (10 minutes default)
TIMEOUT=600
timeout "$TIMEOUT" $INTERPRETER "$SCRIPT_PATH" 2>&1 | tee -a "$LOG_FILE"
EXIT_CODE=${PIPESTATUS[0]}
# If timeout killed the process, EXIT_CODE will be 124
if [ $EXIT_CODE -eq 124 ]; then
echo "⏰ TIMEOUT: script exceeded ${TIMEOUT}s limit" | tee -a "$LOG_FILE"
fi
echo "" | tee -a "$LOG_FILE"
if [ $EXIT_CODE -eq 0 ]; then
echo "✅ VERDICT: PASS" | tee -a "$LOG_FILE"
exit 0
else
echo "🔴 VERDICT: FAIL (exit code $EXIT_CODE)" | tee -a "$LOG_FILE"
# Send alert (Zulip DM to user 9 + stream agent-hub topic alerts-infra)
# Using the same alert path as other monitors
ALERT_MSG="🔴 Contract $CONTRACT_NAME failed (exit $EXIT_CODE). Log: $LOG_FILE"
ALERT_SENT=false
# Take credentials from environment (ZULIP_API_KEY required)
ZULIP_API_URL="${ZULIP_API_URL:-https://chat.sysloggh.net/api/v1}"
ZULIP_API_KEY="${ZULIP_API_KEY:-}"
ZULIP_USER="${ZULIP_USER:-abiba-bot@chat.sysloggh.net}"
if [ -n "$ZULIP_API_KEY" ] && command -v curl &> /dev/null; then
# DM to user 9
DM_EXIT=0
curl -sf -X POST "${ZULIP_API_URL}/messages" \
-u "${ZULIP_USER}:${ZULIP_API_KEY}" \
-d "type=private" \
-d "to=9" \
-d "content=${ALERT_MSG}" > /dev/null 2>&1 || DM_EXIT=$?
# Stream agent-hub topic alerts-infra
STREAM_EXIT=0
curl -sf -X POST "${ZULIP_API_URL}/messages" \
-u "${ZULIP_USER}:${ZULIP_API_KEY}" \
-d "type=stream" \
-d "to=agent-hub" \
-d "topic=alerts-infra" \
-d "content=${ALERT_MSG}" > /dev/null 2>&1 || STREAM_EXIT=$?
if [ $DM_EXIT -eq 0 ] || [ $STREAM_EXIT -eq 0 ]; then
ALERT_SENT=true
else
echo "$(date -u '+%Y-%m-%dT%H:%M:%SZ') ALERT FAILURE: DM exit=$DM_EXIT, stream exit=$STREAM_EXIT" >> "$LOG_FILE"
fi
else
echo "$(date -u '+%Y-%m-%dT%H:%M:%SZ') ALERT SKIPPED: no ZULIP_API_KEY or curl" >> "$LOG_FILE"
fi
exit 1
fi