diff --git a/agent-health-check.prose.md b/agent-health-check.prose.md index ea36f55..3b3c83a 100644 --- a/agent-health-check.prose.md +++ b/agent-health-check.prose.md @@ -38,8 +38,7 @@ Runs every 4 hours (2, 6, 10, 14, 18, 22 UTC at :35) via cron (`35 2,6,10,14,18, - ct_liveness: map of CT → active status - config_integrity: map of config file → valid/invalid -(## Execution -) +## Execution **Execution model**: This contract is executed by a host-scheduled cron job (see `scripts/contract-run.sh`). The cron job runs the monitoring script directly on the target host and appends the result to `/var/log/contract-runs/.log`. An diff --git a/disk-gc-threat-response.prose.md b/disk-gc-threat-response.prose.md index 49023e6..7dd46dd 100644 --- a/disk-gc-threat-response.prose.md +++ b/disk-gc-threat-response.prose.md @@ -159,8 +159,7 @@ from the `report_only_guests` YAML block above. - `timeout`: 300 seconds per CT (GC may take time on large docker hosts) - `retry`: 2 attempts for SSH failures before marking a CT unreachable -(## Execution -) +## Execution **Execution model**: This contract is executed by a host-scheduled cron job (see `scripts/contract-run.sh`). The cron job runs the monitoring script directly on the target host and appends the result to `/var/log/contract-runs/.log`. An diff --git a/infrastructure-monitoring.prose.md b/infrastructure-monitoring.prose.md index bb5db97..a67629c 100644 --- a/infrastructure-monitoring.prose.md +++ b/infrastructure-monitoring.prose.md @@ -101,8 +101,7 @@ GPU .8 (RTX 3090) GPU .110 (RTX 5070) GPU .15 (Strix Halo) - LiteLLM metrics (requests, tokens, latency, errors) visible alongside GPU metrics - Stack persists across reboots (systemd for exporters, Docker restart policy) -(## Execution -) +## Execution **Execution model**: This contract is executed by a host-scheduled cron job (see `scripts/contract-run.sh`). The cron job runs the monitoring script directly on the target host and appends the result to `/var/log/contract-runs/.log`. An diff --git a/litellm-health.prose.md b/litellm-health.prose.md index af7c5aa..8f347cd 100644 --- a/litellm-health.prose.md +++ b/litellm-health.prose.md @@ -110,8 +110,7 @@ contracts — read them there. Do not re-add retired names (`gemma-4-12b`, `gpu- | harness-grafana | grafana/grafana | :3000→:3001 | /api/health | | harness-prometheus | prom/prometheus | :9090 | /-/healthy | -(## Execution -) +## Execution **Execution model**: This contract is executed by a host-scheduled cron job (see `scripts/contract-run.sh`). The cron job runs the monitoring script directly on the target host and appends the result to `/var/log/contract-runs/.log`. An diff --git a/pm2-self-heal.prose.md b/pm2-self-heal.prose.md index c386a44..1336f70 100644 --- a/pm2-self-heal.prose.md +++ b/pm2-self-heal.prose.md @@ -50,8 +50,7 @@ description: > This was the root cause of the 18-restart accumulation. Threshold raised and PM2 counter reset on 2026-06-28. -(## Execution -) +## Execution **Execution model**: This contract is executed by a host-scheduled cron job (see `scripts/contract-run.sh`). The cron job runs the monitoring script directly on the target host and appends the result to `/var/log/contract-runs/.log`. An diff --git a/scripts/contract-run.sh b/scripts/contract-run.sh index 0203f52..7276486 100755 --- a/scripts/contract-run.sh +++ b/scripts/contract-run.sh @@ -2,7 +2,11 @@ # contract-run.sh — Deterministic contract execution from machine scheduler # # Takes a contract name, resolves its script, runs it with timeout, -# logs output to /var/log/contract-runs/, and alerts on failure. +# logs output to $CONTRACT_RUN_LOG_DIR (default: /var/log/contract-runs/), +# and alerts on failure. +# +# Environment: +# CONTRACT_RUN_LOG_DIR Override the log directory (default: /var/log/contract-runs) # # Usage: bash scripts/contract-run.sh # @@ -24,7 +28,7 @@ set -uo pipefail CONTRACT_NAME="$1" SCRIPTS_DIR="$(cd "$(dirname "$0")" && pwd)" -LOG_DIR="/var/log/contract-runs" +LOG_DIR="${CONTRACT_RUN_LOG_DIR:-/var/log/contract-runs}" TIMESTAMP=$(date -u '+%Y%m%d-%H%M%S') LOG_FILE="${LOG_DIR}/${CONTRACT_NAME}-${TIMESTAMP}.log" @@ -90,30 +94,38 @@ else # Send alert (Zulip DM to user 9 + stream agent-hub topic alerts-infra) # Using the same alert path as other monitors ALERT_MSG="🔴 Contract $CONTRACT_NAME failed (exit $EXIT_CODE). Log: $LOG_FILE" + ALERT_SENT=false - # Try to send via the existing alert mechanism - if command -v curl &> /dev/null; then - # Zulip DM to user 9 - ZULIP_API_URL="http://192.168.68.117/api/v1" - ZULIP_API_KEY=$(cat /root/.config/zulip-api-key 2>/dev/null || echo "") - - if [ -n "$ZULIP_API_KEY" ]; then - curl -s -X POST "${ZULIP_API_URL}/messages" \ - -u "user:${ZULIP_API_KEY}" \ - -d "type=private" \ - -d "to=9" \ - -d "content=${ALERT_MSG}" > /dev/null 2>&1 - fi + # Take credentials from environment (ZULIP_API_KEY required) + ZULIP_API_URL="${ZULIP_API_URL:-https://chat.sysloggh.net/api/v1}" + ZULIP_API_KEY="${ZULIP_API_KEY:-}" + ZULIP_USER="${ZULIP_USER:-abiba-bot@chat.sysloggh.net}" + + if [ -n "$ZULIP_API_KEY" ] && command -v curl &> /dev/null; then + # DM to user 9 + DM_EXIT=0 + curl -s -X POST "${ZULIP_API_URL}/messages" \ + -u "${ZULIP_USER}:${ZULIP_API_KEY}" \ + -d "type=private" \ + -d "to=9" \ + -d "content=${ALERT_MSG}" > /dev/null 2>&1 || DM_EXIT=$? # Stream agent-hub topic alerts-infra - if [ -n "$ZULIP_API_KEY" ]; then - curl -s -X POST "${ZULIP_API_URL}/messages" \ - -u "user:${ZULIP_API_KEY}" \ - -d "type=stream" \ - -d "to=agent-hub" \ - -d "topic=alerts-infra" \ - -d "content=${ALERT_MSG}" > /dev/null 2>&1 + STREAM_EXIT=0 + curl -s -X POST "${ZULIP_API_URL}/messages" \ + -u "${ZULIP_USER}:${ZULIP_API_KEY}" \ + -d "type=stream" \ + -d "to=agent-hub" \ + -d "topic=alerts-infra" \ + -d "content=${ALERT_MSG}" > /dev/null 2>&1 || STREAM_EXIT=$? + + if [ $DM_EXIT -eq 0 ] || [ $STREAM_EXIT -eq 0 ]; then + ALERT_SENT=true + else + echo "$(date -u '+%Y-%m-%dT%H:%M:%SZ') ALERT FAILURE: DM exit=$DM_EXIT, stream exit=$STREAM_EXIT" >> "$LOG_FILE" fi + else + echo "$(date -u '+%Y-%m-%dT%H:%M:%SZ') ALERT SKIPPED: no ZULIP_API_KEY or curl" >> "$LOG_FILE" fi exit 1 diff --git a/zulip-health.prose.md b/zulip-health.prose.md index 0e0202d..fabe836 100644 --- a/zulip-health.prose.md +++ b/zulip-health.prose.md @@ -125,8 +125,7 @@ grep -c "async def edit_message" ~/.hermes/plugins/*/zulip*/adapter.py - **On `zulip-status` command**: Run on-demand and report to user - **On critical alert**: Escalate to relay message immediately, don't wait for schedule -(## Execution -) +## Execution **Execution model**: This contract is executed by a host-scheduled cron job (see `scripts/contract-run.sh`). The cron job runs the monitoring script directly on the target host and appends the result to `/var/log/contract-runs/.log`. An