PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 9s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 10s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s
relay-785 / self-heal-contract-stale-script-pointers-20260924. gpu/ had been silent since 2026-09-14T18:02:03Z (12 days) and pm2/ had never posted a run. FINDING - the gpu-self-heal executor was never lost, only its schedule was. The brief concluded the mechanism was gone. It is not: on CT 116 /opt/inference-harness/scripts/gpu-self-heal.py exists (mtime 2026-09-21 11:43), is self-posting via gitea-logger.sh, and runs cleanly - I ran it and it posted health-logs/gpu/gpu-self-heal-20260926-144618.json. What was missing was the cron entry, dropped during a CT 116 /etc/cron.d rework on 2026-09-21. So this is a schedule restoration, not a resurrection. DECISIONS * gpu-self-heal -> RESTORE. Schedule re-added as /etc/cron.d/gpu-self-heal on CT 116, matching the observed historical cadence (:02 past 0/6/12/18). Proven with a REAL unattended cron fire (temporary */1 entry, removed after): gpu-self-heal-20260926-144802.json committed 14:48:03Z. The contract now names the executor, schedule, log and posting, which it previously did not. * pm2 health-logs posting -> RETIRE the claim. pm2-self-heal.prose.md required appending to health-logs/pm2/ as a 'hard rule', but scripts/pm2-self-heal.sh contains no Gitea or push code and never did; the directory has held only its init commit since 2026-07-28. Corrected to point at the contract-runner's durable per-run logs and failure note instead of adding a second, redundant posting path. DEAD-MAN'S-SWITCH - scripts/health-log-freshness.py + contract. Absence of logs must raise an alarm, and the producer cannot raise it, so this runs on CT 100, a different host from the producers, and fails when the newest health-logs/gpu entry is older than 12h (litellm 18h). Verified it would have caught the real gap: evaluated at 2026-09-20 the newest gpu/ entry was 126h old against a 12h limit -> STALE. Bugs found and fixed while testing, each caught by a test that bit: * the documented HEALTH_LOG_MAX_AGE_* override was never implemented; * a Gitea password held under GITEA_TOKEN was sent as an API token -> HTTP 401; now every candidate auth is tried and the first that works is used; * the ~/.git-credentials fallback filtered on GITEA_URL's host, which missed whenever GITEA_URL pointed at the internal IP -> zero candidates. Exit codes: 0 fresh, 1 stale-or-unreadable, 2 cannot run. A directory that cannot be read is a failure, never a skip. Evidence: live PASS; stale override names the directory and producer; no credential -> exit 2; whole thing runs green through contract-run.sh. prose-lint: PASSED.
81 lines
3.0 KiB
Markdown
81 lines
3.0 KiB
Markdown
---
|
|
kind: function
|
|
name: health-log-freshness
|
|
description: >
|
|
Dead-man's-switch for the health-logs posting jobs. Fails when the newest
|
|
commit in a watched SyslogSolution/health-logs directory is older than that
|
|
directory's threshold.
|
|
|
|
Exists because absence of logs raises no alarm: the gpu/ directory went silent
|
|
for 12 days (2026-09-14T18:02:03Z to 2026-09-26) and nothing noticed, because
|
|
the only thing that would have noticed was the job that had stopped. The check
|
|
therefore runs on CT 100, a DIFFERENT host from the producers on CT 116, so a
|
|
dead producer host or a deleted schedule still raises an alarm.
|
|
|
|
Watched: gpu/ (12h, producer 2 */6 * * *), litellm/ (18h, producer 0 */6 * * *).
|
|
Not watched: pm2/ — that contract's health-logs claim was retired 2026-09-26;
|
|
see pm2-self-heal.prose.md step 6.
|
|
|
|
Verified 2026-09-26: would have caught the real gap (at 2026-09-20 the newest
|
|
gpu/ entry was 126h old against a 12h limit).
|
|
|
|
version: 1.0.0
|
|
---
|
|
|
|
## Execution
|
|
|
|
Host-scheduled on CT 100 via `/etc/cron.d/contract-runner`:
|
|
|
|
```
|
|
20 */4 * * * root CONTRACT_RUN_LOG_DIR=/var/log/contract-runs /bin/bash /opt/contract-runner/scripts/contract-run.sh health-log-freshness >/dev/null 2>&1 || /root/abiba-workspace/bin/fm-inbox.sh note "contract-runner: health-log-freshness FAILED - see /var/log/contract-runs/" >/dev/null 2>&1
|
|
```
|
|
|
|
Every 4 hours, offset to `:20` to avoid the existing `:05`/`:15`/`:35` slots.
|
|
|
|
## Output shape
|
|
|
|
```
|
|
Health-log freshness — dead-man's-switch
|
|
==================================================================
|
|
✅ health-logs/gpu/ newest 2026-09-26T14:48:03Z (0.02h old, limit 12.0h)
|
|
last commit: gpu: gpu-self-heal-20260926-144802.json
|
|
producer: gpu-self-heal.py, CT116 cron 2 */6 * * *
|
|
==================================================================
|
|
VERDICT: PASS — every watched health-log directory is advancing
|
|
```
|
|
|
|
`--json` emits `{checked: {...}, failures: [...]}`.
|
|
|
|
## Exit codes
|
|
|
|
| exit | meaning |
|
|
| --- | --- |
|
|
| 0 | every watched directory is advancing |
|
|
| 1 | at least one is stale, or could not be read |
|
|
| 2 | the check could not run (no Gitea credential) |
|
|
|
|
A directory that **cannot be read** is a failure, not a skip: unreadable and
|
|
stopped are indistinguishable from the outside.
|
|
|
|
## Configuration
|
|
|
|
| variable | default | meaning |
|
|
| --- | --- | --- |
|
|
| `GITEA_URL` | `https://git.sysloggh.net` | Gitea base URL |
|
|
| `GITEA_TOKEN` / `GITEA_PAT` | — | API token; falls back to basic auth from `~/.git-credentials` |
|
|
| `HEALTH_LOG_MAX_AGE_GPU` | `12` | hours |
|
|
| `HEALTH_LOG_MAX_AGE_LITELLM` | `18` | hours |
|
|
|
|
## Thresholds
|
|
|
|
Sized for the producer cadence plus one missed run, so a single blip does not
|
|
page but a genuine stop does:
|
|
|
|
* `gpu/` — 6 h cadence, 12 h limit;
|
|
* `litellm/` — 6 h cadence, 18 h limit (proven healthy; a looser bound avoids noise).
|
|
|
|
## Maintains
|
|
|
|
- health-logs-gpu-freshness: { status: "ok|stale", last_check: timestamp }
|
|
- health-logs-litellm-freshness: { status: "ok|stale", last_check: timestamp }
|