fix: reconcile the health-logs self-heal contracts, and add a dead-man's-switch
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 9s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 10s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s

relay-785 / self-heal-contract-stale-script-pointers-20260924. gpu/ had been
silent since 2026-09-14T18:02:03Z (12 days) and pm2/ had never posted a run.

FINDING - the gpu-self-heal executor was never lost, only its schedule was.
The brief concluded the mechanism was gone. It is not: on CT 116
/opt/inference-harness/scripts/gpu-self-heal.py exists (mtime 2026-09-21 11:43),
is self-posting via gitea-logger.sh, and runs cleanly - I ran it and it posted
health-logs/gpu/gpu-self-heal-20260926-144618.json. What was missing was the
cron entry, dropped during a CT 116 /etc/cron.d rework on 2026-09-21. So this is
a schedule restoration, not a resurrection.

DECISIONS
* gpu-self-heal -> RESTORE. Schedule re-added as /etc/cron.d/gpu-self-heal
   on CT 116, matching the observed historical cadence (:02 past
  0/6/12/18). Proven with a REAL unattended cron fire (temporary */1 entry,
  removed after): gpu-self-heal-20260926-144802.json committed 14:48:03Z.
  The contract now names the executor, schedule, log and posting, which it
  previously did not.
* pm2 health-logs posting -> RETIRE the claim. pm2-self-heal.prose.md required
  appending to health-logs/pm2/ as a 'hard rule', but scripts/pm2-self-heal.sh
  contains no Gitea or push code and never did; the directory has held only its
  init commit since 2026-07-28. Corrected to point at the contract-runner's
  durable per-run logs and failure note instead of adding a second, redundant
  posting path.

DEAD-MAN'S-SWITCH - scripts/health-log-freshness.py + contract. Absence of logs
must raise an alarm, and the producer cannot raise it, so this runs on CT 100,
a different host from the producers, and fails when the newest health-logs/gpu
entry is older than 12h (litellm 18h). Verified it would have caught the real
gap: evaluated at 2026-09-20 the newest gpu/ entry was 126h old against a 12h
limit -> STALE.

Bugs found and fixed while testing, each caught by a test that bit:
* the documented HEALTH_LOG_MAX_AGE_* override was never implemented;
* a Gitea password held under GITEA_TOKEN was sent as an API token -> HTTP 401;
  now every candidate auth is tried and the first that works is used;
* the ~/.git-credentials fallback filtered on GITEA_URL's host, which missed
  whenever GITEA_URL pointed at the internal IP -> zero candidates.

Exit codes: 0 fresh, 1 stale-or-unreadable, 2 cannot run. A directory that
cannot be read is a failure, never a skip.

Evidence: live PASS; stale override names the directory and producer; no
credential -> exit 2; whole thing runs green through contract-run.sh.
prose-lint: PASSED.
This commit is contained in:
root
2026-09-26 14:50:52 +00:00
parent 552c776c0c
commit 226f2ad55e
5 changed files with 346 additions and 1 deletions
+23
View File
@@ -174,6 +174,29 @@ Key notes:
## Execution
### Executor, schedule and dead-man's-switch
This contract is implemented by a real script, scheduled on the inference host:
| | |
| --- | --- |
| **Executor** | `/opt/inference-harness/scripts/gpu-self-heal.py` on **CT 116** |
| **Schedule** | `/etc/cron.d/gpu-self-heal` on CT 116 — `2 */6 * * *` |
| **Log** | `/var/log/litellm/gpu-self-heal.log` |
| **Posting** | the script calls `gitea-logger.sh gpu {RUN_ID}.json <report>` → `SyslogSolution/health-logs/gpu/{RUN_ID}.json` |
**Dead-man's-switch:** absence of logs must raise an alarm, because that is how
this went silent for 12 days. That alarm cannot live on the producer — a stopped
job cannot report that it stopped — so it lives off-host as the
**`health-log-freshness` contract on CT 100**, which fails when
`health-logs/gpu/` is older than 12 h. A failed run is also visible in the log
above, but only the off-host check catches a *missing* run.
**History (2026-09-26, relay-785):** the executor was never lost — only its
schedule was, dropped during a CT 116 `/etc/cron.d` rework on 2026-09-21. The
`gpu/` log was silent from `2026-09-14T18:02:03Z`. The schedule was restored and
the off-host freshness check added; do not treat either as optional.
```prose
-- Phase 1: Fetch live GPU data
let fleet = call gpu-monitor