Closes relay-785-health-logs-review-20260926 / folds into self-heal-contract-stale-script-pointers-20260924.
gpu/ had been silent since 2026-09-14T18:02:03Z (12 days) and pm2/ had never posted a run.
Finding: the gpu-self-heal executor was never lost — only its schedule was
The brief concluded the mechanism was gone. It is not. On CT 116, /opt/inference-harness/scripts/gpu-self-heal.py exists (mtime 2026-09-21 11:43), is self-posting via gitea-logger.sh, and runs cleanly — I ran it and it posted health-logs/gpu/gpu-self-heal-20260926-144618.json. What was missing was the cron entry, dropped during a CT 116 /etc/cron.d rework on 2026-09-21.
So this is a schedule restoration, not a resurrection.
Decisions
gpu-self-heal → RESTORE
The contract previously described the logic but named no executor and no schedule. It now does:
Executor
/opt/inference-harness/scripts/gpu-self-heal.py on CT 116
Cadence matches the observed historical pattern (:02 past 0/6/12/18). Proven with a real unattended cron fire (temporary */1 entry, removed after the run): gpu-self-heal-20260926-144802.json, committed 2026-09-26T14:48:03Z.
pm2 health-logs posting → RETIRE the claim
pm2-self-heal.prose.md required appending to health-logs/pm2/{timestamp}.md and called it a "hard rule". But scripts/pm2-self-heal.sh contains no Gitea or git-push code — it never did — and the directory has held only its init commit since 2026-07-28. Corrected to point at the contract-runner's durable per-run logs plus its failure note, rather than adding a second, redundant posting path. Retiring a false coverage claim is the honest fix; implementing a duplicate posting would add work without adding a signal.
Dead-man's-switch
Absence of logs must raise an alarm, and the producer cannot raise it — a stopped job cannot report that it stopped. So scripts/health-log-freshness.py runs on CT 100, a different host from the producers, and fails when the newest entry in a watched directory is older than its threshold (gpu/ 12h, litellm/ 18h — cadence plus one missed run).
Would it have caught the real gap?
newest gpu/ entry 2026-09-14T18:02:03Z
evaluated at 2026-09-15T06:00:00Z -> age 12.0h | limit 12h | fresh
evaluated at 2026-09-20T00:00:00Z -> age 126.0h | limit 12h | STALE -> ALERT
evaluated at 2026-09-26T00:00:00Z -> age 270.0h | limit 12h | STALE -> ALERT
Yes — it would have flagged within about 12–18 hours of the stop.
Bugs found and fixed while testing (each caught by a test that bit)
the documented HEALTH_LOG_MAX_AGE_* override was never implemented;
a Gitea password held under GITEA_TOKEN was sent as an API token → HTTP 401; now every candidate auth is tried and the first that works is used;
the ~/.git-credentials fallback filtered on GITEA_URL's host, which missed entirely whenever GITEA_URL pointed at the internal IP → zero candidates.
Evidence
A: no GITEA_TOKEN -> VERDICT: PASS EXIT=0
B: GITEA_TOKEN=password -> VERDICT: PASS EXIT=0 (recovers from the 401)
C: no credential at all -> "no Gitea credential available" EXIT=2
D: HEALTH_LOG_MAX_AGE_GPU=0.0001 -> STALE reported for gpu/, naming the producer
through contract-run.sh -> ✅ VERDICT: PASS
prose-lint -> ✅ LINT PASSED (19 warning(s))
Exit codes: 0 fresh, 1 stale-or-unreadable, 2 cannot run. A directory that cannot be read is a failure, never a skip — unreadable and stopped are indistinguishable from the outside.
One step lands with the merge
The watchdog's cron line is documented in health-log-freshness.prose.md (20 */4 * * *, offset from the existing :05/:15/:35 slots) but I did not add it to CT 100 yet, because /opt/contract-runner will not have the script until this merges — enabling it now would raise a false FAILED note every 4 hours during review. It goes in as part of the merge-time sequence already documented in docs/contract-execution-pinning.md.
Closes `relay-785-health-logs-review-20260926` / folds into `self-heal-contract-stale-script-pointers-20260924`.
`gpu/` had been silent since `2026-09-14T18:02:03Z` (12 days) and `pm2/` had never posted a run.
## Finding: the gpu-self-heal executor was never lost — only its schedule was
The brief concluded the mechanism was gone. **It is not.** On CT 116, `/opt/inference-harness/scripts/gpu-self-heal.py` exists (mtime 2026-09-21 11:43), is self-posting via `gitea-logger.sh`, and runs cleanly — I ran it and it posted `health-logs/gpu/gpu-self-heal-20260926-144618.json`. What was missing was the **cron entry**, dropped during a CT 116 `/etc/cron.d` rework on 2026-09-21.
So this is a **schedule restoration**, not a resurrection.
## Decisions
### gpu-self-heal → RESTORE
The contract previously described the logic but named no executor and no schedule. It now does:
| | |
| --- | --- |
| Executor | `/opt/inference-harness/scripts/gpu-self-heal.py` on CT 116 |
| Schedule | `/etc/cron.d/gpu-self-heal` — `2 */6 * * *` |
| Log | `/var/log/litellm/gpu-self-heal.log` |
| Posting | `gitea-logger.sh gpu {RUN_ID}.json` → `health-logs/gpu/` |
Cadence matches the observed historical pattern (`:02` past 0/6/12/18). Proven with a **real unattended cron fire** (temporary `*/1` entry, removed after the run): `gpu-self-heal-20260926-144802.json`, committed `2026-09-26T14:48:03Z`.
### pm2 health-logs posting → RETIRE the claim
`pm2-self-heal.prose.md` required appending to `health-logs/pm2/{timestamp}.md` and called it a **"hard rule"**. But `scripts/pm2-self-heal.sh` contains **no Gitea or git-push code** — it never did — and the directory has held only its init commit since 2026-07-28. Corrected to point at the contract-runner's durable per-run logs plus its failure note, rather than adding a second, redundant posting path. Retiring a false coverage claim is the honest fix; implementing a duplicate posting would add work without adding a signal.
## Dead-man's-switch
Absence of logs must raise an alarm, and **the producer cannot raise it** — a stopped job cannot report that it stopped. So `scripts/health-log-freshness.py` runs on **CT 100, a different host from the producers**, and fails when the newest entry in a watched directory is older than its threshold (`gpu/` 12h, `litellm/` 18h — cadence plus one missed run).
Would it have caught the real gap?
```
newest gpu/ entry 2026-09-14T18:02:03Z
evaluated at 2026-09-15T06:00:00Z -> age 12.0h | limit 12h | fresh
evaluated at 2026-09-20T00:00:00Z -> age 126.0h | limit 12h | STALE -> ALERT
evaluated at 2026-09-26T00:00:00Z -> age 270.0h | limit 12h | STALE -> ALERT
```
**Yes** — it would have flagged within about 12–18 hours of the stop.
## Bugs found and fixed while testing (each caught by a test that bit)
1. the documented `HEALTH_LOG_MAX_AGE_*` override was **never implemented**;
2. a Gitea **password held under `GITEA_TOKEN`** was sent as an API token → HTTP 401; now every candidate auth is tried and the first that works is used;
3. the `~/.git-credentials` fallback filtered on `GITEA_URL`'s host, which missed entirely whenever `GITEA_URL` pointed at the internal IP → zero candidates.
## Evidence
```
A: no GITEA_TOKEN -> VERDICT: PASS EXIT=0
B: GITEA_TOKEN=password -> VERDICT: PASS EXIT=0 (recovers from the 401)
C: no credential at all -> "no Gitea credential available" EXIT=2
D: HEALTH_LOG_MAX_AGE_GPU=0.0001 -> STALE reported for gpu/, naming the producer
through contract-run.sh -> ✅ VERDICT: PASS
prose-lint -> ✅ LINT PASSED (19 warning(s))
```
Exit codes: `0` fresh, `1` stale-or-unreadable, `2` cannot run. **A directory that cannot be read is a failure, never a skip** — unreadable and stopped are indistinguishable from the outside.
## One step lands with the merge
The watchdog's cron line is documented in `health-log-freshness.prose.md` (`20 */4 * * *`, offset from the existing `:05`/`:15`/`:35` slots) but I did **not** add it to CT 100 yet, because `/opt/contract-runner` will not have the script until this merges — enabling it now would raise a false `FAILED` note every 4 hours during review. It goes in as part of the merge-time sequence already documented in `docs/contract-execution-pinning.md`.
Constraints honoured: CT 105 / kagentz untouched, .129 report-only, the `litellm-health` cron unchanged.
No master push, no merge.
relay-785 / self-heal-contract-stale-script-pointers-20260924. gpu/ had been
silent since 2026-09-14T18:02:03Z (12 days) and pm2/ had never posted a run.
FINDING - the gpu-self-heal executor was never lost, only its schedule was.
The brief concluded the mechanism was gone. It is not: on CT 116
/opt/inference-harness/scripts/gpu-self-heal.py exists (mtime 2026-09-21 11:43),
is self-posting via gitea-logger.sh, and runs cleanly - I ran it and it posted
health-logs/gpu/gpu-self-heal-20260926-144618.json. What was missing was the
cron entry, dropped during a CT 116 /etc/cron.d rework on 2026-09-21. So this is
a schedule restoration, not a resurrection.
DECISIONS
* gpu-self-heal -> RESTORE. Schedule re-added as /etc/cron.d/gpu-self-heal
on CT 116, matching the observed historical cadence (:02 past
0/6/12/18). Proven with a REAL unattended cron fire (temporary */1 entry,
removed after): gpu-self-heal-20260926-144802.json committed 14:48:03Z.
The contract now names the executor, schedule, log and posting, which it
previously did not.
* pm2 health-logs posting -> RETIRE the claim. pm2-self-heal.prose.md required
appending to health-logs/pm2/ as a 'hard rule', but scripts/pm2-self-heal.sh
contains no Gitea or push code and never did; the directory has held only its
init commit since 2026-07-28. Corrected to point at the contract-runner's
durable per-run logs and failure note instead of adding a second, redundant
posting path.
DEAD-MAN'S-SWITCH - scripts/health-log-freshness.py + contract. Absence of logs
must raise an alarm, and the producer cannot raise it, so this runs on CT 100,
a different host from the producers, and fails when the newest health-logs/gpu
entry is older than 12h (litellm 18h). Verified it would have caught the real
gap: evaluated at 2026-09-20 the newest gpu/ entry was 126h old against a 12h
limit -> STALE.
Bugs found and fixed while testing, each caught by a test that bit:
* the documented HEALTH_LOG_MAX_AGE_* override was never implemented;
* a Gitea password held under GITEA_TOKEN was sent as an API token -> HTTP 401;
now every candidate auth is tried and the first that works is used;
* the ~/.git-credentials fallback filtered on GITEA_URL's host, which missed
whenever GITEA_URL pointed at the internal IP -> zero candidates.
Exit codes: 0 fresh, 1 stale-or-unreadable, 2 cannot run. A directory that
cannot be read is a failure, never a skip.
Evidence: live PASS; stale override names the directory and producer; no
credential -> exit 2; whole thing runs green through contract-run.sh.
prose-lint: PASSED.
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
Closes
relay-785-health-logs-review-20260926/ folds intoself-heal-contract-stale-script-pointers-20260924.gpu/had been silent since2026-09-14T18:02:03Z(12 days) andpm2/had never posted a run.Finding: the gpu-self-heal executor was never lost — only its schedule was
The brief concluded the mechanism was gone. It is not. On CT 116,
/opt/inference-harness/scripts/gpu-self-heal.pyexists (mtime 2026-09-21 11:43), is self-posting viagitea-logger.sh, and runs cleanly — I ran it and it postedhealth-logs/gpu/gpu-self-heal-20260926-144618.json. What was missing was the cron entry, dropped during a CT 116/etc/cron.drework on 2026-09-21.So this is a schedule restoration, not a resurrection.
Decisions
gpu-self-heal → RESTORE
The contract previously described the logic but named no executor and no schedule. It now does:
/opt/inference-harness/scripts/gpu-self-heal.pyon CT 116/etc/cron.d/gpu-self-heal—2 */6 * * */var/log/litellm/gpu-self-heal.loggitea-logger.sh gpu {RUN_ID}.json→health-logs/gpu/Cadence matches the observed historical pattern (
:02past 0/6/12/18). Proven with a real unattended cron fire (temporary*/1entry, removed after the run):gpu-self-heal-20260926-144802.json, committed2026-09-26T14:48:03Z.pm2 health-logs posting → RETIRE the claim
pm2-self-heal.prose.mdrequired appending tohealth-logs/pm2/{timestamp}.mdand called it a "hard rule". Butscripts/pm2-self-heal.shcontains no Gitea or git-push code — it never did — and the directory has held only its init commit since 2026-07-28. Corrected to point at the contract-runner's durable per-run logs plus its failure note, rather than adding a second, redundant posting path. Retiring a false coverage claim is the honest fix; implementing a duplicate posting would add work without adding a signal.Dead-man's-switch
Absence of logs must raise an alarm, and the producer cannot raise it — a stopped job cannot report that it stopped. So
scripts/health-log-freshness.pyruns on CT 100, a different host from the producers, and fails when the newest entry in a watched directory is older than its threshold (gpu/12h,litellm/18h — cadence plus one missed run).Would it have caught the real gap?
Yes — it would have flagged within about 12–18 hours of the stop.
Bugs found and fixed while testing (each caught by a test that bit)
HEALTH_LOG_MAX_AGE_*override was never implemented;GITEA_TOKENwas sent as an API token → HTTP 401; now every candidate auth is tried and the first that works is used;~/.git-credentialsfallback filtered onGITEA_URL's host, which missed entirely wheneverGITEA_URLpointed at the internal IP → zero candidates.Evidence
Exit codes:
0fresh,1stale-or-unreadable,2cannot run. A directory that cannot be read is a failure, never a skip — unreadable and stopped are indistinguishable from the outside.One step lands with the merge
The watchdog's cron line is documented in
health-log-freshness.prose.md(20 */4 * * *, offset from the existing:05/:15/:35slots) but I did not add it to CT 100 yet, because/opt/contract-runnerwill not have the script until this merges — enabling it now would raise a falseFAILEDnote every 4 hours during review. It goes in as part of the merge-time sequence already documented indocs/contract-execution-pinning.md.Constraints honoured: CT 105 / kagentz untouched, .129 report-only, the
litellm-healthcron unchanged.No master push, no merge.