--- kind: function name: health-log-freshness description: > Dead-man's-switch for the health-logs posting jobs. Fails when the newest commit in a watched SyslogSolution/health-logs directory is older than that directory's threshold. Exists because absence of logs raises no alarm: the gpu/ directory went silent for 12 days (2026-09-14T18:02:03Z to 2026-09-26) and nothing noticed, because the only thing that would have noticed was the job that had stopped. The check therefore runs on CT 100, a DIFFERENT host from the producers on CT 116, so a dead producer host or a deleted schedule still raises an alarm. Watched: gpu/ (12h, producer 2 */6 * * *), litellm/ (18h, producer 0 */6 * * *). Not watched: pm2/ — that contract's health-logs claim was retired 2026-09-26; see pm2-self-heal.prose.md step 6. Verified 2026-09-26: would have caught the real gap (at 2026-09-20 the newest gpu/ entry was 126h old against a 12h limit). version: 1.0.0 --- ## Execution Host-scheduled on CT 100 via `/etc/cron.d/contract-runner`: ``` 20 */4 * * * root CONTRACT_RUN_LOG_DIR=/var/log/contract-runs /bin/bash /opt/contract-runner/scripts/contract-run.sh health-log-freshness >/dev/null 2>&1 || /root/abiba-workspace/bin/fm-inbox.sh note "contract-runner: health-log-freshness FAILED - see /var/log/contract-runs/" >/dev/null 2>&1 ``` Every 4 hours, offset to `:20` to avoid the existing `:05`/`:15`/`:35` slots. ## Output shape ``` Health-log freshness — dead-man's-switch ================================================================== ✅ health-logs/gpu/ newest 2026-09-26T14:48:03Z (0.02h old, limit 12.0h) last commit: gpu: gpu-self-heal-20260926-144802.json producer: gpu-self-heal.py, CT116 cron 2 */6 * * * ================================================================== VERDICT: PASS — every watched health-log directory is advancing ``` `--json` emits `{checked: {...}, failures: [...]}`. ## Exit codes | exit | meaning | | --- | --- | | 0 | every watched directory is advancing | | 1 | at least one is stale, or could not be read | | 2 | the check could not run (no Gitea credential) | A directory that **cannot be read** is a failure, not a skip: unreadable and stopped are indistinguishable from the outside. ## Configuration | variable | default | meaning | | --- | --- | --- | | `GITEA_URL` | `https://git.sysloggh.net` | Gitea base URL | | `GITEA_TOKEN` / `GITEA_PAT` | — | API token; falls back to basic auth from `~/.git-credentials` | | `HEALTH_LOG_MAX_AGE_GPU` | `12` | hours | | `HEALTH_LOG_MAX_AGE_LITELLM` | `18` | hours | ## Thresholds Sized for the producer cadence plus one missed run, so a single blip does not page but a genuine stop does: * `gpu/` — 6 h cadence, 12 h limit; * `litellm/` — 6 h cadence, 18 h limit (proven healthy; a looser bound avoids noise). ## Maintains - health-logs-gpu-freshness: { status: "ok|stale", last_check: timestamp } - health-logs-litellm-freshness: { status: "ok|stale", last_check: timestamp }