relay-785 / self-heal-contract-stale-script-pointers-20260924. gpu/ had been
silent since 2026-09-14T18:02:03Z (12 days) and pm2/ had never posted a run.
FINDING - the gpu-self-heal executor was never lost, only its schedule was.
The brief concluded the mechanism was gone. It is not: on CT 116
/opt/inference-harness/scripts/gpu-self-heal.py exists (mtime 2026-09-21 11:43),
is self-posting via gitea-logger.sh, and runs cleanly - I ran it and it posted
health-logs/gpu/gpu-self-heal-20260926-144618.json. What was missing was the
cron entry, dropped during a CT 116 /etc/cron.d rework on 2026-09-21. So this is
a schedule restoration, not a resurrection.
DECISIONS
* gpu-self-heal -> RESTORE. Schedule re-added as /etc/cron.d/gpu-self-heal
on CT 116, matching the observed historical cadence (:02 past
0/6/12/18). Proven with a REAL unattended cron fire (temporary */1 entry,
removed after): gpu-self-heal-20260926-144802.json committed 14:48:03Z.
The contract now names the executor, schedule, log and posting, which it
previously did not.
* pm2 health-logs posting -> RETIRE the claim. pm2-self-heal.prose.md required
appending to health-logs/pm2/ as a 'hard rule', but scripts/pm2-self-heal.sh
contains no Gitea or push code and never did; the directory has held only its
init commit since 2026-07-28. Corrected to point at the contract-runner's
durable per-run logs and failure note instead of adding a second, redundant
posting path.
DEAD-MAN'S-SWITCH - scripts/health-log-freshness.py + contract. Absence of logs
must raise an alarm, and the producer cannot raise it, so this runs on CT 100,
a different host from the producers, and fails when the newest health-logs/gpu
entry is older than 12h (litellm 18h). Verified it would have caught the real
gap: evaluated at 2026-09-20 the newest gpu/ entry was 126h old against a 12h
limit -> STALE.
Bugs found and fixed while testing, each caught by a test that bit:
* the documented HEALTH_LOG_MAX_AGE_* override was never implemented;
* a Gitea password held under GITEA_TOKEN was sent as an API token -> HTTP 401;
now every candidate auth is tried and the first that works is used;
* the ~/.git-credentials fallback filtered on GITEA_URL's host, which missed
whenever GITEA_URL pointed at the internal IP -> zero candidates.
Exit codes: 0 fresh, 1 stale-or-unreadable, 2 cannot run. A directory that
cannot be read is a failure, never a skip.
Evidence: live PASS; stale override names the directory and producer; no
credential -> exit 2; whole thing runs green through contract-run.sh.
prose-lint: PASSED.
1. Replace '(## Execution' + ')' with '## Execution' in 6 contract files
2. Make LOG_DIR honor CONTRACT_RUN_LOG_DIR env var (default: /var/log/contract-runs)
3. Fix Zulip alert: correct URL (https://chat.sysloggh.net/api/v1), user
(abiba-bot@chat.sysloggh.net), take ZULIP_API_KEY from environment, and
write alert failures to run log
- Implement four-state PBS GC logic in proxmox-monitor.sh:
- probe-failed: unparseable JSON, store not found, or empty body → FAIL
- running: collection in progress (last-run-endtime absent, upid present) → DO NOT FAIL
- stale: no completed run within 48h → FAIL, naming last completed run age
- healthy: completed within 48h → PASS, naming endtime and pending bytes
- Add tests/test_pbs_gc_states.sh covering all four states
- Proves the test bites on the pre-fix version (5/6 tests fail)
- All 6 tests pass against the fixed version
- Update contract-run.sh to map disk-gc-threat-response -> scripts/disk-gc-scan.py
- Update Execution sections of host-scheduled contracts:
infrastructure-monitoring, zulip-health, litellm-health, agent-health-check,
disk-gc-threat-response, pm2-self-heal
Adding note that execution is host-scheduled via cron, not agent session ack.
- Add zulip-watchdog to Maintains (it's running, infrastructure-monitoring expects it)
- Remove gpu-monitor from PM2 Maintains (it's systemd-only, not PM2-tracked)
- Add Execution steps for all monitored processes (gitea-runner, zulip-watchdog)
- Update alert channel: Telegram is primary, Zulip DM is secondary
- Update script to check all 4 processes (abiba-telegram, abiba-zulip, gitea-runner, zulip-watchdog)
- Add restart count thresholds for all processes
- Update log output to include all process statuses
pm2-self-heal.prose.md:
- Add AS-BUILT note: gpu-monitor is systemd-managed, NOT PM2
- gpu-watchdog is decommissioned and folded into gpu-monitor.service
- gitea-runner is KEPT; abiba-zulip is KEPT (online for days)
- spoton-service was deleted; live PM2 set is 4 processes
- Preserve historical context for crash-loop guard
litellm-health.prose.md:
- Correct Prometheus node coverage: 6 nodes (.4/.5/.6/.9/.12/.15:9100)
- Note .4:9100 is DEAD target (no route, down for weeks)
- Clarify this does not read as 6 healthy nodes
- Remove retired names (qwen3.6-27B-code, qwen3.6-35B-udq4) from live alias claims
- Update Strix Halo model to Carnice-Qwen3.6-MoE-35B-A3B-Q4_K_M.gguf (strix-moe, 256K ctx)
- Fix litellm-health step 7 probe to gpu-vision (monitor key scoped)
- Move qwen3.6-27B-code/35B-udq4 from raw-but-live to non-resolving in audit
- Fold in pm2-self-heal: remove spoton-service (live PM2 set is 4/4)
- Update hermes templates, key enforcement, timeout tables to live names
- pm2-self-heal.sh: restart abiba-telegram when TEL_RESTARTS > 1000 even if 'online'
(catches quiet crash-loops like the 10k-restarts spoton incident); alert includes count.
- pm2-self-heal.prose.md: document the crash-loop guard under Rule 2.
- gpu-fleet.prose.md / gpu-self-heal.prose.md: gpu-dense is Qwen3.8-27B-Uncensored-Q4_K_M
(~16.8GB, alias qwen3.6-27B-code for LiteLLM routing), not SmartCode-Fable-5 (verified live
on .8:8080 Aug 16).
- pm2-self-heal.prose.md: description said 'logs every action to the
knowledge graph', contradicting its own execution step (line ~60) that
says '(not knowledge graph - hard rule)'. Align description to Gitea.
- zulip-self-heal.prose.md: Reporting section said 'every cycle produces
a knowledge graph node'. Contract is RETIRED; remove graph-node directive.
- Both now consistently log to SyslogSolution/health-logs, never the graph.
Logs (LITELLM-HEALTH, GPU-SELF-HEAL, PM2) now pushed to SyslogSolution/health-logs
instead of creating orphan nodes in the shared knowledge graph.
- litellm-self-heal: Phase 4 now calls gitea-logger instead of kg-logger
- gpu-self-heal: Reporting section updated to Gitea path
- pm2-self-heal: Log step redirected to Gitea
- 271 existing orphan nodes remain in graph (no delete tool available)
Root cause: abiba-zulip extension had STUCK_THRESHOLD_MS=30min.
After 30min of inactivity (overnight), it reported stuck=True, which
triggered the health monitor to restart it — cycling restarts up to 18.
Permanent fixes:
1. STUCK_THRESHOLD_MS raised 30min → 4 hours in extension index.js
2. PM2 restart counter reset: delete+re-add abiba-zulip (was 18, now 0)
3. pm2-self-heal.sh alert threshold documented at 30 (implementation already)
4. Prose contract updated with historical note and threshold rationale
- Added --no-color flag to all PM2 commands
- Self-process guard: abiba-zulip is READ-ONLY, never restarted
- Corrected column indices (PID=6, restarts=8, status=9)
- Added pm2-self-heal.sh shell script implementation
Responsibility contract that checks abiba-zulip and abiba-telegram
every 5 minutes. Restarts stopped/errored processes, escalates if
restart count exceeds 5/hour. Logs to knowledge graph and alerts
owner via Zulip DM.