3b74ca28d12b878d6206c761f780c7d0cd37f549
9
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
226f2ad55e |
fix: reconcile the health-logs self-heal contracts, and add a dead-man's-switch
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 9s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 10s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s
relay-785 / self-heal-contract-stale-script-pointers-20260924. gpu/ had been silent since 2026-09-14T18:02:03Z (12 days) and pm2/ had never posted a run. FINDING - the gpu-self-heal executor was never lost, only its schedule was. The brief concluded the mechanism was gone. It is not: on CT 116 /opt/inference-harness/scripts/gpu-self-heal.py exists (mtime 2026-09-21 11:43), is self-posting via gitea-logger.sh, and runs cleanly - I ran it and it posted health-logs/gpu/gpu-self-heal-20260926-144618.json. What was missing was the cron entry, dropped during a CT 116 /etc/cron.d rework on 2026-09-21. So this is a schedule restoration, not a resurrection. DECISIONS * gpu-self-heal -> RESTORE. Schedule re-added as /etc/cron.d/gpu-self-heal on CT 116, matching the observed historical cadence (:02 past 0/6/12/18). Proven with a REAL unattended cron fire (temporary */1 entry, removed after): gpu-self-heal-20260926-144802.json committed 14:48:03Z. The contract now names the executor, schedule, log and posting, which it previously did not. * pm2 health-logs posting -> RETIRE the claim. pm2-self-heal.prose.md required appending to health-logs/pm2/ as a 'hard rule', but scripts/pm2-self-heal.sh contains no Gitea or push code and never did; the directory has held only its init commit since 2026-07-28. Corrected to point at the contract-runner's durable per-run logs and failure note instead of adding a second, redundant posting path. DEAD-MAN'S-SWITCH - scripts/health-log-freshness.py + contract. Absence of logs must raise an alarm, and the producer cannot raise it, so this runs on CT 100, a different host from the producers, and fails when the newest health-logs/gpu entry is older than 12h (litellm 18h). Verified it would have caught the real gap: evaluated at 2026-09-20 the newest gpu/ entry was 126h old against a 12h limit -> STALE. Bugs found and fixed while testing, each caught by a test that bit: * the documented HEALTH_LOG_MAX_AGE_* override was never implemented; * a Gitea password held under GITEA_TOKEN was sent as an API token -> HTTP 401; now every candidate auth is tried and the first that works is used; * the ~/.git-credentials fallback filtered on GITEA_URL's host, which missed whenever GITEA_URL pointed at the internal IP -> zero candidates. Exit codes: 0 fresh, 1 stale-or-unreadable, 2 cannot run. A directory that cannot be read is a failure, never a skip. Evidence: live PASS; stale override names the directory and producer; no credential -> exit 2; whole thing runs green through contract-run.sh. prose-lint: PASSED. |
||
|
|
c5a6dcd42a |
fix(revision-preflight): default to warn, name the refusal class, bound the fetch
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 10s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 12s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Failing after 6s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Skipped
Bounded correction round on PR #134 after a PASS-WITH-FINDINGS review whose
Finding 4 is High. The guard's purpose and its fail-closed fix stand; the
problem was that with 'enforce' as the default it gates EVERY scheduled
contract, and three legitimate states produce a refusal - a clone legitimately
ahead of origin/master mid-review, a detached HEAD, and an offline or failed
fetch - so any of them would turn the fleet's monitoring into withheld
verdicts. That risk outweighs the staleness the guard catches.
1. DEFAULT IS NOW 'warn'. 'enforce' remains available and documented. The
criteria for flipping the default later are written into the doc as a
decision with evidence - a sustained window (30 days / 200+ runs) with zero
mismatch:* and zero cannot-verify:* refusals, no fetch blips, and a pinned
clone demonstrably kept current - explicitly as its own change, not a silent
flip.
2. 'COULD NOT CHECK' IS NOW DISTINGUISHABLE FROM 'THIS COPY IS WRONG'. Every
non-zero exit prints a machine-readable REASON=<class> line:
cannot-verify:fetch-failed | cannot-verify:ref-unresolvable (exit 2)
mismatch:path-absent | mismatch:content
mismatch:detached-head | mismatch:clone-ahead (exit 1)
detached-head and clone-ahead are named separately because they are
legitimate states, far less alarming than a hand-edited file. clone-ahead
requires HEAD to be STRICTLY ahead; an uncommitted edit on a commit that IS
the ref is a plain content mismatch (my own first cut got this wrong and the
new test 7d caught it).
3. THE DEFAULT FETCH IS BOUNDED: --fetch-timeout, default 20s, 0 = unbounded,
and a missing 'timeout' binary is itself a cannot-verify rather than an
unbounded fetch inside a scheduled contract.
4. TEST COVERAGE ADDED for every new class: fetch failure, fetch timeout
(asserted to return promptly under a 1s bound), unresolvable ref, detached
HEAD, clone-ahead, genuine content mismatch, and the contract-run.sh default.
The pre-fix draft fixture comparisons are kept: 31 passed, 0 failed.
5. MERGE-TIME SEQUENCE documented: fast-forward /opt/contract-runner, confirm
clean, prove a contract runs and reports. Baseline recorded as of today -
firstmate has already fast-forwarded it to
|
||
|
|
574cb99d76 |
feat: land the revision-preflight guard, fixed and wired into contract execution
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Failing after 6s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Skipped
A contract verdict is only meaningful if it came from the merged copy. The fleet has been bitten three times on 2026-09-25 (a clone parked on a merged feature branch while executing from another clone; a script copied into the runner clone by hand; a stale local origin/master making an ancestry check report unlanded work). The control for this existed as an untracked draft and protected nobody, because it was entirely fail-open. Defect in the draft, preserved verbatim as tests/fixtures/revision-preflight.prefix.sh: git -C "$CLONE" show "origin/master:$(basename "$SCRIPT")" basename drops the scripts/ prefix, so for any script under scripts/ it queried the repo root, failed, took the "warn but don't block" branch and exited 0 - passing a script that exists in no revision at all. Reproduced: pre-fix + scripts/demo.sh under scripts/ -> 'could not resolve', EXIT=0 pre-fix + a script in no revision -> EXIT=0 Fixed guard (scripts/revision-preflight.sh): * resolves the repo-relative path inside the clone, so scripts/ paths resolve; * FAILS CLOSED - a path absent from the ref, an unresolvable ref, or a failed fetch is a failure, never a warning; * fetches the remote by default, because a stale local ref would otherwise pass a stale script as current; --no-fetch states the assumption instead of hiding it. Wiring (scripts/contract-run.sh): before executing, the wrapper runs the guard against the clone it lives in. Default CONTRACT_REVISION_PREFLIGHT=enforce withholds the verdict, alerts and exits 2 on mismatch; =warn logs and continues; =off skips. Verified live: match -> contract proceeds and PASSes; mismatch -> 'VERDICT WITHHELD', exit 2; =warn -> continues. Pinning (docs/contract-execution-pinning.md): every contract pins the clone contract-run.sh lives in - the deployed runner being /opt/contract-runner on CT 100. Documented that daily-health-digest has no contract file at all, which is why its execution copy was silently operator-chosen. Tests: tests/test_revision_preflight.sh, 15 assertions over a throwaway clone with a real bare remote. It runs the pre-fix draft against the same cases and shows it passing a ghost script, so the tests provably bite. shellcheck: scripts/revision-preflight.sh and the new test are clean. The three findings remaining in contract-run.sh (SC2086 x2, SC2034) are pre-existing and byte-identical on master. |
||
|
|
040fecef3e |
feat: multi-engine search stack + visibility check
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 10s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s
Search stack (192.168.68.7) was effectively Bing-only: google served a JS
shell, duckduckgo CAPTCHA'd from the house egress, and every other shipped
engine returned a silent zero. Upgraded SearXNG to 2026.9.23 (same pinned
digest as the image already pulled by other hosts) which uses browser
impersonation, and routed DuckDuckGo through a VPS forward proxy over the
existing WireGuard tunnel via a per-engine 'network'.
Live result: bing, google cse, brave and yandex contribute on every query;
duckduckgo is best-effort via the datacenter egress.
Adds the visibility leg so a future regression cannot be silent:
scripts/search-stack-check.py
* two fixed queries; FAILS when fewer than two engines contribute,
printing contributing engines and every unresponsive_engines entry
* FAILS when Firecrawl extraction returns empty markdown or errors
* reports silent-zero engines explicitly
scripts/contract-run.sh
* maps search-stack-visibility -> search-stack-check.py
search-stack-visibility.prose.md
* contract text, execution model, pass/fail shapes, residual risk
Scheduled hourly at :15 on CT 100 via /etc/cron.d/contract-runner.
|
||
|
|
c0d04a2c02 |
fix: three wrapper holes in contract-run.sh
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 9s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 12s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 11s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
1. Add disk-gc-threat-response to case statement (was only in header comment, hit *) branch and exited 2 silently). Mapped to scripts/disk-gc-scan.py per contract's Execution section. 2. Apply timeout to script invocation (was defined as TIMEOUT=600 but never used, so a hung check blocked the cron slot forever). Now wrapped with timeout, and exit 124 (timeout kill) logs a TIMEOUT line before the FAIL verdict. 3. Send alert on exit-2 paths (unknown contract and missing script). Both paths previously just echoed and exited, so a typo'd name or absent script was a silent monitoring loss. Now they send the same Zulip DM as a failed check. Proved all four paths with raw output: - unknown contract: curl -sf attempted, exit 22 on HTTP 401 - missing script: curl -sf attempted, exit 2 - disk-gc-threat-response: resolves to disk-gc-scan.py, runs, PASS - stub sleep > timeout: TIMEOUT line logged, exit 1, alert failure recorded |
||
|
|
fda6c844ff |
fix: add -f to curl to fail on HTTP >= 400
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s
Without -f, a rejected credential (HTTP 401) returns curl exit 0, making a failed alert indistinguishable from a successful one. With -f, curl exits non-zero on HTTP >= 400, so DM_EXIT and STREAM_EXIT correctly capture the transmission failure and the run log records it. |
||
|
|
748ea389be |
fix: correct execution headings, variableize LOG_DIR, fix dead alert path
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 11s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 2s
1. Replace '(## Execution' + ')' with '## Execution' in 6 contract files 2. Make LOG_DIR honor CONTRACT_RUN_LOG_DIR env var (default: /var/log/contract-runs) 3. Fix Zulip alert: correct URL (https://chat.sysloggh.net/api/v1), user (abiba-bot@chat.sysloggh.net), take ZULIP_API_KEY from environment, and write alert failures to run log |
||
|
|
c666d3e15c |
feat: implement PBS GC four-state logic and update contracts
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 2s
- Implement four-state PBS GC logic in proxmox-monitor.sh: - probe-failed: unparseable JSON, store not found, or empty body → FAIL - running: collection in progress (last-run-endtime absent, upid present) → DO NOT FAIL - stale: no completed run within 48h → FAIL, naming last completed run age - healthy: completed within 48h → PASS, naming endtime and pending bytes - Add tests/test_pbs_gc_states.sh covering all four states - Proves the test bites on the pre-fix version (5/6 tests fail) - All 6 tests pass against the fixed version - Update contract-run.sh to map disk-gc-threat-response -> scripts/disk-gc-scan.py - Update Execution sections of host-scheduled contracts: infrastructure-monitoring, zulip-health, litellm-health, agent-health-check, disk-gc-threat-response, pm2-self-heal Adding note that execution is host-scheduled via cron, not agent session ack. |
||
|
|
c07aa5e382 |
Add contract-run.sh for machine scheduler execution and fix PBS GC leg
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 15s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s
- Add scripts/contract-run.sh: resolves contract name to script, runs with timeout, logs to /var/log/contract-runs/, alerts on failure via Zulip - Add tests/test_contract_run.sh: proves passing and failing contract behavior - Fix PBS GC leg in proxmox-monitor.sh: simplify logic to check last-run-endtime, use absolute paths for pct and proxmox-backup-manager to avoid PATH issues Part of Task: contract-execution-host-scheduler-20260924 |