Search stack (192.168.68.7) was effectively Bing-only: google served a JS
shell, duckduckgo CAPTCHA'd from the house egress, and every other shipped
engine returned a silent zero. Upgraded SearXNG to 2026.9.23 (same pinned
digest as the image already pulled by other hosts) which uses browser
impersonation, and routed DuckDuckGo through a VPS forward proxy over the
existing WireGuard tunnel via a per-engine 'network'.
Live result: bing, google cse, brave and yandex contribute on every query;
duckduckgo is best-effort via the datacenter egress.
Adds the visibility leg so a future regression cannot be silent:
scripts/search-stack-check.py
* two fixed queries; FAILS when fewer than two engines contribute,
printing contributing engines and every unresponsive_engines entry
* FAILS when Firecrawl extraction returns empty markdown or errors
* reports silent-zero engines explicitly
scripts/contract-run.sh
* maps search-stack-visibility -> search-stack-check.py
search-stack-visibility.prose.md
* contract text, execution model, pass/fail shapes, residual risk
Scheduled hourly at :15 on CT 100 via /etc/cron.d/contract-runner.
1. Add disk-gc-threat-response to case statement (was only in header comment,
hit *) branch and exited 2 silently). Mapped to scripts/disk-gc-scan.py per
contract's Execution section.
2. Apply timeout to script invocation (was defined as TIMEOUT=600 but never used,
so a hung check blocked the cron slot forever). Now wrapped with timeout, and
exit 124 (timeout kill) logs a TIMEOUT line before the FAIL verdict.
3. Send alert on exit-2 paths (unknown contract and missing script). Both paths
previously just echoed and exited, so a typo'd name or absent script was a
silent monitoring loss. Now they send the same Zulip DM as a failed check.
Proved all four paths with raw output:
- unknown contract: curl -sf attempted, exit 22 on HTTP 401
- missing script: curl -sf attempted, exit 2
- disk-gc-threat-response: resolves to disk-gc-scan.py, runs, PASS
- stub sleep > timeout: TIMEOUT line logged, exit 1, alert failure recorded
Without -f, a rejected credential (HTTP 401) returns curl exit 0, making
a failed alert indistinguishable from a successful one. With -f, curl
exits non-zero on HTTP >= 400, so DM_EXIT and STREAM_EXIT correctly
capture the transmission failure and the run log records it.
1. Replace '(## Execution' + ')' with '## Execution' in 6 contract files
2. Make LOG_DIR honor CONTRACT_RUN_LOG_DIR env var (default: /var/log/contract-runs)
3. Fix Zulip alert: correct URL (https://chat.sysloggh.net/api/v1), user
(abiba-bot@chat.sysloggh.net), take ZULIP_API_KEY from environment, and
write alert failures to run log
- Implement four-state PBS GC logic in proxmox-monitor.sh:
- probe-failed: unparseable JSON, store not found, or empty body → FAIL
- running: collection in progress (last-run-endtime absent, upid present) → DO NOT FAIL
- stale: no completed run within 48h → FAIL, naming last completed run age
- healthy: completed within 48h → PASS, naming endtime and pending bytes
- Add tests/test_pbs_gc_states.sh covering all four states
- Proves the test bites on the pre-fix version (5/6 tests fail)
- All 6 tests pass against the fixed version
- Update contract-run.sh to map disk-gc-threat-response -> scripts/disk-gc-scan.py
- Update Execution sections of host-scheduled contracts:
infrastructure-monitoring, zulip-health, litellm-health, agent-health-check,
disk-gc-threat-response, pm2-self-heal
Adding note that execution is host-scheduled via cron, not agent session ack.
- Add scripts/contract-run.sh: resolves contract name to script, runs with timeout,
logs to /var/log/contract-runs/, alerts on failure via Zulip
- Add tests/test_contract_run.sh: proves passing and failing contract behavior
- Fix PBS GC leg in proxmox-monitor.sh: simplify logic to check last-run-endtime,
use absolute paths for pct and proxmox-backup-manager to avoid PATH issues
Part of Task: contract-execution-host-scheduler-20260924