226f2ad55e1bf8485a10cba9c71632a841203c5b
526
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
226f2ad55e |
fix: reconcile the health-logs self-heal contracts, and add a dead-man's-switch
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 9s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 10s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s
relay-785 / self-heal-contract-stale-script-pointers-20260924. gpu/ had been silent since 2026-09-14T18:02:03Z (12 days) and pm2/ had never posted a run. FINDING - the gpu-self-heal executor was never lost, only its schedule was. The brief concluded the mechanism was gone. It is not: on CT 116 /opt/inference-harness/scripts/gpu-self-heal.py exists (mtime 2026-09-21 11:43), is self-posting via gitea-logger.sh, and runs cleanly - I ran it and it posted health-logs/gpu/gpu-self-heal-20260926-144618.json. What was missing was the cron entry, dropped during a CT 116 /etc/cron.d rework on 2026-09-21. So this is a schedule restoration, not a resurrection. DECISIONS * gpu-self-heal -> RESTORE. Schedule re-added as /etc/cron.d/gpu-self-heal on CT 116, matching the observed historical cadence (:02 past 0/6/12/18). Proven with a REAL unattended cron fire (temporary */1 entry, removed after): gpu-self-heal-20260926-144802.json committed 14:48:03Z. The contract now names the executor, schedule, log and posting, which it previously did not. * pm2 health-logs posting -> RETIRE the claim. pm2-self-heal.prose.md required appending to health-logs/pm2/ as a 'hard rule', but scripts/pm2-self-heal.sh contains no Gitea or push code and never did; the directory has held only its init commit since 2026-07-28. Corrected to point at the contract-runner's durable per-run logs and failure note instead of adding a second, redundant posting path. DEAD-MAN'S-SWITCH - scripts/health-log-freshness.py + contract. Absence of logs must raise an alarm, and the producer cannot raise it, so this runs on CT 100, a different host from the producers, and fails when the newest health-logs/gpu entry is older than 12h (litellm 18h). Verified it would have caught the real gap: evaluated at 2026-09-20 the newest gpu/ entry was 126h old against a 12h limit -> STALE. Bugs found and fixed while testing, each caught by a test that bit: * the documented HEALTH_LOG_MAX_AGE_* override was never implemented; * a Gitea password held under GITEA_TOKEN was sent as an API token -> HTTP 401; now every candidate auth is tried and the first that works is used; * the ~/.git-credentials fallback filtered on GITEA_URL's host, which missed whenever GITEA_URL pointed at the internal IP -> zero candidates. Exit codes: 0 fresh, 1 stale-or-unreadable, 2 cannot run. A directory that cannot be read is a failure, never a skip. Evidence: live PASS; stale override names the directory and producer; no credential -> exit 2; whole thing runs green through contract-run.sh. prose-lint: PASSED. |
||
|
|
552c776c0c |
Merge pull request 'feat: land the revision-preflight guard, fixed and wired into contract execution' (#134) from fix/land-revision-preflight-guard-20260925 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 10s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 0s
|
||
|
|
778424acd4 |
Merge remote-tracking branch 'origin/master' into fix/land-revision-preflight-guard-20260925
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 8s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 9s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 4s
|
||
|
|
c5a6dcd42a |
fix(revision-preflight): default to warn, name the refusal class, bound the fetch
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 10s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 12s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Failing after 6s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Skipped
Bounded correction round on PR #134 after a PASS-WITH-FINDINGS review whose
Finding 4 is High. The guard's purpose and its fail-closed fix stand; the
problem was that with 'enforce' as the default it gates EVERY scheduled
contract, and three legitimate states produce a refusal - a clone legitimately
ahead of origin/master mid-review, a detached HEAD, and an offline or failed
fetch - so any of them would turn the fleet's monitoring into withheld
verdicts. That risk outweighs the staleness the guard catches.
1. DEFAULT IS NOW 'warn'. 'enforce' remains available and documented. The
criteria for flipping the default later are written into the doc as a
decision with evidence - a sustained window (30 days / 200+ runs) with zero
mismatch:* and zero cannot-verify:* refusals, no fetch blips, and a pinned
clone demonstrably kept current - explicitly as its own change, not a silent
flip.
2. 'COULD NOT CHECK' IS NOW DISTINGUISHABLE FROM 'THIS COPY IS WRONG'. Every
non-zero exit prints a machine-readable REASON=<class> line:
cannot-verify:fetch-failed | cannot-verify:ref-unresolvable (exit 2)
mismatch:path-absent | mismatch:content
mismatch:detached-head | mismatch:clone-ahead (exit 1)
detached-head and clone-ahead are named separately because they are
legitimate states, far less alarming than a hand-edited file. clone-ahead
requires HEAD to be STRICTLY ahead; an uncommitted edit on a commit that IS
the ref is a plain content mismatch (my own first cut got this wrong and the
new test 7d caught it).
3. THE DEFAULT FETCH IS BOUNDED: --fetch-timeout, default 20s, 0 = unbounded,
and a missing 'timeout' binary is itself a cannot-verify rather than an
unbounded fetch inside a scheduled contract.
4. TEST COVERAGE ADDED for every new class: fetch failure, fetch timeout
(asserted to return promptly under a 1s bound), unresolvable ref, detached
HEAD, clone-ahead, genuine content mismatch, and the contract-run.sh default.
The pre-fix draft fixture comparisons are kept: 31 passed, 0 failed.
5. MERGE-TIME SEQUENCE documented: fast-forward /opt/contract-runner, confirm
clean, prove a contract runs and reports. Baseline recorded as of today -
firstmate has already fast-forwarded it to
|
||
|
|
9faffe4f6b |
Merge pull request 'feat: add the missing daily-health-digest contract + fix the red lint from #133' (#135) from fix/daily-health-digest-contract-20260925 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 12s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 0s
|
||
|
|
cd00cd0475 |
feat: add the missing daily-health-digest contract and register it
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 13s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
Backlog row daily-health-digest-contract-missing-20260925. The digest has been dispatched on a schedule with NO contract file at all - no *daily*.prose.md, absent from contract-registry.yaml, the only reference anywhere being the CT100 cron line. That absence is why the choice of execution copy was silently the operator's, and how a stale clone could run the check unnoticed. New daily-health-digest.prose.md states: * the PINNED execution path /root/abiba-workspace/projects/prose-contracts/ scripts/daily-infra-report.py and the pinned clone - the cron's FM_HOME clone, the only stable non-ephemeral copy; the treehouse clone is a per-agent working copy and must NOT be pinned; * the output shape (all 18 top-level --json keys) and what a healthy run is; * the exit-code semantics AS THEY ACTUALLY BEHAVE, verified case by case: missing PVE_TOKEN or an unreachable probe exits 1 and raises an alert, while a missing EMAIL credential is a deliberate DEGRADED leg that still exits 0 and still produces the report. PROBE_FAILURES and DEGRADED_LEGS are separate lists on purpose and must not be merged; * the email-delivery dependency, that EMAIL_PASSWORD must be a Google app password, that it is failing with 534 5.7.9 as of 2026-09-25, and that a delivery failure is a credential dependency rather than a code defect; * what counts as a failure versus degraded. Registered in contract-registry.yaml (contracts entry plus index.by_category .monitoring and index.by_domain.infrastructure). Verified: YAML parses, 31 contracts, exactly one daily-health-digest entry. |
||
|
|
819cd53bce |
fix(lint): restore the repo secret scan, which PR #133 broke
Master's lint is RED right now, and it is my doing. at |
||
|
|
574cb99d76 |
feat: land the revision-preflight guard, fixed and wired into contract execution
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Failing after 6s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Skipped
A contract verdict is only meaningful if it came from the merged copy. The fleet has been bitten three times on 2026-09-25 (a clone parked on a merged feature branch while executing from another clone; a script copied into the runner clone by hand; a stale local origin/master making an ancestry check report unlanded work). The control for this existed as an untracked draft and protected nobody, because it was entirely fail-open. Defect in the draft, preserved verbatim as tests/fixtures/revision-preflight.prefix.sh: git -C "$CLONE" show "origin/master:$(basename "$SCRIPT")" basename drops the scripts/ prefix, so for any script under scripts/ it queried the repo root, failed, took the "warn but don't block" branch and exited 0 - passing a script that exists in no revision at all. Reproduced: pre-fix + scripts/demo.sh under scripts/ -> 'could not resolve', EXIT=0 pre-fix + a script in no revision -> EXIT=0 Fixed guard (scripts/revision-preflight.sh): * resolves the repo-relative path inside the clone, so scripts/ paths resolve; * FAILS CLOSED - a path absent from the ref, an unresolvable ref, or a failed fetch is a failure, never a warning; * fetches the remote by default, because a stale local ref would otherwise pass a stale script as current; --no-fetch states the assumption instead of hiding it. Wiring (scripts/contract-run.sh): before executing, the wrapper runs the guard against the clone it lives in. Default CONTRACT_REVISION_PREFLIGHT=enforce withholds the verdict, alerts and exits 2 on mismatch; =warn logs and continues; =off skips. Verified live: match -> contract proceeds and PASSes; mismatch -> 'VERDICT WITHHELD', exit 2; =warn -> continues. Pinning (docs/contract-execution-pinning.md): every contract pins the clone contract-run.sh lives in - the deployed runner being /opt/contract-runner on CT 100. Documented that daily-health-digest has no contract file at all, which is why its execution copy was silently operator-chosen. Tests: tests/test_revision_preflight.sh, 15 assertions over a throwaway clone with a real bare remote. It runs the pre-fix draft against the same cases and shows it passing a ghost script, so the tests provably bite. shellcheck: scripts/revision-preflight.sh and the new test are clean. The three findings remaining in contract-run.sh (SC2086 x2, SC2034) are pre-existing and byte-identical on master. |
||
|
|
9d64b0bd66 |
Merge pull request 'fix(daily-infra-report): resolve PVE token from env; make a dead probe non-silent' (#133) from fix/daily-health-digest-pve-token-20260925 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Failing after 7s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Skipped
|
||
|
|
cdc7ad2c79 |
fix(daily-infra-report): resolve PVE token from env; make a dead probe non-silent
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Failing after 10s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 9s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Skipped
The Proxmox leg of the daily digest has been reporting NOTHING while exiting 0.
Root cause: the auth header was a literal placeholder string,
AUTH = "Authorization: PVEAPIToken=«vault: infrastructure/production PVE_API_TOKEN»"
which was sent verbatim. The API rejected it, pve_get() returned None, and the
report rendered node_count=0 / nodes_online=0 with pve_probe_status='unreachable'
while still exiting 0. A monitoring gap that looks like a healthy run.
Also: the vault key is PVE_TOKEN, not PVE_API_TOKEN, so even reading os.environ
by the old name would not have found it.
Fixes:
* pve_auth() resolves the token at call time from PVE_TOKEN (injected by
'infisical run --env=prod'). Nothing is hardcoded; a missing token raises.
* pve_get() builds the command inside its try block, so a missing token degrades
to None instead of escaping as an unhandled exception.
* PROBE_FAILURES records an unreachable node/resources probe. Probe failures are
deliberately separate from DEGRADED_LEGS: a missing credential stays exit 0
(existing intent), but a probe with no data now exits 1 in both the report and
--json paths, so it cannot pass unnoticed.
Measured effect on the live host: pve_probe_status unreachable -> ok,
node_count 0 -> 5, nodes_online 0 -> 5, total_vms 0 -> 22, running_vms 0 -> 22.
Tests: 4 new regression tests; all 4 fail against the pre-fix script and pass
after, and the 3 pre-existing tests still pass (7/7).
Not fixed here (needs the captain): the email leg fails with
'534 5.7.9 Application-specific password required' - EMAIL_PASSWORD in the vault
is not a valid Gmail app password for jtabiri@gmail.com. That is a credential
action, not a code change.
|
||
|
|
5409dfd73a |
Merge pull request 'feat: multi-engine search stack + visibility check' (#132) from fix/search-stack-multi-engine-20260925 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 13s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 1s
|
||
|
|
8b2eba4f7a |
docs+check: correct DDG status, explain silent-zero semantics, credit google cse
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 11s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 8s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
Follow-up corrections after review: 1. DuckDuckGo is NOT fixed. The VPS fallback egress has since been flagged by DuckDuckGo too (HTTP 202 + challenge markers), so it reports CAPTCHA on both paths. The contract and script docstring now say so instead of claiming a fix that had already expired. It stays enabled as best-effort coverage so a recovery shows up as a contribution. 2. The relationship between silent zeros and the verdict is now explicit in both the script output and the contract: an enabled expected engine that contributes zero with no error is REPORTED, not fatal. Only the <SEARCH_CHECK_MIN_ENGINES> floor and the extraction leg fail the run. This is deliberate - de-duplication and query-shape make a zero non-probative. 3. Recorded that 'google cse' uses a THIRD PARTY's public search-engine id hardcoded in the SearXNG build, not a key we own; its quota and availability are outside our control, and our own free key would need a wrapper (not built). VPS forward proxy is now a real service: /opt/fwd-proxy docker compose with restart: unless-stopped, a healthy healthcheck, and Docker enabled at boot. |
||
|
|
040fecef3e |
feat: multi-engine search stack + visibility check
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 10s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s
Search stack (192.168.68.7) was effectively Bing-only: google served a JS
shell, duckduckgo CAPTCHA'd from the house egress, and every other shipped
engine returned a silent zero. Upgraded SearXNG to 2026.9.23 (same pinned
digest as the image already pulled by other hosts) which uses browser
impersonation, and routed DuckDuckGo through a VPS forward proxy over the
existing WireGuard tunnel via a per-engine 'network'.
Live result: bing, google cse, brave and yandex contribute on every query;
duckduckgo is best-effort via the datacenter egress.
Adds the visibility leg so a future regression cannot be silent:
scripts/search-stack-check.py
* two fixed queries; FAILS when fewer than two engines contribute,
printing contributing engines and every unresponsive_engines entry
* FAILS when Firecrawl extraction returns empty markdown or errors
* reports silent-zero engines explicitly
scripts/contract-run.sh
* maps search-stack-visibility -> search-stack-check.py
search-stack-visibility.prose.md
* contract text, execution model, pass/fail shapes, residual risk
Scheduled hourly at :15 on CT 100 via /etc/cron.d/contract-runner.
|
||
|
|
483a66b7b4 |
Merge pull request 'Fix PBS GC monitor false positive and add contract-run.sh wrapper' (#131) from fix/contract-run-pbs-gc-20260924 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 12s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 13s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 9s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 4s
|
||
|
|
c0d04a2c02 |
fix: three wrapper holes in contract-run.sh
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 9s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 12s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 11s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
1. Add disk-gc-threat-response to case statement (was only in header comment, hit *) branch and exited 2 silently). Mapped to scripts/disk-gc-scan.py per contract's Execution section. 2. Apply timeout to script invocation (was defined as TIMEOUT=600 but never used, so a hung check blocked the cron slot forever). Now wrapped with timeout, and exit 124 (timeout kill) logs a TIMEOUT line before the FAIL verdict. 3. Send alert on exit-2 paths (unknown contract and missing script). Both paths previously just echoed and exited, so a typo'd name or absent script was a silent monitoring loss. Now they send the same Zulip DM as a failed check. Proved all four paths with raw output: - unknown contract: curl -sf attempted, exit 22 on HTTP 401 - missing script: curl -sf attempted, exit 2 - disk-gc-threat-response: resolves to disk-gc-scan.py, runs, PASS - stub sleep > timeout: TIMEOUT line logged, exit 1, alert failure recorded |
||
|
|
fda6c844ff |
fix: add -f to curl to fail on HTTP >= 400
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s
Without -f, a rejected credential (HTTP 401) returns curl exit 0, making a failed alert indistinguishable from a successful one. With -f, curl exits non-zero on HTTP >= 400, so DM_EXIT and STREAM_EXIT correctly capture the transmission failure and the run log records it. |
||
|
|
748ea389be |
fix: correct execution headings, variableize LOG_DIR, fix dead alert path
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 11s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 2s
1. Replace '(## Execution' + ')' with '## Execution' in 6 contract files 2. Make LOG_DIR honor CONTRACT_RUN_LOG_DIR env var (default: /var/log/contract-runs) 3. Fix Zulip alert: correct URL (https://chat.sysloggh.net/api/v1), user (abiba-bot@chat.sysloggh.net), take ZULIP_API_KEY from environment, and write alert failures to run log |
||
|
|
f16a890d0e |
fix: make test_pbs_gc_states.sh self-contained with inline SSH replacement
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 8s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 2s
|
||
|
|
c666d3e15c |
feat: implement PBS GC four-state logic and update contracts
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 2s
- Implement four-state PBS GC logic in proxmox-monitor.sh: - probe-failed: unparseable JSON, store not found, or empty body → FAIL - running: collection in progress (last-run-endtime absent, upid present) → DO NOT FAIL - stale: no completed run within 48h → FAIL, naming last completed run age - healthy: completed within 48h → PASS, naming endtime and pending bytes - Add tests/test_pbs_gc_states.sh covering all four states - Proves the test bites on the pre-fix version (5/6 tests fail) - All 6 tests pass against the fixed version - Update contract-run.sh to map disk-gc-threat-response -> scripts/disk-gc-scan.py - Update Execution sections of host-scheduled contracts: infrastructure-monitoring, zulip-health, litellm-health, agent-health-check, disk-gc-threat-response, pm2-self-heal Adding note that execution is host-scheduled via cron, not agent session ack. |
||
|
|
c07aa5e382 |
Add contract-run.sh for machine scheduler execution and fix PBS GC leg
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 15s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s
- Add scripts/contract-run.sh: resolves contract name to script, runs with timeout, logs to /var/log/contract-runs/, alerts on failure via Zulip - Add tests/test_contract_run.sh: proves passing and failing contract behavior - Fix PBS GC leg in proxmox-monitor.sh: simplify logic to check last-run-endtime, use absolute paths for pct and proxmox-backup-manager to avoid PATH issues Part of Task: contract-execution-host-scheduler-20260924 |
||
|
|
87205e6fdb |
Merge pull request 'feat(security): commit-time secret guard that FAILS the build on a committed credential' (#129) from fm/commit-time-secret-guard-20260917 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 13s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 1s
|
||
|
|
8f1e5eebc4 |
feat(security): commit-time secret guard that FAILS the build on a committed credential
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 18s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
The 2026-09-17 purge removed six live credentials that had sat in this repo for weeks, several in .md prose. Nothing blocked that class of commit, so a warning in a stream nobody reads was the only signal. This adds a guard that fails the build instead of warning. Guard - scripts/secret-scan.sh: bash + coreutils + grep/sed/awk + git only (the Gitea Actions runner executes job steps inside the runner container — BusyBox grep, no node/python). Modes: --tree (git-tracked, default), --path DIR (no git), --staged (pre-commit), --diff REF. Exit 1 on a finding, 2 on config error. - scripts/secret-patterns.tsv: checked-in pattern list — sk-, sk-or-v1-, sk_live_, literal Bearer tokens, PVEAPIToken=, raw Authorization values, PEM private-key blocks, prose credential lines, and password/api_key/secret/token assignments carrying a literal value. Prose is scanned exactly like code. - scripts/secret-allowlist.tsv: one entry per deliberate synthetic example, each with a reason. A missing reason is a hard error (fail closed). The 2026-09-17 purge's `«vault: ...»` placeholders are listed explicitly rather than filtered by a general "vault"/"synthetic" rule, so a new occurrence still needs a reviewed, reasoned entry. - A small inert-value classifier drops env refs, paths, dotted code access, variable names and right-truncated redactions; it does not know the words "synthetic"/"example", so a fabrication is always an explicit exception. - Findings are printed with the credential masked; a scan never echoes a full secret into the log. Wiring - .gitea/workflows/pr-pipeline.yaml lint job: explicit "Committed-credential scan" step plus the self-test. A finding fails the required `pr-pipeline / lint` context, which the merge gate depends on. - scripts/prose-lint.sh (the local gate): a "Secret scan" section, so `bash scripts/prose-lint.sh` before pushing is equivalent to CI. Tests - tests/test_secret_scan.sh: 20 cases. Plants pattern-matching fixtures in temp trees (outside every allowlisted path) and asserts the guard FAILS, including the --staged commit-time path; asserts the tree is quiet; asserts allowlisted text at an unlisted path still fails (path-explicit, not word-based); asserts a reasonless allowlist entry exits 2. Verified: guard run against 8245716^ (the pre-fix revision, before the purge) fails on the real OpenRouter/LiteLLM/Zulip/Proxmox/Stirling credentials; guard run over the current tree is clean. |
||
|
|
bb1b65340e |
Merge pull request #128 from decisions-2026-08-03-rebased
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 8s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 11s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 1s
|
||
|
|
9e87927444 |
pm2-self-heal: reconcile contract with live PM2 set and alert channels
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
- Add zulip-watchdog to Maintains (it's running, infrastructure-monitoring expects it) - Remove gpu-monitor from PM2 Maintains (it's systemd-only, not PM2-tracked) - Add Execution steps for all monitored processes (gitea-runner, zulip-watchdog) - Update alert channel: Telegram is primary, Zulip DM is secondary - Update script to check all 4 processes (abiba-telegram, abiba-zulip, gitea-runner, zulip-watchdog) - Add restart count thresholds for all processes - Update log output to include all process statuses |
||
|
|
b0683e9566 | fix: correct logging destination in description (Gitea health-logs, not knowledge graph) | ||
|
|
dd11c8f14f | Apply captain 2026-08-03 decisions: restore abiba-zulip, retire gpu-watchdog, PM2-track gpu-monitor | ||
|
|
0732eed329 |
Merge pull request 'fix(daily-infra-report): drop the vestigial Zulip key requirement; fail loudly on a failed send' (#127) from fix/daily-health-digest-remove-vestigial-zulip-20260921 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 1s
|
||
|
|
59ed7cdbf7 |
fix(daily-infra-report): remove vestigial ZULIP_API_KEY requirement, exit non-zero on failed send
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s
1. Remove vestigial ZULIP_API_KEY requirement:
- /api/v1/server_settings is a PUBLIC endpoint (verified HTTP 200 with or without credential)
- No Zulip API key is required for this call
- If a future leg genuinely needs abiba-bot's key, it must prove it with a 200 from
/api/v1/users/me as abiba-bot and label itself degraded when it cannot
- Never fall back to the vault's shared ZULIP_API_KEY
2. Make failed sends exit non-zero:
- A degraded leg (no credential configured) must stay exit 0
- A failed send (attempted and failed) must exit 1
- This distinguishes 'not configured' from 'attempted and failed'
Test evidence:
- No-credential run: exit 0, digest still produced
- Wrong password: exit 1, labelled SMTP error
- grep -n ZULIP_API_KEY: only comment reference remains
|
||
|
|
5d70bbf25b |
Merge pull request 'fix(daily-infra-report): missing credentials degrade one leg instead of blacking out the digest' (#126) from fix/daily-health-digest-degraded-credentials-20260921 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 2s
|
||
|
|
ccc916d1ec |
feat(daily-infra-report): make missing credentials a labelled degraded leg
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
- ZULIP_API_KEY: no longer SystemExit, now reports 'credential-missing: ZULIP_API_KEY' - EMAIL_PASSWORD: no longer sys.exit(1), now appends to DEGRADED_LEGS and returns success - PVE API: fixed None check in storage section - Summary: reports degraded legs before summary This allows the digest to be produced and emailed even when credentials are missing, while still explicitly logging which legs are degraded. Test: empty env produces JSON report + degraded leg labels, no SystemExit. |
||
|
|
48fb263d4b |
Merge pull request 'feat(litellm): align contracts for 7-provider cloud consolidation (2026-09-20)' (#125) from feat/align-litellm-contracts-cloud-20260920 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 15s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 2s
|
||
|
|
64790ebd19 |
feat(litellm): align contracts for 7-provider cloud consolidation (2026-09-20)
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s
- hermes-key-enforcement: add Model Access Tiers (local vs cloud) and cloud-scoping clause
- litellm-api-keys: add Cloud Provider Consolidation section (provider map, vault secrets, access tiers)
- litellm-api-keys: pin standard agent keys to explicit local-only models list; forbid {}/all-proxy-models (Community-edition cloud leak)
- contract-registry: register litellm-api-keys (was unregistered drift)
Additive only. Local prose-lint: PASSED.
|
||
|
|
6fa0a255df |
Merge pull request 'fix(zulip-monitor): add C3 public access path, make run verdict non-optimistic, separate C1/C2/C3' (#124) from fm/kagentz-a2a-outage-masked-as-expected-20260919 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 8s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 2s
|
||
|
|
c72436b406 |
no-mistakes(document): docs: reconcile zulip-monitor status in infrastructure-control
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
|
||
|
|
a17379676d | no-mistakes(document): docs: sync zulip-health contract with C1/C3 and version | ||
|
|
63990b84f7 | no-mistakes(review): Fix review findings: test regression, contract verdict, scope trim | ||
|
|
7a5ddb46a9 | chore: Add local monitoring tooling scripts | ||
|
|
568fec2efa |
test(zulip-kagentz): Replace string-presence tests with behavioural sandbox tests
- C3 502 → INCIDENT - C3 000 → INCIDENT - C1 401 + C3 302 → 0 issues, all healthy - C1 000 → INCIDENT Each test asserts from the run's own log/verdict, not from file text. Prose assertions kept as secondary. Proven to bite: run against pre-fix script (origin/master) shows all 4 behavioural cases fail because C3 leg doesn't exist and Result line doesn't say 'INCIDENT'. |
||
|
|
641a52c6da |
fix(zulip-monitor): Add C3 public access path, make Result line non-optimistic
- (b) Changed Result line to say 'INCIDENT' when ISSUES > 0, '0 issues (all healthy)' when ISSUES = 0 - (c) Documented C1 (no credential needed), C2 (requires LITELLM_KEY) distinction - (d) Added C3 public access path leg for https://kagentz.sysloggh.net/ - C3 treats 200/302/401 as alive, 502/000 as incident - Added tests/test_zulip_kagentz_legs.py to verify all changes |
||
|
|
58033f39c5 |
Merge pull request 'fix: Rule 5 accept canonical internal and public host base_url' (#122) from fix/rule5-canonical-baseurl-20260919 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 8s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 0s
|
||
|
|
6c59988e7f |
Merge pull request 'fix(litellm-health): three-state model verdicts - busy is not a failure' (#123) from fix/litellm-health-busy-vs-down-20260919 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 1s
|
||
|
|
9a2ee6faec |
fix(litellm): Remove duplicate 'host healthy' from busy line
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
The busy line was rendering as: 'busy (completion timed out after retry; host healthy host healthy (200))' because host_detail already contains 'host healthy (200)' and the prefix also said 'host healthy'. Fixed to: 'busy (completion timed out after retry; host healthy (200))' F1 cosmetic fix from PR #123 verify. |
||
|
|
e71ded3c8c |
fix: internal /v1 WARN not FAIL; align prose
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 3s
Rule 5 now treats internal http://192.168.68.116/v1 as non-canonical but working (authenticated via nginx), producing a WARNING instead of a FAILURE. The canonical internal path /litellm/v1 and the public host https://litellm.sysloggh.net/v1 both PASS. Everything else FAILS. Prose aligned: hermes-key-enforcement.prose.md now states the canonical internal form, notes that internal /v1 still works but is flagged as non-canonical (WARN not FAIL), and clarifies that the public host serves /v1 ONLY (404 on /litellm/v1). Corrected the 'unauthenticated path' wording at line 116, which was factually wrong. Tests updated: BASE template uses canonical internal path; new test cases prove canonical /litellm/v1 PASSES, wrong path FAILS, public host PASSES, and internal /v1 WARNS (not FAILS). Fixed backwards comment in test_old_rule5_check_would_fail_canonical. |
||
|
|
a12abbeb14 |
fix(litellm): Implement busy vs down with proper degraded state
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 8s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 8s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
Three states: - healthy: passed, exit 0 (unchanged) - busy (completion timed out after retry AND host /health answered): ⚠️ DEGRADED line, does NOT fail the run, exit 0 - host unreachable or real fault: ❌, exit 1 (unchanged) Summary now reports degraded count: - All pass, no degraded: '✅ All checks passed' - All pass, 1+ degraded: '✅ All checks passed (1 degraded: gpu-dense)' - Some failed: '❌ Some checks failed' or '❌ Some checks failed (1 degraded: ...)' Host health mapping verified: - gpu-dense -> 192.168.68.8:8080/health - gpu-vision -> 192.168.68.110:8080/health - strix-moe -> 192.168.68.15:8080/health |
||
|
|
0b9aebca37 |
fix(litellm): Fix timeout kind reporting + add busy/degraded detection
1. TIMEOUT KIND FIX: run_command returns (1, '', 'TIMEOUT') when its own timeout fires. probe_http now checks for this before falling through to 'curl exit <rc>', so a 30s timeout reports 'timeout after 30s' not 'curl exit 1'. 2. BUSY/DEGRADED DETECTION: After both model probes fail, check the model's host health endpoint (e.g. 192.168.68.8:8080/health for gpu-dense). If the host answers 200, report 'busy (completion timed out after retry; host healthy 200)' — do NOT fail the run on that alone. If the host does not answer, that's a real FAIL. 3. RETRY TIMEOUT INCREASED: Single-host retry timeout raised from 45s to 90s. Worst-case prefill on a single-slot .8 host is ~76s (observed 83K-token prompt at 1078 tok/s), so 90s covers it. New line shapes: - Busy: 'gpu-dense: busy (completion timed out after retry; host healthy 200)' - Real failure: 'probe-failed: gpu-dense timeout after 30s then timeout after 90s (2 attempts)' |
||
|
|
fa458afa26 |
fix: Rule 5 accept canonical internal and public host base_url
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
Rule 5 in audit-hermes-config.py had an inverted check: it expected base_url=http://192.168.68.116/v1, but the contract hermes-key-enforcement.prose.md names http://192.168.68.116/litellm/v1 as CORRECT/CANONICAL in multiple places. The audit script would FAIL a config using the contract's canonical internal path and PASS one using a path the contract does not name. Fix: Rule 5 now accepts the canonical internal base (http://192.168.68.116/litellm/v1) AND the public base (https://litellm.sysloggh.net/v1), and FAILS anything else. The internal nginx serves both /litellm/v1 and /v1; the public host serves /v1 only (per 2026-09-19 probe from CT 116). Tests aligned: BASE template updated to use the canonical internal path, and new test cases added to prove the canonical internal path PASSES, a wrong path FAILS, and the public host path PASSES. |
||
|
|
aa3da83af5 |
Merge pull request 'fix(litellm-health): retry single-host model probes once at a longer timeout' (#121) from fix/litellm-health-retry-timeout-20260919 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 9s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 0s
|
||
|
|
aee2de25ac |
fix(litellm): Report both attempts' failure kinds in probe-failed
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 4s
The failure line now preserves both attempts' failure kinds instead of hardcoding 'timeout after retry, 45s'. If both attempts fail, the report shows: 'probe-failed: <model> <first kind> then <retry kind> (2 attempts)'. This fixes the self-contradictory output when the first attempt timed out but the retry failed with connection refused, and prevents the duration from appearing twice when both attempts were timeouts. Example outputs: - timeout then timeout: 'probe-failed: gpu-dense timeout after 30s then timeout after 45s (2 attempts)' - timeout then refused: 'probe-failed: gpu-dense timeout after 30s then connection refused (2 attempts)' - refused then refused: 'probe-failed: gpu-dense connection refused then connection refused (2 attempts)' |
||
|
|
76653381ec |
fix(litellm): Add retry with longer timeout for single-host model probes
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 11s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s
Single-host models (gpu-dense, gpu-vision, strix-moe) now retry once at 45s on initial 30s timeout failure before declaring probe-failed. This prevents a single transient timeout (cold prefill ~13s or concurrent generation hold) from failing the entire health digest. Evidence: 2026-09-19 ~06:55Z digest failed gpu-dense at 30s; 06:56Z direct probe 200 in 1.04s. The failed-probe-fails-the-run property is preserved: if both attempts fail, the script still exits non-zero with the target and duration named. Closes: daily-health-digest false negative on single transient timeout |
||
|
|
3b32cc9658 |
Merge pull request 'fix(infra-monitoring): probe the real Docker Stats and PVE exporter ports' (#120) from fix/infra-monitoring-probe-ports-20260919 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 3s
|