Commit Graph
527 Commits
Author SHA1 Message Date
abiba-bot 042fd3ccc6 Merge pull request 'fix: reconcile the health-logs self-heal contracts, and add a dead-man's-switch' (#136) from fix/health-log-freshness-watchdog-20260926 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 11s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 1s
2026-09-26 15:16:36 +00:00
root 226f2ad55e fix: reconcile the health-logs self-heal contracts, and add a dead-man's-switch
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 9s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 10s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s
relay-785 / self-heal-contract-stale-script-pointers-20260924. gpu/ had been
silent since 2026-09-14T18:02:03Z (12 days) and pm2/ had never posted a run.

FINDING - the gpu-self-heal executor was never lost, only its schedule was.
The brief concluded the mechanism was gone. It is not: on CT 116
/opt/inference-harness/scripts/gpu-self-heal.py exists (mtime 2026-09-21 11:43),
is self-posting via gitea-logger.sh, and runs cleanly - I ran it and it posted
health-logs/gpu/gpu-self-heal-20260926-144618.json. What was missing was the
cron entry, dropped during a CT 116 /etc/cron.d rework on 2026-09-21. So this is
a schedule restoration, not a resurrection.

DECISIONS
* gpu-self-heal -> RESTORE. Schedule re-added as /etc/cron.d/gpu-self-heal
   on CT 116, matching the observed historical cadence (:02 past
  0/6/12/18). Proven with a REAL unattended cron fire (temporary */1 entry,
  removed after): gpu-self-heal-20260926-144802.json committed 14:48:03Z.
  The contract now names the executor, schedule, log and posting, which it
  previously did not.
* pm2 health-logs posting -> RETIRE the claim. pm2-self-heal.prose.md required
  appending to health-logs/pm2/ as a 'hard rule', but scripts/pm2-self-heal.sh
  contains no Gitea or push code and never did; the directory has held only its
  init commit since 2026-07-28. Corrected to point at the contract-runner's
  durable per-run logs and failure note instead of adding a second, redundant
  posting path.

DEAD-MAN'S-SWITCH - scripts/health-log-freshness.py + contract. Absence of logs
must raise an alarm, and the producer cannot raise it, so this runs on CT 100,
a different host from the producers, and fails when the newest health-logs/gpu
entry is older than 12h (litellm 18h). Verified it would have caught the real
gap: evaluated at 2026-09-20 the newest gpu/ entry was 126h old against a 12h
limit -> STALE.

Bugs found and fixed while testing, each caught by a test that bit:
* the documented HEALTH_LOG_MAX_AGE_* override was never implemented;
* a Gitea password held under GITEA_TOKEN was sent as an API token -> HTTP 401;
  now every candidate auth is tried and the first that works is used;
* the ~/.git-credentials fallback filtered on GITEA_URL's host, which missed
  whenever GITEA_URL pointed at the internal IP -> zero candidates.

Exit codes: 0 fresh, 1 stale-or-unreadable, 2 cannot run. A directory that
cannot be read is a failure, never a skip.

Evidence: live PASS; stale override names the directory and producer; no
credential -> exit 2; whole thing runs green through contract-run.sh.
prose-lint: PASSED.
2026-09-26 14:50:52 +00:00
abiba-bot 552c776c0c Merge pull request 'feat: land the revision-preflight guard, fixed and wired into contract execution' (#134) from fix/land-revision-preflight-guard-20260925 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 10s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 0s
2026-09-25 11:46:29 +00:00
root 778424acd4 Merge remote-tracking branch 'origin/master' into fix/land-revision-preflight-guard-20260925
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 8s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 9s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 4s
2026-09-25 11:29:30 +00:00
root c5a6dcd42a fix(revision-preflight): default to warn, name the refusal class, bound the fetch
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 10s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 12s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Failing after 6s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Skipped
Bounded correction round on PR #134 after a PASS-WITH-FINDINGS review whose
Finding 4 is High. The guard's purpose and its fail-closed fix stand; the
problem was that with 'enforce' as the default it gates EVERY scheduled
contract, and three legitimate states produce a refusal - a clone legitimately
ahead of origin/master mid-review, a detached HEAD, and an offline or failed
fetch - so any of them would turn the fleet's monitoring into withheld
verdicts. That risk outweighs the staleness the guard catches.

1. DEFAULT IS NOW 'warn'. 'enforce' remains available and documented. The
   criteria for flipping the default later are written into the doc as a
   decision with evidence - a sustained window (30 days / 200+ runs) with zero
   mismatch:* and zero cannot-verify:* refusals, no fetch blips, and a pinned
   clone demonstrably kept current - explicitly as its own change, not a silent
   flip.

2. 'COULD NOT CHECK' IS NOW DISTINGUISHABLE FROM 'THIS COPY IS WRONG'. Every
   non-zero exit prints a machine-readable REASON=<class> line:
     cannot-verify:fetch-failed | cannot-verify:ref-unresolvable   (exit 2)
     mismatch:path-absent | mismatch:content
     mismatch:detached-head | mismatch:clone-ahead                (exit 1)
   detached-head and clone-ahead are named separately because they are
   legitimate states, far less alarming than a hand-edited file. clone-ahead
   requires HEAD to be STRICTLY ahead; an uncommitted edit on a commit that IS
   the ref is a plain content mismatch (my own first cut got this wrong and the
   new test 7d caught it).

3. THE DEFAULT FETCH IS BOUNDED: --fetch-timeout, default 20s, 0 = unbounded,
   and a missing 'timeout' binary is itself a cannot-verify rather than an
   unbounded fetch inside a scheduled contract.

4. TEST COVERAGE ADDED for every new class: fetch failure, fetch timeout
   (asserted to return promptly under a 1s bound), unresolvable ref, detached
   HEAD, clone-ahead, genuine content mismatch, and the contract-run.sh default.
   The pre-fix draft fixture comparisons are kept: 31 passed, 0 failed.

5. MERGE-TIME SEQUENCE documented: fast-forward /opt/contract-runner, confirm
   clean, prove a contract runs and reports. Baseline recorded as of today -
   firstmate has already fast-forwarded it to 9faffe4 - with the note that an
   untracked file blocks a fast-forward even when byte-identical.

Live behaviour re-verified on the real runner path:
  default: REASON=mismatch:clone-ahead -> 'continuing because ...=warn' -> VERDICT: PASS, exit 0
  enforce: REASON=mismatch:clone-ahead -> 'VERDICT WITHHELD: mismatch:clone-ahead', exit 2

MANDATORY CHECKS (master went red once from a credential-SHAPED string, so
these are now run on every shippable branch):
  bash scripts/prose-lint.sh        -> LINT PASSED (18 warning(s))
  secret scan                       -> secret scan clean (tree; 34 allowlisted,
                                       24 inert value(s) ignored); No committed credentials
  shellcheck revision-preflight.sh  -> clean
  shellcheck test_revision_preflight.sh -> clean
  shellcheck contract-run.sh        -> SC2034 x1, SC2086 x2 - byte-identical on
                                       master, i.e. pre-existing, none introduced

tests/test_probe_drift.py::test_prose_lint_accepts_report_format_with_provenance
fails both before and after this branch (it runs prose-lint from a temp CWD and
cannot find its sibling secret-scan.sh). Pre-existing, unrelated, not fixed here.
2026-09-25 11:29:15 +00:00
abiba-bot 9faffe4f6b Merge pull request 'feat: add the missing daily-health-digest contract + fix the red lint from #133' (#135) from fix/daily-health-digest-contract-20260925 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 12s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 0s
2026-09-25 11:12:56 +00:00
root cd00cd0475 feat: add the missing daily-health-digest contract and register it
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 13s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
Backlog row daily-health-digest-contract-missing-20260925. The digest has been
dispatched on a schedule with NO contract file at all - no *daily*.prose.md,
absent from contract-registry.yaml, the only reference anywhere being the CT100
cron line. That absence is why the choice of execution copy was silently the
operator's, and how a stale clone could run the check unnoticed.

New daily-health-digest.prose.md states:
* the PINNED execution path /root/abiba-workspace/projects/prose-contracts/
  scripts/daily-infra-report.py and the pinned clone - the cron's FM_HOME clone,
  the only stable non-ephemeral copy; the treehouse clone is a per-agent working
  copy and must NOT be pinned;
* the output shape (all 18 top-level --json keys) and what a healthy run is;
* the exit-code semantics AS THEY ACTUALLY BEHAVE, verified case by case:
  missing PVE_TOKEN or an unreachable probe exits 1 and raises an alert, while
  a missing EMAIL credential is a deliberate DEGRADED leg that still exits 0 and
  still produces the report. PROBE_FAILURES and DEGRADED_LEGS are separate lists
  on purpose and must not be merged;
* the email-delivery dependency, that EMAIL_PASSWORD must be a Google app
  password, that it is failing with 534 5.7.9 as of 2026-09-25, and that a
  delivery failure is a credential dependency rather than a code defect;
* what counts as a failure versus degraded.

Registered in contract-registry.yaml (contracts entry plus index.by_category
.monitoring and index.by_domain.infrastructure). Verified: YAML parses, 31
contracts, exactly one daily-health-digest entry.
2026-09-25 11:08:28 +00:00
root 819cd53bce fix(lint): restore the repo secret scan, which PR #133 broke
Master's lint is RED right now, and it is my doing.

  at 483a66b (before #133): prose-lint -> LINT PASSED
  at 9d64b0b (after  #133): prose-lint -> LINT FAILED, 4 credential-shaped strings

The regression came from #133's new pve_auth() work. The secret scanner flags
the literal header shape PVEAPIToken=<value> (rule proxmox-token), and three
occurrences landed in the tree:

  scripts/daily-infra-report.py  the f-string building the auth header
  tests/test_daily_infra_report.py  a docstring quoting the old placeholder
  tests/test_daily_infra_report.py  a synthetic token in an assertion

Fixed without allowlisting anything, because none of these is a credential:

* the header prefix becomes PVE_AUTH_HEADER = "PVEAPIToken=", a constant ending
  at '=' so the scanner's pattern (which needs a character after '=') cannot
  match, and the f-string no longer contains the literal;
* the test docstring no longer reproduces the old placeholder verbatim;
* the test builds its expected value from the constant plus a local sample
  variable instead of embedding a credential-shaped literal.

Behaviour is unchanged and re-verified: with PVE_TOKEN unset the script still
exits 1 with the probe failure, and with the vault token it still reports
pve_probe_status ok / node_count 5 / nodes_online 5.

  before: LINT FAILED — 4 credential-shaped strings
  after:  LINT PASSED (18 warnings)
2026-09-25 11:08:28 +00:00
root 574cb99d76 feat: land the revision-preflight guard, fixed and wired into contract execution
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Failing after 6s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Skipped
A contract verdict is only meaningful if it came from the merged copy. The
fleet has been bitten three times on 2026-09-25 (a clone parked on a merged
feature branch while executing from another clone; a script copied into the
runner clone by hand; a stale local origin/master making an ancestry check
report unlanded work). The control for this existed as an untracked draft and
protected nobody, because it was entirely fail-open.

Defect in the draft, preserved verbatim as tests/fixtures/revision-preflight.prefix.sh:

  git -C "$CLONE" show "origin/master:$(basename "$SCRIPT")"

basename drops the scripts/ prefix, so for any script under scripts/ it queried
the repo root, failed, took the "warn but don't block" branch and exited 0 -
passing a script that exists in no revision at all. Reproduced:
  pre-fix + scripts/demo.sh under scripts/  -> 'could not resolve', EXIT=0
  pre-fix + a script in no revision          -> EXIT=0

Fixed guard (scripts/revision-preflight.sh):
* resolves the repo-relative path inside the clone, so scripts/ paths resolve;
* FAILS CLOSED - a path absent from the ref, an unresolvable ref, or a failed
  fetch is a failure, never a warning;
* fetches the remote by default, because a stale local ref would otherwise
  pass a stale script as current; --no-fetch states the assumption instead of
  hiding it.

Wiring (scripts/contract-run.sh): before executing, the wrapper runs the guard
against the clone it lives in. Default CONTRACT_REVISION_PREFLIGHT=enforce
withholds the verdict, alerts and exits 2 on mismatch; =warn logs and
continues; =off skips. Verified live: match -> contract proceeds and PASSes;
mismatch -> 'VERDICT WITHHELD', exit 2; =warn -> continues.

Pinning (docs/contract-execution-pinning.md): every contract pins the clone
contract-run.sh lives in - the deployed runner being /opt/contract-runner on
CT 100. Documented that daily-health-digest has no contract file at all, which
is why its execution copy was silently operator-chosen.

Tests: tests/test_revision_preflight.sh, 15 assertions over a throwaway clone
with a real bare remote. It runs the pre-fix draft against the same cases and
shows it passing a ghost script, so the tests provably bite.

shellcheck: scripts/revision-preflight.sh and the new test are clean. The three
findings remaining in contract-run.sh (SC2086 x2, SC2034) are pre-existing and
byte-identical on master.
2026-09-25 11:04:29 +00:00
abiba-bot 9d64b0bd66 Merge pull request 'fix(daily-infra-report): resolve PVE token from env; make a dead probe non-silent' (#133) from fix/daily-health-digest-pve-token-20260925 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Failing after 7s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Skipped
2026-09-25 10:51:41 +00:00
root cdc7ad2c79 fix(daily-infra-report): resolve PVE token from env; make a dead probe non-silent
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Failing after 10s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 9s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Skipped
The Proxmox leg of the daily digest has been reporting NOTHING while exiting 0.

Root cause: the auth header was a literal placeholder string,

    AUTH = "Authorization: PVEAPIToken=«vault: infrastructure/production PVE_API_TOKEN»"

which was sent verbatim. The API rejected it, pve_get() returned None, and the
report rendered node_count=0 / nodes_online=0 with pve_probe_status='unreachable'
while still exiting 0. A monitoring gap that looks like a healthy run.

Also: the vault key is PVE_TOKEN, not PVE_API_TOKEN, so even reading os.environ
by the old name would not have found it.

Fixes:
* pve_auth() resolves the token at call time from PVE_TOKEN (injected by
  'infisical run --env=prod'). Nothing is hardcoded; a missing token raises.
* pve_get() builds the command inside its try block, so a missing token degrades
  to None instead of escaping as an unhandled exception.
* PROBE_FAILURES records an unreachable node/resources probe. Probe failures are
  deliberately separate from DEGRADED_LEGS: a missing credential stays exit 0
  (existing intent), but a probe with no data now exits 1 in both the report and
  --json paths, so it cannot pass unnoticed.

Measured effect on the live host: pve_probe_status unreachable -> ok,
node_count 0 -> 5, nodes_online 0 -> 5, total_vms 0 -> 22, running_vms 0 -> 22.

Tests: 4 new regression tests; all 4 fail against the pre-fix script and pass
after, and the 3 pre-existing tests still pass (7/7).

Not fixed here (needs the captain): the email leg fails with
'534 5.7.9 Application-specific password required' - EMAIL_PASSWORD in the vault
is not a valid Gmail app password for jtabiri@gmail.com. That is a credential
action, not a code change.
2026-09-25 10:34:07 +00:00
abiba-bot 5409dfd73a Merge pull request 'feat: multi-engine search stack + visibility check' (#132) from fix/search-stack-multi-engine-20260925 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 13s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 1s
2026-09-25 01:31:55 +00:00
root 8b2eba4f7a docs+check: correct DDG status, explain silent-zero semantics, credit google cse
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 11s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 8s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
Follow-up corrections after review:

1. DuckDuckGo is NOT fixed. The VPS fallback egress has since been flagged by
   DuckDuckGo too (HTTP 202 + challenge markers), so it reports CAPTCHA on both
   paths. The contract and script docstring now say so instead of claiming a
   fix that had already expired. It stays enabled as best-effort coverage so a
   recovery shows up as a contribution.

2. The relationship between silent zeros and the verdict is now explicit in
   both the script output and the contract: an enabled expected engine that
   contributes zero with no error is REPORTED, not fatal. Only the
   <SEARCH_CHECK_MIN_ENGINES> floor and the extraction leg fail the run. This is
   deliberate - de-duplication and query-shape make a zero non-probative.

3. Recorded that 'google cse' uses a THIRD PARTY's public search-engine id
   hardcoded in the SearXNG build, not a key we own; its quota and availability
   are outside our control, and our own free key would need a wrapper (not
   built).

VPS forward proxy is now a real service: /opt/fwd-proxy docker compose with
restart: unless-stopped, a healthy healthcheck, and Docker enabled at boot.
2026-09-25 01:19:27 +00:00
root 040fecef3e feat: multi-engine search stack + visibility check
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 10s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s
Search stack (192.168.68.7) was effectively Bing-only: google served a JS
shell, duckduckgo CAPTCHA'd from the house egress, and every other shipped
engine returned a silent zero. Upgraded SearXNG to 2026.9.23 (same pinned
digest as the image already pulled by other hosts) which uses browser
impersonation, and routed DuckDuckGo through a VPS forward proxy over the
existing WireGuard tunnel via a per-engine 'network'.

Live result: bing, google cse, brave and yandex contribute on every query;
duckduckgo is best-effort via the datacenter egress.

Adds the visibility leg so a future regression cannot be silent:

  scripts/search-stack-check.py
    * two fixed queries; FAILS when fewer than two engines contribute,
      printing contributing engines and every unresponsive_engines entry
    * FAILS when Firecrawl extraction returns empty markdown or errors
    * reports silent-zero engines explicitly

  scripts/contract-run.sh
    * maps search-stack-visibility -> search-stack-check.py

  search-stack-visibility.prose.md
    * contract text, execution model, pass/fail shapes, residual risk

Scheduled hourly at :15 on CT 100 via /etc/cron.d/contract-runner.
2026-09-25 01:14:03 +00:00
abiba-bot 483a66b7b4 Merge pull request 'Fix PBS GC monitor false positive and add contract-run.sh wrapper' (#131) from fix/contract-run-pbs-gc-20260924 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 12s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 13s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 9s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 4s
2026-09-24 06:07:08 +00:00
root c0d04a2c02 fix: three wrapper holes in contract-run.sh
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 9s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 12s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 11s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
1. Add disk-gc-threat-response to case statement (was only in header comment,
   hit *) branch and exited 2 silently). Mapped to scripts/disk-gc-scan.py per
   contract's Execution section.

2. Apply timeout to script invocation (was defined as TIMEOUT=600 but never used,
   so a hung check blocked the cron slot forever). Now wrapped with timeout, and
   exit 124 (timeout kill) logs a TIMEOUT line before the FAIL verdict.

3. Send alert on exit-2 paths (unknown contract and missing script). Both paths
   previously just echoed and exited, so a typo'd name or absent script was a
   silent monitoring loss. Now they send the same Zulip DM as a failed check.

Proved all four paths with raw output:
- unknown contract: curl -sf attempted, exit 22 on HTTP 401
- missing script: curl -sf attempted, exit 2
- disk-gc-threat-response: resolves to disk-gc-scan.py, runs, PASS
- stub sleep > timeout: TIMEOUT line logged, exit 1, alert failure recorded
2026-09-24 06:01:01 +00:00
root fda6c844ff fix: add -f to curl to fail on HTTP >= 400
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s
Without -f, a rejected credential (HTTP 401) returns curl exit 0, making
a failed alert indistinguishable from a successful one. With -f, curl
exits non-zero on HTTP >= 400, so DM_EXIT and STREAM_EXIT correctly
capture the transmission failure and the run log records it.
2026-09-24 05:27:28 +00:00
root 748ea389be fix: correct execution headings, variableize LOG_DIR, fix dead alert path
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 11s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 2s
1. Replace '(## Execution' + ')' with '## Execution' in 6 contract files
2. Make LOG_DIR honor CONTRACT_RUN_LOG_DIR env var (default: /var/log/contract-runs)
3. Fix Zulip alert: correct URL (https://chat.sysloggh.net/api/v1), user
   (abiba-bot@chat.sysloggh.net), take ZULIP_API_KEY from environment, and
   write alert failures to run log
2026-09-24 05:17:49 +00:00
root f16a890d0e fix: make test_pbs_gc_states.sh self-contained with inline SSH replacement
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 8s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 2s
2026-09-24 05:12:44 +00:00
root c666d3e15c feat: implement PBS GC four-state logic and update contracts
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 2s
- Implement four-state PBS GC logic in proxmox-monitor.sh:
  - probe-failed: unparseable JSON, store not found, or empty body → FAIL
  - running: collection in progress (last-run-endtime absent, upid present) → DO NOT FAIL
  - stale: no completed run within 48h → FAIL, naming last completed run age
  - healthy: completed within 48h → PASS, naming endtime and pending bytes

- Add tests/test_pbs_gc_states.sh covering all four states
  - Proves the test bites on the pre-fix version (5/6 tests fail)
  - All 6 tests pass against the fixed version

- Update contract-run.sh to map disk-gc-threat-response -> scripts/disk-gc-scan.py

- Update Execution sections of host-scheduled contracts:
  infrastructure-monitoring, zulip-health, litellm-health, agent-health-check,
  disk-gc-threat-response, pm2-self-heal
  Adding note that execution is host-scheduled via cron, not agent session ack.
2026-09-24 04:53:46 +00:00
root c07aa5e382 Add contract-run.sh for machine scheduler execution and fix PBS GC leg
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 15s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s
- Add scripts/contract-run.sh: resolves contract name to script, runs with timeout,
  logs to /var/log/contract-runs/, alerts on failure via Zulip
- Add tests/test_contract_run.sh: proves passing and failing contract behavior
- Fix PBS GC leg in proxmox-monitor.sh: simplify logic to check last-run-endtime,
  use absolute paths for pct and proxmox-backup-manager to avoid PATH issues

Part of Task: contract-execution-host-scheduler-20260924
2026-09-24 01:22:24 +00:00
abiba-bot 87205e6fdb Merge pull request 'feat(security): commit-time secret guard that FAILS the build on a committed credential' (#129) from fm/commit-time-secret-guard-20260917 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 13s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 1s
2026-09-22 15:33:16 +00:00
abiba-bot 8f1e5eebc4 feat(security): commit-time secret guard that FAILS the build on a committed credential
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 18s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
The 2026-09-17 purge removed six live credentials that had sat in this repo
for weeks, several in .md prose. Nothing blocked that class of commit, so a
warning in a stream nobody reads was the only signal. This adds a guard that
fails the build instead of warning.

Guard
- scripts/secret-scan.sh: bash + coreutils + grep/sed/awk + git only (the Gitea
  Actions runner executes job steps inside the runner container — BusyBox grep,
  no node/python). Modes: --tree (git-tracked, default), --path DIR (no git),
  --staged (pre-commit), --diff REF. Exit 1 on a finding, 2 on config error.
- scripts/secret-patterns.tsv: checked-in pattern list — sk-, sk-or-v1-,
  sk_live_, literal Bearer tokens, PVEAPIToken=, raw Authorization values, PEM
  private-key blocks, prose credential lines, and password/api_key/secret/token
  assignments carrying a literal value. Prose is scanned exactly like code.
- scripts/secret-allowlist.tsv: one entry per deliberate synthetic example, each
  with a reason. A missing reason is a hard error (fail closed). The 2026-09-17
  purge's `«vault: ...»` placeholders are listed explicitly rather than filtered
  by a general "vault"/"synthetic" rule, so a new occurrence still needs a
  reviewed, reasoned entry.
- A small inert-value classifier drops env refs, paths, dotted code access,
  variable names and right-truncated redactions; it does not know the words
  "synthetic"/"example", so a fabrication is always an explicit exception.
- Findings are printed with the credential masked; a scan never echoes a full
  secret into the log.

Wiring
- .gitea/workflows/pr-pipeline.yaml lint job: explicit "Committed-credential
  scan" step plus the self-test. A finding fails the required
  `pr-pipeline / lint` context, which the merge gate depends on.
- scripts/prose-lint.sh (the local gate): a "Secret scan" section, so
  `bash scripts/prose-lint.sh` before pushing is equivalent to CI.

Tests
- tests/test_secret_scan.sh: 20 cases. Plants pattern-matching fixtures in temp
  trees (outside every allowlisted path) and asserts the guard FAILS, including
  the --staged commit-time path; asserts the tree is quiet; asserts allowlisted
  text at an unlisted path still fails (path-explicit, not word-based); asserts
  a reasonless allowlist entry exits 2.

Verified: guard run against 8245716^ (the pre-fix revision, before the purge)
fails on the real OpenRouter/LiteLLM/Zulip/Proxmox/Stirling credentials; guard
run over the current tree is clean.
2026-09-22 15:04:37 +00:00
root bb1b65340e Merge pull request #128 from decisions-2026-08-03-rebased
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 8s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 11s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 1s
2026-09-22 11:28:33 +00:00
root 9e87927444 pm2-self-heal: reconcile contract with live PM2 set and alert channels
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
- Add zulip-watchdog to Maintains (it's running, infrastructure-monitoring expects it)
- Remove gpu-monitor from PM2 Maintains (it's systemd-only, not PM2-tracked)
- Add Execution steps for all monitored processes (gitea-runner, zulip-watchdog)
- Update alert channel: Telegram is primary, Zulip DM is secondary
- Update script to check all 4 processes (abiba-telegram, abiba-zulip, gitea-runner, zulip-watchdog)
- Add restart count thresholds for all processes
- Update log output to include all process statuses
2026-09-22 11:25:45 +00:00
root b0683e9566 fix: correct logging destination in description (Gitea health-logs, not knowledge graph) 2026-09-22 11:03:20 +00:00
root dd11c8f14f Apply captain 2026-08-03 decisions: restore abiba-zulip, retire gpu-watchdog, PM2-track gpu-monitor 2026-09-22 10:45:44 +00:00
abiba-bot 0732eed329 Merge pull request 'fix(daily-infra-report): drop the vestigial Zulip key requirement; fail loudly on a failed send' (#127) from fix/daily-health-digest-remove-vestigial-zulip-20260921 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 1s
2026-09-21 11:36:48 +00:00
root 59ed7cdbf7 fix(daily-infra-report): remove vestigial ZULIP_API_KEY requirement, exit non-zero on failed send
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s
1. Remove vestigial ZULIP_API_KEY requirement:
   - /api/v1/server_settings is a PUBLIC endpoint (verified HTTP 200 with or without credential)
   - No Zulip API key is required for this call
   - If a future leg genuinely needs abiba-bot's key, it must prove it with a 200 from
     /api/v1/users/me as abiba-bot and label itself degraded when it cannot
   - Never fall back to the vault's shared ZULIP_API_KEY

2. Make failed sends exit non-zero:
   - A degraded leg (no credential configured) must stay exit 0
   - A failed send (attempted and failed) must exit 1
   - This distinguishes 'not configured' from 'attempted and failed'

Test evidence:
- No-credential run: exit 0, digest still produced
- Wrong password: exit 1, labelled SMTP error
- grep -n ZULIP_API_KEY: only comment reference remains
2026-09-21 11:30:42 +00:00
abiba-bot 5d70bbf25b Merge pull request 'fix(daily-infra-report): missing credentials degrade one leg instead of blacking out the digest' (#126) from fix/daily-health-digest-degraded-credentials-20260921 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 2s
2026-09-21 11:26:32 +00:00
root ccc916d1ec feat(daily-infra-report): make missing credentials a labelled degraded leg
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
- ZULIP_API_KEY: no longer SystemExit, now reports 'credential-missing: ZULIP_API_KEY'
- EMAIL_PASSWORD: no longer sys.exit(1), now appends to DEGRADED_LEGS and returns success
- PVE API: fixed None check in storage section
- Summary: reports degraded legs before summary

This allows the digest to be produced and emailed even when credentials are missing,
while still explicitly logging which legs are degraded.

Test: empty env produces JSON report + degraded leg labels, no SystemExit.
2026-09-21 11:00:25 +00:00
abiba-bot 48fb263d4b Merge pull request 'feat(litellm): align contracts for 7-provider cloud consolidation (2026-09-20)' (#125) from feat/align-litellm-contracts-cloud-20260920 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 15s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 2s
2026-09-21 08:24:13 +00:00
agent-zero 64790ebd19 feat(litellm): align contracts for 7-provider cloud consolidation (2026-09-20)
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s
- hermes-key-enforcement: add Model Access Tiers (local vs cloud) and cloud-scoping clause
- litellm-api-keys: add Cloud Provider Consolidation section (provider map, vault secrets, access tiers)
- litellm-api-keys: pin standard agent keys to explicit local-only models list; forbid {}/all-proxy-models (Community-edition cloud leak)
- contract-registry: register litellm-api-keys (was unregistered drift)

Additive only. Local prose-lint: PASSED.
2026-09-20 10:53:33 -04:00
abiba-bot 6fa0a255df Merge pull request 'fix(zulip-monitor): add C3 public access path, make run verdict non-optimistic, separate C1/C2/C3' (#124) from fm/kagentz-a2a-outage-masked-as-expected-20260919 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 8s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 2s
2026-09-19 22:50:06 +00:00
root c72436b406 no-mistakes(document): docs: reconcile zulip-monitor status in infrastructure-control
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
2026-09-19 22:40:25 +00:00
root a17379676d no-mistakes(document): docs: sync zulip-health contract with C1/C3 and version 2026-09-19 22:38:53 +00:00
root 63990b84f7 no-mistakes(review): Fix review findings: test regression, contract verdict, scope trim 2026-09-19 22:31:25 +00:00
root 7a5ddb46a9 chore: Add local monitoring tooling scripts 2026-09-19 22:27:19 +00:00
root 568fec2efa test(zulip-kagentz): Replace string-presence tests with behavioural sandbox tests
- C3 502 → INCIDENT
- C3 000 → INCIDENT
- C1 401 + C3 302 → 0 issues, all healthy
- C1 000 → INCIDENT

Each test asserts from the run's own log/verdict, not from file text.
Prose assertions kept as secondary.

Proven to bite: run against pre-fix script (origin/master) shows all 4
behavioural cases fail because C3 leg doesn't exist and Result line
doesn't say 'INCIDENT'.
2026-09-19 22:25:03 +00:00
root 641a52c6da fix(zulip-monitor): Add C3 public access path, make Result line non-optimistic
- (b) Changed Result line to say 'INCIDENT' when ISSUES > 0, '0 issues (all healthy)' when ISSUES = 0
- (c) Documented C1 (no credential needed), C2 (requires LITELLM_KEY) distinction
- (d) Added C3 public access path leg for https://kagentz.sysloggh.net/
- C3 treats 200/302/401 as alive, 502/000 as incident
- Added tests/test_zulip_kagentz_legs.py to verify all changes
2026-09-19 22:14:08 +00:00
abiba-bot 58033f39c5 Merge pull request 'fix: Rule 5 accept canonical internal and public host base_url' (#122) from fix/rule5-canonical-baseurl-20260919 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 8s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 0s
2026-09-19 11:57:35 +00:00
abiba-bot 6c59988e7f Merge pull request 'fix(litellm-health): three-state model verdicts - busy is not a failure' (#123) from fix/litellm-health-busy-vs-down-20260919 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 1s
2026-09-19 11:57:25 +00:00
root 9a2ee6faec fix(litellm): Remove duplicate 'host healthy' from busy line
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
The busy line was rendering as:
  'busy (completion timed out after retry; host healthy host healthy (200))'

because host_detail already contains 'host healthy (200)' and the prefix
also said 'host healthy'. Fixed to:
  'busy (completion timed out after retry; host healthy (200))'

F1 cosmetic fix from PR #123 verify.
2026-09-19 11:53:32 +00:00
root e71ded3c8c fix: internal /v1 WARN not FAIL; align prose
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 3s
Rule 5 now treats internal http://192.168.68.116/v1 as non-canonical
but working (authenticated via nginx), producing a WARNING instead of a
FAILURE. The canonical internal path /litellm/v1 and the public host
https://litellm.sysloggh.net/v1 both PASS. Everything else FAILS.

Prose aligned: hermes-key-enforcement.prose.md now states the canonical
internal form, notes that internal /v1 still works but is flagged as
non-canonical (WARN not FAIL), and clarifies that the public host serves
/v1 ONLY (404 on /litellm/v1). Corrected the 'unauthenticated path'
wording at line 116, which was factually wrong.

Tests updated: BASE template uses canonical internal path; new test
cases prove canonical /litellm/v1 PASSES, wrong path FAILS, public host
PASSES, and internal /v1 WARNS (not FAILS). Fixed backwards comment in
test_old_rule5_check_would_fail_canonical.
2026-09-19 11:48:48 +00:00
root a12abbeb14 fix(litellm): Implement busy vs down with proper degraded state
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 8s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 8s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
Three states:
- healthy: passed, exit 0 (unchanged)
- busy (completion timed out after retry AND host /health answered):
  ⚠️ DEGRADED line, does NOT fail the run, exit 0
- host unreachable or real fault: ❌, exit 1 (unchanged)

Summary now reports degraded count:
- All pass, no degraded: '✅ All checks passed'
- All pass, 1+ degraded: '✅ All checks passed (1 degraded: gpu-dense)'
- Some failed: '❌ Some checks failed' or '❌ Some checks failed (1 degraded: ...)'

Host health mapping verified:
- gpu-dense -> 192.168.68.8:8080/health
- gpu-vision -> 192.168.68.110:8080/health
- strix-moe -> 192.168.68.15:8080/health
2026-09-19 11:43:39 +00:00
root 0b9aebca37 fix(litellm): Fix timeout kind reporting + add busy/degraded detection
1. TIMEOUT KIND FIX: run_command returns (1, '', 'TIMEOUT') when its own
   timeout fires. probe_http now checks for this before falling through to
   'curl exit <rc>', so a 30s timeout reports 'timeout after 30s' not
   'curl exit 1'.

2. BUSY/DEGRADED DETECTION: After both model probes fail, check the
   model's host health endpoint (e.g. 192.168.68.8:8080/health for
   gpu-dense). If the host answers 200, report 'busy (completion timed
   out after retry; host healthy 200)' — do NOT fail the run on that
   alone. If the host does not answer, that's a real FAIL.

3. RETRY TIMEOUT INCREASED: Single-host retry timeout raised from 45s to
   90s. Worst-case prefill on a single-slot .8 host is ~76s (observed
   83K-token prompt at 1078 tok/s), so 90s covers it.

New line shapes:
- Busy: 'gpu-dense: busy (completion timed out after retry; host healthy 200)'
- Real failure: 'probe-failed: gpu-dense timeout after 30s then timeout after 90s (2 attempts)'
2026-09-19 11:42:36 +00:00
root fa458afa26 fix: Rule 5 accept canonical internal and public host base_url
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
Rule 5 in audit-hermes-config.py had an inverted check: it expected
base_url=http://192.168.68.116/v1, but the contract hermes-key-enforcement.prose.md
names http://192.168.68.116/litellm/v1 as CORRECT/CANONICAL in multiple places.
The audit script would FAIL a config using the contract's canonical internal path
and PASS one using a path the contract does not name.

Fix: Rule 5 now accepts the canonical internal base (http://192.168.68.116/litellm/v1)
AND the public base (https://litellm.sysloggh.net/v1), and FAILS anything else.
The internal nginx serves both /litellm/v1 and /v1; the public host serves /v1 only
(per 2026-09-19 probe from CT 116).

Tests aligned: BASE template updated to use the canonical internal path, and new
test cases added to prove the canonical internal path PASSES, a wrong path FAILS,
and the public host path PASSES.
2026-09-19 11:37:14 +00:00
abiba-bot aa3da83af5 Merge pull request 'fix(litellm-health): retry single-host model probes once at a longer timeout' (#121) from fix/litellm-health-retry-timeout-20260919 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 9s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 0s
2026-09-19 10:57:39 +00:00
root aee2de25ac fix(litellm): Report both attempts' failure kinds in probe-failed
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 4s
The failure line now preserves both attempts' failure kinds instead of
hardcoding 'timeout after retry, 45s'. If both attempts fail, the report
shows: 'probe-failed: <model> <first kind> then <retry kind> (2 attempts)'.

This fixes the self-contradictory output when the first attempt timed out
but the retry failed with connection refused, and prevents the duration
from appearing twice when both attempts were timeouts.

Example outputs:
- timeout then timeout: 'probe-failed: gpu-dense timeout after 30s then timeout after 45s (2 attempts)'
- timeout then refused: 'probe-failed: gpu-dense timeout after 30s then connection refused (2 attempts)'
- refused then refused: 'probe-failed: gpu-dense connection refused then connection refused (2 attempts)'
2026-09-19 10:53:44 +00:00
root 76653381ec fix(litellm): Add retry with longer timeout for single-host model probes
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 11s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s
Single-host models (gpu-dense, gpu-vision, strix-moe) now retry once at
45s on initial 30s timeout failure before declaring probe-failed. This
prevents a single transient timeout (cold prefill ~13s or concurrent
generation hold) from failing the entire health digest.

Evidence: 2026-09-19 ~06:55Z digest failed gpu-dense at 30s; 06:56Z
direct probe 200 in 1.04s.

The failed-probe-fails-the-run property is preserved: if both attempts
fail, the script still exits non-zero with the target and duration named.

Closes: daily-health-digest false negative on single transient timeout
2026-09-19 10:43:54 +00:00