Commit Graph
43 Commits
Author SHA1 Message Date
root e0c9852de8 no-mistakes(document): Document llmuser SSH user for .8 GPU health probe
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 11s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 12s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 24s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 23s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 2s
2026-09-28 07:22:13 +00:00
root 96769a103f no-mistakes(review): sync tanko CT 112 mapping to minipve across consumers
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 12s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 9s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 17s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 3s
2026-09-28 06:47:46 +00:00
root e42b970dec fix(audit-hermes): handle fallback_providers as list or dict
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 17s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s
The audit assumed fallback_providers was always a dict (single provider).
Two live agents (koby, koonimo) carry it as a LIST of dicts (one entry per
fallback), so the script crashed with:

    File "audit-hermes-config.py", line 211, in audit
        fb.get("provider") == "deepseek",
    AttributeError: 'list' object has no attribute 'get'

Both are REAL agent configs, so this is not a malformed-input case — the
script simply could not audit two of the four agents it exists to audit.

Fix:
- Normalize fallback_providers to a list of entries (dict → [dict], list → list)
- Apply the existing checks to each entry
- A malformed entry (not a mapping) produces a reported VIOLATION naming the
  offending entry, NOT an uncaught exception

Adds regression test using the real failing shape (list of dicts) and proves
it bites against the pre-fix revision.

Real audit results after fix:
- mumuni: FAIL — 7 violations
- tanko: FAIL — 21 violations
- koby: FAIL — 16 violations (previously crashed)
- koonimo: FAIL — 10 violations (previously crashed)

No agent configs were changed. No existing rules were relaxed.
2026-09-27 11:17:18 +00:00
root 778424acd4 Merge remote-tracking branch 'origin/master' into fix/land-revision-preflight-guard-20260925
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 8s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 9s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 4s
2026-09-25 11:29:30 +00:00
root c5a6dcd42a fix(revision-preflight): default to warn, name the refusal class, bound the fetch
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 10s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 12s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Failing after 6s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Skipped
Bounded correction round on PR #134 after a PASS-WITH-FINDINGS review whose
Finding 4 is High. The guard's purpose and its fail-closed fix stand; the
problem was that with 'enforce' as the default it gates EVERY scheduled
contract, and three legitimate states produce a refusal - a clone legitimately
ahead of origin/master mid-review, a detached HEAD, and an offline or failed
fetch - so any of them would turn the fleet's monitoring into withheld
verdicts. That risk outweighs the staleness the guard catches.

1. DEFAULT IS NOW 'warn'. 'enforce' remains available and documented. The
   criteria for flipping the default later are written into the doc as a
   decision with evidence - a sustained window (30 days / 200+ runs) with zero
   mismatch:* and zero cannot-verify:* refusals, no fetch blips, and a pinned
   clone demonstrably kept current - explicitly as its own change, not a silent
   flip.

2. 'COULD NOT CHECK' IS NOW DISTINGUISHABLE FROM 'THIS COPY IS WRONG'. Every
   non-zero exit prints a machine-readable REASON=<class> line:
     cannot-verify:fetch-failed | cannot-verify:ref-unresolvable   (exit 2)
     mismatch:path-absent | mismatch:content
     mismatch:detached-head | mismatch:clone-ahead                (exit 1)
   detached-head and clone-ahead are named separately because they are
   legitimate states, far less alarming than a hand-edited file. clone-ahead
   requires HEAD to be STRICTLY ahead; an uncommitted edit on a commit that IS
   the ref is a plain content mismatch (my own first cut got this wrong and the
   new test 7d caught it).

3. THE DEFAULT FETCH IS BOUNDED: --fetch-timeout, default 20s, 0 = unbounded,
   and a missing 'timeout' binary is itself a cannot-verify rather than an
   unbounded fetch inside a scheduled contract.

4. TEST COVERAGE ADDED for every new class: fetch failure, fetch timeout
   (asserted to return promptly under a 1s bound), unresolvable ref, detached
   HEAD, clone-ahead, genuine content mismatch, and the contract-run.sh default.
   The pre-fix draft fixture comparisons are kept: 31 passed, 0 failed.

5. MERGE-TIME SEQUENCE documented: fast-forward /opt/contract-runner, confirm
   clean, prove a contract runs and reports. Baseline recorded as of today -
   firstmate has already fast-forwarded it to 9faffe4 - with the note that an
   untracked file blocks a fast-forward even when byte-identical.

Live behaviour re-verified on the real runner path:
  default: REASON=mismatch:clone-ahead -> 'continuing because ...=warn' -> VERDICT: PASS, exit 0
  enforce: REASON=mismatch:clone-ahead -> 'VERDICT WITHHELD: mismatch:clone-ahead', exit 2

MANDATORY CHECKS (master went red once from a credential-SHAPED string, so
these are now run on every shippable branch):
  bash scripts/prose-lint.sh        -> LINT PASSED (18 warning(s))
  secret scan                       -> secret scan clean (tree; 34 allowlisted,
                                       24 inert value(s) ignored); No committed credentials
  shellcheck revision-preflight.sh  -> clean
  shellcheck test_revision_preflight.sh -> clean
  shellcheck contract-run.sh        -> SC2034 x1, SC2086 x2 - byte-identical on
                                       master, i.e. pre-existing, none introduced

tests/test_probe_drift.py::test_prose_lint_accepts_report_format_with_provenance
fails both before and after this branch (it runs prose-lint from a temp CWD and
cannot find its sibling secret-scan.sh). Pre-existing, unrelated, not fixed here.
2026-09-25 11:29:15 +00:00
root 819cd53bce fix(lint): restore the repo secret scan, which PR #133 broke
Master's lint is RED right now, and it is my doing.

  at 483a66b (before #133): prose-lint -> LINT PASSED
  at 9d64b0b (after  #133): prose-lint -> LINT FAILED, 4 credential-shaped strings

The regression came from #133's new pve_auth() work. The secret scanner flags
the literal header shape PVEAPIToken=<value> (rule proxmox-token), and three
occurrences landed in the tree:

  scripts/daily-infra-report.py  the f-string building the auth header
  tests/test_daily_infra_report.py  a docstring quoting the old placeholder
  tests/test_daily_infra_report.py  a synthetic token in an assertion

Fixed without allowlisting anything, because none of these is a credential:

* the header prefix becomes PVE_AUTH_HEADER = "PVEAPIToken=", a constant ending
  at '=' so the scanner's pattern (which needs a character after '=') cannot
  match, and the f-string no longer contains the literal;
* the test docstring no longer reproduces the old placeholder verbatim;
* the test builds its expected value from the constant plus a local sample
  variable instead of embedding a credential-shaped literal.

Behaviour is unchanged and re-verified: with PVE_TOKEN unset the script still
exits 1 with the probe failure, and with the vault token it still reports
pve_probe_status ok / node_count 5 / nodes_online 5.

  before: LINT FAILED — 4 credential-shaped strings
  after:  LINT PASSED (18 warnings)
2026-09-25 11:08:28 +00:00
root 574cb99d76 feat: land the revision-preflight guard, fixed and wired into contract execution
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Failing after 6s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Skipped
A contract verdict is only meaningful if it came from the merged copy. The
fleet has been bitten three times on 2026-09-25 (a clone parked on a merged
feature branch while executing from another clone; a script copied into the
runner clone by hand; a stale local origin/master making an ancestry check
report unlanded work). The control for this existed as an untracked draft and
protected nobody, because it was entirely fail-open.

Defect in the draft, preserved verbatim as tests/fixtures/revision-preflight.prefix.sh:

  git -C "$CLONE" show "origin/master:$(basename "$SCRIPT")"

basename drops the scripts/ prefix, so for any script under scripts/ it queried
the repo root, failed, took the "warn but don't block" branch and exited 0 -
passing a script that exists in no revision at all. Reproduced:
  pre-fix + scripts/demo.sh under scripts/  -> 'could not resolve', EXIT=0
  pre-fix + a script in no revision          -> EXIT=0

Fixed guard (scripts/revision-preflight.sh):
* resolves the repo-relative path inside the clone, so scripts/ paths resolve;
* FAILS CLOSED - a path absent from the ref, an unresolvable ref, or a failed
  fetch is a failure, never a warning;
* fetches the remote by default, because a stale local ref would otherwise
  pass a stale script as current; --no-fetch states the assumption instead of
  hiding it.

Wiring (scripts/contract-run.sh): before executing, the wrapper runs the guard
against the clone it lives in. Default CONTRACT_REVISION_PREFLIGHT=enforce
withholds the verdict, alerts and exits 2 on mismatch; =warn logs and
continues; =off skips. Verified live: match -> contract proceeds and PASSes;
mismatch -> 'VERDICT WITHHELD', exit 2; =warn -> continues.

Pinning (docs/contract-execution-pinning.md): every contract pins the clone
contract-run.sh lives in - the deployed runner being /opt/contract-runner on
CT 100. Documented that daily-health-digest has no contract file at all, which
is why its execution copy was silently operator-chosen.

Tests: tests/test_revision_preflight.sh, 15 assertions over a throwaway clone
with a real bare remote. It runs the pre-fix draft against the same cases and
shows it passing a ghost script, so the tests provably bite.

shellcheck: scripts/revision-preflight.sh and the new test are clean. The three
findings remaining in contract-run.sh (SC2086 x2, SC2034) are pre-existing and
byte-identical on master.
2026-09-25 11:04:29 +00:00
root cdc7ad2c79 fix(daily-infra-report): resolve PVE token from env; make a dead probe non-silent
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Failing after 10s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 9s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Skipped
The Proxmox leg of the daily digest has been reporting NOTHING while exiting 0.

Root cause: the auth header was a literal placeholder string,

    AUTH = "Authorization: PVEAPIToken=«vault: infrastructure/production PVE_API_TOKEN»"

which was sent verbatim. The API rejected it, pve_get() returned None, and the
report rendered node_count=0 / nodes_online=0 with pve_probe_status='unreachable'
while still exiting 0. A monitoring gap that looks like a healthy run.

Also: the vault key is PVE_TOKEN, not PVE_API_TOKEN, so even reading os.environ
by the old name would not have found it.

Fixes:
* pve_auth() resolves the token at call time from PVE_TOKEN (injected by
  'infisical run --env=prod'). Nothing is hardcoded; a missing token raises.
* pve_get() builds the command inside its try block, so a missing token degrades
  to None instead of escaping as an unhandled exception.
* PROBE_FAILURES records an unreachable node/resources probe. Probe failures are
  deliberately separate from DEGRADED_LEGS: a missing credential stays exit 0
  (existing intent), but a probe with no data now exits 1 in both the report and
  --json paths, so it cannot pass unnoticed.

Measured effect on the live host: pve_probe_status unreachable -> ok,
node_count 0 -> 5, nodes_online 0 -> 5, total_vms 0 -> 22, running_vms 0 -> 22.

Tests: 4 new regression tests; all 4 fail against the pre-fix script and pass
after, and the 3 pre-existing tests still pass (7/7).

Not fixed here (needs the captain): the email leg fails with
'534 5.7.9 Application-specific password required' - EMAIL_PASSWORD in the vault
is not a valid Gmail app password for jtabiri@gmail.com. That is a credential
action, not a code change.
2026-09-25 10:34:07 +00:00
abiba-bot 483a66b7b4 Merge pull request 'Fix PBS GC monitor false positive and add contract-run.sh wrapper' (#131) from fix/contract-run-pbs-gc-20260924 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 12s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 13s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 9s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 4s
2026-09-24 06:07:08 +00:00
root f16a890d0e fix: make test_pbs_gc_states.sh self-contained with inline SSH replacement
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 8s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 2s
2026-09-24 05:12:44 +00:00
root c666d3e15c feat: implement PBS GC four-state logic and update contracts
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 2s
- Implement four-state PBS GC logic in proxmox-monitor.sh:
  - probe-failed: unparseable JSON, store not found, or empty body → FAIL
  - running: collection in progress (last-run-endtime absent, upid present) → DO NOT FAIL
  - stale: no completed run within 48h → FAIL, naming last completed run age
  - healthy: completed within 48h → PASS, naming endtime and pending bytes

- Add tests/test_pbs_gc_states.sh covering all four states
  - Proves the test bites on the pre-fix version (5/6 tests fail)
  - All 6 tests pass against the fixed version

- Update contract-run.sh to map disk-gc-threat-response -> scripts/disk-gc-scan.py

- Update Execution sections of host-scheduled contracts:
  infrastructure-monitoring, zulip-health, litellm-health, agent-health-check,
  disk-gc-threat-response, pm2-self-heal
  Adding note that execution is host-scheduled via cron, not agent session ack.
2026-09-24 04:53:46 +00:00
root c07aa5e382 Add contract-run.sh for machine scheduler execution and fix PBS GC leg
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 15s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s
- Add scripts/contract-run.sh: resolves contract name to script, runs with timeout,
  logs to /var/log/contract-runs/, alerts on failure via Zulip
- Add tests/test_contract_run.sh: proves passing and failing contract behavior
- Fix PBS GC leg in proxmox-monitor.sh: simplify logic to check last-run-endtime,
  use absolute paths for pct and proxmox-backup-manager to avoid PATH issues

Part of Task: contract-execution-host-scheduler-20260924
2026-09-24 01:22:24 +00:00
abiba-bot 8f1e5eebc4 feat(security): commit-time secret guard that FAILS the build on a committed credential
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 18s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
The 2026-09-17 purge removed six live credentials that had sat in this repo
for weeks, several in .md prose. Nothing blocked that class of commit, so a
warning in a stream nobody reads was the only signal. This adds a guard that
fails the build instead of warning.

Guard
- scripts/secret-scan.sh: bash + coreutils + grep/sed/awk + git only (the Gitea
  Actions runner executes job steps inside the runner container — BusyBox grep,
  no node/python). Modes: --tree (git-tracked, default), --path DIR (no git),
  --staged (pre-commit), --diff REF. Exit 1 on a finding, 2 on config error.
- scripts/secret-patterns.tsv: checked-in pattern list — sk-, sk-or-v1-,
  sk_live_, literal Bearer tokens, PVEAPIToken=, raw Authorization values, PEM
  private-key blocks, prose credential lines, and password/api_key/secret/token
  assignments carrying a literal value. Prose is scanned exactly like code.
- scripts/secret-allowlist.tsv: one entry per deliberate synthetic example, each
  with a reason. A missing reason is a hard error (fail closed). The 2026-09-17
  purge's `«vault: ...»` placeholders are listed explicitly rather than filtered
  by a general "vault"/"synthetic" rule, so a new occurrence still needs a
  reviewed, reasoned entry.
- A small inert-value classifier drops env refs, paths, dotted code access,
  variable names and right-truncated redactions; it does not know the words
  "synthetic"/"example", so a fabrication is always an explicit exception.
- Findings are printed with the credential masked; a scan never echoes a full
  secret into the log.

Wiring
- .gitea/workflows/pr-pipeline.yaml lint job: explicit "Committed-credential
  scan" step plus the self-test. A finding fails the required
  `pr-pipeline / lint` context, which the merge gate depends on.
- scripts/prose-lint.sh (the local gate): a "Secret scan" section, so
  `bash scripts/prose-lint.sh` before pushing is equivalent to CI.

Tests
- tests/test_secret_scan.sh: 20 cases. Plants pattern-matching fixtures in temp
  trees (outside every allowlisted path) and asserts the guard FAILS, including
  the --staged commit-time path; asserts the tree is quiet; asserts allowlisted
  text at an unlisted path still fails (path-explicit, not word-based); asserts
  a reasonless allowlist entry exits 2.

Verified: guard run against 8245716^ (the pre-fix revision, before the purge)
fails on the real OpenRouter/LiteLLM/Zulip/Proxmox/Stirling credentials; guard
run over the current tree is clean.
2026-09-22 15:04:37 +00:00
root 63990b84f7 no-mistakes(review): Fix review findings: test regression, contract verdict, scope trim 2026-09-19 22:31:25 +00:00
root 568fec2efa test(zulip-kagentz): Replace string-presence tests with behavioural sandbox tests
- C3 502 → INCIDENT
- C3 000 → INCIDENT
- C1 401 + C3 302 → 0 issues, all healthy
- C1 000 → INCIDENT

Each test asserts from the run's own log/verdict, not from file text.
Prose assertions kept as secondary.

Proven to bite: run against pre-fix script (origin/master) shows all 4
behavioural cases fail because C3 leg doesn't exist and Result line
doesn't say 'INCIDENT'.
2026-09-19 22:25:03 +00:00
root 641a52c6da fix(zulip-monitor): Add C3 public access path, make Result line non-optimistic
- (b) Changed Result line to say 'INCIDENT' when ISSUES > 0, '0 issues (all healthy)' when ISSUES = 0
- (c) Documented C1 (no credential needed), C2 (requires LITELLM_KEY) distinction
- (d) Added C3 public access path leg for https://kagentz.sysloggh.net/
- C3 treats 200/302/401 as alive, 502/000 as incident
- Added tests/test_zulip_kagentz_legs.py to verify all changes
2026-09-19 22:14:08 +00:00
root e71ded3c8c fix: internal /v1 WARN not FAIL; align prose
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 3s
Rule 5 now treats internal http://192.168.68.116/v1 as non-canonical
but working (authenticated via nginx), producing a WARNING instead of a
FAILURE. The canonical internal path /litellm/v1 and the public host
https://litellm.sysloggh.net/v1 both PASS. Everything else FAILS.

Prose aligned: hermes-key-enforcement.prose.md now states the canonical
internal form, notes that internal /v1 still works but is flagged as
non-canonical (WARN not FAIL), and clarifies that the public host serves
/v1 ONLY (404 on /litellm/v1). Corrected the 'unauthenticated path'
wording at line 116, which was factually wrong.

Tests updated: BASE template uses canonical internal path; new test
cases prove canonical /litellm/v1 PASSES, wrong path FAILS, public host
PASSES, and internal /v1 WARNS (not FAILS). Fixed backwards comment in
test_old_rule5_check_would_fail_canonical.
2026-09-19 11:48:48 +00:00
root fa458afa26 fix: Rule 5 accept canonical internal and public host base_url
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
Rule 5 in audit-hermes-config.py had an inverted check: it expected
base_url=http://192.168.68.116/v1, but the contract hermes-key-enforcement.prose.md
names http://192.168.68.116/litellm/v1 as CORRECT/CANONICAL in multiple places.
The audit script would FAIL a config using the contract's canonical internal path
and PASS one using a path the contract does not name.

Fix: Rule 5 now accepts the canonical internal base (http://192.168.68.116/litellm/v1)
AND the public base (https://litellm.sysloggh.net/v1), and FAILS anything else.
The internal nginx serves both /litellm/v1 and /v1; the public host serves /v1 only
(per 2026-09-19 probe from CT 116).

Tests aligned: BASE template updated to use the canonical internal path, and new
test cases added to prove the canonical internal path PASSES, a wrong path FAILS,
and the public host path PASSES.
2026-09-19 11:37:14 +00:00
root 8210fd905c fix: align contracts to 4-name LiteLLM registry (2026-09-12)
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 1s
- Remove retired names (qwen3.6-27B-code, qwen3.6-35B-udq4) from live alias claims
- Update Strix Halo model to Carnice-Qwen3.6-MoE-35B-A3B-Q4_K_M.gguf (strix-moe, 256K ctx)
- Fix litellm-health step 7 probe to gpu-vision (monitor key scoped)
- Move qwen3.6-27B-code/35B-udq4 from raw-but-live to non-resolving in audit
- Fold in pm2-self-heal: remove spoton-service (live PM2 set is 4/4)
- Update hermes templates, key enforcement, timeout tables to live names
2026-09-12 21:39:54 +00:00
root 6272d29978 no-mistakes(review): Fail closed on empty report-only exclusion list 2026-09-12 18:55:34 +00:00
root 4bd6132cf5 no-mistakes(review): Reject report-only exclusion entries lacking identity keys 2026-09-12 18:51:42 +00:00
root 6d65cba064 no-mistakes(review): Canonicalize guest identities in GC report-only gate 2026-09-12 18:47:17 +00:00
root de1428b4ae no-mistakes(review): Harden GC gate key aliases, fail closed, fix baseline 2026-09-12 18:42:36 +00:00
root 4ea2d0309f no-mistakes(review): Fix report-only gate, dedupe list, correct fleet map 2026-09-12 18:38:19 +00:00
abiba 7b8cc5f9ac fix(disk-gc): hard guest-level report-only gate for CT 111/.129; correct stale fleet map
CT 111 (tdunna, 192.168.68.129) is Theo's box and is report-only per the captain
(2026-08-17, re-confirmed 2026-09-10). The contract defined AMBER as 'GC scheduled
for next run' and its Execution loop called gc-executor for EVERY threat with no
guest-level exclusion - so a single AMBER reading there would have scheduled apt
clean / journal vacuum / log+tmp deletion against someone else's box. The only
marker was frontmatter report_only_agents, which names an AGENT while the scan unit
is a GUEST.

- Gate the Execution loop on a guest/host-keyed report_only_guests block (guest id,
  hostname and IP all match); an excluded guest is alerted and skipped, so no
  gc-executor call is constructed for it at any level.
- Carry the ruling in the contract body next to the loop, not only in frontmatter.
- scripts/disk-gc-plan.py: executable planner that reads the contract's
  authoritative exclusion block and emits the action plan; tests/ covers it.
- Correct the stale fleet map against pvesh /cluster/resources: CT 105 -> amdpve
  (was minipve), CT 111 -> storepve (was amdpve), add guests 118/119/120, and fix
  the '15 CTs' counts (20 guests: 17 LXC + 3 QEMU VMs).
- scripts/pct-run.sh: same stale map (105/111 wrong node, 120 missing) - this is
  why pct-run 111/105 failed.
- Fold in disk-gc-ct100-probe-gap-20260911: CT 100 verifiably works through
  pct-run now that the map is correct; documented that the scanner must probe it
  like any other guest, never via a local-only path (the scanner runs inside CT 100).
2026-09-12 18:33:02 +00:00
root 3d55764799 no-mistakes(review): Split retired-alias audit into fail vs warn; fix key claim 2026-09-12 16:41:35 +00:00
root ce48070f21 no-mistakes(review): Derive model fields recursively; fix historical latency and key claims 2026-09-12 16:33:58 +00:00
root 9e581ab203 no-mistakes(review): Complete retired-alias field coverage; fix misleading example labels 2026-09-12 16:26:14 +00:00
root 9cac3589cf no-mistakes(review): Fix compression example alias; make retired aliases fail audit 2026-09-12 16:19:43 +00:00
abiba 221f9f79f3 fix(audit): reject retired aliases; sweep gpu-light/gemma-4-12b to gpu-vision
audit-hermes-config.py Rule 8 required auxiliary.vision.model and
auxiliary.web_extract.model to equal the retired 'gpu-light', so a config
adopting the live canonical 'gpu-vision' FAILED our own audit - the audit was
enforcing a dead alias (400 Invalid model name). Rule 8 now requires
gpu-vision; retired names gpu-light/crew-auto join the raw-name rejection set;
the guidance message names the live aliases.

Sweep of the remaining references: gpu-self-heal stops canonicalizing
gpu-light; hermes-config-template, hermes-agent-baseline, hermes-key-enforcement,
inference-optimization, litellm-client-timeouts and gpu-fleet now use the live
gpu-vision alias. Where a file restated model/rpm/weight/fallback state it now
points at CT 116 /opt/inference-harness/litellm_config.yaml instead of
duplicating it. koby's .129 config is report-only and recorded, not edited.

Adds tests/test_audit_hermes_config_alias.py: executes the audit CLI and asserts
gpu-vision passes while gpu-light and gemma-4-12b fail.
2026-09-12 16:11:46 +00:00
root 5f582e2c9c no-mistakes(review): fix duplicate-000 probe capture at assignment boundary 2026-09-12 14:23:35 +00:00
root b1b3b4c010 no-mistakes(review): fix kagentz A2A status logging, tests, and contract retirement 2026-09-12 14:19:13 +00:00
root 287657a77a Merge remote-tracking branch 'origin/master' into update/docker-ecosystems-20260908
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
2026-09-10 23:01:32 +00:00
root 21f9073e0b fix(daily-infra-report): use pve_probe_status in render + add regression tests (PR #64 round 2) 2026-09-10 23:01:18 +00:00
abiba-bot 8a4ee4cf5a no-mistakes(review): Pin Mumuni removal behaviorally and fix digest import bug 2026-09-10 10:18:38 +00:00
abiba 952fca9c92 fix(monitoring): retire the Mumuni leg — stop probing her decommissioned .24 deployment
Captain ruling 2026-09-10: Mumuni moved off this host onto her own container
(kagentz CT 105 on minipve, 192.168.68.14, dedicated `hermes` user) and is
monitored from her side. This host must not monitor anything Mumuni.

The stale probes fired false alerts repeatedly:
  * scripts/zulip-monitor.sh ssh'd to root@192.168.68.24 for the
    decommissioned deployment's ~/.hermes/gateway_state.json, read "unknown"
    on every run, and posted a 🔴 "Mumuni (Hermes) Zulip state: unknown" DM +
    #agent-hub stream alert each cycle.
  * scripts/daily-infra-report.py published a matching "mumuni:unknown" row in
    every digest.

Changes:
  * zulip-monitor.sh: delete the "Platform B: Hermes (Mumuni)" leg and its
    notify; keep the Zulip-server, Platform A pi/Abiba (bridge), Platform B
    Tanko and Platform C Agent Zero legs. A comment records why the leg is
    retired so it is not re-added. Also fixes SC2155 so shellcheck is clean.
  * daily-infra-report.py: delete the .24 ~/.hermes/gateway_state.json agent
    probe, its hermes --version probe, and the now-dead mumuni render branch.
    Abiba (CT 100, its own .24 address) and Tanko legs unchanged; the PVE API
    token name is untouched.
  * zulip-health.prose.md (v3.1.0): drop the Mumuni-only B4 gateway-process,
    B5 heartbeat and B6 response-delivery steps and the stale 192.168.68.24
    references; state explicitly that Mumuni is not monitored from this host.
    Tanko/Agent-Zero/bridge steps retained.
  * agent-health-check.py: correct the v2 changelog roster comment that still
    placed mumuni at .24/CT100. No behavior change — the mumuni probe was
    already absent from the AGENTS dict; v5 changelog notes the correction.

Tests: tests/test_mumuni_monitor_removal.py pins the removal structurally and
behaviorally — the shipped zulip-monitor.sh is run in a sandbox (only its LOG
constant rewritten) with stub ssh/curl on PATH; the ssh stub records every
target host, so "never reaches .24" and "no Mumuni notify even on the alert
path" are asserted from observed behavior. A mutation check (re-inject the old
leg) fails the suite, so the guarantee is not vacuous.
2026-09-10 10:11:19 +00:00
abiba-bot bbdf6c1249 no-mistakes(review): Restore quiet-mode silence; keep provenance on failure alert 2026-09-10 01:52:38 +00:00
abiba-bot 1974959cc9 no-mistakes(review): Gate wrapper checks on executed infisical; scope liveness guide 2026-09-10 01:46:06 +00:00
abiba-bot c59c9fb174 no-mistakes(review): Ignore commented infisical paths; normalize probe-model tests 2026-09-10 01:41:27 +00:00
abiba-bot 194e256ac5 no-mistakes(review): Harden health-check provenance, infisical verification, report-only JSON, tests 2026-09-10 01:35:44 +00:00
abiba-bot c66d9c1e20 fix: probe-drift round 2 — stale monitoring expectations (health check, PVE API, GPU port 80, report provenance)
Second probe-drift correction pass after #65/#66/#68. All four legs were stale
consumer expectations, not live faults.

1. scripts/agent-health-check.py (v4)
   - abiba declared pi-only runtime (harness purge): Hermes-era gateway, config
     and wrapper legs are skipped instead of failing.
   - koby declared report_only (captain ruling 2026-08-17, Rule 17): every koby
     leg is detected and reported, never counted as a fleet failure or repaired.
   - koby's CT 111 mapping corrected to storepve (.6); the old amdpve mapping
     made `pct status 111` fail and read as ct-unreachable.
   - wrapper check no longer FAILs .env-based wrappers that legitimately never
     invoke infisical (koonimo).
   - keys load in main() (load_agent_keys) so the module is importable/testable.
   - every run prints absolute execution provenance (script + cwd), in the header
     and in --json.
   Before: 6 FAILURE(S). After: 0 failures, koby reported read-only.

2. infrastructure-monitoring.prose.md
   - PVE API probe repointed from CT 116 (no pveproxy, 000) to the five real
     nodes on https://<node>:8006/api2/json/version, all 401 = alive.
   - any-HTTP-response liveness rule added (401/3xx alive; 000/timeout = DOWN).
   - LiteLLM health documented as 301 -> /litellm/health/liveliness, not bare 200.

3. gpu-monitor.prose.md
   - GPU health probes on :8080 (or router /health/unified); bare port 80 on a
     GPU host is forbidden (no listener -> false DEGRADED).
   - router /health/unified 301 -> /gpu/gpu-data documented as alive.
   - port-discipline + liveness rule + direct-fallback execution step.

4. Report provenance (all contracts)
   - docs/AUTHORING-GUIDE.md documents the rule; scripts/prose-lint.sh enforces
     that any **Report format** contract states an absolute path (pwd -P).
   - provenance added to gpu-monitor, infrastructure-monitoring, proxmox-monitor.

Tests: tests/test_probe_drift.py (23 passed); prose-lint.sh clean + shellcheck
clean; CI frontmatter validation passes. Evidence with per-leg before/after and
absolute paths: docs/probe-drift-round2-evidence.md.
2026-09-10 01:23:54 +00:00
root f57923b4fa no-mistakes(document): test shellcheck hygiene; flagged monitor contract doc drift
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Failing after 4s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Skipped
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Skipped
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Skipped
2026-09-09 10:26:35 +00:00
root 7ed2e4e923 fix(zulip-monitor): Abiba leg reads nested zulip.connected; probe failures never restart
The monitor's Abiba leg parsed the :9200/health payload at the top level
(d.get('connected',False)) while the pi Zulip extension serves connection
state NESTED at zulip.connected / zulip.last_error (verified against the
extension's startHealthServer handler and the live endpoint). PI_CONNECTED
was therefore always False and every run took the DISCONNECTED path,
restarting a healthy bot: pm2 restarts=8 with the process created
2026-09-09T09:35:09Z, four ❌ Abiba verdicts today (04:23/05:35/06:55/09:35
UTC) and zero ✅, while the Zulip server answered HTTP 200 and the bot kept
heartbeating. The watchdog was the fault, not the connection.

Changes (Abiba leg only; every other leg byte-identical):
- Read the real nested shape: zulip.connected, zulip.last_error and
  zulip.messages_processed. The old retry_count branch is DROPPED — the
  payload exposes no retry counter (the extension keeps retryCount internal
  and never serialises it), so the branch is fabricated and cannot stay.
- Fail-safe restart decision: a fetch error (HTTP 000), non-2xx response,
  empty/unparseable body, or payload missing a boolean zulip.connected is a
  clearly-labelled PROBE FAILURE (🟠 alert + ⚠️ log line naming reason and
  HTTP code) and NEVER calls pm2 restart. pm2 restart runs only on
  affirmative zulip.connected=false (🔴/❌ path unchanged in wording).
  connected=true with last_error keeps the degraded 🟡 warn-no-restart path.

Regression test (new tests/zulip-monitor-abiba.sh, 36 checks): extracts the
real Abiba leg from between the # -- abiba-leg-start/-end markers in the
shipped script and executes it verbatim with stubbed curl/notify/pm2 against
fixtures of the real payload shape — asserts connected=true → no restart,
connected=false → restart, and empty/garbage/missing-key/non-boolean/HTTP
000/HTTP 500 bodies → probe-failure alert with zero restarts. The suite
fails loudly on the pre-fix base and on a marker-intact top-level-parse
variant, so this class of bug cannot return silently.
2026-09-09 09:51:46 +00:00