Commit Graph
529 Commits
Author SHA1 Message Date
root 39209c7ac9 fix: untrack host-disk-bands.json and document gitignored status
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 8s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s
The state file is runtime state (rewritten every scan), so tracking it in git
means:
- every executor's clone becomes permanently dirty after one run
- a scan in one clone produces a merge conflict with a scan in another
- the committed baseline can be stale in a way nobody notices

Added to .gitignore and removed from the index. Contract updated to say
'the state file lives at <abs path> and is gitignored runtime state - the
scanner creates it on first run'.
2026-09-16 00:45:46 +00:00
root 8ed3b9c606 fix: implement host-filesystem state file with transition detection
The contract said the state file was written after every scan, but the script
had no state-file logic at all. This PR adds:

1. Host filesystem scanning (probe_host_filesystems) - probes df on all PVE nodes
2. State file I/O (read_state_file/write_state_file) - absolute path from script location
3. Band classification (classify_band) - HOST-WARN/AMBER/RED thresholds
4. Transition detection (detect_transitions) - alerts on escalation/recovery
5. CLI flags (--hosts-only, --guests-only) to control which parts run

The contract now specifies the state file path resolves from the script's own
location (not CWD-relative), so two different execution contexts cannot write
to two different places.
2026-09-16 00:41:34 +00:00
abiba-bot 6c616a9e58 Merge pull request 'fix(disk-gc): host filesystem bands, named volumes, report-only, and state-change alerts' (#105) from fix/host-filesystem-thresholds-20260915 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 1s
2026-09-15 14:40:20 +00:00
root b9b1712ac6 fix: make host escalations state-change driven
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
A volume alerts ONCE when it enters a higher band (GREEN->WARN, WARN->AMBER,
AMBER->RED) and ONCE when it drops back down (recovery notice). While it stays
in the same band, it is reported in the scan output only — no DM, no channel
alert. This stops the same 96% easystore2 from re-DMing the owner on every 6h
scan.

State lives in a small JSON state file (state/host-disk-bands.json), keyed by
host/volume -> last-seen band. The scanner reads the prior band, compares to the
current band, and DMs only on a transition; the state file is written after
every scan. Chosen over a periodic digest because the scan already runs every
6h and a transition is genuinely new, actionable state.

First-run behavior: when the state file does not yet exist, the current band of
every volume is recorded as baseline WITHOUT alerting — a first run would
otherwise DM every already-elevated volume at once.

Report-only restriction and volume-naming output kept exactly as-is.
2026-09-15 14:25:49 +00:00
root 6fb411613e fix: add host filesystem thresholds to disk-gc contract
Add separate threat bands for HOST filesystems (distinct from guest bands):
- HOST-WARN at 85%: name volume + % + absolute free space in scan output
- HOST-AMBER at 90%: flag for owner attention, Zulip DM
- HOST-RED at 95%: flag for immediate owner attention, Zulip DM + channel alert

Volume naming rule: every host line MUST name the volume and what lives on it.
Action classes by volume type:
- host-root: near full = real risk (backup staging, thin-pool metadata)
- media (/media/*): near full = capacity decision for owner, never auto-delete
- pbs-datastore (tank): near full = breaks Proxmox Backup Server

Report-only restriction: no automatic deletion of media or datastore content ever.

Justification (measured 2026-09-15): storepve /media/easystore2 at 96% was
reported but never banded or acted on. Two incidents this weekend showed the
host filesystem is the thing that breaks, not the guest's.

Added HOST-WARN/AMBER/RED alert templates.
Added report-only execution rule for host filesystems.
2026-09-15 14:15:53 +00:00
abiba-bot 4bdd88613b Merge pull request 'docs: pm2/spoton AS-BUILT correction and the real Prometheus node coverage' (#104) from fix/contract-corrections-20260915 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 2s
2026-09-15 13:14:39 +00:00
root bc7a55122f fix: pm2 contract corrections
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
pm2-self-heal.prose.md:
- Add AS-BUILT note: gpu-monitor is systemd-managed, NOT PM2
- gpu-watchdog is decommissioned and folded into gpu-monitor.service
- gitea-runner is KEPT; abiba-zulip is KEPT (online for days)
- spoton-service was deleted; live PM2 set is 4 processes
- Preserve historical context for crash-loop guard

litellm-health.prose.md:
- Correct Prometheus node coverage: 6 nodes (.4/.5/.6/.9/.12/.15:9100)
- Note .4:9100 is DEAD target (no route, down for weeks)
- Clarify this does not read as 6 healthy nodes
2026-09-15 13:05:32 +00:00
abiba-bot 713b9ce80c Merge pull request 'Remove client deliverable from the contracts repo (process fix)' (#77) from cleanup/remove-scot-deliverable-20260912 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 1s
2026-09-15 12:57:59 +00:00
abiba-bot b451d6f81a Merge branch 'master' into cleanup/remove-scot-deliverable-20260912
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 9s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
2026-09-15 12:53:37 +00:00
abiba-bot fe6eb35291 Merge pull request 'ci: make PR validation unconditional - the paths filter silently skipped whole classes of pull requests' (#103) from fm/ci-paths-filter-skips-deliverables-prs-20260915 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 2s
2026-09-15 12:53:20 +00:00
abiba-bot 40cf057370 Merge pull request 'fix(backups): document the thin-pool headroom and staging-directory preconditions, with the incidents that justify them' (#97) from fix/backup-safety-preconditions-20260915 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 1s
2026-09-15 12:19:19 +00:00
abiba-bot 5053a33e2c ci: make PR validation trigger unconditional (remove paths filter)
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 16s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
The pr-pipeline workflow filtered both push and pull_request on
paths (**.prose.md, scripts/**.sh, **.yaml, **.yml). A PR whose diff
touched none of those paths — e.g. PR #77, deliverables/-only —
produced no Gitea Actions run at all, so validate/lint/ai-review and
the merge gate were silently skipped.

Remove the paths filter from both triggers and record in the file
that the trigger is intentionally unfiltered. No job, step, needs,
if or command is changed.
2026-09-15 12:14:40 +00:00
root b4b5321011 fix: add hostname resolution warning and fix dmsetup field documentation
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 8s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s
Add warning that bare PVE hostnames (acerpve, amdpve, etc.) resolve to VPS
via *.dns.sysloggh.net wildcard, not to actual nodes. List IP addresses:
- acerpve 192.168.68.9
- amdpve 192.168.68.15
- storepve 192.168.68.6
- minipve 192.168.68.12
- ocupve 192.168.68.5

Update acerpve example in backup preflight to include address (192.168.68.9).

Fix dmsetup comment to show full field order:
=start =length =thin-pool =transaction-id
=metadata_used/metadata_total =data_used/data_total
remaining fields are flags

Make it clear lvs command is the primary source for percentages, dmsetup is only for error-state check.

Incidents now include addresses: acerpve (192.168.68.9) and amdpve (192.168.68.15).
2026-09-15 12:09:59 +00:00
root 69940bc9eb fix: correct backup preflight commands - lvs pve/data and dmsetup field documentation
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 2s
Fix two errors in the PREFLIGHT section (measured on acerpve 2026-09-15):
1. lvs -o ... pve/data (not pve-data-tpool) - this is the PRIMARY check that yields percentages directly
   - Quote the acerpve example: data 29.95% 1.22% <816.21g
2. dmsetup status pve-data-tpool - document fields correctly:
   -  = transaction ID (99), NOT data_percent
   -  = metadata used/total blocks
   -  = data used/total sectors
   - Show how to derive percentages if needed
3. Keep the error-state check (grep -q 'Error|Fail') - this is how the incident presented

Everything else stays: 1777 tmpdir requirement with EACCES symptom, incidents as rationale,
GPU-host fact, --output-format json rule, honest note that metadata/snapshot pressure is unproven.
2026-09-15 11:42:10 +00:00
root 65eaffe1c6 fix: add backup safety preconditions - thin-pool headroom and tmpdir 1777
Add documented preflight checks for VM/CT backups on LVM thin-pool hosts:
- dmsetup status pve-data-tpool + lvs to verify data_percent < 90% and metadata_percent < 70%
- Exit 1 if pool shows Error/Fail state (takes down entire VG including host root)
- tmpdir must be mode 1777 (world-traversable) for vzdump archive step
- --output-format json for tasks started from truncating shells
- Document two incidents: acerpve thin-pool VM 101 (twice on 2026-09-13) and amdpve 0700 tmpdir (2026-09-14)
- Note metadata/snapshot-pressure hypothesis is UNPROVEN; preflight is the control
- Document GPU-host fact: VM 101 (llm-gpu) and VM 103 (ocu-llm) have no scheduled backup
2026-09-15 11:37:41 +00:00
abiba-bot b386bd0c19 Merge pull request 'fix(keys): correct the key-lifecycle contracts to the measured truth' (#95) from fix/key-expiry-enforcement-20260915 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 17s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 3s
2026-09-15 05:37:45 +00:00
root 8a4dd08b05 fix: capture 401/403 response body and key alias for credential faults
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s
- get_response_body() returns first 200 chars of response body (single line)
- On 401/403 model probe: report code + body + key_alias
- Monitor key alias: monitor-20260813 (from /etc/litellm-monitor.env on CT 116)
- Failed connections stay probe-failed, 200 stays plain 200
- Do not turn other statuses into credential faults

Signed-off-by: Abiba
2026-09-15 05:11:07 +00:00
root 88b6decb31 fix: remove false Infisical claim - master key NOT in infrastructure project
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
- Replace Infisical retrieval path with proven docker exec + .env note
- State explicitly that master key is NOT in Infisical project=infrastructure
- Keep the live-key check and never-trust-a-literal instruction
- All other corrections from PR #95 preserved

Signed-off-by: Abiba
2026-09-15 04:42:31 +00:00
root 7bf9f78fc6 fix: gpu-dense probe timeout handling - report probe-failed with kind, not service verdict
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
- probe_http now returns (code, failure_kind) tuple
- Model probes report 'probe-failed: <model> <kind> (Ns timeout)' on 000
- Do not assert a service verdict from a failed probe
- 30s timeout for single-host aliases (RTX 3090 needs long warmup/prefill)
- 60s timeout for syslog-auto pool alias with retry on 000

Signed-off-by: Abiba
2026-09-15 04:22:46 +00:00
root dd9e68329e fix: correct infisical --plain flag and use verified localhost:4000 endpoint
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 2s
- Replace broken --plain flag (prints nothing on CLI 0.43.110) with awk parsing
- Note that --plain is broken so nobody fixes it back
- Replace unproven nginx path with verified direct endpoint http://127.0.0.1:4000/key/list

Signed-off-by: Abiba
2026-09-15 03:05:30 +00:00
root cd9ec6a0df fix: correct key-lifecycle contracts to match measured reality
- hermes-key-enforcement.prose.md:
  - State that expiry must be set EXPLICITLY at creation with duration
  - Record that config default is NOT honoured by LiteLLM 1.99.1
  - Describe daily audit as AUDIT-ONLY (reports non-expiring and soon-to-expire)
  - State that renewal is NOT implemented
  - Document exclusions: abiba-pi and all crewmate keys stay WITHOUT expiry
  - koby is report-only

- litellm-api-keys.prose.md:
  - Replace literal master key with retrieval path (docker exec + infisical)
  - State that literal values must never be trusted again (key rotates)
  - Add live-key check (200 from /key/list)

Signed-off-by: Abiba
2026-09-15 02:59:19 +00:00
abiba-bot 1d20fbaa7f Merge pull request 'fix(agent-health): gateway liveness reports a probe failure as a probe failure, never as an agent outage' (#93) from fix/agent-health-gateway-leg-deterministic-20260914 into master 2026-09-14 16:45:25 +00:00
root 2f961d7e7a fix: agent-health-check gateway leg — deterministic probe-failed reporting
Per defect report 1154.msg (agent-health-gateway-leg-flap-20260913):
1. Add _ssh_retry() helper with one retry at longer timeout (25s)
2. Name probe target explicitly: ssh {user}@{host} <command>
3. Print probe-failed when first attempt fails, then retry
4. Only declare gateway-down after retry fails
5. Make Koby report-only explicit in output

Changes:
- check_agents(): all gateway probes now use _ssh_retry()
- All output lines name the probe target (ssh host:port)
- Koby's report-only status is explicit in output
- Never print bare "gateway down" — always name target and failure kind

Verified: koonimo shows "probe-failed" on first attempt (transient SSH),
retries at 25s, succeeds, reports ✅ koonimo: gw=running
2026-09-14 15:21:36 +00:00
abiba-bot 4fe4f3621d Merge pull request 'fix(monitoring): probe precision - any HTTP status means alive, failed probes never become service verdicts' (#92) from fix/probe-precision-20260914 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 1s
2026-09-14 15:07:35 +00:00
root 86d2987ad8 fix: probe precision — add retry + probe-failed reporting to zulip-health
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
Per defect report 1150.msg:
- Add standing probe rules section (2026-09-14)
- Step 1 (Zulip API): retry once at 25s on 000, print target + code
- Step 2 (Platform A): retry once at 25s on 000, print target + code, note must run on Abiba host
- Apply same shape as infrastructure-monitoring: any HTTP status = ALIVE; only 000/timeout = probe-failed

Verified: all probes now return real HTTP codes (Zulip API 200, Platform A 200, Tanko 401, Agent Zero 401)
2026-09-14 14:53:57 +00:00
root efe9381283 fix: probe precision — correct GPU exporter path, Grafana port, add retry + probe-failed reporting
Per defect report 1150.msg:
- GPU exporters: probe /metrics (Prometheus scrape target), not bare /
- Grafana: correct port 3001 (not 3000)
- All probes: print target name + full URL + HTTP code
- All probes: retry once at 25s on 000/timeout
- Apply standing rules: any HTTP status = ALIVE; only 000/timeout/refused = probe-failed

Verified: all probes now return real HTTP codes (GPU 200, Grafana 200, Router 200, LiteLLM 301)
2026-09-14 14:44:37 +00:00
abiba-bot 6a1c4967db Merge pull request 'fix(disk-gc): deterministic reachability verdict and correctly labelled disk figures' (#91) from fix/disk-gc-probe-deterministic-20260914 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 11s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 1s
2026-09-14 14:43:43 +00:00
root 5ed6f8179c fix: add os import for HELPER_PCT_RUN.readable() check
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 3s
2026-09-14 14:14:15 +00:00
root b9322973ce fix disk-gc-scan: CWD independence (FIX A) and df column parsing (FIX B) 2026-09-14 14:05:06 +00:00
root fb1916707d document deterministic disk-gc-scan.py in contract 2026-09-14 13:20:26 +00:00
mumuni-bot 0c7d7be5ad ci: re-trigger pipeline (no statuses recorded on 9312c1a8 head) 2026-09-14 13:19:33 +00:00
root d28df4f4de add disk-gc-scan.py: deterministic fleet disk probe with per-guest access methods 2026-09-14 13:14:10 +00:00
abiba-bot cb26ee06d6 Merge pull request 'fix(hermes): separate POLICY observations from FAULT findings in the audit contracts' (#90) from fix/hermes-violation-classification-20260914 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 1s
2026-09-14 12:50:15 +00:00
root 9a789ab76d fix(hermes): separate policy observations from fault findings
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 3s
Contract design defect: 'uses a non-harness provider' (POLICY) and
'cannot authenticate' (FAULT) were printed as the same violation class.
A policy observation must never be phrased as if the agent were broken.

Changes:
1. Added 'Violation Classification' section to all three contracts
2. Separated POLICY (observation only) from FAULT (requires request-level evidence)
3. Rules:
   - Do NOT infer runtime credential resolution from config text alone
   - Require request-level evidence before calling a FAULT: observed auth failure
     or absence of successful calls
   - If calls are succeeding, output is 'POLICY: uses <provider> directly; calls
     succeeding' - not a violation
   - State what you OBSERVED, not what the field implies

Files changed (3):
- hermes-key-enforcement.prose.md
- hermes-config-template.prose.md
- hermes-agent-baseline.prose.md
2026-09-14 12:32:06 +00:00
abiba-bot 71ceda0042 Merge pull request 'fix(hermes): reachability verdict must not come from the remote command's exit code' (#89) from fix/hermes-reachability-20260914 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 8s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 8s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 1s
2026-09-14 12:20:09 +00:00
root 4e34b7a2a2 fix(hermes): wire all three contracts to reachability helper
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
The helper was correct but dead code - nothing called it. This commit:
1. Makes the helper runnable standalone: scripts/hermes-reachability-check.sh <host> <pattern> <path>
2. Wires all THREE contracts to it:
   - hermes-key-enforcement.prose.md
   - hermes-config-template.prose.md
   - hermes-agent-baseline.prose.md
3. Each contract now explicitly instructs to run the helper and interpret the three outcomes
4. States that the bug this replaces was deriving reachability from the remote grep's exit code

Files changed (4):
- scripts/hermes-reachability-check.sh (standalone mode added)
- hermes-key-enforcement.prose.md (reachability section added)
- hermes-config-template.prose.md (reachability section added)
- hermes-agent-baseline.prose.md (reachability section added)
2026-09-14 12:01:32 +00:00
root f9f6661dd5 fix(hermes): add shared reachability check helper
Bug: The reachability check used 'ssh ... grep ... || echo unreachable',
which conflated grep's 'no matches found' (exit 1) with SSH failure.
This caused clean hosts to be reported as unreachable.

Fix: Add scripts/hermes-reachability-check.sh with the pattern:
  out=$(ssh -o BatchMode=yes root@HOST "grep ... 2>/dev/null; true")
  if [ $? -ne 0 ]; then verdict="unreachable"
  elif [ -n "$out" ]; then verdict="violation: $out"
  else verdict="compliant"
  fi

This correctly distinguishes:
- SSH failure (connection/auth/route) -> unreachable
- SSH success + grep found matches -> violation
- SSH success + grep found nothing -> compliant

Evidence: 3 of 4 hosts (Tanko, Mumuni, Koonimo) were reported as
'unreachable' when they were actually compliant. Only Koby (.129)
has a real finding (plaintext key in state snapshot).

Used by: hermes-key-enforcement, hermes-config-template,
hermes-agent-baseline contracts.
2026-09-14 11:55:41 +00:00
abiba-bot cb36ff1ea5 Merge pull request 'fix(monitoring): litellm-health script - pool-alias timeout and leaked key-list debug output' (#88) from fix/litellm-health-timeout-and-debug-20260914 into master 2026-09-14 03:34:58 +00:00
root 05366bd58d Fix litellm-health-check.py robustness defects
1. Timeout fix for pool alias (syslog-auto):
   - Single-host aliases (gpu-dense, gpu-vision, strix-moe): 30s timeout
   - Pool alias (syslog-auto): 60s timeout, retry once on 000 before failing
   - Cold first request to pool alias can take ~13s; 10s was too short

2. Remove DEBUG prints from output:
   - Removed 'DEBUG: keylen=...' and 'DEBUG: response=...' lines
   - These leaked key inventory to status logs
   - Success output now shows only counts (e.g., 'Admin Key List: 10 keys')
2026-09-14 03:24:30 +00:00
abiba-bot e648b5ac0e Merge pull request 'feat(monitoring): add scripts/litellm-health-check.py so the litellm-health contract is executed, not improvised' (#87) from fix/litellm-health-executor-script-20260913 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 8s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 2s
2026-09-14 03:20:36 +00:00
root 38d7e8b064 Update litellm-health contract to mandate executor script
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
Add 'Executor Script' section:
- Run scripts/litellm-health-check.py from the clone
- Hand-rolled probes not acceptable substitute
- Backend-edge checks use internal IP 192.168.68.116, not public URL
- Docker Stats fetched from CT 116 host (127.0.0.1:9324/metrics)
- Admin Key List requires proper quoting for SSH commands
2026-09-14 03:15:42 +00:00
root e94fadfa60 Add litellm-health-check.py: standardized health check script for LiteLLM fleet monitoring
Implements all 11 checks from litellm-health contract:
- Liveliness, Containers, Prometheus, Grafana health probes
- Model probes for gpu-dense, gpu-vision, strix-moe, syslog-auto
- Admin API key list (10 keys), GitHub status, Docker Stats metrics

Fixed quoting for SSH commands and response parsing (dict with 'keys' field).
Backend edge uses internal IP 192.168.68.116, not public URL.
Docker Stats fetched from CT 116 host itself (127.0.0.1:9324/metrics).

All 11 checks passing consistently.
2026-09-14 03:15:27 +00:00
abiba-bot ba76f2c7d3 Merge pull request 'fix(contracts): abiba default model row + disk-gc access-method note' (#86) from fix-litellm-health-keylist-20260913 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 10s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 9s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 3s
2026-09-13 19:09:19 +00:00
root 7906b2d52d fix: change abiba default model from deepseek-v4-pro to syslog-auto
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Failing after 13m52s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Skipped
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Skipped
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Skipped
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Skipped
The abiba-zulip-restore.prose.md file incorrectly claimed the default model
was 'deepseek-v4-pro', but /root/.pi/agent/settings.json declares
'defaultModel: syslog-auto' and 'defaultProvider: syslog-harness'.

deepseek-v4-pro is not a servable LiteLLM model name for this fleet.
The only live models are: syslog-auto, gpu-dense, gpu-vision, strix-moe.

Note: The firstmate/pi session talks to DeepSeek through its own
provider (auth.json), NOT through LiteLLM, so no LiteLLM change is
needed to keep firstmate's deepseek usage working.
2026-09-13 18:55:24 +00:00
root c81cf5b6f0 fix: disk-gc-threat-response - correct access methods for kagentz and docker-vm
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 8s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 12s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s
- kagentz (105): use ssh root@kagentz (hostname), NOT pct exec 105 (shows loop0 59G, not real 99G)
- docker-vm (109): use ssh root@192.168.68.7 (correct QEMU VM access)
- abiba (100): use ssh root@abiba (hostname)
- syslog-api (116): use pct-run 116 (correct)

Verified:
  kagentz: 99G 8.9G 86G 10% /
  docker-vm: 158G 17G 135G 11% /
  abiba: 59G 13G 44G 23% /
  syslog-api: 40G 13G 25G 34% /
2026-09-13 04:30:28 +00:00
abiba-bot ef168d9690 Merge pull request 'fix(litellm-health): run the key-list admin call on the gateway host, not inside the container' (#84) from fix-litellm-health-registry-20260912 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 1s
2026-09-13 03:20:04 +00:00
root edcf465831 fix: litellm-health step 8 - run key/list on CT 116 host, not in container
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s
- harness-litellm container has no curl/wget, so docker exec harness-litellm curl returns empty
- Fix: run curl on CT 116 host (ssh root@192.168.68.116 then curl)
- Add admin-call-failed label for empty/unparseable responses
- Verify: 10 keys found (host-side curl), NO-CURL confirmed in container
2026-09-13 03:10:09 +00:00
abiba-bot 89651cf37c Merge pull request 'fix(monitoring): make credential sourcing explicit and fail loudly when missing' (#83) from fix-litellm-health-registry-20260912 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 2s
2026-09-12 23:22:43 +00:00
root 6f40a3be60 fix: address credential-sourcing review findings
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
- infrastructure-monitoring: Correct Zulip key path to /etc/litellm-monitor.env
  (not /etc/zulip-bot.env which doesn't exist); add credential-missing check;
  remove unused monitor_key variable

- litellm-health: Clarify that syslog-auto is a fallback pool alias, not a
  step 7 probe; monitor key must be scoped for all four aliases
2026-09-12 23:17:27 +00:00
root d9368467ff fix: credential-sourcing fix for litellm-health and infrastructure-monitoring
- litellm-health.prose.md: Make monitor key retrieval explicit (ssh from CT 100 to CT 116)
  with executable commands; add syslog-auto alias; document credential-missing failure
  condition (not bare 401 or 0 keys)

- infrastructure-monitoring.prose.md: Fix Zulip POST probe to retrieve keys from CT 116
  via ssh instead of sourcing local env file that doesn't exist on executor host

- Verify model inference probes return 200 for gpu-dense, gpu-vision, strix-moe,
  syslog-auto with corrected credential retrieval
2026-09-12 23:13:24 +00:00