Compare commits

...
Author SHA1 Message Date
abiba-bot 5053a33e2c ci: make PR validation trigger unconditional (remove paths filter)
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 16s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
The pr-pipeline workflow filtered both push and pull_request on
paths (**.prose.md, scripts/**.sh, **.yaml, **.yml). A PR whose diff
touched none of those paths — e.g. PR #77, deliverables/-only —
produced no Gitea Actions run at all, so validate/lint/ai-review and
the merge gate were silently skipped.

Remove the paths filter from both triggers and record in the file
that the trigger is intentionally unfiltered. No job, step, needs,
if or command is changed.
2026-09-15 12:14:40 +00:00
abiba-bot b386bd0c19 Merge pull request 'fix(keys): correct the key-lifecycle contracts to the measured truth' (#95) from fix/key-expiry-enforcement-20260915 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 17s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 3s
2026-09-15 05:37:45 +00:00
root 8a4dd08b05 fix: capture 401/403 response body and key alias for credential faults
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s
- get_response_body() returns first 200 chars of response body (single line)
- On 401/403 model probe: report code + body + key_alias
- Monitor key alias: monitor-20260813 (from /etc/litellm-monitor.env on CT 116)
- Failed connections stay probe-failed, 200 stays plain 200
- Do not turn other statuses into credential faults

Signed-off-by: Abiba
2026-09-15 05:11:07 +00:00
root 88b6decb31 fix: remove false Infisical claim - master key NOT in infrastructure project
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
- Replace Infisical retrieval path with proven docker exec + .env note
- State explicitly that master key is NOT in Infisical project=infrastructure
- Keep the live-key check and never-trust-a-literal instruction
- All other corrections from PR #95 preserved

Signed-off-by: Abiba
2026-09-15 04:42:31 +00:00
root 7bf9f78fc6 fix: gpu-dense probe timeout handling - report probe-failed with kind, not service verdict
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
- probe_http now returns (code, failure_kind) tuple
- Model probes report 'probe-failed: <model> <kind> (Ns timeout)' on 000
- Do not assert a service verdict from a failed probe
- 30s timeout for single-host aliases (RTX 3090 needs long warmup/prefill)
- 60s timeout for syslog-auto pool alias with retry on 000

Signed-off-by: Abiba
2026-09-15 04:22:46 +00:00
root dd9e68329e fix: correct infisical --plain flag and use verified localhost:4000 endpoint
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 2s
- Replace broken --plain flag (prints nothing on CLI 0.43.110) with awk parsing
- Note that --plain is broken so nobody fixes it back
- Replace unproven nginx path with verified direct endpoint http://127.0.0.1:4000/key/list

Signed-off-by: Abiba
2026-09-15 03:05:30 +00:00
root cd9ec6a0df fix: correct key-lifecycle contracts to match measured reality
- hermes-key-enforcement.prose.md:
  - State that expiry must be set EXPLICITLY at creation with duration
  - Record that config default is NOT honoured by LiteLLM 1.99.1
  - Describe daily audit as AUDIT-ONLY (reports non-expiring and soon-to-expire)
  - State that renewal is NOT implemented
  - Document exclusions: abiba-pi and all crewmate keys stay WITHOUT expiry
  - koby is report-only

- litellm-api-keys.prose.md:
  - Replace literal master key with retrieval path (docker exec + infisical)
  - State that literal values must never be trusted again (key rotates)
  - Add live-key check (200 from /key/list)

Signed-off-by: Abiba
2026-09-15 02:59:19 +00:00
abiba-bot 1d20fbaa7f Merge pull request 'fix(agent-health): gateway liveness reports a probe failure as a probe failure, never as an agent outage' (#93) from fix/agent-health-gateway-leg-deterministic-20260914 into master 2026-09-14 16:45:25 +00:00
root 2f961d7e7a fix: agent-health-check gateway leg — deterministic probe-failed reporting
Per defect report 1154.msg (agent-health-gateway-leg-flap-20260913):
1. Add _ssh_retry() helper with one retry at longer timeout (25s)
2. Name probe target explicitly: ssh {user}@{host} <command>
3. Print probe-failed when first attempt fails, then retry
4. Only declare gateway-down after retry fails
5. Make Koby report-only explicit in output

Changes:
- check_agents(): all gateway probes now use _ssh_retry()
- All output lines name the probe target (ssh host:port)
- Koby's report-only status is explicit in output
- Never print bare "gateway down" — always name target and failure kind

Verified: koonimo shows "probe-failed" on first attempt (transient SSH),
retries at 25s, succeeds, reports ✅ koonimo: gw=running
2026-09-14 15:21:36 +00:00
abiba-bot 4fe4f3621d Merge pull request 'fix(monitoring): probe precision - any HTTP status means alive, failed probes never become service verdicts' (#92) from fix/probe-precision-20260914 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 1s
2026-09-14 15:07:35 +00:00
root 86d2987ad8 fix: probe precision — add retry + probe-failed reporting to zulip-health
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
Per defect report 1150.msg:
- Add standing probe rules section (2026-09-14)
- Step 1 (Zulip API): retry once at 25s on 000, print target + code
- Step 2 (Platform A): retry once at 25s on 000, print target + code, note must run on Abiba host
- Apply same shape as infrastructure-monitoring: any HTTP status = ALIVE; only 000/timeout = probe-failed

Verified: all probes now return real HTTP codes (Zulip API 200, Platform A 200, Tanko 401, Agent Zero 401)
2026-09-14 14:53:57 +00:00
root efe9381283 fix: probe precision — correct GPU exporter path, Grafana port, add retry + probe-failed reporting
Per defect report 1150.msg:
- GPU exporters: probe /metrics (Prometheus scrape target), not bare /
- Grafana: correct port 3001 (not 3000)
- All probes: print target name + full URL + HTTP code
- All probes: retry once at 25s on 000/timeout
- Apply standing rules: any HTTP status = ALIVE; only 000/timeout/refused = probe-failed

Verified: all probes now return real HTTP codes (GPU 200, Grafana 200, Router 200, LiteLLM 301)
2026-09-14 14:44:37 +00:00
abiba-bot 6a1c4967db Merge pull request 'fix(disk-gc): deterministic reachability verdict and correctly labelled disk figures' (#91) from fix/disk-gc-probe-deterministic-20260914 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 11s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 1s
2026-09-14 14:43:43 +00:00
root 5ed6f8179c fix: add os import for HELPER_PCT_RUN.readable() check
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 3s
2026-09-14 14:14:15 +00:00
root b9322973ce fix disk-gc-scan: CWD independence (FIX A) and df column parsing (FIX B) 2026-09-14 14:05:06 +00:00
root fb1916707d document deterministic disk-gc-scan.py in contract 2026-09-14 13:20:26 +00:00
root d28df4f4de add disk-gc-scan.py: deterministic fleet disk probe with per-guest access methods 2026-09-14 13:14:10 +00:00
abiba-bot cb26ee06d6 Merge pull request 'fix(hermes): separate POLICY observations from FAULT findings in the audit contracts' (#90) from fix/hermes-violation-classification-20260914 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 1s
2026-09-14 12:50:15 +00:00
root 9a789ab76d fix(hermes): separate policy observations from fault findings
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 3s
Contract design defect: 'uses a non-harness provider' (POLICY) and
'cannot authenticate' (FAULT) were printed as the same violation class.
A policy observation must never be phrased as if the agent were broken.

Changes:
1. Added 'Violation Classification' section to all three contracts
2. Separated POLICY (observation only) from FAULT (requires request-level evidence)
3. Rules:
   - Do NOT infer runtime credential resolution from config text alone
   - Require request-level evidence before calling a FAULT: observed auth failure
     or absence of successful calls
   - If calls are succeeding, output is 'POLICY: uses <provider> directly; calls
     succeeding' - not a violation
   - State what you OBSERVED, not what the field implies

Files changed (3):
- hermes-key-enforcement.prose.md
- hermes-config-template.prose.md
- hermes-agent-baseline.prose.md
2026-09-14 12:32:06 +00:00
abiba-bot 71ceda0042 Merge pull request 'fix(hermes): reachability verdict must not come from the remote command's exit code' (#89) from fix/hermes-reachability-20260914 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 8s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 8s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 1s
2026-09-14 12:20:09 +00:00
root 4e34b7a2a2 fix(hermes): wire all three contracts to reachability helper
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
The helper was correct but dead code - nothing called it. This commit:
1. Makes the helper runnable standalone: scripts/hermes-reachability-check.sh <host> <pattern> <path>
2. Wires all THREE contracts to it:
   - hermes-key-enforcement.prose.md
   - hermes-config-template.prose.md
   - hermes-agent-baseline.prose.md
3. Each contract now explicitly instructs to run the helper and interpret the three outcomes
4. States that the bug this replaces was deriving reachability from the remote grep's exit code

Files changed (4):
- scripts/hermes-reachability-check.sh (standalone mode added)
- hermes-key-enforcement.prose.md (reachability section added)
- hermes-config-template.prose.md (reachability section added)
- hermes-agent-baseline.prose.md (reachability section added)
2026-09-14 12:01:32 +00:00
root f9f6661dd5 fix(hermes): add shared reachability check helper
Bug: The reachability check used 'ssh ... grep ... || echo unreachable',
which conflated grep's 'no matches found' (exit 1) with SSH failure.
This caused clean hosts to be reported as unreachable.

Fix: Add scripts/hermes-reachability-check.sh with the pattern:
  out=$(ssh -o BatchMode=yes root@HOST "grep ... 2>/dev/null; true")
  if [ $? -ne 0 ]; then verdict="unreachable"
  elif [ -n "$out" ]; then verdict="violation: $out"
  else verdict="compliant"
  fi

This correctly distinguishes:
- SSH failure (connection/auth/route) -> unreachable
- SSH success + grep found matches -> violation
- SSH success + grep found nothing -> compliant

Evidence: 3 of 4 hosts (Tanko, Mumuni, Koonimo) were reported as
'unreachable' when they were actually compliant. Only Koby (.129)
has a real finding (plaintext key in state snapshot).

Used by: hermes-key-enforcement, hermes-config-template,
hermes-agent-baseline contracts.
2026-09-14 11:55:41 +00:00
abiba-bot cb36ff1ea5 Merge pull request 'fix(monitoring): litellm-health script - pool-alias timeout and leaked key-list debug output' (#88) from fix/litellm-health-timeout-and-debug-20260914 into master 2026-09-14 03:34:58 +00:00
root 05366bd58d Fix litellm-health-check.py robustness defects
1. Timeout fix for pool alias (syslog-auto):
   - Single-host aliases (gpu-dense, gpu-vision, strix-moe): 30s timeout
   - Pool alias (syslog-auto): 60s timeout, retry once on 000 before failing
   - Cold first request to pool alias can take ~13s; 10s was too short

2. Remove DEBUG prints from output:
   - Removed 'DEBUG: keylen=...' and 'DEBUG: response=...' lines
   - These leaked key inventory to status logs
   - Success output now shows only counts (e.g., 'Admin Key List: 10 keys')
2026-09-14 03:24:30 +00:00
abiba-bot e648b5ac0e Merge pull request 'feat(monitoring): add scripts/litellm-health-check.py so the litellm-health contract is executed, not improvised' (#87) from fix/litellm-health-executor-script-20260913 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 8s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 2s
2026-09-14 03:20:36 +00:00
root 38d7e8b064 Update litellm-health contract to mandate executor script
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
Add 'Executor Script' section:
- Run scripts/litellm-health-check.py from the clone
- Hand-rolled probes not acceptable substitute
- Backend-edge checks use internal IP 192.168.68.116, not public URL
- Docker Stats fetched from CT 116 host (127.0.0.1:9324/metrics)
- Admin Key List requires proper quoting for SSH commands
2026-09-14 03:15:42 +00:00
root e94fadfa60 Add litellm-health-check.py: standardized health check script for LiteLLM fleet monitoring
Implements all 11 checks from litellm-health contract:
- Liveliness, Containers, Prometheus, Grafana health probes
- Model probes for gpu-dense, gpu-vision, strix-moe, syslog-auto
- Admin API key list (10 keys), GitHub status, Docker Stats metrics

Fixed quoting for SSH commands and response parsing (dict with 'keys' field).
Backend edge uses internal IP 192.168.68.116, not public URL.
Docker Stats fetched from CT 116 host itself (127.0.0.1:9324/metrics).

All 11 checks passing consistently.
2026-09-14 03:15:27 +00:00
abiba-bot ba76f2c7d3 Merge pull request 'fix(contracts): abiba default model row + disk-gc access-method note' (#86) from fix-litellm-health-keylist-20260913 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 10s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 9s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 3s
2026-09-13 19:09:19 +00:00
root 7906b2d52d fix: change abiba default model from deepseek-v4-pro to syslog-auto
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Failing after 13m52s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Skipped
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Skipped
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Skipped
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Skipped
The abiba-zulip-restore.prose.md file incorrectly claimed the default model
was 'deepseek-v4-pro', but /root/.pi/agent/settings.json declares
'defaultModel: syslog-auto' and 'defaultProvider: syslog-harness'.

deepseek-v4-pro is not a servable LiteLLM model name for this fleet.
The only live models are: syslog-auto, gpu-dense, gpu-vision, strix-moe.

Note: The firstmate/pi session talks to DeepSeek through its own
provider (auth.json), NOT through LiteLLM, so no LiteLLM change is
needed to keep firstmate's deepseek usage working.
2026-09-13 18:55:24 +00:00
root c81cf5b6f0 fix: disk-gc-threat-response - correct access methods for kagentz and docker-vm
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 8s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 12s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s
- kagentz (105): use ssh root@kagentz (hostname), NOT pct exec 105 (shows loop0 59G, not real 99G)
- docker-vm (109): use ssh root@192.168.68.7 (correct QEMU VM access)
- abiba (100): use ssh root@abiba (hostname)
- syslog-api (116): use pct-run 116 (correct)

Verified:
  kagentz: 99G 8.9G 86G 10% /
  docker-vm: 158G 17G 135G 11% /
  abiba: 59G 13G 44G 23% /
  syslog-api: 40G 13G 25G 34% /
2026-09-13 04:30:28 +00:00
abiba-bot ef168d9690 Merge pull request 'fix(litellm-health): run the key-list admin call on the gateway host, not inside the container' (#84) from fix-litellm-health-registry-20260912 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 1s
2026-09-13 03:20:04 +00:00
root edcf465831 fix: litellm-health step 8 - run key/list on CT 116 host, not in container
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s
- harness-litellm container has no curl/wget, so docker exec harness-litellm curl returns empty
- Fix: run curl on CT 116 host (ssh root@192.168.68.116 then curl)
- Add admin-call-failed label for empty/unparseable responses
- Verify: 10 keys found (host-side curl), NO-CURL confirmed in container
2026-09-13 03:10:09 +00:00
abiba-bot 89651cf37c Merge pull request 'fix(monitoring): make credential sourcing explicit and fail loudly when missing' (#83) from fix-litellm-health-registry-20260912 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 2s
2026-09-12 23:22:43 +00:00
root 6f40a3be60 fix: address credential-sourcing review findings
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
- infrastructure-monitoring: Correct Zulip key path to /etc/litellm-monitor.env
  (not /etc/zulip-bot.env which doesn't exist); add credential-missing check;
  remove unused monitor_key variable

- litellm-health: Clarify that syslog-auto is a fallback pool alias, not a
  step 7 probe; monitor key must be scoped for all four aliases
2026-09-12 23:17:27 +00:00
root d9368467ff fix: credential-sourcing fix for litellm-health and infrastructure-monitoring
- litellm-health.prose.md: Make monitor key retrieval explicit (ssh from CT 100 to CT 116)
  with executable commands; add syslog-auto alias; document credential-missing failure
  condition (not bare 401 or 0 keys)

- infrastructure-monitoring.prose.md: Fix Zulip POST probe to retrieve keys from CT 116
  via ssh instead of sourcing local env file that doesn't exist on executor host

- Verify model inference probes return 200 for gpu-dense, gpu-vision, strix-moe,
  syslog-auto with corrected credential retrieval
2026-09-12 23:13:24 +00:00
abiba-bot e38598eea4 Merge pull request 'fix: restore per-host probe coverage + sweep residual retired names' (#82) from fix-litellm-health-registry-20260912 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 1s
2026-09-12 22:35:55 +00:00
root 876b011359 no-mistakes(document): Align Rule 9 max_context_window rationale with pool floor
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 8s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
2026-09-12 22:21:57 +00:00
root d23cce89e1 no-mistakes(document): Fix stale retired-name and GPU-context contradictions in contracts 2026-09-12 22:20:02 +00:00
root 9ada2b23c7 no-mistakes(review): Align Rule 8 max_context_window rationale with pool floor 2026-09-12 22:11:40 +00:00
root bd0065bb31 no-mistakes(review): Fix monitor-key scope, Strix context, retired delegation names 2026-09-12 22:02:58 +00:00
root f99f7e1e34 fix: restore per-host probe coverage + sweep residual retired names
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
A. litellm-health step 7: restore .8 → gpu-dense probe (was duplicated
   to gpu-vision after 8210fd9). Add note documenting monitor key scope
   gap for gpu-dense.
B. Sweep remaining retired names presented as usable:
   - README.md:91: qwen3.6-27B-code → gpu-dense in runnable example
   - gpu-fleet.prose.md:14: qwen3.6-35B-udq4 → Carnice-Qwen3.6-MoE...
   - infrastructure-control.prose.md:226: qwen3.6-35B-udq4 → strix-moe
   - proxmox-monitor.prose.md:90: qwen3.6-35B-udq4 → strix-moe
C. Audit test: 10/10 passed (retired raw names now hard-fail)

Fix-forward from 8210fd9 (direct master push).
2026-09-12 21:56:24 +00:00
root 8210fd905c fix: align contracts to 4-name LiteLLM registry (2026-09-12)
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 1s
- Remove retired names (qwen3.6-27B-code, qwen3.6-35B-udq4) from live alias claims
- Update Strix Halo model to Carnice-Qwen3.6-MoE-35B-A3B-Q4_K_M.gguf (strix-moe, 256K ctx)
- Fix litellm-health step 7 probe to gpu-vision (monitor key scoped)
- Move qwen3.6-27B-code/35B-udq4 from raw-but-live to non-resolving in audit
- Fold in pm2-self-heal: remove spoton-service (live PM2 set is 4/4)
- Update hermes templates, key enforcement, timeout tables to live names
2026-09-12 21:39:54 +00:00
abiba-bot b3644f0292 Merge pull request 'fix(disk-gc): hard guest-level report-only gate for CT 111/.129 + correct stale fleet map' (#81) from fix/disk-gc-report-only-129 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 8s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 1s
2026-09-12 19:13:45 +00:00
25 changed files with 1160 additions and 209 deletions
+9 -10
View File
@@ -1,19 +1,18 @@
name: PR Pipeline — Authorize → Validate → Review → Merge
# TRIGGER IS INTENTIONALLY UNFILTERED — DO NOT RE-ADD A `paths:` FILTER.
#
# This workflow previously carried `paths: ['**.prose.md', 'scripts/**.sh',
# '**.yaml', '**.yml']` on both `push` and `pull_request`. Any PR whose diff
# touched none of those patterns (for example a `deliverables/`-only PR, or a
# `scripts/*.py` / `bin/*` change) therefore produced NO Gitea Actions run at
# all: validation, lint, ai-review and the merge gate were silently skipped.
# Validation must run for every pull request and every push to master, so the
# trigger is deliberately unconditional.
on:
push:
branches: [master]
paths:
- '**.prose.md'
- 'scripts/**.sh'
- '**.yaml'
- '**.yml'
pull_request:
types: [opened, synchronize, reopened]
paths:
- '**.prose.md'
- 'scripts/**.sh'
- '**.yaml'
- '**.yml'
jobs:
auth:
+1 -1
View File
@@ -88,7 +88,7 @@ prose run memory-audit-maintenance memory_threshold=90 verify_configs=true
prose run hermes-config-template agent_name=syslog-devops default_model=claude-sonnet-4
# Configure an agent with a different auxiliary model
prose run hermes-config-template agent_name=syslog-code default_model=qwen3.6-27B-code auxiliary_model=gpu-vision
prose run hermes-config-template agent_name=syslog-code default_model=gpu-dense auxiliary_model=gpu-vision
```
### Option B: Manual Execution
+1 -1
View File
@@ -34,7 +34,7 @@ verification, and DM loopback testing.
| @all-bots user ID | 20 | ✅ (config, verified by API at runtime) |
| PM2 process name | abiba-zulip | ✅ |
| Provider | syslog-harness (http://192.168.68.116/v1) | ✅ |
| Default model | deepseek-v4-pro | ✅ (settings.json) |
| Default model | syslog-auto | ✅ (settings.json) |
## Architecture
+2 -3
View File
@@ -163,7 +163,7 @@ def audit(path):
check(
comp.get("max_context_window") == 131072,
"Rule 9",
f"compression.max_context_window must be 131072 (got {comp.get('max_context_window')!r}) — matches 128K GPU capacity",
f"compression.max_context_window must be 131072 (got {comp.get('max_context_window')!r}) — syslog-auto pool floor (NVIDIA hosts 128K; Strix Halo 256K)",
)
# --- Rule 10: Default Model Must Be syslog-auto ---
@@ -245,11 +245,10 @@ def audit(path):
"gemma-4-12b": "gpu-vision",
"crew-auto": "syslog-auto (its 64K cap is retired; no cap in force)",
"ornith-1.0-35b": "strix-moe",
}
raw_but_live = {
"qwen3.6-27B-code": "gpu-dense",
"qwen3.6-35B-udq4": "strix-moe",
}
raw_but_live = {}
for field_path, value in _iter_model_values(cfg):
if value in non_resolving:
check(
+27
View File
@@ -84,6 +84,32 @@ and escalation trail.
- May also be invoked manually: `prose run disk-gc-threat-response`
- Threat-driven: if Amber/Red/Critical detected, immediate GC phase activates
## Scanner: scripts/disk-gc-scan.py
The fleet scan is executed by `scripts/disk-gc-scan.py`, which makes reachability
verdicts deterministic:
1. **Retry on failure:** Each probe retries once before declaring a guest unreachable.
2. **Named probe target:** Every rendered line names the guest, CT id, node, and
access method actually used.
3. **Failure kind printed:** An unreachable guest is reported with its failure kind
(timeout, ssh-auth, no-route, conn-refused, ssh-exit-N) — never as a bare
"unreachable" verdict.
4. **Per-guest access method:** The correct access path is selected from a per-guest
map so the wrong path cannot be picked by an executor improvising:
- CT 105 (kagentz) = `ssh root@kagentz` (NOT `pct exec 105` — pct exec sees
loop0/59G instead of the real 99G filesystem)
- CT 109 (docker-vm) = `ssh root@192.168.68.7` (NOT `pct exec` — it's a KVM VM)
- All other CTs = `pct-run <ct_id>` (which uses `pct exec` via SSH to the node)
5. **Every figure traces to a named probe:** The scan output prints the exact command
that produced each disk figure, so two different guests can never render
identical numbers without the probe commands proving it.
Run: `python3 scripts/disk-gc-scan.py` (or `--json` for machine-readable output).
The scan feeds into `scripts/disk-gc-plan.py`, which applies the report-only gate
from the `report_only_guests` YAML block above.
## Shape
- `self`: scan all CTs via Proxmox API + SSH exec, trigger GC, alert
@@ -357,3 +383,4 @@ one-off GPU builds. No automated post-migration cleanup was in place.
> container has no `pct` binary.
>
> **KVM VM:** CT 109 (docker-vm) is a QEMU VM, not LXC — access via SSH .7.
> **NOTE:** For kagentz (CT 105), use `ssh root@kagentz` (hostname), NOT `pct exec 105` — `pct exec 105` shows loop0 (59G) while `ssh root@kagentz` shows the real filesystem (99G). For docker-vm (CT 109), use `ssh root@192.168.68.7`, not `pct exec`.
+15 -15
View File
@@ -8,12 +8,12 @@ description: >
UPDATED 2026-07-15: Stable role-based aliases introduced: strix-moe,
gpu-dense, gpu-vision (gpu-light was superseded by gpu-vision on 2026-09-12).
These never change — only the underlying model does.
Strix Halo: strix-moe → unsloth/Qwen3.6-35B-A3B-MTP (UD-Q4_K_M, 22GB).
Strix Halo: strix-moe → Carnice-Qwen3.6-MoE-35B-A3B-Q4_K_M.gguf (22GB, 256K ctx).
RTX 5070: gpu-vision — IQ4_NL + MTP draft (~122 tok/s, 2x faster).
UPDATED 2026-07-17: Context reduced fleet-wide from 256K to 128K for stability.
Strix Halo model swapped to qwen3.6-35B-udq4 (22GB, strix-moe alias).
Instability observed near 100K at 256K (now all GPUs at 128K). 128K is the stable ceiling.
For larger context needs → fall back to external providers (deepseek).
UPDATED 2026-07-17: NVIDIA host context reduced from 256K to 128K for stability.
Strix Halo model: Carnice-Qwen3.6-MoE-35B-A3B-Q4_K_M.gguf (alias strix-moe, 256K context).
Strix Halo runs 256K (n_ctx 262144, --kv-unified); RTX 3090 and RTX 5070 remain at 128K.
For >128K on NVIDIA hosts → fall back to external providers (deepseek).
VRAM headroom improved: RTX 3090 ~70%, RTX 5070 ~65%.
agent: abiba
triggers:
@@ -91,8 +91,8 @@ Single source of truth for models, aliases, rpm caps, weights and fallback chain
CT 116 `/opt/inference-harness/litellm_config.yaml`. Do not duplicate those values in
contracts — read them there.
**Backward compatibility**: Old model-specific names (qwen3.6-27B-code, qwen3.5-9b-it) still work
but are deprecated for agent configs. Only the stable aliases survive model swaps. `gemma-4-12b`
**No backward compatibility**: Old model-specific names (qwen3.6-27B-code, qwen3.5-9b-it) are retired as of 2026-09-12
and no longer resolve; do not use them in agent configs. Only the stable aliases survive model swaps. `gemma-4-12b`
is retired and returns 400 `Invalid model name`.
## Routing Configuration (LiteLLM — July 2026)
@@ -190,7 +190,7 @@ If no SSH access, send Zulip DM via abiba-bot with vault update instructions.
| `/root/scripts/gpu-saturation-watchdog.py` | pi (.24) | Auto-restart stuck llama-server |
| `/root/dashboard/gpu-fleet.html` | pi (.24) | Live HTML dashboard |
| `/etc/systemd/system/llama-server.service` | .8, .110 | llama-server daemons (Nvidia GPUs) |
| `/etc/systemd/system/strix-server.service` | .15 (amdpve) | llama-server daemon (Vulkan, Strix Halo) running unsloth/Qwen3.6-35B-A3B-MTP-GGUF. Note: `llama-server.service` and `llama-server@.service` are **masked** on .15 to prevent port 8080 collisions. |
| `/etc/systemd/system/strix-server.service` | .15 (amdpve) | llama-server daemon (Vulkan, Strix Halo) running Carnice-Qwen3.6-MoE-35B-A3B-Q4_K_M.gguf (256K ctx). Note: `llama-server.service` and `llama-server@.service` are **masked** on .15 to prevent port 8080 collisions. |
## Prometheus & Grafana
@@ -207,7 +207,7 @@ If no SSH access, send Zulip DM via abiba-bot with vault update instructions.
- **Router startup race**: Compose router.py doesn't call load_roster(). Reload thread sleeps 30s first.
Fix: trigger roster reload via SSH after restart, or rebuild image with startup load_roster().
- **LiteLLM /metrics**: Requires auth. Prometheus uses `/health/liveliness` as workaround.
- **Strix Halo GPU**: Vulkan is the working backend (ROCm/HIP path abandoned — HSA runtime blocked on Debian 13). Build at `/root/llama.cpp/build-vk/`, commit `4fc4ec5` (2026-07-01), ggml 0.15.3 shared-lib arch. Mesa RADV 25.0.7, KHR_coopmat fast path active. ~70 tok/s gen, 532 tok/s prompt. Service: `strix-server.service` on port 8080, model: `qwen3.6-35B-udq4`, alias `strix-moe`, 128K context, flash-attn + q4 KV, multimodal (mmproj loaded).
- **Strix Halo GPU**: Vulkan is the working backend (ROCm/HIP path abandoned — HSA runtime blocked on Debian 13). Build at `/root/llama.cpp/build-vk/`, commit `4fc4ec5` (2026-07-01), ggml 0.15.3 shared-lib arch. Mesa RADV 25.0.7, KHR_coopmat fast path active. ~70 tok/s gen, 532 tok/s prompt. Service: `strix-server.service` on port 8080, model: `Carnice-Qwen3.6-MoE-35B-A3B-Q4_K_M.gguf`, alias `strix-moe`, 256K context (n_ctx 262144), --parallel 2 --kv-unified, flash-attn + q4 KV, multimodal (mmproj loaded).
- **Port conflict detection (2026-07-05)**: All 3 GPU wrappers now detect ghost processes squatting port 8080 before starting. `.8` and `.110` use inline pre-start check in `llama-wrapper.sh`; `.15` uses `/usr/local/bin/port-cleanup.sh` ExecStartPre. Replaces the blanket `pkill -9 -x llama-server` on .15 which would kill ALL llama-server instances regardless of port. Ghost detection was the root cause of .8 crash-looping for 27+ restarts (stale pid 25836 squatting 8080 after OOM kill).
- **Strix Halo thermal safeguard (2026-07-02)**: `strix-server.service` has `-n 8192` (hard generation cap per request). Without it, `--predict` defaults to -1 (infinity) — a runaway request from .123 (old Mumuni CT114 — now inside Abiba CT100 at .24) decoded 39,868 tokens over 24 min, pushing Tctl to 98°C (crit 89.8°C) and throttling 70→29 t/s. The cap bounds worst-case generation to ~5 min. Do NOT remove `-n` without a replacement ceiling. Sustained load hits ~84°C even at 92s; the APU is fanless/low-flow. Clients MUST also set `max_tokens`.
- **Port 8080 firewall**: amdpve iptables restricts 8080 to 192.168.68.116 (LiteLLM/router host) only. All inbound connections are from .116 (LiteLLM proxied via nginx). Localhost curls hang (SYN dropped). Always test from .116.
@@ -223,10 +223,10 @@ If no SSH access, send Zulip DM via abiba-bot with vault update instructions.
| GPU | Model | Gen tok/s | Prompt tok/s | Baseline | Context |
|-----|-------|-----------|--------------|----------|---------|
| RTX 3090 (.8) | Qwen3.8-27B-Uncensored-Q4_K_M | **TBD** | — | — | **128K** |
| Strix Halo (.15) | qwen3.6-35B-udq4 | **65** | 140 | — | **128K** |
| Strix Halo (.15) | Carnice-Qwen3.6-MoE-35B-A3B-Q4_K_M.gguf (strix-moe) | **65** | 140 | — | **256K** |
Benchmarks from 2026-07-17. Strix Halo model: qwen3.6-35B-udq4. RTX 5070 MTP provides 2.7x speedup over pre-upgrade 70 tok/s.
All 3 GPUs now at 128K context (2026-07-17, reduced from 256K for stability).
Benchmarks from 2026-07-17. Strix Halo model: Carnice-Qwen3.6-MoE-35B-A3B-Q4_K_M.gguf (alias strix-moe), n_ctx 262144, --parallel 2 --kv-unified. RTX 5070 MTP provides 2.7x speedup over pre-upgrade 70 tok/s.
GPU contexts: RTX 3090 (.8) and RTX 5070 (.110) at 128K; Strix Halo (.15) at 256K (2026-09-12).
Benchmarks run through LiteLLM proxy (192.168.68.116:4000) every 5 minutes.
Degradation alerts fire at 30% (warning) and 50% (critical) below baseline.
@@ -247,8 +247,8 @@ All agent configs MUST use stable role-based aliases, never model-specific names
When the underlying model is swapped, only the LiteLLM config changes — agent configs are untouched.
### Context Windows
- RTX 3090: **128K** (reduced from 256K 2026-07-17) | RTX 5070: **128K** (reduced from 256K) | Strix Halo: **128K**
- **All agents**: 128K ceiling — stable margin. For >128K workloads, use external providers (deepseek)
- RTX 3090: **128K** (reduced from 256K 2026-07-17) | RTX 5070: **128K** (reduced from 256K) | Strix Halo: **256K** (2026-09-12)
- **Agents via `syslog-auto`**: 128K ceiling — the pool's safe floor (NVIDIA hosts are 128K). For >128K workloads, use external providers (deepseek)
- Compression threshold 0.65 (audit Rule 9): fires at ~85K (~43K headroom before 128K ceiling)
- **Pi agents (Abiba)**: `compaction.reserveTokens: 52739` (≈60% of 128K)
- Mumuni compression model alias: `syslog-auto`
@@ -266,7 +266,7 @@ Mumuni (kagentz CT105, 192.168.68.14 — migrated from CT100 2026-08-29) is the
| `aux.vision.model` | `gpu-vision` | Vision tasks (RTX 5070) |
| `aux.web_extract.model` | `gpu-vision` | Web extraction |
| `delegation.model` | `gpu-dense` | Sub-agent reasoning (RTX 3090) |
| `context.max_context_window` | 131072 (128K) | Reduced from 256K 2026-07-17 — stable 128K ceiling |
| `context.max_context_window` | 131072 (128K) | Conservative `syslog-auto` pool floor (NVIDIA hosts 128K; Strix Halo 256K) |
| `compression.threshold` | 0.65 | Rule 9: triggers at ~85K for a 128K window |
| `compression.target_ratio` | 0.3 | Compresses to ~38K |
| `compression.protect_last_n` | 40 | Preserves last 40 messages |
+4 -4
View File
@@ -48,11 +48,11 @@ depends_on:
| Alias | GPU | Host | Model | VRAM | Ctx | tok/s | Role |
|-------|-----|------|-------|------|-----|-------|------|
| `strix-moe` | Strix Halo 64GB | ct15 (.15:8080) | qwen3.6-35B-udq4 | ~10/64GB (16%) | 128K | 62.9 | Compression, summarization, long docs |
| `strix-moe` | Strix Halo 64GB | ct15 (.15:8080) | Carnice-Qwen3.6-MoE-35B-A3B-Q4_K_M.gguf | ~10/64GB (16%) | 256K | 62.9 | Compression, summarization, long docs |
Key notes:
- All models use direct GPU routing via LiteLLM (`api_key: not-needed`). Router (port 9000) was decommissioned 2026-09-11 and is NOT in the inference path.
- Stable aliases (gpu-dense, gpu-vision, strix-moe) from gpu-fleet are the canonical names for agent configs. Model-specific names still work but are deprecated. The retired names `gpu-light` and `gemma-4-12b` were superseded by `gpu-vision` on 2026-09-12 and no longer resolve (400 `Invalid model name`).
- Stable aliases (gpu-dense, gpu-vision, strix-moe) from gpu-fleet are the canonical names for agent configs — use these, not model-specific names. The retired names `gpu-light` and `gemma-4-12b` were superseded by `gpu-vision` on 2026-09-12 and no longer resolve (400 `Invalid model name`).
- The RTX 5070 is the fastest endpoint per token — `gpu-vision` is its canonical alias. Route vision/web/light work there first. The RTX 5070 model is multimodal (image+text).
- Strix Halo is 62.9 tok/s (89% of 70.5 baseline) — below optimal but stable. Check for competing workloads.
- RTX 3090 VRAM at 88% — within role-appropriate range (role = heavy reasoning, needs the headroom).
@@ -142,10 +142,10 @@ Key notes:
- **Escalate**: If Tier 2 triggers and temp still rising after 5 min → possible hardware failure
### Rule 9: Context Window Optimization
- **Detect**: Benchmark tok/s vs baseline for each GPU at current context (all 128K)
- **Detect**: Benchmark tok/s vs baseline for each GPU at its current context
- RTX 3090 (128K ctx, ThinkingCap): baseline 74.8 tok/s — currently at 74.9 (100%)
- RTX 5070 (128K ctx, HauhauCS QAT): baseline 165.2 tok/s — currently at 169.6 (103%)
- Strix Halo (128K ctx, qwen3.6-35B-udq4): baseline 70.5 tok/s — currently at 62.9 (89%)
- Strix Halo (256K ctx, Carnice-Qwen3.6-MoE-35B-A3B-Q4_K_M.gguf): baseline 70.5 tok/s — currently at 62.9 (89%)
- **Fix**:
- If tok/s > baseline → context has headroom, consider increasing
- If tok/s < 90% baseline → reduce context by 25% and retest
+40 -3
View File
@@ -5,12 +5,30 @@ version: 1.0.0
description: >
Canonical known-good baseline for all Syslog Hermes agents. Captures the exact
configuration state, keys, workarounds, and audit procedure. When an agent's
configuration goes sideways, restore from this baseline. Last verified 2026-07-16. All GPUs 128K context (reduced from 256K for stability Jul 2026) (RTX 3090 .8, RTX 5070 .110, Strix Halo .15). Parallel 1 fleet-wide (Strix Halo handles compression solo).
configuration goes sideways, restore from this baseline. Last verified 2026-07-16. RTX 3090/5070 at 128K (reduced from 256K for stability Jul 2026); Strix Halo at 256K (2026-09-12) (RTX 3090 .8, RTX 5070 .110, Strix Halo .15). Parallel 1 fleet-wide (Strix Halo handles compression solo).
author: Abiba (pi agent)
---
# Hermes Agent Baseline — Canonical Good State
## Reachability Detection
Before checking agent baseline, verify the host is reachable and can be audited. Use the shared reachability helper from the clone root:
```bash
# Run on each host to check reachability (Tanko, Mumuni, Koonimo, Koby)
scripts/hermes-reachability-check.sh <host> "api_key:" "/root/.hermes/config.yaml"
# Example: scripts/hermes-reachability-check.sh 192.168.68.122 "api_key:" "/root/.hermes/config.yaml"
# Expected outcomes:
# - UNREACHABLE: SSH connection failed (host is down)
# - VIOLATION: SSH succeeded and found matches (report the finding)
# - COMPLIANT: SSH succeeded and found no matches (no api_key in config)
#
# NOTE: The bug this replaces was deriving reachability from the remote grep's exit code.
# The correct pattern: remote side always succeeds (grep ...; true), so ssh status = connection only.
```
## Quick Restore
```bash
@@ -103,6 +121,26 @@ auxiliary:
timeout: 120
```
## Violation Classification
When reporting findings, separate POLICY observations from FAULT findings:
### POLICY (observation only, not a fault)
- Agent uses a non-internal-harness provider (e.g., direct DeepSeek, Tencent, OpenRouter)
- Config text has a field that looks unusual but the agent's calls are succeeding
- Example: "POLICY: Koonimo uses deepseek directly; calls succeeding in last hour"
### FAULT (requires request-level evidence)
- Agent's calls are failing with auth errors (401/403 in logs)
- Agent's config has no valid API key AND calls are failing
- Example: "FAULT: Koby's LiteLLM key expired; 401 observed at 2026-09-14 11:42:00"
### Rules
1. Do NOT infer the runtime's credential resolution from config text alone.
2. Require request-level evidence before calling something a FAULT: an observed auth failure in the agent's log, or the absence of successful calls in the window.
3. If calls are succeeding, the correct output is "POLICY: uses <provider> directly; calls succeeding" - not a violation.
4. State what you OBSERVED, not what the field implies.
## Known Bug: `api_key_env` Ignored by Auxiliary Client
**Bug location**: `agent/auxiliary_client.py` → `_resolve_task_provider_model()` (line ~5478)
@@ -200,8 +238,7 @@ the registry before applying. Key is injected via `infisical run --` wrapper at
{ "id": "syslog-auto" },
{ "id": "strix-moe" },
{ "id": "gpu-dense" },
{ "id": "gpu-vision" },
{ "id": "qwen3.6-27B-code" }
{ "id": "gpu-vision" }
]
}
}
+53 -14
View File
@@ -19,6 +19,24 @@ description: >
- agent_keys: map (see Agent Keys section)
- infra_endpoints_verified: array
## Reachability Detection
Before auditing the config template, verify the host is reachable and can be checked. Use the shared reachability helper from the clone root:
```bash
# Run on each host to check reachability (Tanko, Mumuni, Koonimo, Koby)
scripts/hermes-reachability-check.sh <host> "base_url:" "/root/.hermes/config.yaml"
# Example: scripts/hermes-reachability-check.sh 192.168.68.122 "base_url:" "/root/.hermes/config.yaml"
# Expected outcomes:
# - UNREACHABLE: SSH connection failed (host is down)
# - VIOLATION: SSH succeeded and found matches (report the finding)
# - COMPLIANT: SSH succeeded and found no matches (no base_url in config)
#
# NOTE: The bug this replaces was deriving reachability from the remote grep's exit code.
# The correct pattern: remote side always succeeds (grep ...; true), so ssh status = connection only.
```
## Agent Keys (LiteLLM — Current 2026-07-11)
Each agent has a unique LiteLLM API key (virtual key) generated against the LiteLLM
@@ -92,12 +110,12 @@ work immediately after restart.
```yaml
# ─── Model Selection ───
model:
default: <agent_model> # e.g., strix-moe, qwen3.6-27B-code, syslog-auto
default: <agent_model> # e.g., strix-moe, gpu-dense, syslog-auto
provider: harness
base_url: http://192.168.68.116/litellm/v1 # Rule 5 (2026-08-09): canonical authenticated path; /v1 also OK
api_key_env: LITELLM_API_KEY # Injected via infisical run -- wrapper
max_tokens: 4096 # ⚠️ CRITICAL: Prevents unbounded generation
context_length: 131072 # For syslog-auto (all GPUs at 128K for stability).
context_length: 131072 # Conservative floor for syslog-auto (NVIDIA hosts 128K; Strix Halo 256K).
# ⚠️ MANDATORY: Hermes probes unknown models from 256K
# and falls back to 256K when /v1/models lacks a context
# field (llama-server does). Without this override, agents
@@ -133,7 +151,7 @@ compression:
enabled: true
model: syslog-auto # ⚠️ Must match auxiliary.compression.model. Stable alias (gpu-fleet § Stable Role-Based Aliases). NOT ornith-1.0-35b (LiteLLM does not serve that name).
provider: harness
max_context_window: 131072 # MUST match actual GPU capacity. All 3 GPUs are 128K (Jul 17).
max_context_window: 131072 # MUST stay at the syslog-auto pool floor: NVIDIA hosts are 128K, Strix Halo 256K (2026-09-12).
threshold: 0.65 # Fires at ~170K for 262K window, ~85K for 128K
target_ratio: 0.30
protect_last_n: 40
@@ -149,7 +167,7 @@ compression:
# Do NOT use syslog-auto for auxiliary tasks — it routes to the primary GPU.
# gpu-vision = RTX 5070 (12B), freeing the Strix Halo for agent reasoning.
# Heavy aux (delegation, x_search) use gpu-dense (RTX 3090) instead.
# NEVER use raw model names (e.g. qwen3.6-27B-code, qwen3.6-35B-udq4; gemma-4-12b is retired
# NEVER use retired model names (qwen3.6-27B-code, qwen3.6-35B-udq4; gemma-4-12b is retired
# and no longer resolves) in agent configs — use the stable aliases so model swaps don't break agents.
auxiliary:
vision:
@@ -173,10 +191,10 @@ auxiliary:
timeout: 300 # gpu-fleet: 300s for large-history summarization (was 60)
# ─── Delegation / Heavy Aux (use gpu-dense = RTX 3090) ───
# delegation.model and x_search.model use gpu-dense (NOT raw qwen3.6-27B-code).
# delegation.model and x_search.model use gpu-dense (NOT retired raw name).
delegation:
model: gpu-dense # stable alias for RTX 3090 (was raw qwen3.6-27B-code)
model: gpu-dense # stable alias for RTX 3090
provider: harness
base_url: http://192.168.68.116/litellm/v1 # Rule 5 (2026-08-09): canonical authenticated; /v1 also OK
api_key_env: LITELLM_API_KEY
@@ -199,6 +217,26 @@ When LiteLLM keys are regenerated (e.g., after infrastructure changes):
3. **After update**: Restart Hermes on the agent host
4. **Verify**: `curl -H "Authorization: Bearer sk-<KEY>" http://192.168.68.116/v1/models`
## Violation Classification
When reporting findings, separate POLICY observations from FAULT findings:
### POLICY (observation only, not a fault)
- Agent uses a non-internal-harness provider (e.g., direct DeepSeek, Tencent, OpenRouter)
- Config text has a field that looks unusual but the agent's calls are succeeding
- Example: "POLICY: Koonimo uses deepseek directly; calls succeeding in last hour"
### FAULT (requires request-level evidence)
- Agent's calls are failing with auth errors (401/403 in logs)
- Agent's config has no valid API key AND calls are failing
- Example: "FAULT: Koby's LiteLLM key expired; 401 observed at 2026-09-14 11:42:00"
### Rules
1. Do NOT infer the runtime's credential resolution from config text alone.
2. Require request-level evidence before calling something a FAULT: an observed auth failure in the agent's log, or the absence of successful calls in the window.
3. If calls are succeeding, the correct output is "POLICY: uses <provider> directly; calls succeeding" - not a violation.
4. State what you OBSERVED, not what the field implies.
## Configuration Rules
### Rule 1: Shared Infra Is Locked
@@ -253,7 +291,7 @@ The following MUST be identical across ALL profiles:
### Rule 7: Auxiliary Model Consistency (UPDATED 2026-07-16)
- Vision and web_extract use `gpu-vision` (RTX 5070 — 12GB, vision-optimized)
- Compression uses `syslog-auto` — the Strix Halo weighted pool (64GB, 128K ctx, compression-optimized); do NOT pin `compression.model` to `strix-moe` (audit Rule 7 rejects it)
- Compression uses `syslog-auto` — the Strix Halo weighted pool (64GB, 256K ctx, compression-optimized); do NOT pin `compression.model` to `strix-moe` (audit Rule 7 rejects it)
- **`ornith-1.0-35b` is NOT a valid compression model name** — LiteLLM does not serve it
(do not restate the served model list here — CT 116 `/opt/inference-harness/litellm_config.yaml`
is the single source of truth for models, aliases, weights and fallbacks). Old configs with `ornith-1.0-35b` cause 403/model-not-found on compression calls.
@@ -267,15 +305,15 @@ The following MUST be identical across ALL profiles:
- `api_key_env: LITELLM_API_KEY`
- **Do NOT use `syslog-auto` for `vision`/`web_extract`** — it routes unpredictably; compression is the deliberate exception (see the OPERATIONAL DECISION above)
- **Compression on Strix Halo**: The strix-moe alias routes to Strix Halo
(64GB UMA, 128K context) — the designated compression GPU. This frees the
(64GB UMA, 256K context) — the designated compression GPU. This frees the
RTX 5070 for vision and web search, and the RTX 3090 for heavy reasoning.
- The `compression:` block's `model` MUST match `auxiliary: compression: model`
- The `compression: max_context_window: 131072` MUST match actual GPU capacity (128K)
- The `compression: max_context_window: 131072` MUST stay at the syslog-auto pool floor (NVIDIA hosts 128K; Strix Halo 256K)
### Rule 8: GPU Workload Distribution (UPDATED 2026-07-16)
- **RTX 3090 (24GB, 128K ctx, qwen3.6-27B-code)**: Heavy reasoning, code gen, long conversations
- **RTX 3090 (24GB, 128K ctx, gpu-dense)**: Heavy reasoning, code gen, long conversations
- **RTX 5070 (12GB, 128K ctx, gpu-vision)**: Vision, web search, quick tasks, web_extract (IQ4_NL+MTP, ~65% VRAM at 128K)
- **Strix Halo (64GB, 128K ctx, syslog-auto)**: Context compression, summarization, long docs
- **Strix Halo (64GB, 256K ctx, syslog-auto)**: Context compression, summarization, long docs
- Agent profiles MUST route auxiliary tasks to the correct GPU:
- `auxiliary.vision.model: gpu-vision` (RTX 5070)
- `auxiliary.web_extract.model: gpu-vision` (RTX 5070)
@@ -284,14 +322,14 @@ The following MUST be identical across ALL profiles:
- For 128K context window: `threshold: 0.65` (fires at ~85K tokens)
- Do NOT use `threshold: 0.25` — this fires at 65K, causing premature context loss
- Do NOT use `threshold: 0.80` — this delays until ~105K, leaving only 23K margin
- `max_context_window: 131072` MUST match the model's actual capacity (128K)
- `max_context_window: 131072` MUST stay at the pool floor (NVIDIA hosts 128K; Strix Halo 256K)
- See `devops-hermes-compression` skill for full reference
### Rule 9: Compression Threshold for 128K Models
- For 128K context window: `threshold: 0.65` (fires at ~85K tokens)
- Do NOT use `threshold: 0.25` — this fires at 65K, causing premature context loss
- Do NOT use `threshold: 0.80` — this delays until ~105K, leaving only 23K margin
- `max_context_window: 131072` MUST match the model's actual capacity (all GPUs = 128K)
- `max_context_window: 131072` MUST stay at the pool floor (NVIDIA hosts 128K; Strix Halo 256K)
- See `devops-hermes-compression` skill for full reference
### Rule 10: Default Model Must Be `syslog-auto` (All Agents)
@@ -322,7 +360,8 @@ When an agent shows "context issues" (premature compression, 401s, 504s, DeepSee
verify ALL FOUR of these against the live config. They are the only root causes found in production:
1. **max_context_window correct?** — BOTH `compression.max_context_window` AND `context.max_context_window`
MUST be `131072` (all GPUs are 128K). A value of `262144` causes instability near 100K and must NOT be used.
MUST be `131072` (the syslog-auto pool floor: NVIDIA hosts are 128K; Strix Halo is 256K).
A `262144` client window can route to a 128K NVIDIA host and fail, so it must NOT be used.
~83K instead of ~170K. Check: `grep -n max_context_window ~/.hermes/config.yaml`
2. **base_url uses authenticated path?** — `custom_providers[0].base_url`, `delegation.base_url`,
and ALL `auxiliary.*.base_url` MUST be `http://192.168.68.116/litellm/v1` (Rule 5, canonical)
+43 -3
View File
@@ -117,6 +117,44 @@ model:
api_key_env: LITELLM_API_KEY
```
## Reachability Detection
Before checking for hardcoded keys, verify the host is reachable and can be audited. Use the shared reachability helper from the clone root:
```bash
# Run on each host to check reachability (Tanko, Mumuni, Koonimo, Koby)
scripts/hermes-reachability-check.sh <host> "api_key: sk-" "/root/.hermes/"
# Example: scripts/hermes-reachability-check.sh 192.168.68.122 "api_key: sk-" "/root/.hermes/"
# Expected outcomes:
# - UNREACHABLE: SSH connection failed (host is down)
# - VIOLATION: SSH succeeded and found matches (report the finding)
# - COMPLIANT: SSH succeeded and found no matches (no hardcoded keys in config)
#
# NOTE: The bug this replaces was deriving reachability from the remote grep's exit code.
# The correct pattern: remote side always succeeds (grep ...; true), so ssh status = connection only.
```
## Violation Classification
When reporting findings, separate POLICY observations from FAULT findings:
### POLICY (observation only, not a fault)
- Agent uses a non-internal-harness provider (e.g., direct DeepSeek, Tencent, OpenRouter)
- Config text has a field that looks unusual but the agent's calls are succeeding
- Example: "POLICY: Koonimo uses deepseek directly; calls succeeding in last hour"
### FAULT (requires request-level evidence)
- Agent's calls are failing with auth errors (401/403 in logs)
- Agent's config has no valid API key AND calls are failing
- Example: "FAULT: Koby's LiteLLM key expired; 401 observed at 2026-09-14 11:42:00"
### Rules
1. Do NOT infer the runtime's credential resolution from config text alone.
2. Require request-level evidence before calling something a FAULT: an observed auth failure in the agent's log, or the absence of successful calls in the window.
3. If calls are succeeding, the correct output is "POLICY: uses <provider> directly; calls succeeding" - not a violation.
4. State what you OBSERVED, not what the field implies.
## Detection Query
Run on any Hermes host to detect violations:
@@ -169,9 +207,11 @@ The agent picks up the new key via `infisical run --` at gateway startup.
**Keys are permanent and use bare agent name aliases.**
- **Duration**: `null` — keys never expire. NOT enforced today: CT 116 `litellm_config.yaml` has no `default_key_generate_params` block, and a key generated with no explicit models comes back with an empty models list. OPEN policy question: should agent keys expire by default? (captain security-policy decision, raised separately.)
- **Duration**: `null` — keys never expire by default. **Expiry must be set EXPLICITLY at creation** with the `duration` parameter (e.g., `90d` for 90 days). The 90-day default is the standard; however, the config default is **NOT honoured** by LiteLLM 1.99.1 (verified on CT 116: a key generated with no explicit duration returns `expires=null`). This has been recorded in `/opt/inference-harness/litellm_config.yaml` to prevent re-filing as a bug.
- **Daily Audit**: A daily audit job runs at 00:00 UTC (`/usr/local/bin/litellm-key-renewal-ct116.sh`, cron 00:00). It is **AUDIT-ONLY** and does not perform renewal. It lists every key, reports those with no expiry and those inside a 14-day warning window, explicitly EXCLUDES `abiba-pi` and `koby` (report-only, and .129 must never be touched), and logs `RENEWAL-REQUIRED-BUT-NOT-PERFORMED + NO KEY WAS CHANGED` when renewal is skipped. **Renewal is NOT implemented** — keys must not be rotated until delivery (vault injection + consumer verification) exists and is proven end-to-end.
- **Exclusions**: `abiba-pi` and every firstmate/secondmate/crewmate key stay **WITHOUT an expiry** until a proven renewal path exists. `koby` is **report-only** (never touched). These exclusions are enforced by the audit job.
- **Alias convention**: bare agent name only (e.g., `tanko`, `mumuni`, `koby`, `koonimo`). No dates, no versions. The alias IS the identity.
- **Rotation triggers**: compromise, personnel departure, or quarterly security hygiene. NOT calendar-driven.
- **Rotation triggers**: compromise, personnel departure, or quarterly security hygiene. NOT calendar-driven. Manual rotation is permitted only when the renewal delivery path is proven and verified on a throwaway consumer before production use.
- **Max budget**: $100 per key (config default).
```yaml
@@ -182,7 +222,7 @@ The agent picks up the new key via `infisical run --` at gateway startup.
# re-verify before applying.
litellm_settings:
default_key_generate_params:
models: ["syslog-auto", "qwen3.6-27B-code", "gpu-vision"]
models: ["syslog-auto", "gpu-dense", "gpu-vision", "strix-moe"]
duration: null # ← permanent
max_budget: 100
metadata:
+3 -3
View File
@@ -4,7 +4,7 @@ kind: responsibility
description: >
Optimizes the full Syslog inference stack — LiteLLM routing weights, GPU model
assignments, agent context management, and prompt caching — to reduce response
times to sub-15s average. All GPUs now at 128K context (stable ceiling).
times to sub-15s average. NVIDIA GPUs at 128K context; Strix Halo at 256K (2026-09-12).
id: 067NC6KP02RG60S50M40E30928
---
@@ -63,7 +63,7 @@ prefill time at 532 tok/s. Fix context first, routing second.
- **Cache everything repeated**: system prompts, skill docs, AGENTS.md — these
never change between turns. Single-digit cache hit rate is unacceptable.
- **Lower context ceiling**: 128K window is the stable ceiling for agent conversations.
GPUs reduced from 256K to 128K (2026-07-17). 128K window should compact at 85K (0.65 threshold). For larger contexts, route to external providers.
GPUs reduced from 256K to 128K (2026-07-17) for the NVIDIA hosts; Strix Halo runs 256K (2026-09-12). 128K window should compact at 85K (0.65 threshold). For larger contexts, route to external providers.
### Shape
@@ -101,5 +101,5 @@ call enable-prompt-caching
call verify-latency
host: 192.168.68.116
models: [syslog-auto, qwen3.6-27B-code, gpu-vision]
models: [syslog-auto, gpu-dense, gpu-vision, strix-moe]
```
+1 -1
View File
@@ -223,7 +223,7 @@ description: >
**Prometheus targets**:
- 192.168.68.8:9400 (RTX 3090 — qwen)
- 192.168.68.110:9400 (RTX 5070 — gpu-vision)
- 192.168.68.15:9400 (Strix Halo — qwen3.6-35B-udq4)
- 192.168.68.15:9400 (Strix Halo — strix-moe)
- harness-litellm:4000 (LiteLLM health)
### Ecosystem C: Netbird (72.61.0.17 — Hostinger srv1079750.hstgr.cloud)
+127 -98
View File
@@ -17,7 +17,7 @@ description: >
⚠️ This contract is target-state aspirational — but GPU export + alerting
are now as-built (verified 2026-08-09).
As-built GPU monitoring is via gpu-monitor contract (port 9100 poll).
version: 1.0.0
version: 1.0.1
---
## Architecture
@@ -120,127 +120,156 @@ the any-HTTP rule. On those — the authenticated Zulip POST and the router
`/health` — an unexpected status (`401`/`403` from a bad or missing credential,
`5xx`, or anything other than the expected `200`) is an **ALERT**, not "alive".
**STANDING PROBE RULES (2026-09-14, from defect report 1150.msg):**
1. **Any HTTP status means ALIVE.** 200, 301, 302, 401, 403, 404 all prove the
service answered — report the code, never "down". A redirect is not a failure.
Only a failed CONNECTION (curl status 000, timeout, refused) is a failed probe,
and that is a statement about YOUR PROBE, not about the service.
2. **A failed probe is never a service verdict.** Print
`probe-failed: <target> <kind>` naming the exact URL/host/port and the failure
kind (timeout, refused, no-route, dns), retry once at a longer timeout, and only
then report. Apply the same shape as scripts/disk-gc-scan.py.
3. **Say which probe produced each number.** "Grafana: 000" is unusable;
"Grafana http://192.168.68.116:3001/api/health -> connection timeout after 10s
(retried at 25s: also timeout)" is actionable.
### check-health
**RUN LIVE, NEVER ECHO — every dispatch must execute the probes below with real tool calls; never repeat a prior report unless a live probe fails.**
**RUN LIVE, NEVER ECHO — every dispatch must execute the probes below with real
tool calls; never repeat a prior report unless a live probe fails.**
**PROBE SHAPE (per standing rules above):**
- Every probe prints the target name + URL + HTTP code (or failure kind)
- Retry once on connection failure at longer timeout
- Any HTTP status = ALIVE; only 000/timeout/refused = probe-failed
- Report the actual probe command and its result, not a summary verdict
```bash
# Provenance — run first; paste the absolute path into the report
pwd -P
# Zulip API health (POST ping)
source /etc/litellm-monitor.env
ZULIP_USER="abiba-bot@chat.sysloggh.net"
curl -s -o /dev/null -w '%{http_code}' -X POST https://chat.sysloggh.net/api/v1/messages -u "${ZULIP_USER}:${ZULIP_BOT_KEY}"
# Expected: 200 (HTTP 000 = unreachable/cache)
# ============================================================
# 1. ZULIP API HEALTH (POST ping) — bare-200 probe
# ============================================================
# NOTE: /etc/litellm-monitor.env exists only on CT 116, retrieve keys from CT 116 via:
zulip_key=$(ssh root@192.168.68.116 "grep ZULIP_BOT_KEY /etc/litellm-monitor.env | cut -d= -f2")
if [ -z "$zulip_key" ]; then
echo "credential-missing: ZULIP_BOT_KEY not found in /etc/litellm-monitor.env"
else
ZULIP_USER="abiba-bot@chat.sysloggh.net"
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 -X POST https://chat.sysloggh.net/api/v1/messages -u "${ZULIP_USER}:${zulip_key}")
echo "Zulip API https://chat.sysloggh.net/api/v1/messages -> $code"
# Expected: 200 (bare-200 probe; any other status is an ALERT)
fi
# PM2 process health
# ============================================================
# 2. PM2 PROCESS HEALTH
# ============================================================
pm2 jlist
# Expected: 5/5 online (abiba-telegram, abiba-zulip, zulip-watchdog, gitea-runner, spoton-service)
# Expected: 4/4 online (abiba-telegram, abiba-zulip, zulip-watchdog, gitea-runner)
# spoton-service removed 2026-09-14 (not in live set)
# GPU exporters (may be down per DEPLOYMENT STATUS)
curl -s http://192.168.68.8:9400/metrics && echo " - OK" || echo " - FAIL"
curl -s http://192.168.68.110:9400/metrics && echo " - OK" || echo " - FAIL"
curl -s http://192.168.68.15:9400/metrics && echo " - OK" || echo " - FAIL"
# ============================================================
# 3. GPU EXPORTERS — any-HTTP probe (metrics endpoint)
# ============================================================
# Probe /metrics (the Prometheus scrape target), not bare /
for host in 192.168.68.8 192.168.68.110 192.168.68.15; do
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 "http://$host:9400/metrics")
if [ "$code" == "000" ]; then
# Retry with longer timeout
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 25 "http://$host:9400/metrics")
echo "GPU exporter http://$host:9400/metrics -> probe-failed: timeout (retried at 25s: still $code)"
else
echo "GPU exporter http://$host:9400/metrics -> $code"
fi
done
# Expected: 200 on all 3 hosts (RTX 3090, RTX 5070, Strix Halo)
# Router health (via nginx on port 80)
curl -s -o /dev/null -w '%{http_code}' http://192.168.68.116/health
# Expected: 200 (Router is up and responding)
# ============================================================
# 4. ROUTER HEALTH (via nginx on port 80) — bare-200 probe
# ============================================================
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 http://192.168.68.116/health)
if [ "$code" == "000" ]; then
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 25 http://192.168.68.116/health)
echo "Router http://192.168.68.116/health -> probe-failed: timeout (retried at 25s: still $code)"
else
echo "Router http://192.168.68.116/health -> $code"
fi
# Expected: 200 (bare-200 probe; any other status is an ALERT)
# LiteLLM health (via nginx on port 80)
curl -s -o /dev/null -w '%{http_code}' http://192.168.68.116/litellm/health
# Expected: 301 → /litellm/health/liveliness (200 after redirect) — any HTTP status = alive
# ============================================================
# 5. LITELLM HEALTH (via nginx on port 80) — any-HTTP probe
# ============================================================
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 http://192.168.68.116/litellm/health)
if [ "$code" == "000" ]; then
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 25 http://192.168.68.116/litellm/health)
echo "LiteLLM http://192.168.68.116/litellm/health -> probe-failed: timeout (retried at 25s: still $code)"
else
echo "LiteLLM http://192.168.68.116/litellm/health -> $code"
fi
# Expected: 301 → /litellm/health/liveliness (any HTTP status = ALIVE)
# PVE API liveness — probe the REAL PVE nodes on :8006, never the monitoring
# host CT 116. CT 116 runs no pveproxy, so probing it on :8006 returns 000 —
# that was the stale-vantage bug this replaces (CT 116 is the monitoring host,
# not a cluster node). Unauthenticated GET answers 401 while the API is ALIVE
# by design. Alive = ANY HTTP status (401 is the EXPECTED healthy response);
# DOWN = connection refused (000) or timeout only.
# ============================================================
# 6. PVE API LIVENESS — any-HTTP probe (auth-gated)
# ============================================================
# Probe the REAL PVE nodes on :8006, never the monitoring host CT 116.
for node in 192.168.68.9 192.168.68.5 192.168.68.15 192.168.68.6 192.168.68.12; do
printf '%s:8006 -> %s\n' "$node" \
"$(curl -sk -o /dev/null -w '%{http_code}' --connect-timeout 5 "https://$node:8006/api2/json/version")"
code=$(curl -sk -o /dev/null -w '%{http_code}' --connect-timeout 10 "https://$node:8006/api2/json/version")
if [ "$code" == "000" ]; then
code=$(curl -sk -o /dev/null -w '%{http_code}' --connect-timeout 25 "https://$node:8006/api2/json/version")
echo "PVE API https://$node:8006/api2/json/version -> probe-failed: timeout (retried at 25s: still $code)"
else
echo "PVE API https://$node:8006/api2/json/version -> $code"
fi
done
# Expected: 401 on every node (acerpve .9, ocupve .5, amdpve .15, storepve .6, minipve .12)
# A node answering 000/timeout is DOWN — flag that node. 401 is NOT a fault.
# 401 is the EXPECTED healthy response (auth-gated); 000/timeout = DOWN
# Prometheus targets
curl -s http://192.168.68.116:9090/api/v1/targets | jq '.data.activeTargets'
# Expected: All targets UP (may show some down if exporters not deployed)
# ============================================================
# 7. PROMETHEUS TARGETS — bare-200 probe
# ============================================================
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 http://192.168.68.116:9090/api/v1/targets)
if [ "$code" == "000" ]; then
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 25 http://192.168.68.116:9090/api/v1/targets)
echo "Prometheus http://192.168.68.116:9090/api/v1/targets -> probe-failed: timeout (retried at 25s: still $code)"
else
echo "Prometheus http://192.168.68.116:9090/api/v1/targets -> $code"
fi
# Expected: 200 (bare-200 probe; any other status is an ALERT)
# Grafana health
curl -s http://192.168.68.116:3001/api/health | jq '{status, version}'
# Expected: {"status":"ok","version":"..."}
# ============================================================
# 8. GRAFANA HEALTH — any-HTTP probe
# ============================================================
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 http://192.168.68.116:3001/api/health)
if [ "$code" == "000" ]; then
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 25 http://192.168.68.116:3001/api/health)
echo "Grafana http://192.168.68.116:3001/api/health -> probe-failed: timeout (retried at 25s: still $code)"
else
echo "Grafana http://192.168.68.116:3001/api/health -> $code"
fi
# Expected: 200 (any HTTP status = ALIVE; 000/timeout = DOWN)
# LiteLLM metrics (Prometheus endpoint)
curl -s http://192.168.68.116:4000/metrics | head -20
# Expected: Prometheus-formatted metrics output
# ============================================================
# 9. LITELLM METRICS (Prometheus endpoint) — any-HTTP probe
# ============================================================
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 http://192.168.68.116:4000/metrics)
if [ "$code" == "000" ]; then
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 25 http://192.168.68.116:4000/metrics)
echo "LiteLLM metrics http://192.168.68.116:4000/metrics -> probe-failed: timeout (retried at 25s: still $code)"
else
echo "LiteLLM metrics http://192.168.68.116:4000/metrics -> $code"
fi
# Expected: 200 (any HTTP status = ALIVE; 000/timeout = DOWN)
```
**Report format**: Begin every report with the **absolute path the probe executed
from** (`pwd -P`, or the script's absolute path) so a stale-consumer report is
distinguishable from a real fault at read time. Summarize actual results from
each probe. Apply the any-HTTP-response liveness rule ONLY to the auth-gated PVE
API and LiteLLM endpoints above: only connection-refused (`000`) or timeout is
DOWN; empty output is a warning. For probes whose expected result is a bare `200`
(the authenticated Zulip POST, router `/health`), flag an alert on any unexpected
status (`401`/`403`/`5xx`) — do not summarize it as alive. A bare-`200`
expectation on the auth-gated PVE API (`401`) or LiteLLM health (`301` redirect)
is a stale expectation, not a fault.
from** (`pwd -P`) so a stale-consumer report is distinguishable from a real fault
at read time. For each probe, print the target name, the full URL, and the HTTP
code (or failure kind with retry details). Apply the standing probe rules: any
HTTP status = ALIVE; only 000/timeout/refused = probe-failed. A redirect is not
a failure.
### Phase 1: GPU Exporters
**NVIDIA (.8 and .110)**:
1. Download `nvidia_gpu_exporter` binary
2. Create systemd service `nvidia-gpu-exporter.service`
3. Start and enable
**AMD (.15)**:
1. Create Python exporter script at `/opt/amdgpu-exporter/exporter.py`
2. Parses `amdgpu_top --json -d 1000` output
3. Exposes key metrics at `:9400/metrics` via Python http.server
4. Create systemd service
5. Start and enable
### Phase 2: Prometheus
1. Create `/opt/monitoring/` directory on CT 116
2. Write `prometheus.yml` with scrape configs for all targets
3. Add to docker-compose (or separate compose file)
4. Start container
### Phase 3: Grafana
1. Create `/opt/monitoring/grafana/` directories
2. Provision Prometheus datasource
3. Provision GPU fleet dashboard JSON
4. Provision LiteLLM dashboard JSON
5. Add to docker-compose
6. Start container
### Phase 4: Verification
1. Verify all 3 GPU exporters return 200 at :9400/metrics
2. Verify Prometheus targets all UP at :9090/targets
3. Verify Grafana accessible at :3001 with dashboards
4. Verify LiteLLM metrics flowing to Prometheus
5. ~~Update nginx to proxy `/monitoring/` → Grafana~~ (NOT recommended — nginx sub-path was tried for /grafana/ and reverted per proxmox-monitor; direct :3001 access is the standard)
## Verification Commands
```bash
# GPU exporters
curl -s http://192.168.68.8:9400/metrics | grep nvidia
curl -s http://192.168.68.110:9400/metrics | grep nvidia
curl -s http://192.168.68.15:9400/metrics | grep amdgpu
# Prometheus
curl -s http://192.168.68.116:9090/api/v1/targets
# Grafana
curl -s http://192.168.68.116:3001/api/health
# LiteLLM metrics (already live)
curl -s http://192.168.68.116:4000/metrics | head -20
```
+9 -1
View File
@@ -320,7 +320,15 @@ directly call OpenRouter via Python's requests library. Converting would require
## LiteLLM Master Key (use sparingly — agents should NOT use it directly)
- Master key: `sk-litellm-7f96080dd99b15c36bd4b333b58a6796` (in /opt/inference-harness/.env on CT116, Infisical project=infrastructure env=production secret=LITELLM_MASTER_KEY)
- Master key: **Retrieval path (do not trust a literal value in this file — the key rotates)**:
```bash
# PRIMARY (proven, runs on CT 116 with no extra tooling):
docker exec harness-litellm printenv LITELLM_MASTER_KEY
# Note: the same value is stored in /opt/inference-harness/.env on CT 116 (verified matching)
# The master key is NOT in the Infisical vault (project=infrastructure env=production does not contain it)
# Prove a key is live with a 200 from /key/list on the CT 116 host (the container has no curl):
curl -s -H "Authorization: Bearer <key>" http://127.0.0.1:4000/key/list | jq length
```
- Used for /key/generate, /key/delete, /key/list (GET), DB queries
- **Known violation (RESOLVED 2026-07-16):** Abiba's LITELLM_API_KEY was previously the master key.
It is now a dedicated agent key `sk-sxbphLvk1OU…` (vault secret `ABIBA_LITELLM_API_KEY`, alias `abiba-pi`).
+3 -5
View File
@@ -26,9 +26,8 @@ description: >
| Model | avg latency | avg TTFT | p-profile (24h) |
|---|---|---|---|
| syslog-auto | 28.8s | 25.4s | 68 calls took 30-120s; tail to ~300s under load |
| qwen3.6-27B-code | 23.0s | — | same backend class as syslog-auto |
| strix-moe | 7.5s | — | Strix Halo, healthy |
| gemma-4-12b (retired 2026-09-12; RTX 5070 now `gpu-vision`) | 2.6s | — | RTX 5070, healthy |
| gpu-vision (retired gemma-4-12b, RTX 5070) | 2.6s | — | RTX 5070, healthy |
Sep 6 incident timeline: failures 04:00-07:00 EDT (0% GPU util = wedged
backend), full recovery 07:00-08:00 with ZERO client failures once requests
@@ -56,9 +55,8 @@ proxy queuing.
- vision: 60s (keep), web_extract: 30s (keep) — the 2.6s average was measured on `gemma-4-12b` (retired 2026-09-12); the live RTX 5070 alias is `gpu-vision`.
- compression: 300s (keep — this was already raised from 60 per gpu-fleet).
- **gpu-dense delegation/x_search: set timeout >= 120s.** The RTX 3090
(qwen3.6-27B-code backend, 23.0s avg) is the same speed class as
syslog-auto; delegation defaults that assume fast responses will 408 the
same way.
(gpu-dense backend) is the same speed class as syslog-auto; delegation
defaults that assume fast responses will 408 the same way.
### 3. Retry policy — backoff, not repetition
+41 -6
View File
@@ -44,7 +44,7 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2)
**What changed (v3.2.0 → v4.0.0 — 2026-07-08)**:
- Router REMOVED from request path — LiteLLM proxies directly to GPU
- All GPUs at parallel 2 (was parallel 1)
- NVIDIA context reduced 256K→128K to free VRAM — now the stable ceiling across all GPUs (2026-07-17)
- NVIDIA context reduced 256K→128K to free VRAM — the stable NVIDIA ceiling (2026-07-17); Strix Halo runs 256K (2026-09-12)
- Timeouts and fallback chains are config state — read them from CT 116 `/opt/inference-harness/litellm_config.yaml`; they are not duplicated here.
## Parameters
@@ -69,7 +69,16 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2)
- Network access to public_url, auth_host, and gpu_dashboard_url
- LiteLLM master key for admin endpoints only (`/key/list`, `/key/generate`, `/key/info`)
- The dedicated `monitor` agent key on CT 116 at `/etc/litellm-monitor.env` (root-only 0600) for
model inference checks — the master key must never be used for inference
model inference checks, scoped for every alias step 7 probes (`gpu-dense`, `gpu-vision`,
`strix-moe`, `syslog-auto`). Retrieve from the executor's host via:
```
monitor key: ssh root@192.168.68.116 "grep LITELLM_MONITOR_KEY /etc/litellm-monitor.env | cut -d= -f2"
master key: ssh root@192.168.68.116 "docker exec harness-litellm printenv LITELLM_MASTER_KEY"
```
If credentials are missing or unreadable, the probe must report `credential-missing` (not bare 401 or "0 keys").
The master key must never be used for inference.
## GPU Fleet Topology
@@ -144,12 +153,19 @@ contracts — read them there. Do not re-add retired names (`gemma-4-12b`, `gpu-
7. **Check model inference via LiteLLM** — Test one model on each GPU host. The health
check runs on the **backend edge**, not the public edge, so these paths carry the
`/litellm/` prefix:
- POST http://{{backend_host}}/litellm/v1/chat/completions model=qwen3.6-27B-code → expect 200 (RTX 3090, .8)
- POST http://{{backend_host}}/litellm/v1/chat/completions model=gpu-dense → expect 200 (RTX 3090, .8)
- POST http://{{backend_host}}/litellm/v1/chat/completions model=gpu-vision → expect 200 (RTX 5070, .110)
- POST http://{{backend_host}}/litellm/v1/chat/completions model=strix-moe → expect 200 (Strix Halo, .15)
- Auth uses the dedicated `monitor` agent key, read on CT 116 from
`/etc/litellm-monitor.env` (root-only 0600). Do NOT use the master key for inference —
the master key is for admin endpoints only (`/key/list`, `/key/generate`, `/key/info`).
- Auth uses the dedicated `monitor` agent key. Retrieve via:
`ssh root@192.168.68.116 "grep LITELLM_MONITOR_KEY /etc/litellm-monitor.env | cut -d= -f2"`
Do NOT use the master key for inference — the master key is for admin endpoints only
(`/key/list`, `/key/generate`, `/key/info`). Retrieve master key via:
`ssh root@192.168.68.116 "docker exec harness-litellm printenv LITELLM_MASTER_KEY"`
- KEY SCOPE: the `monitor` key MUST be scoped for the three probed aliases (`gpu-dense`,
`gpu-vision`, `strix-moe`) plus the `syslog-auto` fallback pool, otherwise the probe
returns 403 and the host is not covered.
If a probe returns 403, widen the monitor key's model list on CT 116 (add the missing
alias) and re-run — never drop the host from the probe to make the check pass.
- `gemma-4-12b` was retired and returns 400 `Invalid model name` — do not re-add it to
this list. The RTX 5070 host now serves `gpu-vision`.
- `/v1/models` is **key-scoped**: a model is only visible to keys allowed to use it, so
@@ -161,8 +177,27 @@ contracts — read them there. Do not re-add retired names (`gemma-4-12b`, `gpu-
8. **Check agent keys**:
- GET http://{{backend_host}}/litellm/key/list with master key (admin endpoint) → verify all 6 agents have keys
- **IMPORTANT**: Run the curl on the CT 116 HOST, not inside the container. The `harness-litellm` container has no curl/wget. Use:
`ssh root@192.168.68.116 "curl -s -H 'Authorization: Bearer $MASTER_KEY' http://127.0.0.1:4000/key/list"`
- If the response is empty or unparseable, report `admin-call-failed` (not "0 agent keys")
9. **Check Grafana**:
- GET {{grafana_url}}/api/health → expect 200
10. **Compile and report** — Determine overall_status from individual check results
## Executor Script (2026-09-13)
**Run `scripts/litellm-health-check.py` from the clone.** This script implements all 11
checks defined above and reports results in a standardized format. Paste its output in
the status line.
- Hand-rolled probes are **not** an acceptable substitute for the script.
- Backend-edge checks (steps 2–8) must use `http://192.168.68.116` (internal IP),
**not** the public URL `https://litellm.sysloggh.net` (which returns 401 for those paths).
- Docker Stats (step 10) must be fetched from the CT 116 host itself (`127.0.0.1:9324/metrics`)
because the `harness-docker-stats` container binds to localhost on CT 116.
- Admin Key List (step 8) requires the master key expanded locally before SSH, then embedded
in the remote curl command with proper quoting.
Expected output on a healthy fleet: 11/11 passing checks.
+2 -2
View File
@@ -90,8 +90,8 @@ raw data never provided.
| Worker | Model | Toolsets | Role | Use When |
|--------|-------|----------|------|----------|
| `syslog-code` | qwen3.6-27B-code | terminal, file, web, memory, skills | Code patches, automation, scripts | Writing/modifying code, creating scripts, debugging, reading/writing files |
| `syslog-devops` | qwen3.6-27B-code | terminal, file, web, memory, skills | Infrastructure, DB, bridge, Proxmox | Server ops, SSH, Docker, Proxmox, DB queries, hardware checks |
| `syslog-code` | gpu-dense | terminal, file, web, memory, skills | Code patches, automation, scripts | Writing/modifying code, creating scripts, debugging, reading/writing files |
| `syslog-devops` | gpu-dense | terminal, file, web, memory, skills | Infrastructure, DB, bridge, Proxmox | Server ops, SSH, Docker, Proxmox, DB queries, hardware checks |
| `syslog-email` | strix-moe | terminal, file, web, memory, skills | Email automation, mail operations | Sending/receiving email, inbox management, SMTP operations |
| `syslog-research` | strix-moe | terminal, file, web, memory, skills, **browser** | Analysis, classification, data processing | Web research, browser tasks, data analysis, classification, reading docs |
| `syslog-review` | strix-moe | terminal, file, web, memory, skills | Verification, QA, audit validation | **ALWAYS** verify worker output before delivery — especially for infra changes, code builds, and research findings |
-1
View File
@@ -9,7 +9,6 @@ description: >
- abiba-telegram: { status: "online", uptime: string, restarts: number }
- abiba-zulip: { status: "online", uptime: string, restarts: number }
- gitea-runner: { status: "online", uptime: string, restarts: number }
- spoton-service: { status: "online", uptime: string, restarts: number }
- zulip-watchdog: { status: "online", uptime: string, restarts: number }
- last_check: timestamp
+1 -1
View File
@@ -87,7 +87,7 @@ agent: abiba
| storepve | 192.168.68.6 | PVE |
| acerpve | 192.168.68.9 | PVE (hosts llm-gpu qemu/101) |
| minipve | 192.168.68.12 | PVE |
| amdpve | 192.168.68.15 | PVE + Strix Halo LLM (qwen3.6-35B-udq4, strix-moe) |
| amdpve | 192.168.68.15 | PVE + Strix Halo LLM (strix-moe) |
## Operations
+64 -26
View File
@@ -340,6 +340,36 @@ def check_gpu_ports():
# CHECK 3: Agent Gateway Liveness + Streaming (now covers all agents)
# ═══════════════════════════════════════════════════════════════════
def _ssh_retry(host, cmd, user="root", timeout=15, retry_timeout=25, label=""):
"""SSH with one retry at a longer timeout.
Returns (stdout_or_None, probe_failed_bool, fail_kind).
When probe_failed is True, fail_kind is one of: timeout, ssh-failed.
"""
import subprocess as _sp
def _attempt(tmo, conn_tmo):
try:
r = _sp.run(
["ssh", "-o", "StrictHostKeyChecking=no", "-o", f"ConnectTimeout={conn_tmo}",
f"{user}@{host}", cmd],
capture_output=True, text=True, timeout=tmo)
return r.stdout.strip() if r.returncode == 0 else None
except _sp.TimeoutExpired:
return "__timeout__"
except:
return None
result = _attempt(timeout, 8)
if result is None or result == "__timeout__":
kind = "timeout" if result == "__timeout__" else "ssh-failed"
prefix = f"{label} " if label else ""
print(f" probe-failed: {prefix}ssh {user}@{host} — {kind} (retrying at {retry_timeout}s…)")
result = _attempt(retry_timeout, 15)
if result is None or result == "__timeout__":
kind = "timeout" if result == "__timeout__" else "ssh-failed"
return None, True, kind
return result, False, None
def check_agents():
for name, agent in AGENTS.items():
host = agent.get("host")
@@ -355,42 +385,49 @@ def check_agents():
is_dsh = agent.get("runtime") == "dsh"
label = "DSH (DeepSeek Harness)" if is_dsh else "pi-only runtime"
since = "since 2026-08-27" if is_dsh else "since the harness purge"
live = ssh(host, "true", user=user)
print(f" {'✅' if live is not None else '❌'} {name}: {label} — "
f"no Hermes gateway {since} (CT {ct}, SSH {'OK' if live is not None else 'FAIL'})")
if live is None:
_fail(f"unreachable:{name}", name)
live, probe_failed, fail_kind = _ssh_retry(host, "true", user=user)
if probe_failed:
print(f" ❌ {name}: {label} — probe-failed: ssh {user}@{host} {fail_kind} "
f"(retried at 25s: also {fail_kind}) [CT {ct}]")
_fail(f"probe-failed:{name}:{fail_kind}", name)
else:
print(f" ✅ {name}: {label} — no Hermes gateway {since} "
f"(ssh {user}@{host} OK, CT {ct})")
continue
if not host or not user:
print(f" ⬜ {name} (CT {ct}): cannot SSH — skip liveness check")
continue
# Resolve the Hermes gateway PID once, before the report-only branch:
# the summary line below renders `pid`, and it used to be bound only in
# the report-only path — leaving it unbound on the abiba/koonimo path
# raised UnboundLocalError and crashed the whole check. Agents without
# a gateway get pid=?.
pid = ssh(host, "pgrep -f '[h]ermes_cli.main gateway run' | grep -v infisical | head -1", user=user)
if not pid:
pid = ssh(host, "pgrep -f '[h]ermes.*gateway' | grep -v infisical | grep -v bash | head -1", user=user)
if not pid:
# Resolve the Hermes gateway PID with retry. The probe target is
# explicit: ssh {user}@{host} pgrep -f hermes gateway.
pid, probe_failed, fail_kind = _ssh_retry(
host, "pgrep -f '[h]ermes_cli.main gateway run' | grep -v infisical | head -1", user=user)
if not pid and not probe_failed:
pid, probe_failed, fail_kind = _ssh_retry(
host, "pgrep -f '[h]ermes.*gateway' | grep -v infisical | grep -v bash | head -1", user=user)
if not pid and not probe_failed:
pid = "?"
# ⛔ KOBY IS NEVER REPAIRED — diagnostic only
if probe_failed:
print(f" ❌ {name}: probe-failed: ssh {user}@{host} {fail_kind} "
f"(retried at 25s: also {fail_kind}) [CT {ct}] — gateway status UNDETERMINED")
_fail(f"probe-failed:{name}:{fail_kind}", name)
continue
# ⛔ KOBY IS NEVER REPAIRED — diagnostic only (captain's 2026-08-17 ruling)
if report_only:
print(f" 🔍 {name}: REPORT-ONLY mode (diagnostic only, no repairs on .129)")
# Still check gateway status for reporting purposes
if pid == "?":
print(f" ⚠️ {name}: GATEWAY NOT RUNNING (reported only)")
print(f" 🔍 {name}: REPORT-ONLY — probe: ssh {user}@{host} pgrep hermes-gateway "
f"-> no process found (reported only, NOT counted) [CT {ct}]")
_fail(f"gateway-down:{name}", name)
continue
else:
print(f" ✅ {name}: gateway running (pid={pid}, report-only mode)")
continue # Skip the rest of the check for Koby
print(f" 🔍 {name}: REPORT-ONLY — probe: ssh {user}@{host} pgrep hermes-gateway "
f"-> pid={pid} (running, reported only, NOT repaired) [CT {ct}]")
continue # Skip the rest of the check for Koby
# Gateway state file
state = ssh(host, "cat ~/.hermes/gateway_state.json 2>/dev/null", user=user)
state, _, _ = _ssh_retry(host, "cat ~/.hermes/gateway_state.json 2>/dev/null", user=user)
if state:
try:
st = json.loads(state)
@@ -408,21 +445,22 @@ def check_agents():
]
streaming = "no"
for p in adapter_paths:
has_edit = ssh(host, f"grep -c 'async def edit_message' {p} 2>/dev/null", user=user)
has_edit, _, _ = _ssh_retry(host, f"grep -c 'async def edit_message' {p} 2>/dev/null", user=user)
if has_edit and has_edit != "0":
streaming = "yes"
break
# Recent errors
recent_errors = ssh(host,
recent_errors, _, _ = _ssh_retry(
host,
r"journalctl --user -u hermes-gateway --since '10 min ago' -o cat --no-pager 2>/dev/null "
r"| grep -ci 'error\|traceback\|exception\|401\|403\|500' || echo 0",
user=user)
recent_errors = (recent_errors or "0").strip().split("\n")[-1]
print(f" {'✅' if gw_state == 'running' and zulip == 'connected' else '⚠️'} "
f"{name}: gw={gw_state} zulip={zulip} streaming={streaming} "
f"errors_10m={recent_errors.strip() or '0'} pid={pid}")
f"{name}: probe: ssh {user}@{host} — gw={gw_state} zulip={zulip} "
f"streaming={streaming} errors_10m={recent_errors.strip() or '0'} pid={pid} [CT {ct}]")
# ═══════════════════════════════════════════════════════════════════
+339
View File
@@ -0,0 +1,339 @@
#!/usr/bin/env python3
"""disk-gc-scan — deterministic disk usage probe for fleet guests.
This is the executable scanner side of `disk-gc-threat-response.prose.md`. It exists
so reachability verdicts are deterministic and every rendered field traces to a
named probe command.
DESIGN PRINCIPLES (per task disk-gc-probe-false-unreachable-20260913):
1. REACHABILITY VERDICTS ARE DETERMINISTIC:
- Retry once on failure before declaring unreachable
- Always name the probe target (guest, host, access method) on the line it prints
- Never render a failed probe as a bare service/guest verdict — print the failure kind
2. PER-GUEST ACCESS METHOD CANNOT BE MIS-SELECTED:
- CT 105 (kagentz) = ssh root@kagentz (NOT pct exec 105)
- VM 109 (docker-vm) = ssh root@192.168.68.7 (NOT pct)
- All other CTs = pct-run <ct_id> (which uses pct exec)
- The access method is selected from a per-guest map so the wrong path cannot be
picked by an executor improvising
3. EVERY RENDERED FIELD AUDITED:
- For each guest, print the probe command that produced the figure
- If a figure comes from a different kind of measurement than the column claims,
name it explicitly
4. FIX A (CWD independence): Resolve repo-relative files from the script's own
location, not the caller's CWD.
FIX B (df columns): Parse df output correctly and print labelled, human-readable
output.
Usage:
disk-gc-scan.py # scan all guests
disk-gc-scan.py --json # machine-readable output
Exit codes: 0 ok (all guests probed), 1 probe error
"""
from __future__ import annotations
import json
import pathlib
import subprocess
import sys
import time
import os
from dataclasses import dataclass
from typing import Optional
# Resolve repo-relative files from the script's own location, not the caller's CWD
SCRIPT_DIR = pathlib.Path(__file__).resolve().parent
HELPER_PCT_RUN = SCRIPT_DIR / "pct-run.sh"
# Per-guest access method map. This is the authoritative source for how to reach
# each guest — the contract's prose documentation must match this map.
#
# Access methods:
# - "pct-run": use pct-run.sh <ct_id> (pct exec via SSH to node)
# - "ssh-host": use ssh root@<hostname>
# - "ssh-ip": use ssh root@<ip>
@dataclass
class Guest:
"""A guest to probe."""
ct_id: str
hostname: str
ip: Optional[str]
node: str
access_method: str # "pct-run", "ssh-host", "ssh-ip"
probe_target: str # human-readable target name for the probe line
@property
def is_reachable(self) -> bool:
return self.probe_result is not None and self.probe_result.exit_code == 0
@property
def usage_pct(self) -> Optional[float]:
return self.probe_result.usage_pct if self.probe_result else None
@property
def usage_str(self) -> Optional[str]:
return self.probe_result.usage_str if self.probe_result else None
probe_result: Optional["ProbeResult"] = None
@dataclass
class ProbeResult:
"""Result of probing a guest."""
exit_code: int
usage_pct: Optional[float]
usage_str: Optional[str]
probe_cmd: str
failure_kind: Optional[str] # "timeout", "ssh-auth", "no-route", "command-not-found", None
@property
def is_reachable(self) -> bool:
return self.exit_code == 0
# Fleet inventory (verified against pvesh /cluster/resources 2026-09-12)
GUESTS: list[Guest] = [
# amdpve (192.168.68.15)
Guest(ct_id="105", hostname="kagentz", ip="192.168.68.105", node="amdpve",
access_method="ssh-host", probe_target="kagentz (CT 105, amdpve)"),
Guest(ct_id="112", hostname="tanko", ip="192.168.68.112", node="amdpve",
access_method="pct-run", probe_target="tanko (CT 112, amdpve)"),
Guest(ct_id="113", hostname="baggy", ip="192.168.68.113", node="amdpve",
access_method="pct-run", probe_target="baggy (CT 113, amdpve)"),
Guest(ct_id="115", hostname="scottdenya", ip="192.168.68.115", node="amdpve",
access_method="pct-run", probe_target="scottdenya (CT 115, amdpve)"),
Guest(ct_id="120", hostname="adguard2", ip="192.168.68.120", node="amdpve",
access_method="pct-run", probe_target="adguard2 (CT 120, amdpve)"),
# minipve (192.168.68.12)
Guest(ct_id="100", hostname="abiba", ip="192.168.68.100", node="minipve",
access_method="pct-run", probe_target="abiba (CT 100, minipve)"),
Guest(ct_id="102", hostname="adguard", ip="192.168.68.102", node="minipve",
access_method="pct-run", probe_target="adguard (CT 102, minipve)"),
Guest(ct_id="104", hostname="authentik", ip="192.168.68.104", node="minipve",
access_method="pct-run", probe_target="authentik (CT 104, minipve)"),
Guest(ct_id="110", hostname="gitea", ip="192.168.68.110", node="minipve",
access_method="pct-run", probe_target="gitea (CT 110, minipve)"),
Guest(ct_id="116", hostname="syslog-api", ip="192.168.68.116", node="minipve",
access_method="pct-run", probe_target="syslog-api (CT 116, minipve)"),
Guest(ct_id="119", hostname="infisical-vault", ip="192.168.68.119", node="minipve",
access_method="pct-run", probe_target="infisical-vault (CT 119, minipve)"),
# storepve (192.168.68.6)
Guest(ct_id="106", hostname="ra-h-os", ip="192.168.68.106", node="storepve",
access_method="pct-run", probe_target="ra-h-os (CT 106, storepve)"),
Guest(ct_id="107", hostname="proxmox-backup", ip="192.168.68.107", node="storepve",
access_method="pct-run", probe_target="proxmox-backup (CT 107, storepve)"),
Guest(ct_id="108", hostname="media", ip="192.168.68.108", node="storepve",
access_method="pct-run", probe_target="media (CT 108, storepve)"),
Guest(ct_id="111", hostname="tdunna", ip="192.168.68.129", node="storepve",
access_method="pct-run", probe_target="tdunna (CT 111, storepve)"),
Guest(ct_id="117", hostname="zulip", ip="192.168.68.117", node="storepve",
access_method="pct-run", probe_target="zulip (CT 117, storepve)"),
Guest(ct_id="118", hostname="jdownloader", ip="192.168.68.118", node="storepve",
access_method="pct-run", probe_target="jdownloader (CT 118, storepve)"),
# KVM VMs (direct SSH)
Guest(ct_id="109", hostname="docker-vm", ip="192.168.68.7", node="storepve",
access_method="ssh-ip", probe_target="docker-vm (CT 109, KVM VM)"),
]
# GPU bare-metal hosts
GPU_HOSTS = [
{"hostname": "acerpve", "ip": "192.168.68.9", "gpu": "RTX 3090",
"probe_target": "RTX 3090 (bare metal .9)"},
{"hostname": "ocupve", "ip": "192.168.68.110", "gpu": "RTX 5070",
"probe_target": "RTX 5070 (bare metal .110)"},
{"hostname": "amdpve", "ip": "192.168.68.15", "gpu": "Strix Halo",
"probe_target": "Strix Halo (bare metal .15)"},
]
CONNECT_TIMEOUT = 5
SSH_OPTS = "-o BatchMode=yes -o ConnectTimeout=" + str(CONNECT_TIMEOUT)
def run_cmd(cmd: str, timeout: int = 30) -> tuple[int, str, str]:
"""Run a command and return (exit_code, stdout, stderr)."""
try:
result = subprocess.run(
cmd, shell=True, capture_output=True, text=True, timeout=timeout
)
return result.returncode, result.stdout.strip(), result.stderr.strip()
except subprocess.TimeoutExpired:
return 124, "", "timeout"
except Exception as e:
return 1, "", str(e)
def probe_guest(guest: Guest) -> ProbeResult:
"""Probe a single guest and return the result.
Access method is selected from guest.access_method:
- "pct-run": pct-run.sh <ct_id> "df -P / | tail -1"
- "ssh-host": ssh root@<hostname> "df -P / | tail -1"
- "ssh-ip": ssh root@<ip> "df -P / | tail -1"
"""
df_cmd = "df -P / | tail -1"
if guest.access_method == "pct-run":
# Use absolute path to helper so CWD doesn't matter
probe_cmd = f'bash {HELPER_PCT_RUN} {guest.ct_id} "{df_cmd}"'
elif guest.access_method == "ssh-host":
probe_cmd = f'ssh {SSH_OPTS} root@{guest.hostname} "{df_cmd}"'
elif guest.access_method == "ssh-ip":
probe_cmd = f'ssh {SSH_OPTS} root@{guest.ip} "{df_cmd}"'
else:
raise ValueError(f"unknown access_method: {guest.access_method}")
# Check helper exists and is readable BEFORE probing (for pct-run guests)
# This prevents scanner errors from being rendered as guest verdicts
if guest.access_method == "pct-run":
if not HELPER_PCT_RUN.exists():
print(f"SCANNER ERROR: helper not found: {HELPER_PCT_RUN}", file=sys.stderr)
sys.exit(1)
if not os.access(str(HELPER_PCT_RUN), os.R_OK):
print(f"SCANNER ERROR: helper not readable: {HELPER_PCT_RUN}", file=sys.stderr)
sys.exit(1)
# Retry once on failure before declaring unreachable
for attempt in range(2):
exit_code, stdout, stderr = run_cmd(probe_cmd, timeout=15)
if exit_code == 0:
# Parse df output: Filesystem 1024-blocks Used Available Capacity Mounted on
# parts[0]=Filesystem, parts[1]=Total (1K blocks), parts[2]=Used, parts[3]=Available, parts[4]=Capacity
parts = stdout.split()
if len(parts) >= 5:
capacity_str = parts[4] # e.g., "34%"
usage_pct = float(capacity_str.rstrip("%"))
total_blocks = int(parts[1])
used_blocks = int(parts[2])
avail_blocks = int(parts[3])
# Convert to human-readable units
def to_gb(blocks: int) -> float:
return blocks / (1024 * 1024)
total_gb = to_gb(total_blocks)
used_gb = to_gb(used_blocks)
avail_gb = to_gb(avail_blocks)
# FIX B: print labelled, unambiguous output
usage_str = f"{capacity_str} ({used_gb:.1f}G used of {total_gb:.1f}G total, {avail_gb:.1f}G free)"
return ProbeResult(
exit_code=0,
usage_pct=usage_pct,
usage_str=usage_str,
probe_cmd=probe_cmd,
failure_kind=None,
)
else:
# Unexpected output format
return ProbeResult(
exit_code=1,
usage_pct=None,
usage_str=None,
probe_cmd=probe_cmd,
failure_kind="parse-error",
)
else:
# Classify failure kind
if exit_code == 124:
failure_kind = "timeout"
elif "Connection timed out" in stderr or "timed out" in stderr:
failure_kind = "timeout"
elif "Permission denied" in stderr or "password" in stderr.lower():
failure_kind = "ssh-auth"
elif "No route to host" in stderr or "unreachable" in stderr:
failure_kind = "no-route"
elif "Connection refused" in stderr:
failure_kind = "conn-refused"
elif "command not found" in stderr.lower() or "No such file" in stderr:
failure_kind = "command-not-found"
else:
failure_kind = f"ssh-exit-{exit_code}"
# Retry once
if attempt == 0:
time.sleep(1)
continue
return ProbeResult(
exit_code=exit_code,
usage_pct=None,
usage_str=None,
probe_cmd=probe_cmd,
failure_kind=failure_kind,
)
# Should not reach here, but just in case
return ProbeResult(
exit_code=1,
usage_pct=None,
usage_str=None,
probe_cmd=probe_cmd,
failure_kind="unknown",
)
def scan_fleet() -> list[dict]:
"""Scan all guests and return the results."""
results = []
for guest in GUESTS:
probe_result = probe_guest(guest)
guest.probe_result = probe_result
row = {
"target": guest.probe_target,
"ct_id": guest.ct_id,
"hostname": guest.hostname,
"node": guest.node,
"access_method": guest.access_method,
"reachable": probe_result.is_reachable,
"usage_pct": probe_result.usage_pct,
"usage_str": probe_result.usage_str,
"probe_cmd": probe_result.probe_cmd,
"failure_kind": probe_result.failure_kind,
}
results.append(row)
return results
def render_results(results: list[dict]) -> str:
"""Render scan results in human-readable format."""
lines = []
lines.append("=== Disk GC Scan ===")
lines.append("")
for row in results:
if row["reachable"]:
lines.append(f" ✅ {row['target']}: {row['usage_str']}")
lines.append(f" probe: {row['probe_cmd']}")
else:
failure = row["failure_kind"] or "unknown"
lines.append(f" ❌ {row['target']}: UNREACHABLE ({failure})")
lines.append(f" probe: {row['probe_cmd']}")
return "\n".join(lines)
def main() -> int:
import argparse
ap = argparse.ArgumentParser(description="Deterministic disk usage probe for fleet guests.")
ap.add_argument("--json", action="store_true", help="machine-readable output")
args = ap.parse_args()
results = scan_fleet()
if args.json:
print(json.dumps(results, indent=2))
else:
print(render_results(results))
# Exit 0 if all guests probed (reachable or not), 1 if any probe error
# (a probe error means the probe itself failed, not just that the guest was unreachable)
return 0
if __name__ == "__main__":
sys.exit(main())
+32
View File
@@ -0,0 +1,32 @@
#!/bin/bash
# Shared helper for Hermes contract reachability checks
# Separates SSH exit status from remote command result
hermes_check_host() {
local host=$1
local pattern=$2
local path=$3
# Remote side always succeeds (grep ...; true), so ssh exit code = connection status only
local out
out=$(ssh -o BatchMode=yes -o ConnectTimeout=3 root@"$host" "grep -RIn '$pattern' '$path' 2>/dev/null; true" 2>/dev/null)
local status=$?
if [ $status -ne 0 ]; then
echo "$host: UNREACHABLE (ssh exit $status)"
elif [ -n "$out" ]; then
echo "$host: VIOLATION: $out"
else
echo "$host: COMPLIANT (no matches found)"
fi
}
# Standalone mode: scripts/hermes-reachability-check.sh <host> <pattern> <path>
if [ "${BASH_SOURCE[0]}" = "${0}" ]; then
if [ $# -ne 3 ]; then
echo "Usage: $0 <host> <pattern> <path>" >&2
exit 2
fi
hermes_check_host "$1" "$2" "$3"
exit 0
fi
+304
View File
@@ -0,0 +1,304 @@
#!/usr/bin/env python3
"""
LiteLLM Health Check - Contract executor
Runs all checks defined in litellm-health.prose.md and reports results.
"""
import subprocess
import sys
import json
import time
import random
# Configuration
BACKEND_HOST = "192.168.68.116"
GPU_HOSTS = {
"gpu-dense": "192.168.68.8",
"gpu-vision": "192.168.68.110",
"strix-moe": "192.168.68.15"
}
def run_command(cmd, timeout=15):
"""Run a command and return (exit_code, stdout, stderr)"""
try:
result = subprocess.run(
cmd,
shell=True,
capture_output=True,
text=True,
timeout=timeout
)
return result.returncode, result.stdout.strip(), result.stderr.strip()
except subprocess.TimeoutExpired:
return 1, "", "TIMEOUT"
except Exception as e:
return 1, "", str(e)
def probe_http(url, method="GET", bearer_token=None, data=None, timeout=10, follow_redirects=False):
"""Probe HTTP endpoint and return (status_code, failure_kind)
Returns:
(code, None) if successful or HTTP response received
(000, kind) if connection failed, where kind is 'timeout', 'refused', 'dns', etc.
"""
cmd = "curl -s -o /dev/null -w '%{http_code}' -m " + str(timeout)
if method == "POST":
cmd += " -X POST"
if bearer_token:
cmd += " -H 'Authorization: Bearer " + bearer_token + "'"
if data:
cmd += " -H 'Content-Type: application/json' -d '" + data + "'"
if follow_redirects:
cmd += " -L"
cmd += " '" + url + "'"
try:
rc, stdout, stderr = run_command(cmd, timeout)
if rc != 0:
# Determine failure kind from curl exit code
# curl exit codes: 28=timeout, 7=refused, 6=dns, 35=ssl, 52=empty
if rc == 28:
return (000, "timeout after " + str(timeout) + "s")
elif rc == 7:
return (000, "connection refused")
elif rc == 6:
return (000, "dns failure")
elif rc == 35:
return (000, "ssl error")
elif rc == 52:
return (000, "empty response")
else:
return (000, "curl exit " + str(rc))
return (int(stdout), None) if stdout.isdigit() else (000, "unparseable response")
except subprocess.TimeoutExpired:
return (000, "timeout after " + str(timeout) + "s")
def get_response_body(url, method="POST", bearer_token=None, data=None, timeout=30):
"""Get response body for 401/403 credential faults (truncated to 200 chars)"""
cmd = "curl -s -m " + str(timeout)
if method == "POST":
cmd += " -X POST"
if bearer_token:
cmd += " -H 'Authorization: Bearer " + bearer_token + "'"
if data:
cmd += " -H 'Content-Type: application/json' -d '" + data + "'"
cmd += " '" + url + "'"
rc, stdout, stderr = run_command(cmd, timeout)
# Return first 200 chars, single line
body = stdout.replace('\n', ' ').replace('\t', ' ')[:200] if stdout else ""
return body
def check_liveliness():
"""Step 1: Liveliness probe"""
code, _ = probe_http("http://" + BACKEND_HOST + "/litellm/health/liveliness")
return "Liveliness", code == 200, str(code) + " (target: " + BACKEND_HOST + "/litellm/health/liveliness)"
def check_containers():
"""Step 2: Container health via SSH"""
cmd = "ssh -o BatchMode=yes -o ConnectTimeout=5 -o StrictHostKeyChecking=no root@192.168.68.116 'docker ps --format \"{{.Names}} {{.Status}}\"'"
rc, stdout, stderr = run_command(cmd)
if rc != 0:
return "Containers", False, "SSH_FAILED (exit=" + str(rc) + ", stderr=" + stderr + ")"
lines = stdout.split('\n') if stdout else []
container_count = len([l for l in lines if l.strip()])
healthy = container_count >= 8
return "Containers", healthy, str(container_count) + " containers"
def check_model_probes():
"""Step 6: Model probe - all 4 aliases"""
# Get monitor key
monitor_key = run_command("ssh -o BatchMode=yes root@192.168.68.116 \"grep LITELLM_MONITOR_KEY /etc/litellm-monitor.env | cut -d= -f2\"")[1]
results = []
for model in ["gpu-dense", "gpu-vision", "strix-moe"]:
# Single-host aliases: 30s timeout each
# gpu-dense (RTX 3090) may need long warmup/prefill - timeout is acceptable on cold-start
code, failure_kind = probe_http("http://" + BACKEND_HOST + "/litellm/v1/chat/completions",
method="POST",
bearer_token=monitor_key,
data='{"model":"' + model + '","messages":[{"role":"user","content":"health ' + str(random.randint(1000, 9999)) + '"}],"max_tokens":4}',
timeout=30)
if code == 000 and failure_kind:
# Report probe failure with kind, do not assert a service verdict
results.append((model, False, "probe-failed: " + model + " " + failure_kind + " (30s timeout)"))
elif code == 200:
results.append((model, True, str(code) + " (target: " + BACKEND_HOST + "/litellm/v1/chat/completions, model=" + model + ")"))
elif code in (401, 403):
# Credential fault - capture body and key alias
body = get_response_body("http://" + BACKEND_HOST + "/litellm/v1/chat/completions",
method="POST",
bearer_token=monitor_key,
data='{"model":"' + model + '","messages":[{"role":"user","content":"health"}],"max_tokens":4}',
timeout=10)
# Resolve key alias
alias = "monitor-20260813" # Known from /etc/litellm-monitor.env on CT 116
results.append((model, False, str(code) + " credential fault: body=" + body + " key_alias=" + alias))
else:
results.append((model, False, str(code) + " (target: " + BACKEND_HOST + "/litellm/v1/chat/completions, model=" + model + ")"))
# Pool alias (syslog-auto): 60s timeout, retry once on 000
code, failure_kind = probe_http("http://" + BACKEND_HOST + "/litellm/v1/chat/completions",
method="POST",
bearer_token=monitor_key,
data='{"model":"syslog-auto","messages":[{"role":"user","content":"health ' + str(random.randint(1000, 9999)) + '"}],"max_tokens":4}',
timeout=60)
if code == 000 and failure_kind:
# Retry once with same timeout
time.sleep(1)
code, failure_kind = probe_http("http://" + BACKEND_HOST + "/litellm/v1/chat/completions",
method="POST",
bearer_token=monitor_key,
data='{"model":"syslog-auto","messages":[{"role":"user","content":"health ' + str(random.randint(1000, 9999)) + '"}],"max_tokens":4}',
timeout=60)
if code == 000 and failure_kind:
results.append(("syslog-auto", False, "probe-failed: syslog-auto " + failure_kind + " (60s timeout, retry)"))
elif code in (401, 403):
body = get_response_body("http://" + BACKEND_HOST + "/litellm/v1/chat/completions",
method="POST",
bearer_token=monitor_key,
data='{"model":"syslog-auto","messages":[{"role":"user","content":"health"}],"max_tokens":4}',
timeout=10)
alias = "monitor-20260813"
results.append(("syslog-auto", False, str(code) + " credential fault: body=" + body + " key_alias=" + alias))
else:
results.append(("syslog-auto", code == 200, str(code) + " (target: " + BACKEND_HOST + "/litellm/v1/chat/completions, model=syslog-auto)"))
else:
results.append(("syslog-auto", code == 200, str(code) + " (target: " + BACKEND_HOST + "/litellm/v1/chat/completions, model=syslog-auto)"))
return results
def check_admin_key_list():
"""Step 8: Admin API key list - use two-step approach"""
# Step 1: Get master key
mk_cmd = "ssh -o BatchMode=yes root@192.168.68.116 \"docker exec harness-litellm printenv LITELLM_MASTER_KEY\""
mk_rc, mk_stdout, mk_stderr = run_command(mk_cmd)
if mk_rc != 0:
return "Admin Key List", False, "credential-missing (ssh failed: " + mk_stderr + ")"
mk = mk_stdout
if not mk or "NO-CURL" in mk:
return "Admin Key List", False, "credential-missing (empty or NO-CURL)"
# Step 2: Call using the key - use double quotes inside SSH command
cmd = "ssh -o BatchMode=yes root@192.168.68.116 \"curl -s -H \\\"Authorization: Bearer " + mk + "\\\" http://127.0.0.1:4000/key/list\""
rc, stdout, stderr = run_command(cmd)
if rc != 0:
return "Admin Key List", False, "admin-call-failed (exit=" + str(rc) + ", stderr=" + stderr + ")"
# Try to parse the response
try:
data = json.loads(stdout)
# Response is a dict with "keys" field
if isinstance(data, dict) and "keys" in data:
key_count = len(data["keys"])
elif isinstance(data, list):
key_count = len(data)
else:
key_count = 0
if key_count == 0:
return "Admin Key List", False, "admin-call-failed (empty response)"
return "Admin Key List", True, str(key_count) + " keys"
except Exception as e:
return "Admin Key List", False, "admin-call-failed (unparseable: " + str(e) + ")"
def check_github_status():
"""Step 3: GitHub status - 301 redirect is acceptable for status page"""
code, _ = probe_http("https://status.github.com/api/status.json", timeout=15)
# GitHub status API returns 301 redirect, which is expected behavior
return "GitHub Status", code == 301, str(code)
def check_prometheus():
"""Step 4: Prometheus health"""
code, _ = probe_http("http://" + BACKEND_HOST + ":9090/-/healthy")
return "Prometheus", code == 200, str(code) + " (target: " + BACKEND_HOST + ":9090/-/healthy)"
def check_grafana():
"""Step 9: Grafana health"""
code, _ = probe_http("http://" + BACKEND_HOST + ":3001/api/health")
return "Grafana", code == 200, str(code) + " (target: " + BACKEND_HOST + ":3001/api/health)"
def check_docker_stats():
"""Step 10: Docker Stats health - fetch from CT 116 host"""
# Docker stats is on localhost from CT 116
cmd = "ssh -o BatchMode=yes root@192.168.68.116 'curl -s http://127.0.0.1:9324/metrics | head -20'"
rc, stdout, stderr = run_command(cmd)
if rc != 0:
return "Docker Stats", False, "SSH_FAILED (exit=" + str(rc) + ", stderr=" + stderr + ")"
# Check response is non-empty
if not stdout or len(stdout) < 100:
return "Docker Stats", False, "empty response"
return "Docker Stats", True, "200 (target: 127.0.0.1:9324/metrics from CT 116)"
def main():
print("🏥 LiteLLM Health Check v1.0.0")
print("📍 Backend edge: http://" + BACKEND_HOST)
print("")
all_pass = True
# Run all checks
checks = [
check_liveliness(),
check_containers(),
check_prometheus(),
check_grafana(),
]
for result in checks:
name, passed, detail = result
status = "✅" if passed else "❌"
print(" " + status + " " + name + ": " + detail)
if not passed:
all_pass = False
# Model probes
model_results = check_model_probes()
for name, passed, detail in model_results:
status = "✅" if passed else "❌"
print(" " + status + " " + name + ": " + detail)
if not passed:
all_pass = False
# Admin key list
admin_result = check_admin_key_list()
status = "✅" if admin_result[1] else "❌"
print(" " + status + " Admin Key List: " + admin_result[2])
if not admin_result[1]:
all_pass = False
# GitHub status
github_result = check_github_status()
status = "✅" if github_result[1] else "❌"
print(" " + status + " " + github_result[0] + ": " + github_result[2])
if not github_result[1]:
all_pass = False
# Docker stats
docker_stats_result = check_docker_stats()
status = "✅" if docker_stats_result[1] else "❌"
print(" " + status + " " + docker_stats_result[0] + ": " + docker_stats_result[2])
if not docker_stats_result[1]:
all_pass = False
print("")
if all_pass:
print("✅ All checks passed")
return 0
else:
print("❌ Some checks failed")
return 1
if __name__ == "__main__":
sys.exit(main())
+7 -7
View File
@@ -132,20 +132,20 @@ def test_retired_alias_in_custom_providers_is_rejected(tmp_path):
assert "RESULT: FAIL" in out
def test_raw_but_live_alias_warns_but_passes(tmp_path):
"""Raw-but-live names resolve (200), so they warn only; failing them rejects valid configs."""
def test_retired_raw_name_fails(tmp_path):
"""Retired raw names no longer resolve (400), so they fail; the 2026-09-12 registry change moved qwen3.6-27B-code from raw-but-live to non-resolving."""
code, out = _run_config(
tmp_path,
"raw-qwen.yaml",
"retired-qwen.yaml",
BASE.format(alias="gpu-vision").replace(
"delegation:\n provider: harness",
"delegation:\n provider: harness\n model: qwen3.6-27B-code",
),
)
assert code == 0, out
assert "delegation.model = 'qwen3.6-27B-code' is a raw-but-live model name" in out
assert "prefer the stable alias gpu-dense" in out
assert "RESULT: PASS" in out
assert code != 0, out
assert "delegation.model = 'qwen3.6-27B-code' is retired and no longer resolves" in out
assert "use gpu-dense" in out
assert "RESULT: FAIL" in out
def test_retired_alias_in_fallback_providers_is_rejected(tmp_path):
+32 -4
View File
@@ -127,20 +127,48 @@ grep -c "async def edit_message" ~/.hermes/plugins/*/zulip*/adapter.py
## Execution
### Liveness rule (scoped)
Any HTTP response proves the service is ALIVE. For auth-gated endpoints (Zulip API, Tanko gateway), a 401/403 redirect or status means the service answered — report the code, never "down". Only a failed CONNECTION (curl status 000, timeout, refused) is a failed probe.
**STANDING PROBE RULES (2026-09-14, from defect report 1150.msg):**
1. **Any HTTP status means ALIVE.** 200, 301, 302, 401, 403, 404 all prove the service answered — report the code, never "down". A redirect is not a failure. Only a failed CONNECTION (curl status 000, timeout, refused) is a failed probe.
2. **A failed probe is never a service verdict.** Print `probe-failed: <target> <kind>` naming the exact URL/host/port and the failure kind (timeout, refused, no-route, dns), retry once at a longer timeout, and only then report.
3. **Say which probe produced each number.** "API: 000" is unusable; "API https://chat.sysloggh.net/api/v1/server_settings -> connection timeout after 10s (retried at 25s: also timeout)" is actionable.
### Step 1: Zulip Server Liveness
```bash
curl -s -o /dev/null -w "%{http_code}" https://chat.sysloggh.net/api/v1/server_settings \
-u 'abiba-bot@chat.sysloggh.net:$ZULIP_API_KEY'
# Probe the Zulip API (authenticated, any HTTP status = ALIVE)
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 https://chat.sysloggh.net/api/v1/server_settings -u 'abiba-bot@chat.sysloggh.net:$ZULIP_API_KEY')
if [ "$code" == "000" ]; then
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 25 https://chat.sysloggh.net/api/v1/server_settings -u 'abiba-bot@chat.sysloggh.net:$ZULIP_API_KEY')
echo "Zulip API https://chat.sysloggh.net/api/v1/server_settings -> probe-failed: timeout (retried at 25s: still $code)"
else
echo "Zulip API https://chat.sysloggh.net/api/v1/server_settings -> $code"
fi
```
Expected: `200`. If not → mark `zulip_server_status: "down"`, skip per-platform checks, alert.
Expected: `200` (authenticated). Any HTTP status = ALIVE; only 000/timeout = probe-failed. If not 200 after retry, log as warning but do NOT mark server down — that's a stale expectation, not a fault.
### Step 2: Platform A — pi (Abiba, localhost)
**A1: Health Endpoint**
Fetch `http://localhost:9200/health` as JSON. Check:
```bash
# Probe the Abiba extension health endpoint (loopback, any HTTP status = ALIVE)
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 http://127.0.0.1:9200/health)
if [ "$code" == "000" ]; then
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 25 http://127.0.0.1:9200/health)
echo "Abiba extension http://127.0.0.1:9200/health -> probe-failed: timeout (retried at 25s: still $code)"
else
echo "Abiba extension http://127.0.0.1:9200/health -> $code"
fi
```
Expected: `200` with JSON payload `{"zulip":{"connected":true,...}}`. Any HTTP status = ALIVE; only 000/timeout = probe-failed. **NOTE: This probe MUST run on the Abiba host (CT 100) where 127.0.0.1:9200 is the extension. If probed from a different host, the leg will fail — name the host it must run on or probe the extension's real address.**
Check the JSON payload:
| Field | Healthy | Critical |
|-------|---------|----------|