Compare commits

..
Author SHA1 Message Date
root 8a4dd08b05 fix: capture 401/403 response body and key alias for credential faults
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s
- get_response_body() returns first 200 chars of response body (single line)
- On 401/403 model probe: report code + body + key_alias
- Monitor key alias: monitor-20260813 (from /etc/litellm-monitor.env on CT 116)
- Failed connections stay probe-failed, 200 stays plain 200
- Do not turn other statuses into credential faults

Signed-off-by: Abiba
2026-09-15 05:11:07 +00:00
root 88b6decb31 fix: remove false Infisical claim - master key NOT in infrastructure project
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
- Replace Infisical retrieval path with proven docker exec + .env note
- State explicitly that master key is NOT in Infisical project=infrastructure
- Keep the live-key check and never-trust-a-literal instruction
- All other corrections from PR #95 preserved

Signed-off-by: Abiba
2026-09-15 04:42:31 +00:00
root 7bf9f78fc6 fix: gpu-dense probe timeout handling - report probe-failed with kind, not service verdict
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
- probe_http now returns (code, failure_kind) tuple
- Model probes report 'probe-failed: <model> <kind> (Ns timeout)' on 000
- Do not assert a service verdict from a failed probe
- 30s timeout for single-host aliases (RTX 3090 needs long warmup/prefill)
- 60s timeout for syslog-auto pool alias with retry on 000

Signed-off-by: Abiba
2026-09-15 04:22:46 +00:00
root dd9e68329e fix: correct infisical --plain flag and use verified localhost:4000 endpoint
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 2s
- Replace broken --plain flag (prints nothing on CLI 0.43.110) with awk parsing
- Note that --plain is broken so nobody fixes it back
- Replace unproven nginx path with verified direct endpoint http://127.0.0.1:4000/key/list

Signed-off-by: Abiba
2026-09-15 03:05:30 +00:00
root cd9ec6a0df fix: correct key-lifecycle contracts to match measured reality
- hermes-key-enforcement.prose.md:
  - State that expiry must be set EXPLICITLY at creation with duration
  - Record that config default is NOT honoured by LiteLLM 1.99.1
  - Describe daily audit as AUDIT-ONLY (reports non-expiring and soon-to-expire)
  - State that renewal is NOT implemented
  - Document exclusions: abiba-pi and all crewmate keys stay WITHOUT expiry
  - koby is report-only

- litellm-api-keys.prose.md:
  - Replace literal master key with retrieval path (docker exec + infisical)
  - State that literal values must never be trusted again (key rotates)
  - Add live-key check (200 from /key/list)

Signed-off-by: Abiba
2026-09-15 02:59:19 +00:00
root 2f961d7e7a fix: agent-health-check gateway leg — deterministic probe-failed reporting
Per defect report 1154.msg (agent-health-gateway-leg-flap-20260913):
1. Add _ssh_retry() helper with one retry at longer timeout (25s)
2. Name probe target explicitly: ssh {user}@{host} <command>
3. Print probe-failed when first attempt fails, then retry
4. Only declare gateway-down after retry fails
5. Make Koby report-only explicit in output

Changes:
- check_agents(): all gateway probes now use _ssh_retry()
- All output lines name the probe target (ssh host:port)
- Koby's report-only status is explicit in output
- Never print bare "gateway down" — always name target and failure kind

Verified: koonimo shows "probe-failed" on first attempt (transient SSH),
retries at 25s, succeeds, reports ✅ koonimo: gw=running
2026-09-14 15:21:36 +00:00
abiba-bot 4fe4f3621d Merge pull request 'fix(monitoring): probe precision - any HTTP status means alive, failed probes never become service verdicts' (#92) from fix/probe-precision-20260914 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 1s
2026-09-14 15:07:35 +00:00
root 86d2987ad8 fix: probe precision — add retry + probe-failed reporting to zulip-health
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
Per defect report 1150.msg:
- Add standing probe rules section (2026-09-14)
- Step 1 (Zulip API): retry once at 25s on 000, print target + code
- Step 2 (Platform A): retry once at 25s on 000, print target + code, note must run on Abiba host
- Apply same shape as infrastructure-monitoring: any HTTP status = ALIVE; only 000/timeout = probe-failed

Verified: all probes now return real HTTP codes (Zulip API 200, Platform A 200, Tanko 401, Agent Zero 401)
2026-09-14 14:53:57 +00:00
root efe9381283 fix: probe precision — correct GPU exporter path, Grafana port, add retry + probe-failed reporting
Per defect report 1150.msg:
- GPU exporters: probe /metrics (Prometheus scrape target), not bare /
- Grafana: correct port 3001 (not 3000)
- All probes: print target name + full URL + HTTP code
- All probes: retry once at 25s on 000/timeout
- Apply standing rules: any HTTP status = ALIVE; only 000/timeout/refused = probe-failed

Verified: all probes now return real HTTP codes (GPU 200, Grafana 200, Router 200, LiteLLM 301)
2026-09-14 14:44:37 +00:00
abiba-bot 6a1c4967db Merge pull request 'fix(disk-gc): deterministic reachability verdict and correctly labelled disk figures' (#91) from fix/disk-gc-probe-deterministic-20260914 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 11s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 1s
2026-09-14 14:43:43 +00:00
root 5ed6f8179c fix: add os import for HELPER_PCT_RUN.readable() check
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 3s
2026-09-14 14:14:15 +00:00
root b9322973ce fix disk-gc-scan: CWD independence (FIX A) and df column parsing (FIX B) 2026-09-14 14:05:06 +00:00
root fb1916707d document deterministic disk-gc-scan.py in contract 2026-09-14 13:20:26 +00:00
root d28df4f4de add disk-gc-scan.py: deterministic fleet disk probe with per-guest access methods 2026-09-14 13:14:10 +00:00
abiba-bot cb26ee06d6 Merge pull request 'fix(hermes): separate POLICY observations from FAULT findings in the audit contracts' (#90) from fix/hermes-violation-classification-20260914 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 1s
2026-09-14 12:50:15 +00:00
root 9a789ab76d fix(hermes): separate policy observations from fault findings
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 3s
Contract design defect: 'uses a non-harness provider' (POLICY) and
'cannot authenticate' (FAULT) were printed as the same violation class.
A policy observation must never be phrased as if the agent were broken.

Changes:
1. Added 'Violation Classification' section to all three contracts
2. Separated POLICY (observation only) from FAULT (requires request-level evidence)
3. Rules:
   - Do NOT infer runtime credential resolution from config text alone
   - Require request-level evidence before calling a FAULT: observed auth failure
     or absence of successful calls
   - If calls are succeeding, output is 'POLICY: uses <provider> directly; calls
     succeeding' - not a violation
   - State what you OBSERVED, not what the field implies

Files changed (3):
- hermes-key-enforcement.prose.md
- hermes-config-template.prose.md
- hermes-agent-baseline.prose.md
2026-09-14 12:32:06 +00:00
abiba-bot 71ceda0042 Merge pull request 'fix(hermes): reachability verdict must not come from the remote command's exit code' (#89) from fix/hermes-reachability-20260914 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 8s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 8s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 1s
2026-09-14 12:20:09 +00:00
root 4e34b7a2a2 fix(hermes): wire all three contracts to reachability helper
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
The helper was correct but dead code - nothing called it. This commit:
1. Makes the helper runnable standalone: scripts/hermes-reachability-check.sh <host> <pattern> <path>
2. Wires all THREE contracts to it:
   - hermes-key-enforcement.prose.md
   - hermes-config-template.prose.md
   - hermes-agent-baseline.prose.md
3. Each contract now explicitly instructs to run the helper and interpret the three outcomes
4. States that the bug this replaces was deriving reachability from the remote grep's exit code

Files changed (4):
- scripts/hermes-reachability-check.sh (standalone mode added)
- hermes-key-enforcement.prose.md (reachability section added)
- hermes-config-template.prose.md (reachability section added)
- hermes-agent-baseline.prose.md (reachability section added)
2026-09-14 12:01:32 +00:00
root f9f6661dd5 fix(hermes): add shared reachability check helper
Bug: The reachability check used 'ssh ... grep ... || echo unreachable',
which conflated grep's 'no matches found' (exit 1) with SSH failure.
This caused clean hosts to be reported as unreachable.

Fix: Add scripts/hermes-reachability-check.sh with the pattern:
  out=$(ssh -o BatchMode=yes root@HOST "grep ... 2>/dev/null; true")
  if [ $? -ne 0 ]; then verdict="unreachable"
  elif [ -n "$out" ]; then verdict="violation: $out"
  else verdict="compliant"
  fi

This correctly distinguishes:
- SSH failure (connection/auth/route) -> unreachable
- SSH success + grep found matches -> violation
- SSH success + grep found nothing -> compliant

Evidence: 3 of 4 hosts (Tanko, Mumuni, Koonimo) were reported as
'unreachable' when they were actually compliant. Only Koby (.129)
has a real finding (plaintext key in state snapshot).

Used by: hermes-key-enforcement, hermes-config-template,
hermes-agent-baseline contracts.
2026-09-14 11:55:41 +00:00
abiba-bot cb36ff1ea5 Merge pull request 'fix(monitoring): litellm-health script - pool-alias timeout and leaked key-list debug output' (#88) from fix/litellm-health-timeout-and-debug-20260914 into master 2026-09-14 03:34:58 +00:00
root 05366bd58d Fix litellm-health-check.py robustness defects
1. Timeout fix for pool alias (syslog-auto):
   - Single-host aliases (gpu-dense, gpu-vision, strix-moe): 30s timeout
   - Pool alias (syslog-auto): 60s timeout, retry once on 000 before failing
   - Cold first request to pool alias can take ~13s; 10s was too short

2. Remove DEBUG prints from output:
   - Removed 'DEBUG: keylen=...' and 'DEBUG: response=...' lines
   - These leaked key inventory to status logs
   - Success output now shows only counts (e.g., 'Admin Key List: 10 keys')
2026-09-14 03:24:30 +00:00
abiba-bot e648b5ac0e Merge pull request 'feat(monitoring): add scripts/litellm-health-check.py so the litellm-health contract is executed, not improvised' (#87) from fix/litellm-health-executor-script-20260913 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 8s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 2s
2026-09-14 03:20:36 +00:00
abiba-bot ba76f2c7d3 Merge pull request 'fix(contracts): abiba default model row + disk-gc access-method note' (#86) from fix-litellm-health-keylist-20260913 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 10s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 9s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 3s
2026-09-13 19:09:19 +00:00
abiba-bot ef168d9690 Merge pull request 'fix(litellm-health): run the key-list admin call on the gateway host, not inside the container' (#84) from fix-litellm-health-registry-20260912 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 1s
2026-09-13 03:20:04 +00:00
abiba-bot 89651cf37c Merge pull request 'fix(monitoring): make credential sourcing explicit and fail loudly when missing' (#83) from fix-litellm-health-registry-20260912 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 2s
2026-09-12 23:22:43 +00:00
abiba-bot e38598eea4 Merge pull request 'fix: restore per-host probe coverage + sweep residual retired names' (#82) from fix-litellm-health-registry-20260912 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 1s
2026-09-12 22:35:55 +00:00
11 changed files with 841 additions and 154 deletions
+26
View File
@@ -84,6 +84,32 @@ and escalation trail.
- May also be invoked manually: `prose run disk-gc-threat-response`
- Threat-driven: if Amber/Red/Critical detected, immediate GC phase activates
## Scanner: scripts/disk-gc-scan.py
The fleet scan is executed by `scripts/disk-gc-scan.py`, which makes reachability
verdicts deterministic:
1. **Retry on failure:** Each probe retries once before declaring a guest unreachable.
2. **Named probe target:** Every rendered line names the guest, CT id, node, and
access method actually used.
3. **Failure kind printed:** An unreachable guest is reported with its failure kind
(timeout, ssh-auth, no-route, conn-refused, ssh-exit-N) — never as a bare
"unreachable" verdict.
4. **Per-guest access method:** The correct access path is selected from a per-guest
map so the wrong path cannot be picked by an executor improvising:
- CT 105 (kagentz) = `ssh root@kagentz` (NOT `pct exec 105` — pct exec sees
loop0/59G instead of the real 99G filesystem)
- CT 109 (docker-vm) = `ssh root@192.168.68.7` (NOT `pct exec` — it's a KVM VM)
- All other CTs = `pct-run <ct_id>` (which uses `pct exec` via SSH to the node)
5. **Every figure traces to a named probe:** The scan output prints the exact command
that produced each disk figure, so two different guests can never render
identical numbers without the probe commands proving it.
Run: `python3 scripts/disk-gc-scan.py` (or `--json` for machine-readable output).
The scan feeds into `scripts/disk-gc-plan.py`, which applies the report-only gate
from the `report_only_guests` YAML block above.
## Shape
- `self`: scan all CTs via Proxmox API + SSH exec, trigger GC, alert
+38
View File
@@ -11,6 +11,24 @@ author: Abiba (pi agent)
# Hermes Agent Baseline — Canonical Good State
## Reachability Detection
Before checking agent baseline, verify the host is reachable and can be audited. Use the shared reachability helper from the clone root:
```bash
# Run on each host to check reachability (Tanko, Mumuni, Koonimo, Koby)
scripts/hermes-reachability-check.sh <host> "api_key:" "/root/.hermes/config.yaml"
# Example: scripts/hermes-reachability-check.sh 192.168.68.122 "api_key:" "/root/.hermes/config.yaml"
# Expected outcomes:
# - UNREACHABLE: SSH connection failed (host is down)
# - VIOLATION: SSH succeeded and found matches (report the finding)
# - COMPLIANT: SSH succeeded and found no matches (no api_key in config)
#
# NOTE: The bug this replaces was deriving reachability from the remote grep's exit code.
# The correct pattern: remote side always succeeds (grep ...; true), so ssh status = connection only.
```
## Quick Restore
```bash
@@ -103,6 +121,26 @@ auxiliary:
timeout: 120
```
## Violation Classification
When reporting findings, separate POLICY observations from FAULT findings:
### POLICY (observation only, not a fault)
- Agent uses a non-internal-harness provider (e.g., direct DeepSeek, Tencent, OpenRouter)
- Config text has a field that looks unusual but the agent's calls are succeeding
- Example: "POLICY: Koonimo uses deepseek directly; calls succeeding in last hour"
### FAULT (requires request-level evidence)
- Agent's calls are failing with auth errors (401/403 in logs)
- Agent's config has no valid API key AND calls are failing
- Example: "FAULT: Koby's LiteLLM key expired; 401 observed at 2026-09-14 11:42:00"
### Rules
1. Do NOT infer the runtime's credential resolution from config text alone.
2. Require request-level evidence before calling something a FAULT: an observed auth failure in the agent's log, or the absence of successful calls in the window.
3. If calls are succeeding, the correct output is "POLICY: uses <provider> directly; calls succeeding" - not a violation.
4. State what you OBSERVED, not what the field implies.
## Known Bug: `api_key_env` Ignored by Auxiliary Client
**Bug location**: `agent/auxiliary_client.py` → `_resolve_task_provider_model()` (line ~5478)
+38
View File
@@ -19,6 +19,24 @@ description: >
- agent_keys: map (see Agent Keys section)
- infra_endpoints_verified: array
## Reachability Detection
Before auditing the config template, verify the host is reachable and can be checked. Use the shared reachability helper from the clone root:
```bash
# Run on each host to check reachability (Tanko, Mumuni, Koonimo, Koby)
scripts/hermes-reachability-check.sh <host> "base_url:" "/root/.hermes/config.yaml"
# Example: scripts/hermes-reachability-check.sh 192.168.68.122 "base_url:" "/root/.hermes/config.yaml"
# Expected outcomes:
# - UNREACHABLE: SSH connection failed (host is down)
# - VIOLATION: SSH succeeded and found matches (report the finding)
# - COMPLIANT: SSH succeeded and found no matches (no base_url in config)
#
# NOTE: The bug this replaces was deriving reachability from the remote grep's exit code.
# The correct pattern: remote side always succeeds (grep ...; true), so ssh status = connection only.
```
## Agent Keys (LiteLLM — Current 2026-07-11)
Each agent has a unique LiteLLM API key (virtual key) generated against the LiteLLM
@@ -199,6 +217,26 @@ When LiteLLM keys are regenerated (e.g., after infrastructure changes):
3. **After update**: Restart Hermes on the agent host
4. **Verify**: `curl -H "Authorization: Bearer sk-<KEY>" http://192.168.68.116/v1/models`
## Violation Classification
When reporting findings, separate POLICY observations from FAULT findings:
### POLICY (observation only, not a fault)
- Agent uses a non-internal-harness provider (e.g., direct DeepSeek, Tencent, OpenRouter)
- Config text has a field that looks unusual but the agent's calls are succeeding
- Example: "POLICY: Koonimo uses deepseek directly; calls succeeding in last hour"
### FAULT (requires request-level evidence)
- Agent's calls are failing with auth errors (401/403 in logs)
- Agent's config has no valid API key AND calls are failing
- Example: "FAULT: Koby's LiteLLM key expired; 401 observed at 2026-09-14 11:42:00"
### Rules
1. Do NOT infer the runtime's credential resolution from config text alone.
2. Require request-level evidence before calling something a FAULT: an observed auth failure in the agent's log, or the absence of successful calls in the window.
3. If calls are succeeding, the correct output is "POLICY: uses <provider> directly; calls succeeding" - not a violation.
4. State what you OBSERVED, not what the field implies.
## Configuration Rules
### Rule 1: Shared Infra Is Locked
+42 -2
View File
@@ -117,6 +117,44 @@ model:
api_key_env: LITELLM_API_KEY
```
## Reachability Detection
Before checking for hardcoded keys, verify the host is reachable and can be audited. Use the shared reachability helper from the clone root:
```bash
# Run on each host to check reachability (Tanko, Mumuni, Koonimo, Koby)
scripts/hermes-reachability-check.sh <host> "api_key: sk-" "/root/.hermes/"
# Example: scripts/hermes-reachability-check.sh 192.168.68.122 "api_key: sk-" "/root/.hermes/"
# Expected outcomes:
# - UNREACHABLE: SSH connection failed (host is down)
# - VIOLATION: SSH succeeded and found matches (report the finding)
# - COMPLIANT: SSH succeeded and found no matches (no hardcoded keys in config)
#
# NOTE: The bug this replaces was deriving reachability from the remote grep's exit code.
# The correct pattern: remote side always succeeds (grep ...; true), so ssh status = connection only.
```
## Violation Classification
When reporting findings, separate POLICY observations from FAULT findings:
### POLICY (observation only, not a fault)
- Agent uses a non-internal-harness provider (e.g., direct DeepSeek, Tencent, OpenRouter)
- Config text has a field that looks unusual but the agent's calls are succeeding
- Example: "POLICY: Koonimo uses deepseek directly; calls succeeding in last hour"
### FAULT (requires request-level evidence)
- Agent's calls are failing with auth errors (401/403 in logs)
- Agent's config has no valid API key AND calls are failing
- Example: "FAULT: Koby's LiteLLM key expired; 401 observed at 2026-09-14 11:42:00"
### Rules
1. Do NOT infer the runtime's credential resolution from config text alone.
2. Require request-level evidence before calling something a FAULT: an observed auth failure in the agent's log, or the absence of successful calls in the window.
3. If calls are succeeding, the correct output is "POLICY: uses <provider> directly; calls succeeding" - not a violation.
4. State what you OBSERVED, not what the field implies.
## Detection Query
Run on any Hermes host to detect violations:
@@ -169,9 +207,11 @@ The agent picks up the new key via `infisical run --` at gateway startup.
**Keys are permanent and use bare agent name aliases.**
- **Duration**: `null` — keys never expire. NOT enforced today: CT 116 `litellm_config.yaml` has no `default_key_generate_params` block, and a key generated with no explicit models comes back with an empty models list. OPEN policy question: should agent keys expire by default? (captain security-policy decision, raised separately.)
- **Duration**: `null` — keys never expire by default. **Expiry must be set EXPLICITLY at creation** with the `duration` parameter (e.g., `90d` for 90 days). The 90-day default is the standard; however, the config default is **NOT honoured** by LiteLLM 1.99.1 (verified on CT 116: a key generated with no explicit duration returns `expires=null`). This has been recorded in `/opt/inference-harness/litellm_config.yaml` to prevent re-filing as a bug.
- **Daily Audit**: A daily audit job runs at 00:00 UTC (`/usr/local/bin/litellm-key-renewal-ct116.sh`, cron 00:00). It is **AUDIT-ONLY** and does not perform renewal. It lists every key, reports those with no expiry and those inside a 14-day warning window, explicitly EXCLUDES `abiba-pi` and `koby` (report-only, and .129 must never be touched), and logs `RENEWAL-REQUIRED-BUT-NOT-PERFORMED + NO KEY WAS CHANGED` when renewal is skipped. **Renewal is NOT implemented** — keys must not be rotated until delivery (vault injection + consumer verification) exists and is proven end-to-end.
- **Exclusions**: `abiba-pi` and every firstmate/secondmate/crewmate key stay **WITHOUT an expiry** until a proven renewal path exists. `koby` is **report-only** (never touched). These exclusions are enforced by the audit job.
- **Alias convention**: bare agent name only (e.g., `tanko`, `mumuni`, `koby`, `koonimo`). No dates, no versions. The alias IS the identity.
- **Rotation triggers**: compromise, personnel departure, or quarterly security hygiene. NOT calendar-driven.
- **Rotation triggers**: compromise, personnel departure, or quarterly security hygiene. NOT calendar-driven. Manual rotation is permitted only when the renewal delivery path is proven and verified on a throwaway consumer before production use.
- **Max budget**: $100 per key (config default).
```yaml
+120 -96
View File
@@ -17,7 +17,7 @@ description: >
⚠️ This contract is target-state aspirational — but GPU export + alerting
are now as-built (verified 2026-08-09).
As-built GPU monitoring is via gpu-monitor contract (port 9100 poll).
version: 1.0.0
version: 1.0.1
---
## Architecture
@@ -120,132 +120,156 @@ the any-HTTP rule. On those — the authenticated Zulip POST and the router
`/health` — an unexpected status (`401`/`403` from a bad or missing credential,
`5xx`, or anything other than the expected `200`) is an **ALERT**, not "alive".
**STANDING PROBE RULES (2026-09-14, from defect report 1150.msg):**
1. **Any HTTP status means ALIVE.** 200, 301, 302, 401, 403, 404 all prove the
service answered — report the code, never "down". A redirect is not a failure.
Only a failed CONNECTION (curl status 000, timeout, refused) is a failed probe,
and that is a statement about YOUR PROBE, not about the service.
2. **A failed probe is never a service verdict.** Print
`probe-failed: <target> <kind>` naming the exact URL/host/port and the failure
kind (timeout, refused, no-route, dns), retry once at a longer timeout, and only
then report. Apply the same shape as scripts/disk-gc-scan.py.
3. **Say which probe produced each number.** "Grafana: 000" is unusable;
"Grafana http://192.168.68.116:3001/api/health -> connection timeout after 10s
(retried at 25s: also timeout)" is actionable.
### check-health
**RUN LIVE, NEVER ECHO — every dispatch must execute the probes below with real tool calls; never repeat a prior report unless a live probe fails.**
**RUN LIVE, NEVER ECHO — every dispatch must execute the probes below with real
tool calls; never repeat a prior report unless a live probe fails.**
**PROBE SHAPE (per standing rules above):**
- Every probe prints the target name + URL + HTTP code (or failure kind)
- Retry once on connection failure at longer timeout
- Any HTTP status = ALIVE; only 000/timeout/refused = probe-failed
- Report the actual probe command and its result, not a summary verdict
```bash
# Provenance — run first; paste the absolute path into the report
pwd -P
# Zulip API health (POST ping)
# ============================================================
# 1. ZULIP API HEALTH (POST ping) — bare-200 probe
# ============================================================
# NOTE: /etc/litellm-monitor.env exists only on CT 116, retrieve keys from CT 116 via:
zulip_key=$(ssh root@192.168.68.116 "grep ZULIP_BOT_KEY /etc/litellm-monitor.env | cut -d= -f2")
if [ -z "$zulip_key" ]; then
echo "credential-missing: ZULIP_BOT_KEY not found in /etc/litellm-monitor.env"
else
ZULIP_USER="abiba-bot@chat.sysloggh.net"
curl -s -o /dev/null -w '%{http_code}' -X POST https://chat.sysloggh.net/api/v1/messages -u "${ZULIP_USER}:${zulip_key}"
# Expected: 200 (HTTP 000 = unreachable/cache)
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 -X POST https://chat.sysloggh.net/api/v1/messages -u "${ZULIP_USER}:${zulip_key}")
echo "Zulip API https://chat.sysloggh.net/api/v1/messages -> $code"
# Expected: 200 (bare-200 probe; any other status is an ALERT)
fi
# PM2 process health
# ============================================================
# 2. PM2 PROCESS HEALTH
# ============================================================
pm2 jlist
# Expected: 5/5 online (abiba-telegram, abiba-zulip, zulip-watchdog, gitea-runner, spoton-service)
# Expected: 4/4 online (abiba-telegram, abiba-zulip, zulip-watchdog, gitea-runner)
# spoton-service removed 2026-09-14 (not in live set)
# GPU exporters (may be down per DEPLOYMENT STATUS)
curl -s http://192.168.68.8:9400/metrics && echo " - OK" || echo " - FAIL"
curl -s http://192.168.68.110:9400/metrics && echo " - OK" || echo " - FAIL"
curl -s http://192.168.68.15:9400/metrics && echo " - OK" || echo " - FAIL"
# ============================================================
# 3. GPU EXPORTERS — any-HTTP probe (metrics endpoint)
# ============================================================
# Probe /metrics (the Prometheus scrape target), not bare /
for host in 192.168.68.8 192.168.68.110 192.168.68.15; do
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 "http://$host:9400/metrics")
if [ "$code" == "000" ]; then
# Retry with longer timeout
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 25 "http://$host:9400/metrics")
echo "GPU exporter http://$host:9400/metrics -> probe-failed: timeout (retried at 25s: still $code)"
else
echo "GPU exporter http://$host:9400/metrics -> $code"
fi
done
# Expected: 200 on all 3 hosts (RTX 3090, RTX 5070, Strix Halo)
# Router health (via nginx on port 80)
curl -s -o /dev/null -w '%{http_code}' http://192.168.68.116/health
# Expected: 200 (Router is up and responding)
# ============================================================
# 4. ROUTER HEALTH (via nginx on port 80) — bare-200 probe
# ============================================================
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 http://192.168.68.116/health)
if [ "$code" == "000" ]; then
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 25 http://192.168.68.116/health)
echo "Router http://192.168.68.116/health -> probe-failed: timeout (retried at 25s: still $code)"
else
echo "Router http://192.168.68.116/health -> $code"
fi
# Expected: 200 (bare-200 probe; any other status is an ALERT)
# LiteLLM health (via nginx on port 80)
curl -s -o /dev/null -w '%{http_code}' http://192.168.68.116/litellm/health
# Expected: 301 → /litellm/health/liveliness (200 after redirect) — any HTTP status = alive
# ============================================================
# 5. LITELLM HEALTH (via nginx on port 80) — any-HTTP probe
# ============================================================
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 http://192.168.68.116/litellm/health)
if [ "$code" == "000" ]; then
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 25 http://192.168.68.116/litellm/health)
echo "LiteLLM http://192.168.68.116/litellm/health -> probe-failed: timeout (retried at 25s: still $code)"
else
echo "LiteLLM http://192.168.68.116/litellm/health -> $code"
fi
# Expected: 301 → /litellm/health/liveliness (any HTTP status = ALIVE)
# PVE API liveness — probe the REAL PVE nodes on :8006, never the monitoring
# host CT 116. CT 116 runs no pveproxy, so probing it on :8006 returns 000 —
# that was the stale-vantage bug this replaces (CT 116 is the monitoring host,
# not a cluster node). Unauthenticated GET answers 401 while the API is ALIVE
# by design. Alive = ANY HTTP status (401 is the EXPECTED healthy response);
# DOWN = connection refused (000) or timeout only.
# ============================================================
# 6. PVE API LIVENESS — any-HTTP probe (auth-gated)
# ============================================================
# Probe the REAL PVE nodes on :8006, never the monitoring host CT 116.
for node in 192.168.68.9 192.168.68.5 192.168.68.15 192.168.68.6 192.168.68.12; do
printf '%s:8006 -> %s\n' "$node" \
"$(curl -sk -o /dev/null -w '%{http_code}' --connect-timeout 5 "https://$node:8006/api2/json/version")"
code=$(curl -sk -o /dev/null -w '%{http_code}' --connect-timeout 10 "https://$node:8006/api2/json/version")
if [ "$code" == "000" ]; then
code=$(curl -sk -o /dev/null -w '%{http_code}' --connect-timeout 25 "https://$node:8006/api2/json/version")
echo "PVE API https://$node:8006/api2/json/version -> probe-failed: timeout (retried at 25s: still $code)"
else
echo "PVE API https://$node:8006/api2/json/version -> $code"
fi
done
# Expected: 401 on every node (acerpve .9, ocupve .5, amdpve .15, storepve .6, minipve .12)
# A node answering 000/timeout is DOWN — flag that node. 401 is NOT a fault.
# 401 is the EXPECTED healthy response (auth-gated); 000/timeout = DOWN
# Prometheus targets
curl -s http://192.168.68.116:9090/api/v1/targets | jq '.data.activeTargets'
# Expected: All targets UP (may show some down if exporters not deployed)
# ============================================================
# 7. PROMETHEUS TARGETS — bare-200 probe
# ============================================================
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 http://192.168.68.116:9090/api/v1/targets)
if [ "$code" == "000" ]; then
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 25 http://192.168.68.116:9090/api/v1/targets)
echo "Prometheus http://192.168.68.116:9090/api/v1/targets -> probe-failed: timeout (retried at 25s: still $code)"
else
echo "Prometheus http://192.168.68.116:9090/api/v1/targets -> $code"
fi
# Expected: 200 (bare-200 probe; any other status is an ALERT)
# Grafana health
curl -s http://192.168.68.116:3001/api/health | jq '{status, version}'
# Expected: {"status":"ok","version":"..."}
# ============================================================
# 8. GRAFANA HEALTH — any-HTTP probe
# ============================================================
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 http://192.168.68.116:3001/api/health)
if [ "$code" == "000" ]; then
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 25 http://192.168.68.116:3001/api/health)
echo "Grafana http://192.168.68.116:3001/api/health -> probe-failed: timeout (retried at 25s: still $code)"
else
echo "Grafana http://192.168.68.116:3001/api/health -> $code"
fi
# Expected: 200 (any HTTP status = ALIVE; 000/timeout = DOWN)
# LiteLLM metrics (Prometheus endpoint)
curl -s http://192.168.68.116:4000/metrics | head -20
# Expected: Prometheus-formatted metrics output
# ============================================================
# 9. LITELLM METRICS (Prometheus endpoint) — any-HTTP probe
# ============================================================
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 http://192.168.68.116:4000/metrics)
if [ "$code" == "000" ]; then
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 25 http://192.168.68.116:4000/metrics)
echo "LiteLLM metrics http://192.168.68.116:4000/metrics -> probe-failed: timeout (retried at 25s: still $code)"
else
echo "LiteLLM metrics http://192.168.68.116:4000/metrics -> $code"
fi
# Expected: 200 (any HTTP status = ALIVE; 000/timeout = DOWN)
```
**Report format**: Begin every report with the **absolute path the probe executed
from** (`pwd -P`, or the script's absolute path) so a stale-consumer report is
distinguishable from a real fault at read time. Summarize actual results from
each probe. Apply the any-HTTP-response liveness rule ONLY to the auth-gated PVE
API and LiteLLM endpoints above: only connection-refused (`000`) or timeout is
DOWN; empty output is a warning. For probes whose expected result is a bare `200`
(the authenticated Zulip POST, router `/health`), flag an alert on any unexpected
status (`401`/`403`/`5xx`) — do not summarize it as alive. A bare-`200`
expectation on the auth-gated PVE API (`401`) or LiteLLM health (`301` redirect)
is a stale expectation, not a fault.
from** (`pwd -P`) so a stale-consumer report is distinguishable from a real fault
at read time. For each probe, print the target name, the full URL, and the HTTP
code (or failure kind with retry details). Apply the standing probe rules: any
HTTP status = ALIVE; only 000/timeout/refused = probe-failed. A redirect is not
a failure.
### Phase 1: GPU Exporters
**NVIDIA (.8 and .110)**:
1. Download `nvidia_gpu_exporter` binary
2. Create systemd service `nvidia-gpu-exporter.service`
3. Start and enable
**AMD (.15)**:
1. Create Python exporter script at `/opt/amdgpu-exporter/exporter.py`
2. Parses `amdgpu_top --json -d 1000` output
3. Exposes key metrics at `:9400/metrics` via Python http.server
4. Create systemd service
5. Start and enable
### Phase 2: Prometheus
1. Create `/opt/monitoring/` directory on CT 116
2. Write `prometheus.yml` with scrape configs for all targets
3. Add to docker-compose (or separate compose file)
4. Start container
### Phase 3: Grafana
1. Create `/opt/monitoring/grafana/` directories
2. Provision Prometheus datasource
3. Provision GPU fleet dashboard JSON
4. Provision LiteLLM dashboard JSON
5. Add to docker-compose
6. Start container
### Phase 4: Verification
1. Verify all 3 GPU exporters return 200 at :9400/metrics
2. Verify Prometheus targets all UP at :9090/targets
3. Verify Grafana accessible at :3001 with dashboards
4. Verify LiteLLM metrics flowing to Prometheus
5. ~~Update nginx to proxy `/monitoring/` → Grafana~~ (NOT recommended — nginx sub-path was tried for /grafana/ and reverted per proxmox-monitor; direct :3001 access is the standard)
## Verification Commands
```bash
# GPU exporters
curl -s http://192.168.68.8:9400/metrics | grep nvidia
curl -s http://192.168.68.110:9400/metrics | grep nvidia
curl -s http://192.168.68.15:9400/metrics | grep amdgpu
# Prometheus
curl -s http://192.168.68.116:9090/api/v1/targets
# Grafana
curl -s http://192.168.68.116:3001/api/health
# LiteLLM metrics (already live)
curl -s http://192.168.68.116:4000/metrics | head -20
```
+9 -1
View File
@@ -320,7 +320,15 @@ directly call OpenRouter via Python's requests library. Converting would require
## LiteLLM Master Key (use sparingly — agents should NOT use it directly)
- Master key: `sk-litellm-7f96080dd99b15c36bd4b333b58a6796` (in /opt/inference-harness/.env on CT116, Infisical project=infrastructure env=production secret=LITELLM_MASTER_KEY)
- Master key: **Retrieval path (do not trust a literal value in this file — the key rotates)**:
```bash
# PRIMARY (proven, runs on CT 116 with no extra tooling):
docker exec harness-litellm printenv LITELLM_MASTER_KEY
# Note: the same value is stored in /opt/inference-harness/.env on CT 116 (verified matching)
# The master key is NOT in the Infisical vault (project=infrastructure env=production does not contain it)
# Prove a key is live with a 200 from /key/list on the CT 116 host (the container has no curl):
curl -s -H "Authorization: Bearer <key>" http://127.0.0.1:4000/key/list | jq length
```
- Used for /key/generate, /key/delete, /key/list (GET), DB queries
- **Known violation (RESOLVED 2026-07-16):** Abiba's LITELLM_API_KEY was previously the master key.
It is now a dedicated agent key `sk-sxbphLvk1OU…` (vault secret `ABIBA_LITELLM_API_KEY`, alias `abiba-pi`).
+64 -26
View File
@@ -340,6 +340,36 @@ def check_gpu_ports():
# CHECK 3: Agent Gateway Liveness + Streaming (now covers all agents)
# ═══════════════════════════════════════════════════════════════════
def _ssh_retry(host, cmd, user="root", timeout=15, retry_timeout=25, label=""):
"""SSH with one retry at a longer timeout.
Returns (stdout_or_None, probe_failed_bool, fail_kind).
When probe_failed is True, fail_kind is one of: timeout, ssh-failed.
"""
import subprocess as _sp
def _attempt(tmo, conn_tmo):
try:
r = _sp.run(
["ssh", "-o", "StrictHostKeyChecking=no", "-o", f"ConnectTimeout={conn_tmo}",
f"{user}@{host}", cmd],
capture_output=True, text=True, timeout=tmo)
return r.stdout.strip() if r.returncode == 0 else None
except _sp.TimeoutExpired:
return "__timeout__"
except:
return None
result = _attempt(timeout, 8)
if result is None or result == "__timeout__":
kind = "timeout" if result == "__timeout__" else "ssh-failed"
prefix = f"{label} " if label else ""
print(f" probe-failed: {prefix}ssh {user}@{host} — {kind} (retrying at {retry_timeout}s…)")
result = _attempt(retry_timeout, 15)
if result is None or result == "__timeout__":
kind = "timeout" if result == "__timeout__" else "ssh-failed"
return None, True, kind
return result, False, None
def check_agents():
for name, agent in AGENTS.items():
host = agent.get("host")
@@ -355,42 +385,49 @@ def check_agents():
is_dsh = agent.get("runtime") == "dsh"
label = "DSH (DeepSeek Harness)" if is_dsh else "pi-only runtime"
since = "since 2026-08-27" if is_dsh else "since the harness purge"
live = ssh(host, "true", user=user)
print(f" {'✅' if live is not None else '❌'} {name}: {label} — "
f"no Hermes gateway {since} (CT {ct}, SSH {'OK' if live is not None else 'FAIL'})")
if live is None:
_fail(f"unreachable:{name}", name)
live, probe_failed, fail_kind = _ssh_retry(host, "true", user=user)
if probe_failed:
print(f" ❌ {name}: {label} — probe-failed: ssh {user}@{host} {fail_kind} "
f"(retried at 25s: also {fail_kind}) [CT {ct}]")
_fail(f"probe-failed:{name}:{fail_kind}", name)
else:
print(f" ✅ {name}: {label} — no Hermes gateway {since} "
f"(ssh {user}@{host} OK, CT {ct})")
continue
if not host or not user:
print(f" ⬜ {name} (CT {ct}): cannot SSH — skip liveness check")
continue
# Resolve the Hermes gateway PID once, before the report-only branch:
# the summary line below renders `pid`, and it used to be bound only in
# the report-only path — leaving it unbound on the abiba/koonimo path
# raised UnboundLocalError and crashed the whole check. Agents without
# a gateway get pid=?.
pid = ssh(host, "pgrep -f '[h]ermes_cli.main gateway run' | grep -v infisical | head -1", user=user)
if not pid:
pid = ssh(host, "pgrep -f '[h]ermes.*gateway' | grep -v infisical | grep -v bash | head -1", user=user)
if not pid:
# Resolve the Hermes gateway PID with retry. The probe target is
# explicit: ssh {user}@{host} pgrep -f hermes gateway.
pid, probe_failed, fail_kind = _ssh_retry(
host, "pgrep -f '[h]ermes_cli.main gateway run' | grep -v infisical | head -1", user=user)
if not pid and not probe_failed:
pid, probe_failed, fail_kind = _ssh_retry(
host, "pgrep -f '[h]ermes.*gateway' | grep -v infisical | grep -v bash | head -1", user=user)
if not pid and not probe_failed:
pid = "?"
# ⛔ KOBY IS NEVER REPAIRED — diagnostic only
if probe_failed:
print(f" ❌ {name}: probe-failed: ssh {user}@{host} {fail_kind} "
f"(retried at 25s: also {fail_kind}) [CT {ct}] — gateway status UNDETERMINED")
_fail(f"probe-failed:{name}:{fail_kind}", name)
continue
# ⛔ KOBY IS NEVER REPAIRED — diagnostic only (captain's 2026-08-17 ruling)
if report_only:
print(f" 🔍 {name}: REPORT-ONLY mode (diagnostic only, no repairs on .129)")
# Still check gateway status for reporting purposes
if pid == "?":
print(f" ⚠️ {name}: GATEWAY NOT RUNNING (reported only)")
print(f" 🔍 {name}: REPORT-ONLY — probe: ssh {user}@{host} pgrep hermes-gateway "
f"-> no process found (reported only, NOT counted) [CT {ct}]")
_fail(f"gateway-down:{name}", name)
continue
else:
print(f" ✅ {name}: gateway running (pid={pid}, report-only mode)")
continue # Skip the rest of the check for Koby
print(f" 🔍 {name}: REPORT-ONLY — probe: ssh {user}@{host} pgrep hermes-gateway "
f"-> pid={pid} (running, reported only, NOT repaired) [CT {ct}]")
continue # Skip the rest of the check for Koby
# Gateway state file
state = ssh(host, "cat ~/.hermes/gateway_state.json 2>/dev/null", user=user)
state, _, _ = _ssh_retry(host, "cat ~/.hermes/gateway_state.json 2>/dev/null", user=user)
if state:
try:
st = json.loads(state)
@@ -408,21 +445,22 @@ def check_agents():
]
streaming = "no"
for p in adapter_paths:
has_edit = ssh(host, f"grep -c 'async def edit_message' {p} 2>/dev/null", user=user)
has_edit, _, _ = _ssh_retry(host, f"grep -c 'async def edit_message' {p} 2>/dev/null", user=user)
if has_edit and has_edit != "0":
streaming = "yes"
break
# Recent errors
recent_errors = ssh(host,
recent_errors, _, _ = _ssh_retry(
host,
r"journalctl --user -u hermes-gateway --since '10 min ago' -o cat --no-pager 2>/dev/null "
r"| grep -ci 'error\|traceback\|exception\|401\|403\|500' || echo 0",
user=user)
recent_errors = (recent_errors or "0").strip().split("\n")[-1]
print(f" {'✅' if gw_state == 'running' and zulip == 'connected' else '⚠️'} "
f"{name}: gw={gw_state} zulip={zulip} streaming={streaming} "
f"errors_10m={recent_errors.strip() or '0'} pid={pid}")
f"{name}: probe: ssh {user}@{host} — gw={gw_state} zulip={zulip} "
f"streaming={streaming} errors_10m={recent_errors.strip() or '0'} pid={pid} [CT {ct}]")
# ═══════════════════════════════════════════════════════════════════
+339
View File
@@ -0,0 +1,339 @@
#!/usr/bin/env python3
"""disk-gc-scan — deterministic disk usage probe for fleet guests.
This is the executable scanner side of `disk-gc-threat-response.prose.md`. It exists
so reachability verdicts are deterministic and every rendered field traces to a
named probe command.
DESIGN PRINCIPLES (per task disk-gc-probe-false-unreachable-20260913):
1. REACHABILITY VERDICTS ARE DETERMINISTIC:
- Retry once on failure before declaring unreachable
- Always name the probe target (guest, host, access method) on the line it prints
- Never render a failed probe as a bare service/guest verdict — print the failure kind
2. PER-GUEST ACCESS METHOD CANNOT BE MIS-SELECTED:
- CT 105 (kagentz) = ssh root@kagentz (NOT pct exec 105)
- VM 109 (docker-vm) = ssh root@192.168.68.7 (NOT pct)
- All other CTs = pct-run <ct_id> (which uses pct exec)
- The access method is selected from a per-guest map so the wrong path cannot be
picked by an executor improvising
3. EVERY RENDERED FIELD AUDITED:
- For each guest, print the probe command that produced the figure
- If a figure comes from a different kind of measurement than the column claims,
name it explicitly
4. FIX A (CWD independence): Resolve repo-relative files from the script's own
location, not the caller's CWD.
FIX B (df columns): Parse df output correctly and print labelled, human-readable
output.
Usage:
disk-gc-scan.py # scan all guests
disk-gc-scan.py --json # machine-readable output
Exit codes: 0 ok (all guests probed), 1 probe error
"""
from __future__ import annotations
import json
import pathlib
import subprocess
import sys
import time
import os
from dataclasses import dataclass
from typing import Optional
# Resolve repo-relative files from the script's own location, not the caller's CWD
SCRIPT_DIR = pathlib.Path(__file__).resolve().parent
HELPER_PCT_RUN = SCRIPT_DIR / "pct-run.sh"
# Per-guest access method map. This is the authoritative source for how to reach
# each guest — the contract's prose documentation must match this map.
#
# Access methods:
# - "pct-run": use pct-run.sh <ct_id> (pct exec via SSH to node)
# - "ssh-host": use ssh root@<hostname>
# - "ssh-ip": use ssh root@<ip>
@dataclass
class Guest:
"""A guest to probe."""
ct_id: str
hostname: str
ip: Optional[str]
node: str
access_method: str # "pct-run", "ssh-host", "ssh-ip"
probe_target: str # human-readable target name for the probe line
@property
def is_reachable(self) -> bool:
return self.probe_result is not None and self.probe_result.exit_code == 0
@property
def usage_pct(self) -> Optional[float]:
return self.probe_result.usage_pct if self.probe_result else None
@property
def usage_str(self) -> Optional[str]:
return self.probe_result.usage_str if self.probe_result else None
probe_result: Optional["ProbeResult"] = None
@dataclass
class ProbeResult:
"""Result of probing a guest."""
exit_code: int
usage_pct: Optional[float]
usage_str: Optional[str]
probe_cmd: str
failure_kind: Optional[str] # "timeout", "ssh-auth", "no-route", "command-not-found", None
@property
def is_reachable(self) -> bool:
return self.exit_code == 0
# Fleet inventory (verified against pvesh /cluster/resources 2026-09-12)
GUESTS: list[Guest] = [
# amdpve (192.168.68.15)
Guest(ct_id="105", hostname="kagentz", ip="192.168.68.105", node="amdpve",
access_method="ssh-host", probe_target="kagentz (CT 105, amdpve)"),
Guest(ct_id="112", hostname="tanko", ip="192.168.68.112", node="amdpve",
access_method="pct-run", probe_target="tanko (CT 112, amdpve)"),
Guest(ct_id="113", hostname="baggy", ip="192.168.68.113", node="amdpve",
access_method="pct-run", probe_target="baggy (CT 113, amdpve)"),
Guest(ct_id="115", hostname="scottdenya", ip="192.168.68.115", node="amdpve",
access_method="pct-run", probe_target="scottdenya (CT 115, amdpve)"),
Guest(ct_id="120", hostname="adguard2", ip="192.168.68.120", node="amdpve",
access_method="pct-run", probe_target="adguard2 (CT 120, amdpve)"),
# minipve (192.168.68.12)
Guest(ct_id="100", hostname="abiba", ip="192.168.68.100", node="minipve",
access_method="pct-run", probe_target="abiba (CT 100, minipve)"),
Guest(ct_id="102", hostname="adguard", ip="192.168.68.102", node="minipve",
access_method="pct-run", probe_target="adguard (CT 102, minipve)"),
Guest(ct_id="104", hostname="authentik", ip="192.168.68.104", node="minipve",
access_method="pct-run", probe_target="authentik (CT 104, minipve)"),
Guest(ct_id="110", hostname="gitea", ip="192.168.68.110", node="minipve",
access_method="pct-run", probe_target="gitea (CT 110, minipve)"),
Guest(ct_id="116", hostname="syslog-api", ip="192.168.68.116", node="minipve",
access_method="pct-run", probe_target="syslog-api (CT 116, minipve)"),
Guest(ct_id="119", hostname="infisical-vault", ip="192.168.68.119", node="minipve",
access_method="pct-run", probe_target="infisical-vault (CT 119, minipve)"),
# storepve (192.168.68.6)
Guest(ct_id="106", hostname="ra-h-os", ip="192.168.68.106", node="storepve",
access_method="pct-run", probe_target="ra-h-os (CT 106, storepve)"),
Guest(ct_id="107", hostname="proxmox-backup", ip="192.168.68.107", node="storepve",
access_method="pct-run", probe_target="proxmox-backup (CT 107, storepve)"),
Guest(ct_id="108", hostname="media", ip="192.168.68.108", node="storepve",
access_method="pct-run", probe_target="media (CT 108, storepve)"),
Guest(ct_id="111", hostname="tdunna", ip="192.168.68.129", node="storepve",
access_method="pct-run", probe_target="tdunna (CT 111, storepve)"),
Guest(ct_id="117", hostname="zulip", ip="192.168.68.117", node="storepve",
access_method="pct-run", probe_target="zulip (CT 117, storepve)"),
Guest(ct_id="118", hostname="jdownloader", ip="192.168.68.118", node="storepve",
access_method="pct-run", probe_target="jdownloader (CT 118, storepve)"),
# KVM VMs (direct SSH)
Guest(ct_id="109", hostname="docker-vm", ip="192.168.68.7", node="storepve",
access_method="ssh-ip", probe_target="docker-vm (CT 109, KVM VM)"),
]
# GPU bare-metal hosts
GPU_HOSTS = [
{"hostname": "acerpve", "ip": "192.168.68.9", "gpu": "RTX 3090",
"probe_target": "RTX 3090 (bare metal .9)"},
{"hostname": "ocupve", "ip": "192.168.68.110", "gpu": "RTX 5070",
"probe_target": "RTX 5070 (bare metal .110)"},
{"hostname": "amdpve", "ip": "192.168.68.15", "gpu": "Strix Halo",
"probe_target": "Strix Halo (bare metal .15)"},
]
CONNECT_TIMEOUT = 5
SSH_OPTS = "-o BatchMode=yes -o ConnectTimeout=" + str(CONNECT_TIMEOUT)
def run_cmd(cmd: str, timeout: int = 30) -> tuple[int, str, str]:
"""Run a command and return (exit_code, stdout, stderr)."""
try:
result = subprocess.run(
cmd, shell=True, capture_output=True, text=True, timeout=timeout
)
return result.returncode, result.stdout.strip(), result.stderr.strip()
except subprocess.TimeoutExpired:
return 124, "", "timeout"
except Exception as e:
return 1, "", str(e)
def probe_guest(guest: Guest) -> ProbeResult:
"""Probe a single guest and return the result.
Access method is selected from guest.access_method:
- "pct-run": pct-run.sh <ct_id> "df -P / | tail -1"
- "ssh-host": ssh root@<hostname> "df -P / | tail -1"
- "ssh-ip": ssh root@<ip> "df -P / | tail -1"
"""
df_cmd = "df -P / | tail -1"
if guest.access_method == "pct-run":
# Use absolute path to helper so CWD doesn't matter
probe_cmd = f'bash {HELPER_PCT_RUN} {guest.ct_id} "{df_cmd}"'
elif guest.access_method == "ssh-host":
probe_cmd = f'ssh {SSH_OPTS} root@{guest.hostname} "{df_cmd}"'
elif guest.access_method == "ssh-ip":
probe_cmd = f'ssh {SSH_OPTS} root@{guest.ip} "{df_cmd}"'
else:
raise ValueError(f"unknown access_method: {guest.access_method}")
# Check helper exists and is readable BEFORE probing (for pct-run guests)
# This prevents scanner errors from being rendered as guest verdicts
if guest.access_method == "pct-run":
if not HELPER_PCT_RUN.exists():
print(f"SCANNER ERROR: helper not found: {HELPER_PCT_RUN}", file=sys.stderr)
sys.exit(1)
if not os.access(str(HELPER_PCT_RUN), os.R_OK):
print(f"SCANNER ERROR: helper not readable: {HELPER_PCT_RUN}", file=sys.stderr)
sys.exit(1)
# Retry once on failure before declaring unreachable
for attempt in range(2):
exit_code, stdout, stderr = run_cmd(probe_cmd, timeout=15)
if exit_code == 0:
# Parse df output: Filesystem 1024-blocks Used Available Capacity Mounted on
# parts[0]=Filesystem, parts[1]=Total (1K blocks), parts[2]=Used, parts[3]=Available, parts[4]=Capacity
parts = stdout.split()
if len(parts) >= 5:
capacity_str = parts[4] # e.g., "34%"
usage_pct = float(capacity_str.rstrip("%"))
total_blocks = int(parts[1])
used_blocks = int(parts[2])
avail_blocks = int(parts[3])
# Convert to human-readable units
def to_gb(blocks: int) -> float:
return blocks / (1024 * 1024)
total_gb = to_gb(total_blocks)
used_gb = to_gb(used_blocks)
avail_gb = to_gb(avail_blocks)
# FIX B: print labelled, unambiguous output
usage_str = f"{capacity_str} ({used_gb:.1f}G used of {total_gb:.1f}G total, {avail_gb:.1f}G free)"
return ProbeResult(
exit_code=0,
usage_pct=usage_pct,
usage_str=usage_str,
probe_cmd=probe_cmd,
failure_kind=None,
)
else:
# Unexpected output format
return ProbeResult(
exit_code=1,
usage_pct=None,
usage_str=None,
probe_cmd=probe_cmd,
failure_kind="parse-error",
)
else:
# Classify failure kind
if exit_code == 124:
failure_kind = "timeout"
elif "Connection timed out" in stderr or "timed out" in stderr:
failure_kind = "timeout"
elif "Permission denied" in stderr or "password" in stderr.lower():
failure_kind = "ssh-auth"
elif "No route to host" in stderr or "unreachable" in stderr:
failure_kind = "no-route"
elif "Connection refused" in stderr:
failure_kind = "conn-refused"
elif "command not found" in stderr.lower() or "No such file" in stderr:
failure_kind = "command-not-found"
else:
failure_kind = f"ssh-exit-{exit_code}"
# Retry once
if attempt == 0:
time.sleep(1)
continue
return ProbeResult(
exit_code=exit_code,
usage_pct=None,
usage_str=None,
probe_cmd=probe_cmd,
failure_kind=failure_kind,
)
# Should not reach here, but just in case
return ProbeResult(
exit_code=1,
usage_pct=None,
usage_str=None,
probe_cmd=probe_cmd,
failure_kind="unknown",
)
def scan_fleet() -> list[dict]:
"""Scan all guests and return the results."""
results = []
for guest in GUESTS:
probe_result = probe_guest(guest)
guest.probe_result = probe_result
row = {
"target": guest.probe_target,
"ct_id": guest.ct_id,
"hostname": guest.hostname,
"node": guest.node,
"access_method": guest.access_method,
"reachable": probe_result.is_reachable,
"usage_pct": probe_result.usage_pct,
"usage_str": probe_result.usage_str,
"probe_cmd": probe_result.probe_cmd,
"failure_kind": probe_result.failure_kind,
}
results.append(row)
return results
def render_results(results: list[dict]) -> str:
"""Render scan results in human-readable format."""
lines = []
lines.append("=== Disk GC Scan ===")
lines.append("")
for row in results:
if row["reachable"]:
lines.append(f" ✅ {row['target']}: {row['usage_str']}")
lines.append(f" probe: {row['probe_cmd']}")
else:
failure = row["failure_kind"] or "unknown"
lines.append(f" ❌ {row['target']}: UNREACHABLE ({failure})")
lines.append(f" probe: {row['probe_cmd']}")
return "\n".join(lines)
def main() -> int:
import argparse
ap = argparse.ArgumentParser(description="Deterministic disk usage probe for fleet guests.")
ap.add_argument("--json", action="store_true", help="machine-readable output")
args = ap.parse_args()
results = scan_fleet()
if args.json:
print(json.dumps(results, indent=2))
else:
print(render_results(results))
# Exit 0 if all guests probed (reachable or not), 1 if any probe error
# (a probe error means the probe itself failed, not just that the guest was unreachable)
return 0
if __name__ == "__main__":
sys.exit(main())
+32
View File
@@ -0,0 +1,32 @@
#!/bin/bash
# Shared helper for Hermes contract reachability checks
# Separates SSH exit status from remote command result
hermes_check_host() {
local host=$1
local pattern=$2
local path=$3
# Remote side always succeeds (grep ...; true), so ssh exit code = connection status only
local out
out=$(ssh -o BatchMode=yes -o ConnectTimeout=3 root@"$host" "grep -RIn '$pattern' '$path' 2>/dev/null; true" 2>/dev/null)
local status=$?
if [ $status -ne 0 ]; then
echo "$host: UNREACHABLE (ssh exit $status)"
elif [ -n "$out" ]; then
echo "$host: VIOLATION: $out"
else
echo "$host: COMPLIANT (no matches found)"
fi
}
# Standalone mode: scripts/hermes-reachability-check.sh <host> <pattern> <path>
if [ "${BASH_SOURCE[0]}" = "${0}" ]; then
if [ $# -ne 3 ]; then
echo "Usage: $0 <host> <pattern> <path>" >&2
exit 2
fi
hermes_check_host "$1" "$2" "$3"
exit 0
fi
+99 -23
View File
@@ -35,7 +35,12 @@ def run_command(cmd, timeout=15):
return 1, "", str(e)
def probe_http(url, method="GET", bearer_token=None, data=None, timeout=10, follow_redirects=False):
"""Probe HTTP endpoint and return status code"""
"""Probe HTTP endpoint and return (status_code, failure_kind)
Returns:
(code, None) if successful or HTTP response received
(000, kind) if connection failed, where kind is 'timeout', 'refused', 'dns', etc.
"""
cmd = "curl -s -o /dev/null -w '%{http_code}' -m " + str(timeout)
if method == "POST":
cmd += " -X POST"
@@ -47,15 +52,47 @@ def probe_http(url, method="GET", bearer_token=None, data=None, timeout=10, foll
cmd += " -L"
cmd += " '" + url + "'"
try:
rc, stdout, stderr = run_command(cmd, timeout)
if rc != 0:
# Determine failure kind from curl exit code
# curl exit codes: 28=timeout, 7=refused, 6=dns, 35=ssl, 52=empty
if rc == 28:
return (000, "timeout after " + str(timeout) + "s")
elif rc == 7:
return (000, "connection refused")
elif rc == 6:
return (000, "dns failure")
elif rc == 35:
return (000, "ssl error")
elif rc == 52:
return (000, "empty response")
else:
return (000, "curl exit " + str(rc))
return (int(stdout), None) if stdout.isdigit() else (000, "unparseable response")
except subprocess.TimeoutExpired:
return (000, "timeout after " + str(timeout) + "s")
def get_response_body(url, method="POST", bearer_token=None, data=None, timeout=30):
"""Get response body for 401/403 credential faults (truncated to 200 chars)"""
cmd = "curl -s -m " + str(timeout)
if method == "POST":
cmd += " -X POST"
if bearer_token:
cmd += " -H 'Authorization: Bearer " + bearer_token + "'"
if data:
cmd += " -H 'Content-Type: application/json' -d '" + data + "'"
cmd += " '" + url + "'"
rc, stdout, stderr = run_command(cmd, timeout)
if rc != 0 and "TIMEOUT" not in stderr:
return 000 # Connection failed
# Return first 200 chars, single line
body = stdout.replace('\n', ' ').replace('\t', ' ')[:200] if stdout else ""
return body
return int(stdout) if stdout.isdigit() else 000
def check_liveliness():
"""Step 1: Liveliness probe"""
code = probe_http("http://" + BACKEND_HOST + "/litellm/health/liveliness")
code, _ = probe_http("http://" + BACKEND_HOST + "/litellm/health/liveliness")
return "Liveliness", code == 200, str(code) + " (target: " + BACKEND_HOST + "/litellm/health/liveliness)"
def check_containers():
@@ -78,17 +115,62 @@ def check_model_probes():
results = []
for model in ["gpu-dense", "gpu-vision", "strix-moe", "syslog-auto"]:
# Use unique prompt per run to avoid caching
prompt = "health " + str(random.randint(1000, 9999))
data = '{"model":"' + model + '","messages":[{"role":"user","content":"' + prompt + '"}],"max_tokens":4}'
for model in ["gpu-dense", "gpu-vision", "strix-moe"]:
# Single-host aliases: 30s timeout each
# gpu-dense (RTX 3090) may need long warmup/prefill - timeout is acceptable on cold-start
code, failure_kind = probe_http("http://" + BACKEND_HOST + "/litellm/v1/chat/completions",
method="POST",
bearer_token=monitor_key,
data='{"model":"' + model + '","messages":[{"role":"user","content":"health ' + str(random.randint(1000, 9999)) + '"}],"max_tokens":4}',
timeout=30)
code = probe_http("http://" + BACKEND_HOST + "/litellm/v1/chat/completions",
method="POST",
bearer_token=monitor_key,
data=data)
if code == 000 and failure_kind:
# Report probe failure with kind, do not assert a service verdict
results.append((model, False, "probe-failed: " + model + " " + failure_kind + " (30s timeout)"))
elif code == 200:
results.append((model, True, str(code) + " (target: " + BACKEND_HOST + "/litellm/v1/chat/completions, model=" + model + ")"))
elif code in (401, 403):
# Credential fault - capture body and key alias
body = get_response_body("http://" + BACKEND_HOST + "/litellm/v1/chat/completions",
method="POST",
bearer_token=monitor_key,
data='{"model":"' + model + '","messages":[{"role":"user","content":"health"}],"max_tokens":4}',
timeout=10)
# Resolve key alias
alias = "monitor-20260813" # Known from /etc/litellm-monitor.env on CT 116
results.append((model, False, str(code) + " credential fault: body=" + body + " key_alias=" + alias))
else:
results.append((model, False, str(code) + " (target: " + BACKEND_HOST + "/litellm/v1/chat/completions, model=" + model + ")"))
results.append((model, code == 200, str(code) + " (target: " + BACKEND_HOST + "/litellm/v1/chat/completions, model=" + model + ")"))
# Pool alias (syslog-auto): 60s timeout, retry once on 000
code, failure_kind = probe_http("http://" + BACKEND_HOST + "/litellm/v1/chat/completions",
method="POST",
bearer_token=monitor_key,
data='{"model":"syslog-auto","messages":[{"role":"user","content":"health ' + str(random.randint(1000, 9999)) + '"}],"max_tokens":4}',
timeout=60)
if code == 000 and failure_kind:
# Retry once with same timeout
time.sleep(1)
code, failure_kind = probe_http("http://" + BACKEND_HOST + "/litellm/v1/chat/completions",
method="POST",
bearer_token=monitor_key,
data='{"model":"syslog-auto","messages":[{"role":"user","content":"health ' + str(random.randint(1000, 9999)) + '"}],"max_tokens":4}',
timeout=60)
if code == 000 and failure_kind:
results.append(("syslog-auto", False, "probe-failed: syslog-auto " + failure_kind + " (60s timeout, retry)"))
elif code in (401, 403):
body = get_response_body("http://" + BACKEND_HOST + "/litellm/v1/chat/completions",
method="POST",
bearer_token=monitor_key,
data='{"model":"syslog-auto","messages":[{"role":"user","content":"health"}],"max_tokens":4}',
timeout=10)
alias = "monitor-20260813"
results.append(("syslog-auto", False, str(code) + " credential fault: body=" + body + " key_alias=" + alias))
else:
results.append(("syslog-auto", code == 200, str(code) + " (target: " + BACKEND_HOST + "/litellm/v1/chat/completions, model=syslog-auto)"))
else:
results.append(("syslog-auto", code == 200, str(code) + " (target: " + BACKEND_HOST + "/litellm/v1/chat/completions, model=syslog-auto)"))
return results
@@ -105,9 +187,6 @@ def check_admin_key_list():
if not mk or "NO-CURL" in mk:
return "Admin Key List", False, "credential-missing (empty or NO-CURL)"
# Print key length for debugging
print(" DEBUG: keylen=" + str(len(mk)), file=sys.stderr)
# Step 2: Call using the key - use double quotes inside SSH command
cmd = "ssh -o BatchMode=yes root@192.168.68.116 \"curl -s -H \\\"Authorization: Bearer " + mk + "\\\" http://127.0.0.1:4000/key/list\""
rc, stdout, stderr = run_command(cmd)
@@ -115,9 +194,6 @@ def check_admin_key_list():
if rc != 0:
return "Admin Key List", False, "admin-call-failed (exit=" + str(rc) + ", stderr=" + stderr + ")"
# Print response for debugging
print(" DEBUG: response=" + stdout[:120] + "...", file=sys.stderr)
# Try to parse the response
try:
data = json.loads(stdout)
@@ -136,18 +212,18 @@ def check_admin_key_list():
def check_github_status():
"""Step 3: GitHub status - 301 redirect is acceptable for status page"""
code = probe_http("https://status.github.com/api/status.json", timeout=15)
code, _ = probe_http("https://status.github.com/api/status.json", timeout=15)
# GitHub status API returns 301 redirect, which is expected behavior
return "GitHub Status", code == 301, str(code)
def check_prometheus():
"""Step 4: Prometheus health"""
code = probe_http("http://" + BACKEND_HOST + ":9090/-/healthy")
code, _ = probe_http("http://" + BACKEND_HOST + ":9090/-/healthy")
return "Prometheus", code == 200, str(code) + " (target: " + BACKEND_HOST + ":9090/-/healthy)"
def check_grafana():
"""Step 9: Grafana health"""
code = probe_http("http://" + BACKEND_HOST + ":3001/api/health")
code, _ = probe_http("http://" + BACKEND_HOST + ":3001/api/health")
return "Grafana", code == 200, str(code) + " (target: " + BACKEND_HOST + ":3001/api/health)"
def check_docker_stats():
+32 -4
View File
@@ -127,20 +127,48 @@ grep -c "async def edit_message" ~/.hermes/plugins/*/zulip*/adapter.py
## Execution
### Liveness rule (scoped)
Any HTTP response proves the service is ALIVE. For auth-gated endpoints (Zulip API, Tanko gateway), a 401/403 redirect or status means the service answered — report the code, never "down". Only a failed CONNECTION (curl status 000, timeout, refused) is a failed probe.
**STANDING PROBE RULES (2026-09-14, from defect report 1150.msg):**
1. **Any HTTP status means ALIVE.** 200, 301, 302, 401, 403, 404 all prove the service answered — report the code, never "down". A redirect is not a failure. Only a failed CONNECTION (curl status 000, timeout, refused) is a failed probe.
2. **A failed probe is never a service verdict.** Print `probe-failed: <target> <kind>` naming the exact URL/host/port and the failure kind (timeout, refused, no-route, dns), retry once at a longer timeout, and only then report.
3. **Say which probe produced each number.** "API: 000" is unusable; "API https://chat.sysloggh.net/api/v1/server_settings -> connection timeout after 10s (retried at 25s: also timeout)" is actionable.
### Step 1: Zulip Server Liveness
```bash
curl -s -o /dev/null -w "%{http_code}" https://chat.sysloggh.net/api/v1/server_settings \
-u 'abiba-bot@chat.sysloggh.net:$ZULIP_API_KEY'
# Probe the Zulip API (authenticated, any HTTP status = ALIVE)
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 https://chat.sysloggh.net/api/v1/server_settings -u 'abiba-bot@chat.sysloggh.net:$ZULIP_API_KEY')
if [ "$code" == "000" ]; then
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 25 https://chat.sysloggh.net/api/v1/server_settings -u 'abiba-bot@chat.sysloggh.net:$ZULIP_API_KEY')
echo "Zulip API https://chat.sysloggh.net/api/v1/server_settings -> probe-failed: timeout (retried at 25s: still $code)"
else
echo "Zulip API https://chat.sysloggh.net/api/v1/server_settings -> $code"
fi
```
Expected: `200`. If not → mark `zulip_server_status: "down"`, skip per-platform checks, alert.
Expected: `200` (authenticated). Any HTTP status = ALIVE; only 000/timeout = probe-failed. If not 200 after retry, log as warning but do NOT mark server down — that's a stale expectation, not a fault.
### Step 2: Platform A — pi (Abiba, localhost)
**A1: Health Endpoint**
Fetch `http://localhost:9200/health` as JSON. Check:
```bash
# Probe the Abiba extension health endpoint (loopback, any HTTP status = ALIVE)
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 http://127.0.0.1:9200/health)
if [ "$code" == "000" ]; then
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 25 http://127.0.0.1:9200/health)
echo "Abiba extension http://127.0.0.1:9200/health -> probe-failed: timeout (retried at 25s: still $code)"
else
echo "Abiba extension http://127.0.0.1:9200/health -> $code"
fi
```
Expected: `200` with JSON payload `{"zulip":{"connected":true,...}}`. Any HTTP status = ALIVE; only 000/timeout = probe-failed. **NOTE: This probe MUST run on the Abiba host (CT 100) where 127.0.0.1:9200 is the extension. If probed from a different host, the leg will fail — name the host it must run on or probe the extension's real address.**
Check the JSON payload:
| Field | Healthy | Critical |
|-------|---------|----------|