Compare commits

..
Author SHA1 Message Date
root 5ed6f8179c fix: add os import for HELPER_PCT_RUN.readable() check
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 3s
2026-09-14 14:14:15 +00:00
root b9322973ce fix disk-gc-scan: CWD independence (FIX A) and df column parsing (FIX B) 2026-09-14 14:05:06 +00:00
root fb1916707d document deterministic disk-gc-scan.py in contract 2026-09-14 13:20:26 +00:00
root d28df4f4de add disk-gc-scan.py: deterministic fleet disk probe with per-guest access methods 2026-09-14 13:14:10 +00:00
abiba-bot cb26ee06d6 Merge pull request 'fix(hermes): separate POLICY observations from FAULT findings in the audit contracts' (#90) from fix/hermes-violation-classification-20260914 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 1s
2026-09-14 12:50:15 +00:00
root 9a789ab76d fix(hermes): separate policy observations from fault findings
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 3s
Contract design defect: 'uses a non-harness provider' (POLICY) and
'cannot authenticate' (FAULT) were printed as the same violation class.
A policy observation must never be phrased as if the agent were broken.

Changes:
1. Added 'Violation Classification' section to all three contracts
2. Separated POLICY (observation only) from FAULT (requires request-level evidence)
3. Rules:
   - Do NOT infer runtime credential resolution from config text alone
   - Require request-level evidence before calling a FAULT: observed auth failure
     or absence of successful calls
   - If calls are succeeding, output is 'POLICY: uses <provider> directly; calls
     succeeding' - not a violation
   - State what you OBSERVED, not what the field implies

Files changed (3):
- hermes-key-enforcement.prose.md
- hermes-config-template.prose.md
- hermes-agent-baseline.prose.md
2026-09-14 12:32:06 +00:00
abiba-bot 71ceda0042 Merge pull request 'fix(hermes): reachability verdict must not come from the remote command's exit code' (#89) from fix/hermes-reachability-20260914 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 8s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 8s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 1s
2026-09-14 12:20:09 +00:00
root 4e34b7a2a2 fix(hermes): wire all three contracts to reachability helper
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
The helper was correct but dead code - nothing called it. This commit:
1. Makes the helper runnable standalone: scripts/hermes-reachability-check.sh <host> <pattern> <path>
2. Wires all THREE contracts to it:
   - hermes-key-enforcement.prose.md
   - hermes-config-template.prose.md
   - hermes-agent-baseline.prose.md
3. Each contract now explicitly instructs to run the helper and interpret the three outcomes
4. States that the bug this replaces was deriving reachability from the remote grep's exit code

Files changed (4):
- scripts/hermes-reachability-check.sh (standalone mode added)
- hermes-key-enforcement.prose.md (reachability section added)
- hermes-config-template.prose.md (reachability section added)
- hermes-agent-baseline.prose.md (reachability section added)
2026-09-14 12:01:32 +00:00
root f9f6661dd5 fix(hermes): add shared reachability check helper
Bug: The reachability check used 'ssh ... grep ... || echo unreachable',
which conflated grep's 'no matches found' (exit 1) with SSH failure.
This caused clean hosts to be reported as unreachable.

Fix: Add scripts/hermes-reachability-check.sh with the pattern:
  out=$(ssh -o BatchMode=yes root@HOST "grep ... 2>/dev/null; true")
  if [ $? -ne 0 ]; then verdict="unreachable"
  elif [ -n "$out" ]; then verdict="violation: $out"
  else verdict="compliant"
  fi

This correctly distinguishes:
- SSH failure (connection/auth/route) -> unreachable
- SSH success + grep found matches -> violation
- SSH success + grep found nothing -> compliant

Evidence: 3 of 4 hosts (Tanko, Mumuni, Koonimo) were reported as
'unreachable' when they were actually compliant. Only Koby (.129)
has a real finding (plaintext key in state snapshot).

Used by: hermes-key-enforcement, hermes-config-template,
hermes-agent-baseline contracts.
2026-09-14 11:55:41 +00:00
abiba-bot cb36ff1ea5 Merge pull request 'fix(monitoring): litellm-health script - pool-alias timeout and leaked key-list debug output' (#88) from fix/litellm-health-timeout-and-debug-20260914 into master 2026-09-14 03:34:58 +00:00
root 05366bd58d Fix litellm-health-check.py robustness defects
1. Timeout fix for pool alias (syslog-auto):
   - Single-host aliases (gpu-dense, gpu-vision, strix-moe): 30s timeout
   - Pool alias (syslog-auto): 60s timeout, retry once on 000 before failing
   - Cold first request to pool alias can take ~13s; 10s was too short

2. Remove DEBUG prints from output:
   - Removed 'DEBUG: keylen=...' and 'DEBUG: response=...' lines
   - These leaked key inventory to status logs
   - Success output now shows only counts (e.g., 'Admin Key List: 10 keys')
2026-09-14 03:24:30 +00:00
abiba-bot e648b5ac0e Merge pull request 'feat(monitoring): add scripts/litellm-health-check.py so the litellm-health contract is executed, not improvised' (#87) from fix/litellm-health-executor-script-20260913 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 8s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 2s
2026-09-14 03:20:36 +00:00
root 38d7e8b064 Update litellm-health contract to mandate executor script
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
Add 'Executor Script' section:
- Run scripts/litellm-health-check.py from the clone
- Hand-rolled probes not acceptable substitute
- Backend-edge checks use internal IP 192.168.68.116, not public URL
- Docker Stats fetched from CT 116 host (127.0.0.1:9324/metrics)
- Admin Key List requires proper quoting for SSH commands
2026-09-14 03:15:42 +00:00
root e94fadfa60 Add litellm-health-check.py: standardized health check script for LiteLLM fleet monitoring
Implements all 11 checks from litellm-health contract:
- Liveliness, Containers, Prometheus, Grafana health probes
- Model probes for gpu-dense, gpu-vision, strix-moe, syslog-auto
- Admin API key list (10 keys), GitHub status, Docker Stats metrics

Fixed quoting for SSH commands and response parsing (dict with 'keys' field).
Backend edge uses internal IP 192.168.68.116, not public URL.
Docker Stats fetched from CT 116 host itself (127.0.0.1:9324/metrics).

All 11 checks passing consistently.
2026-09-14 03:15:27 +00:00
abiba-bot ba76f2c7d3 Merge pull request 'fix(contracts): abiba default model row + disk-gc access-method note' (#86) from fix-litellm-health-keylist-20260913 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 10s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 9s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 3s
2026-09-13 19:09:19 +00:00
root 7906b2d52d fix: change abiba default model from deepseek-v4-pro to syslog-auto
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Failing after 13m52s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Skipped
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Skipped
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Skipped
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Skipped
The abiba-zulip-restore.prose.md file incorrectly claimed the default model
was 'deepseek-v4-pro', but /root/.pi/agent/settings.json declares
'defaultModel: syslog-auto' and 'defaultProvider: syslog-harness'.

deepseek-v4-pro is not a servable LiteLLM model name for this fleet.
The only live models are: syslog-auto, gpu-dense, gpu-vision, strix-moe.

Note: The firstmate/pi session talks to DeepSeek through its own
provider (auth.json), NOT through LiteLLM, so no LiteLLM change is
needed to keep firstmate's deepseek usage working.
2026-09-13 18:55:24 +00:00
root c81cf5b6f0 fix: disk-gc-threat-response - correct access methods for kagentz and docker-vm
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 8s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 12s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s
- kagentz (105): use ssh root@kagentz (hostname), NOT pct exec 105 (shows loop0 59G, not real 99G)
- docker-vm (109): use ssh root@192.168.68.7 (correct QEMU VM access)
- abiba (100): use ssh root@abiba (hostname)
- syslog-api (116): use pct-run 116 (correct)

Verified:
  kagentz: 99G 8.9G 86G 10% /
  docker-vm: 158G 17G 135G 11% /
  abiba: 59G 13G 44G 23% /
  syslog-api: 40G 13G 25G 34% /
2026-09-13 04:30:28 +00:00
abiba-bot ef168d9690 Merge pull request 'fix(litellm-health): run the key-list admin call on the gateway host, not inside the container' (#84) from fix-litellm-health-registry-20260912 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 1s
2026-09-13 03:20:04 +00:00
abiba-bot 89651cf37c Merge pull request 'fix(monitoring): make credential sourcing explicit and fail loudly when missing' (#83) from fix-litellm-health-registry-20260912 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 2s
2026-09-12 23:22:43 +00:00
abiba-bot e38598eea4 Merge pull request 'fix: restore per-host probe coverage + sweep residual retired names' (#82) from fix-litellm-health-registry-20260912 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 1s
2026-09-12 22:35:55 +00:00
9 changed files with 767 additions and 1 deletions
+1 -1
View File
@@ -34,7 +34,7 @@ verification, and DM loopback testing.
| @all-bots user ID | 20 | ✅ (config, verified by API at runtime) |
| PM2 process name | abiba-zulip | ✅ |
| Provider | syslog-harness (http://192.168.68.116/v1) | ✅ |
| Default model | deepseek-v4-pro | ✅ (settings.json) |
| Default model | syslog-auto | ✅ (settings.json) |
## Architecture
+27
View File
@@ -84,6 +84,32 @@ and escalation trail.
- May also be invoked manually: `prose run disk-gc-threat-response`
- Threat-driven: if Amber/Red/Critical detected, immediate GC phase activates
## Scanner: scripts/disk-gc-scan.py
The fleet scan is executed by `scripts/disk-gc-scan.py`, which makes reachability
verdicts deterministic:
1. **Retry on failure:** Each probe retries once before declaring a guest unreachable.
2. **Named probe target:** Every rendered line names the guest, CT id, node, and
access method actually used.
3. **Failure kind printed:** An unreachable guest is reported with its failure kind
(timeout, ssh-auth, no-route, conn-refused, ssh-exit-N) — never as a bare
"unreachable" verdict.
4. **Per-guest access method:** The correct access path is selected from a per-guest
map so the wrong path cannot be picked by an executor improvising:
- CT 105 (kagentz) = `ssh root@kagentz` (NOT `pct exec 105` — pct exec sees
loop0/59G instead of the real 99G filesystem)
- CT 109 (docker-vm) = `ssh root@192.168.68.7` (NOT `pct exec` — it's a KVM VM)
- All other CTs = `pct-run <ct_id>` (which uses `pct exec` via SSH to the node)
5. **Every figure traces to a named probe:** The scan output prints the exact command
that produced each disk figure, so two different guests can never render
identical numbers without the probe commands proving it.
Run: `python3 scripts/disk-gc-scan.py` (or `--json` for machine-readable output).
The scan feeds into `scripts/disk-gc-plan.py`, which applies the report-only gate
from the `report_only_guests` YAML block above.
## Shape
- `self`: scan all CTs via Proxmox API + SSH exec, trigger GC, alert
@@ -357,3 +383,4 @@ one-off GPU builds. No automated post-migration cleanup was in place.
> container has no `pct` binary.
>
> **KVM VM:** CT 109 (docker-vm) is a QEMU VM, not LXC — access via SSH .7.
> **NOTE:** For kagentz (CT 105), use `ssh root@kagentz` (hostname), NOT `pct exec 105` — `pct exec 105` shows loop0 (59G) while `ssh root@kagentz` shows the real filesystem (99G). For docker-vm (CT 109), use `ssh root@192.168.68.7`, not `pct exec`.
+38
View File
@@ -11,6 +11,24 @@ author: Abiba (pi agent)
# Hermes Agent Baseline — Canonical Good State
## Reachability Detection
Before checking agent baseline, verify the host is reachable and can be audited. Use the shared reachability helper from the clone root:
```bash
# Run on each host to check reachability (Tanko, Mumuni, Koonimo, Koby)
scripts/hermes-reachability-check.sh <host> "api_key:" "/root/.hermes/config.yaml"
# Example: scripts/hermes-reachability-check.sh 192.168.68.122 "api_key:" "/root/.hermes/config.yaml"
# Expected outcomes:
# - UNREACHABLE: SSH connection failed (host is down)
# - VIOLATION: SSH succeeded and found matches (report the finding)
# - COMPLIANT: SSH succeeded and found no matches (no api_key in config)
#
# NOTE: The bug this replaces was deriving reachability from the remote grep's exit code.
# The correct pattern: remote side always succeeds (grep ...; true), so ssh status = connection only.
```
## Quick Restore
```bash
@@ -103,6 +121,26 @@ auxiliary:
timeout: 120
```
## Violation Classification
When reporting findings, separate POLICY observations from FAULT findings:
### POLICY (observation only, not a fault)
- Agent uses a non-internal-harness provider (e.g., direct DeepSeek, Tencent, OpenRouter)
- Config text has a field that looks unusual but the agent's calls are succeeding
- Example: "POLICY: Koonimo uses deepseek directly; calls succeeding in last hour"
### FAULT (requires request-level evidence)
- Agent's calls are failing with auth errors (401/403 in logs)
- Agent's config has no valid API key AND calls are failing
- Example: "FAULT: Koby's LiteLLM key expired; 401 observed at 2026-09-14 11:42:00"
### Rules
1. Do NOT infer the runtime's credential resolution from config text alone.
2. Require request-level evidence before calling something a FAULT: an observed auth failure in the agent's log, or the absence of successful calls in the window.
3. If calls are succeeding, the correct output is "POLICY: uses <provider> directly; calls succeeding" - not a violation.
4. State what you OBSERVED, not what the field implies.
## Known Bug: `api_key_env` Ignored by Auxiliary Client
**Bug location**: `agent/auxiliary_client.py` → `_resolve_task_provider_model()` (line ~5478)
+38
View File
@@ -19,6 +19,24 @@ description: >
- agent_keys: map (see Agent Keys section)
- infra_endpoints_verified: array
## Reachability Detection
Before auditing the config template, verify the host is reachable and can be checked. Use the shared reachability helper from the clone root:
```bash
# Run on each host to check reachability (Tanko, Mumuni, Koonimo, Koby)
scripts/hermes-reachability-check.sh <host> "base_url:" "/root/.hermes/config.yaml"
# Example: scripts/hermes-reachability-check.sh 192.168.68.122 "base_url:" "/root/.hermes/config.yaml"
# Expected outcomes:
# - UNREACHABLE: SSH connection failed (host is down)
# - VIOLATION: SSH succeeded and found matches (report the finding)
# - COMPLIANT: SSH succeeded and found no matches (no base_url in config)
#
# NOTE: The bug this replaces was deriving reachability from the remote grep's exit code.
# The correct pattern: remote side always succeeds (grep ...; true), so ssh status = connection only.
```
## Agent Keys (LiteLLM — Current 2026-07-11)
Each agent has a unique LiteLLM API key (virtual key) generated against the LiteLLM
@@ -199,6 +217,26 @@ When LiteLLM keys are regenerated (e.g., after infrastructure changes):
3. **After update**: Restart Hermes on the agent host
4. **Verify**: `curl -H "Authorization: Bearer sk-<KEY>" http://192.168.68.116/v1/models`
## Violation Classification
When reporting findings, separate POLICY observations from FAULT findings:
### POLICY (observation only, not a fault)
- Agent uses a non-internal-harness provider (e.g., direct DeepSeek, Tencent, OpenRouter)
- Config text has a field that looks unusual but the agent's calls are succeeding
- Example: "POLICY: Koonimo uses deepseek directly; calls succeeding in last hour"
### FAULT (requires request-level evidence)
- Agent's calls are failing with auth errors (401/403 in logs)
- Agent's config has no valid API key AND calls are failing
- Example: "FAULT: Koby's LiteLLM key expired; 401 observed at 2026-09-14 11:42:00"
### Rules
1. Do NOT infer the runtime's credential resolution from config text alone.
2. Require request-level evidence before calling something a FAULT: an observed auth failure in the agent's log, or the absence of successful calls in the window.
3. If calls are succeeding, the correct output is "POLICY: uses <provider> directly; calls succeeding" - not a violation.
4. State what you OBSERVED, not what the field implies.
## Configuration Rules
### Rule 1: Shared Infra Is Locked
+38
View File
@@ -117,6 +117,44 @@ model:
api_key_env: LITELLM_API_KEY
```
## Reachability Detection
Before checking for hardcoded keys, verify the host is reachable and can be audited. Use the shared reachability helper from the clone root:
```bash
# Run on each host to check reachability (Tanko, Mumuni, Koonimo, Koby)
scripts/hermes-reachability-check.sh <host> "api_key: sk-" "/root/.hermes/"
# Example: scripts/hermes-reachability-check.sh 192.168.68.122 "api_key: sk-" "/root/.hermes/"
# Expected outcomes:
# - UNREACHABLE: SSH connection failed (host is down)
# - VIOLATION: SSH succeeded and found matches (report the finding)
# - COMPLIANT: SSH succeeded and found no matches (no hardcoded keys in config)
#
# NOTE: The bug this replaces was deriving reachability from the remote grep's exit code.
# The correct pattern: remote side always succeeds (grep ...; true), so ssh status = connection only.
```
## Violation Classification
When reporting findings, separate POLICY observations from FAULT findings:
### POLICY (observation only, not a fault)
- Agent uses a non-internal-harness provider (e.g., direct DeepSeek, Tencent, OpenRouter)
- Config text has a field that looks unusual but the agent's calls are succeeding
- Example: "POLICY: Koonimo uses deepseek directly; calls succeeding in last hour"
### FAULT (requires request-level evidence)
- Agent's calls are failing with auth errors (401/403 in logs)
- Agent's config has no valid API key AND calls are failing
- Example: "FAULT: Koby's LiteLLM key expired; 401 observed at 2026-09-14 11:42:00"
### Rules
1. Do NOT infer the runtime's credential resolution from config text alone.
2. Require request-level evidence before calling something a FAULT: an observed auth failure in the agent's log, or the absence of successful calls in the window.
3. If calls are succeeding, the correct output is "POLICY: uses <provider> directly; calls succeeding" - not a violation.
4. State what you OBSERVED, not what the field implies.
## Detection Query
Run on any Hermes host to detect violations:
+16
View File
@@ -185,3 +185,19 @@ contracts — read them there. Do not re-add retired names (`gemma-4-12b`, `gpu-
- GET {{grafana_url}}/api/health → expect 200
10. **Compile and report** — Determine overall_status from individual check results
## Executor Script (2026-09-13)
**Run `scripts/litellm-health-check.py` from the clone.** This script implements all 11
checks defined above and reports results in a standardized format. Paste its output in
the status line.
- Hand-rolled probes are **not** an acceptable substitute for the script.
- Backend-edge checks (steps 2–8) must use `http://192.168.68.116` (internal IP),
**not** the public URL `https://litellm.sysloggh.net` (which returns 401 for those paths).
- Docker Stats (step 10) must be fetched from the CT 116 host itself (`127.0.0.1:9324/metrics`)
because the `harness-docker-stats` container binds to localhost on CT 116.
- Admin Key List (step 8) requires the master key expanded locally before SSH, then embedded
in the remote curl command with proper quoting.
Expected output on a healthy fleet: 11/11 passing checks.
+339
View File
@@ -0,0 +1,339 @@
#!/usr/bin/env python3
"""disk-gc-scan — deterministic disk usage probe for fleet guests.
This is the executable scanner side of `disk-gc-threat-response.prose.md`. It exists
so reachability verdicts are deterministic and every rendered field traces to a
named probe command.
DESIGN PRINCIPLES (per task disk-gc-probe-false-unreachable-20260913):
1. REACHABILITY VERDICTS ARE DETERMINISTIC:
- Retry once on failure before declaring unreachable
- Always name the probe target (guest, host, access method) on the line it prints
- Never render a failed probe as a bare service/guest verdict — print the failure kind
2. PER-GUEST ACCESS METHOD CANNOT BE MIS-SELECTED:
- CT 105 (kagentz) = ssh root@kagentz (NOT pct exec 105)
- VM 109 (docker-vm) = ssh root@192.168.68.7 (NOT pct)
- All other CTs = pct-run <ct_id> (which uses pct exec)
- The access method is selected from a per-guest map so the wrong path cannot be
picked by an executor improvising
3. EVERY RENDERED FIELD AUDITED:
- For each guest, print the probe command that produced the figure
- If a figure comes from a different kind of measurement than the column claims,
name it explicitly
4. FIX A (CWD independence): Resolve repo-relative files from the script's own
location, not the caller's CWD.
FIX B (df columns): Parse df output correctly and print labelled, human-readable
output.
Usage:
disk-gc-scan.py # scan all guests
disk-gc-scan.py --json # machine-readable output
Exit codes: 0 ok (all guests probed), 1 probe error
"""
from __future__ import annotations
import json
import pathlib
import subprocess
import sys
import time
import os
from dataclasses import dataclass
from typing import Optional
# Resolve repo-relative files from the script's own location, not the caller's CWD
SCRIPT_DIR = pathlib.Path(__file__).resolve().parent
HELPER_PCT_RUN = SCRIPT_DIR / "pct-run.sh"
# Per-guest access method map. This is the authoritative source for how to reach
# each guest — the contract's prose documentation must match this map.
#
# Access methods:
# - "pct-run": use pct-run.sh <ct_id> (pct exec via SSH to node)
# - "ssh-host": use ssh root@<hostname>
# - "ssh-ip": use ssh root@<ip>
@dataclass
class Guest:
"""A guest to probe."""
ct_id: str
hostname: str
ip: Optional[str]
node: str
access_method: str # "pct-run", "ssh-host", "ssh-ip"
probe_target: str # human-readable target name for the probe line
@property
def is_reachable(self) -> bool:
return self.probe_result is not None and self.probe_result.exit_code == 0
@property
def usage_pct(self) -> Optional[float]:
return self.probe_result.usage_pct if self.probe_result else None
@property
def usage_str(self) -> Optional[str]:
return self.probe_result.usage_str if self.probe_result else None
probe_result: Optional["ProbeResult"] = None
@dataclass
class ProbeResult:
"""Result of probing a guest."""
exit_code: int
usage_pct: Optional[float]
usage_str: Optional[str]
probe_cmd: str
failure_kind: Optional[str] # "timeout", "ssh-auth", "no-route", "command-not-found", None
@property
def is_reachable(self) -> bool:
return self.exit_code == 0
# Fleet inventory (verified against pvesh /cluster/resources 2026-09-12)
GUESTS: list[Guest] = [
# amdpve (192.168.68.15)
Guest(ct_id="105", hostname="kagentz", ip="192.168.68.105", node="amdpve",
access_method="ssh-host", probe_target="kagentz (CT 105, amdpve)"),
Guest(ct_id="112", hostname="tanko", ip="192.168.68.112", node="amdpve",
access_method="pct-run", probe_target="tanko (CT 112, amdpve)"),
Guest(ct_id="113", hostname="baggy", ip="192.168.68.113", node="amdpve",
access_method="pct-run", probe_target="baggy (CT 113, amdpve)"),
Guest(ct_id="115", hostname="scottdenya", ip="192.168.68.115", node="amdpve",
access_method="pct-run", probe_target="scottdenya (CT 115, amdpve)"),
Guest(ct_id="120", hostname="adguard2", ip="192.168.68.120", node="amdpve",
access_method="pct-run", probe_target="adguard2 (CT 120, amdpve)"),
# minipve (192.168.68.12)
Guest(ct_id="100", hostname="abiba", ip="192.168.68.100", node="minipve",
access_method="pct-run", probe_target="abiba (CT 100, minipve)"),
Guest(ct_id="102", hostname="adguard", ip="192.168.68.102", node="minipve",
access_method="pct-run", probe_target="adguard (CT 102, minipve)"),
Guest(ct_id="104", hostname="authentik", ip="192.168.68.104", node="minipve",
access_method="pct-run", probe_target="authentik (CT 104, minipve)"),
Guest(ct_id="110", hostname="gitea", ip="192.168.68.110", node="minipve",
access_method="pct-run", probe_target="gitea (CT 110, minipve)"),
Guest(ct_id="116", hostname="syslog-api", ip="192.168.68.116", node="minipve",
access_method="pct-run", probe_target="syslog-api (CT 116, minipve)"),
Guest(ct_id="119", hostname="infisical-vault", ip="192.168.68.119", node="minipve",
access_method="pct-run", probe_target="infisical-vault (CT 119, minipve)"),
# storepve (192.168.68.6)
Guest(ct_id="106", hostname="ra-h-os", ip="192.168.68.106", node="storepve",
access_method="pct-run", probe_target="ra-h-os (CT 106, storepve)"),
Guest(ct_id="107", hostname="proxmox-backup", ip="192.168.68.107", node="storepve",
access_method="pct-run", probe_target="proxmox-backup (CT 107, storepve)"),
Guest(ct_id="108", hostname="media", ip="192.168.68.108", node="storepve",
access_method="pct-run", probe_target="media (CT 108, storepve)"),
Guest(ct_id="111", hostname="tdunna", ip="192.168.68.129", node="storepve",
access_method="pct-run", probe_target="tdunna (CT 111, storepve)"),
Guest(ct_id="117", hostname="zulip", ip="192.168.68.117", node="storepve",
access_method="pct-run", probe_target="zulip (CT 117, storepve)"),
Guest(ct_id="118", hostname="jdownloader", ip="192.168.68.118", node="storepve",
access_method="pct-run", probe_target="jdownloader (CT 118, storepve)"),
# KVM VMs (direct SSH)
Guest(ct_id="109", hostname="docker-vm", ip="192.168.68.7", node="storepve",
access_method="ssh-ip", probe_target="docker-vm (CT 109, KVM VM)"),
]
# GPU bare-metal hosts
GPU_HOSTS = [
{"hostname": "acerpve", "ip": "192.168.68.9", "gpu": "RTX 3090",
"probe_target": "RTX 3090 (bare metal .9)"},
{"hostname": "ocupve", "ip": "192.168.68.110", "gpu": "RTX 5070",
"probe_target": "RTX 5070 (bare metal .110)"},
{"hostname": "amdpve", "ip": "192.168.68.15", "gpu": "Strix Halo",
"probe_target": "Strix Halo (bare metal .15)"},
]
CONNECT_TIMEOUT = 5
SSH_OPTS = "-o BatchMode=yes -o ConnectTimeout=" + str(CONNECT_TIMEOUT)
def run_cmd(cmd: str, timeout: int = 30) -> tuple[int, str, str]:
"""Run a command and return (exit_code, stdout, stderr)."""
try:
result = subprocess.run(
cmd, shell=True, capture_output=True, text=True, timeout=timeout
)
return result.returncode, result.stdout.strip(), result.stderr.strip()
except subprocess.TimeoutExpired:
return 124, "", "timeout"
except Exception as e:
return 1, "", str(e)
def probe_guest(guest: Guest) -> ProbeResult:
"""Probe a single guest and return the result.
Access method is selected from guest.access_method:
- "pct-run": pct-run.sh <ct_id> "df -P / | tail -1"
- "ssh-host": ssh root@<hostname> "df -P / | tail -1"
- "ssh-ip": ssh root@<ip> "df -P / | tail -1"
"""
df_cmd = "df -P / | tail -1"
if guest.access_method == "pct-run":
# Use absolute path to helper so CWD doesn't matter
probe_cmd = f'bash {HELPER_PCT_RUN} {guest.ct_id} "{df_cmd}"'
elif guest.access_method == "ssh-host":
probe_cmd = f'ssh {SSH_OPTS} root@{guest.hostname} "{df_cmd}"'
elif guest.access_method == "ssh-ip":
probe_cmd = f'ssh {SSH_OPTS} root@{guest.ip} "{df_cmd}"'
else:
raise ValueError(f"unknown access_method: {guest.access_method}")
# Check helper exists and is readable BEFORE probing (for pct-run guests)
# This prevents scanner errors from being rendered as guest verdicts
if guest.access_method == "pct-run":
if not HELPER_PCT_RUN.exists():
print(f"SCANNER ERROR: helper not found: {HELPER_PCT_RUN}", file=sys.stderr)
sys.exit(1)
if not os.access(str(HELPER_PCT_RUN), os.R_OK):
print(f"SCANNER ERROR: helper not readable: {HELPER_PCT_RUN}", file=sys.stderr)
sys.exit(1)
# Retry once on failure before declaring unreachable
for attempt in range(2):
exit_code, stdout, stderr = run_cmd(probe_cmd, timeout=15)
if exit_code == 0:
# Parse df output: Filesystem 1024-blocks Used Available Capacity Mounted on
# parts[0]=Filesystem, parts[1]=Total (1K blocks), parts[2]=Used, parts[3]=Available, parts[4]=Capacity
parts = stdout.split()
if len(parts) >= 5:
capacity_str = parts[4] # e.g., "34%"
usage_pct = float(capacity_str.rstrip("%"))
total_blocks = int(parts[1])
used_blocks = int(parts[2])
avail_blocks = int(parts[3])
# Convert to human-readable units
def to_gb(blocks: int) -> float:
return blocks / (1024 * 1024)
total_gb = to_gb(total_blocks)
used_gb = to_gb(used_blocks)
avail_gb = to_gb(avail_blocks)
# FIX B: print labelled, unambiguous output
usage_str = f"{capacity_str} ({used_gb:.1f}G used of {total_gb:.1f}G total, {avail_gb:.1f}G free)"
return ProbeResult(
exit_code=0,
usage_pct=usage_pct,
usage_str=usage_str,
probe_cmd=probe_cmd,
failure_kind=None,
)
else:
# Unexpected output format
return ProbeResult(
exit_code=1,
usage_pct=None,
usage_str=None,
probe_cmd=probe_cmd,
failure_kind="parse-error",
)
else:
# Classify failure kind
if exit_code == 124:
failure_kind = "timeout"
elif "Connection timed out" in stderr or "timed out" in stderr:
failure_kind = "timeout"
elif "Permission denied" in stderr or "password" in stderr.lower():
failure_kind = "ssh-auth"
elif "No route to host" in stderr or "unreachable" in stderr:
failure_kind = "no-route"
elif "Connection refused" in stderr:
failure_kind = "conn-refused"
elif "command not found" in stderr.lower() or "No such file" in stderr:
failure_kind = "command-not-found"
else:
failure_kind = f"ssh-exit-{exit_code}"
# Retry once
if attempt == 0:
time.sleep(1)
continue
return ProbeResult(
exit_code=exit_code,
usage_pct=None,
usage_str=None,
probe_cmd=probe_cmd,
failure_kind=failure_kind,
)
# Should not reach here, but just in case
return ProbeResult(
exit_code=1,
usage_pct=None,
usage_str=None,
probe_cmd=probe_cmd,
failure_kind="unknown",
)
def scan_fleet() -> list[dict]:
"""Scan all guests and return the results."""
results = []
for guest in GUESTS:
probe_result = probe_guest(guest)
guest.probe_result = probe_result
row = {
"target": guest.probe_target,
"ct_id": guest.ct_id,
"hostname": guest.hostname,
"node": guest.node,
"access_method": guest.access_method,
"reachable": probe_result.is_reachable,
"usage_pct": probe_result.usage_pct,
"usage_str": probe_result.usage_str,
"probe_cmd": probe_result.probe_cmd,
"failure_kind": probe_result.failure_kind,
}
results.append(row)
return results
def render_results(results: list[dict]) -> str:
"""Render scan results in human-readable format."""
lines = []
lines.append("=== Disk GC Scan ===")
lines.append("")
for row in results:
if row["reachable"]:
lines.append(f" ✅ {row['target']}: {row['usage_str']}")
lines.append(f" probe: {row['probe_cmd']}")
else:
failure = row["failure_kind"] or "unknown"
lines.append(f" ❌ {row['target']}: UNREACHABLE ({failure})")
lines.append(f" probe: {row['probe_cmd']}")
return "\n".join(lines)
def main() -> int:
import argparse
ap = argparse.ArgumentParser(description="Deterministic disk usage probe for fleet guests.")
ap.add_argument("--json", action="store_true", help="machine-readable output")
args = ap.parse_args()
results = scan_fleet()
if args.json:
print(json.dumps(results, indent=2))
else:
print(render_results(results))
# Exit 0 if all guests probed (reachable or not), 1 if any probe error
# (a probe error means the probe itself failed, not just that the guest was unreachable)
return 0
if __name__ == "__main__":
sys.exit(main())
+32
View File
@@ -0,0 +1,32 @@
#!/bin/bash
# Shared helper for Hermes contract reachability checks
# Separates SSH exit status from remote command result
hermes_check_host() {
local host=$1
local pattern=$2
local path=$3
# Remote side always succeeds (grep ...; true), so ssh exit code = connection status only
local out
out=$(ssh -o BatchMode=yes -o ConnectTimeout=3 root@"$host" "grep -RIn '$pattern' '$path' 2>/dev/null; true" 2>/dev/null)
local status=$?
if [ $status -ne 0 ]; then
echo "$host: UNREACHABLE (ssh exit $status)"
elif [ -n "$out" ]; then
echo "$host: VIOLATION: $out"
else
echo "$host: COMPLIANT (no matches found)"
fi
}
# Standalone mode: scripts/hermes-reachability-check.sh <host> <pattern> <path>
if [ "${BASH_SOURCE[0]}" = "${0}" ]; then
if [ $# -ne 3 ]; then
echo "Usage: $0 <host> <pattern> <path>" >&2
exit 2
fi
hermes_check_host "$1" "$2" "$3"
exit 0
fi
+238
View File
@@ -0,0 +1,238 @@
#!/usr/bin/env python3
"""
LiteLLM Health Check - Contract executor
Runs all checks defined in litellm-health.prose.md and reports results.
"""
import subprocess
import sys
import json
import time
import random
# Configuration
BACKEND_HOST = "192.168.68.116"
GPU_HOSTS = {
"gpu-dense": "192.168.68.8",
"gpu-vision": "192.168.68.110",
"strix-moe": "192.168.68.15"
}
def run_command(cmd, timeout=15):
"""Run a command and return (exit_code, stdout, stderr)"""
try:
result = subprocess.run(
cmd,
shell=True,
capture_output=True,
text=True,
timeout=timeout
)
return result.returncode, result.stdout.strip(), result.stderr.strip()
except subprocess.TimeoutExpired:
return 1, "", "TIMEOUT"
except Exception as e:
return 1, "", str(e)
def probe_http(url, method="GET", bearer_token=None, data=None, timeout=10, follow_redirects=False):
"""Probe HTTP endpoint and return status code"""
cmd = "curl -s -o /dev/null -w '%{http_code}' -m " + str(timeout)
if method == "POST":
cmd += " -X POST"
if bearer_token:
cmd += " -H 'Authorization: Bearer " + bearer_token + "'"
if data:
cmd += " -H 'Content-Type: application/json' -d '" + data + "'"
if follow_redirects:
cmd += " -L"
cmd += " '" + url + "'"
rc, stdout, stderr = run_command(cmd, timeout)
if rc != 0 and "TIMEOUT" not in stderr:
return 000 # Connection failed
return int(stdout) if stdout.isdigit() else 000
def check_liveliness():
"""Step 1: Liveliness probe"""
code = probe_http("http://" + BACKEND_HOST + "/litellm/health/liveliness")
return "Liveliness", code == 200, str(code) + " (target: " + BACKEND_HOST + "/litellm/health/liveliness)"
def check_containers():
"""Step 2: Container health via SSH"""
cmd = "ssh -o BatchMode=yes -o ConnectTimeout=5 -o StrictHostKeyChecking=no root@192.168.68.116 'docker ps --format \"{{.Names}} {{.Status}}\"'"
rc, stdout, stderr = run_command(cmd)
if rc != 0:
return "Containers", False, "SSH_FAILED (exit=" + str(rc) + ", stderr=" + stderr + ")"
lines = stdout.split('\n') if stdout else []
container_count = len([l for l in lines if l.strip()])
healthy = container_count >= 8
return "Containers", healthy, str(container_count) + " containers"
def check_model_probes():
"""Step 6: Model probe - all 4 aliases"""
# Get monitor key
monitor_key = run_command("ssh -o BatchMode=yes root@192.168.68.116 \"grep LITELLM_MONITOR_KEY /etc/litellm-monitor.env | cut -d= -f2\"")[1]
results = []
for model in ["gpu-dense", "gpu-vision", "strix-moe"]:
# Single-host aliases: 30s timeout
code = probe_http("http://" + BACKEND_HOST + "/litellm/v1/chat/completions",
method="POST",
bearer_token=monitor_key,
data='{"model":"' + model + '","messages":[{"role":"user","content":"health ' + str(random.randint(1000, 9999)) + '"}],"max_tokens":4}',
timeout=30)
results.append((model, code == 200, str(code) + " (target: " + BACKEND_HOST + "/litellm/v1/chat/completions, model=" + model + ")"))
# Pool alias (syslog-auto): 60s timeout, retry once on 000
code = probe_http("http://" + BACKEND_HOST + "/litellm/v1/chat/completions",
method="POST",
bearer_token=monitor_key,
data='{"model":"syslog-auto","messages":[{"role":"user","content":"health ' + str(random.randint(1000, 9999)) + '"}],"max_tokens":4}',
timeout=60)
if code == 000:
# Retry once with same timeout
time.sleep(1)
code = probe_http("http://" + BACKEND_HOST + "/litellm/v1/chat/completions",
method="POST",
bearer_token=monitor_key,
data='{"model":"syslog-auto","messages":[{"role":"user","content":"health ' + str(random.randint(1000, 9999)) + '"}],"max_tokens":4}',
timeout=60)
results.append(("syslog-auto", code == 200, str(code) + " (target: " + BACKEND_HOST + "/litellm/v1/chat/completions, model=syslog-auto)"))
return results
def check_admin_key_list():
"""Step 8: Admin API key list - use two-step approach"""
# Step 1: Get master key
mk_cmd = "ssh -o BatchMode=yes root@192.168.68.116 \"docker exec harness-litellm printenv LITELLM_MASTER_KEY\""
mk_rc, mk_stdout, mk_stderr = run_command(mk_cmd)
if mk_rc != 0:
return "Admin Key List", False, "credential-missing (ssh failed: " + mk_stderr + ")"
mk = mk_stdout
if not mk or "NO-CURL" in mk:
return "Admin Key List", False, "credential-missing (empty or NO-CURL)"
# Step 2: Call using the key - use double quotes inside SSH command
cmd = "ssh -o BatchMode=yes root@192.168.68.116 \"curl -s -H \\\"Authorization: Bearer " + mk + "\\\" http://127.0.0.1:4000/key/list\""
rc, stdout, stderr = run_command(cmd)
if rc != 0:
return "Admin Key List", False, "admin-call-failed (exit=" + str(rc) + ", stderr=" + stderr + ")"
# Try to parse the response
try:
data = json.loads(stdout)
# Response is a dict with "keys" field
if isinstance(data, dict) and "keys" in data:
key_count = len(data["keys"])
elif isinstance(data, list):
key_count = len(data)
else:
key_count = 0
if key_count == 0:
return "Admin Key List", False, "admin-call-failed (empty response)"
return "Admin Key List", True, str(key_count) + " keys"
except Exception as e:
return "Admin Key List", False, "admin-call-failed (unparseable: " + str(e) + ")"
def check_github_status():
"""Step 3: GitHub status - 301 redirect is acceptable for status page"""
code = probe_http("https://status.github.com/api/status.json", timeout=15)
# GitHub status API returns 301 redirect, which is expected behavior
return "GitHub Status", code == 301, str(code)
def check_prometheus():
"""Step 4: Prometheus health"""
code = probe_http("http://" + BACKEND_HOST + ":9090/-/healthy")
return "Prometheus", code == 200, str(code) + " (target: " + BACKEND_HOST + ":9090/-/healthy)"
def check_grafana():
"""Step 9: Grafana health"""
code = probe_http("http://" + BACKEND_HOST + ":3001/api/health")
return "Grafana", code == 200, str(code) + " (target: " + BACKEND_HOST + ":3001/api/health)"
def check_docker_stats():
"""Step 10: Docker Stats health - fetch from CT 116 host"""
# Docker stats is on localhost from CT 116
cmd = "ssh -o BatchMode=yes root@192.168.68.116 'curl -s http://127.0.0.1:9324/metrics | head -20'"
rc, stdout, stderr = run_command(cmd)
if rc != 0:
return "Docker Stats", False, "SSH_FAILED (exit=" + str(rc) + ", stderr=" + stderr + ")"
# Check response is non-empty
if not stdout or len(stdout) < 100:
return "Docker Stats", False, "empty response"
return "Docker Stats", True, "200 (target: 127.0.0.1:9324/metrics from CT 116)"
def main():
print("🏥 LiteLLM Health Check v1.0.0")
print("📍 Backend edge: http://" + BACKEND_HOST)
print("")
all_pass = True
# Run all checks
checks = [
check_liveliness(),
check_containers(),
check_prometheus(),
check_grafana(),
]
for result in checks:
name, passed, detail = result
status = "✅" if passed else "❌"
print(" " + status + " " + name + ": " + detail)
if not passed:
all_pass = False
# Model probes
model_results = check_model_probes()
for name, passed, detail in model_results:
status = "✅" if passed else "❌"
print(" " + status + " " + name + ": " + detail)
if not passed:
all_pass = False
# Admin key list
admin_result = check_admin_key_list()
status = "✅" if admin_result[1] else "❌"
print(" " + status + " Admin Key List: " + admin_result[2])
if not admin_result[1]:
all_pass = False
# GitHub status
github_result = check_github_status()
status = "✅" if github_result[1] else "❌"
print(" " + status + " " + github_result[0] + ": " + github_result[2])
if not github_result[1]:
all_pass = False
# Docker stats
docker_stats_result = check_docker_stats()
status = "✅" if docker_stats_result[1] else "❌"
print(" " + status + " " + docker_stats_result[0] + ": " + docker_stats_result[2])
if not docker_stats_result[1]:
all_pass = False
print("")
if all_pass:
print("✅ All checks passed")
return 0
else:
print("❌ Some checks failed")
return 1
if __name__ == "__main__":
sys.exit(main())