Compare commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
937afc0baf | ||
|
|
288d613876 | ||
|
|
1d20fbaa7f | ||
|
|
2f961d7e7a | ||
|
|
4fe4f3621d | ||
|
|
86d2987ad8 | ||
|
|
efe9381283 | ||
|
|
6a1c4967db | ||
|
|
5ed6f8179c | ||
|
|
b9322973ce | ||
|
|
fb1916707d | ||
|
|
d28df4f4de | ||
|
|
cb26ee06d6 | ||
|
|
9a789ab76d | ||
|
|
71ceda0042 | ||
|
|
4e34b7a2a2 | ||
|
|
f9f6661dd5 | ||
|
|
cb36ff1ea5 | ||
|
|
05366bd58d | ||
|
|
e648b5ac0e | ||
|
|
38d7e8b064 | ||
|
|
e94fadfa60 | ||
|
|
ba76f2c7d3 | ||
|
|
7906b2d52d | ||
|
|
c81cf5b6f0 | ||
|
|
ef168d9690 | ||
|
|
edcf465831 | ||
|
|
89651cf37c | ||
|
|
6f40a3be60 | ||
|
|
d9368467ff | ||
|
|
e38598eea4 | ||
|
|
876b011359 | ||
|
|
d23cce89e1 | ||
|
|
9ada2b23c7 | ||
|
|
bd0065bb31 | ||
|
|
f99f7e1e34 | ||
|
|
8210fd905c | ||
|
|
b3644f0292 |
@@ -88,7 +88,7 @@ prose run memory-audit-maintenance memory_threshold=90 verify_configs=true
|
||||
prose run hermes-config-template agent_name=syslog-devops default_model=claude-sonnet-4
|
||||
|
||||
# Configure an agent with a different auxiliary model
|
||||
prose run hermes-config-template agent_name=syslog-code default_model=qwen3.6-27B-code auxiliary_model=gpu-vision
|
||||
prose run hermes-config-template agent_name=syslog-code default_model=gpu-dense auxiliary_model=gpu-vision
|
||||
```
|
||||
|
||||
### Option B: Manual Execution
|
||||
|
||||
@@ -34,7 +34,7 @@ verification, and DM loopback testing.
|
||||
| @all-bots user ID | 20 | ✅ (config, verified by API at runtime) |
|
||||
| PM2 process name | abiba-zulip | ✅ |
|
||||
| Provider | syslog-harness (http://192.168.68.116/v1) | ✅ |
|
||||
| Default model | deepseek-v4-pro | ✅ (settings.json) |
|
||||
| Default model | syslog-auto | ✅ (settings.json) |
|
||||
|
||||
## Architecture
|
||||
|
||||
|
||||
@@ -163,7 +163,7 @@ def audit(path):
|
||||
check(
|
||||
comp.get("max_context_window") == 131072,
|
||||
"Rule 9",
|
||||
f"compression.max_context_window must be 131072 (got {comp.get('max_context_window')!r}) — matches 128K GPU capacity",
|
||||
f"compression.max_context_window must be 131072 (got {comp.get('max_context_window')!r}) — syslog-auto pool floor (NVIDIA hosts 128K; Strix Halo 256K)",
|
||||
)
|
||||
|
||||
# --- Rule 10: Default Model Must Be syslog-auto ---
|
||||
@@ -245,11 +245,10 @@ def audit(path):
|
||||
"gemma-4-12b": "gpu-vision",
|
||||
"crew-auto": "syslog-auto (its 64K cap is retired; no cap in force)",
|
||||
"ornith-1.0-35b": "strix-moe",
|
||||
}
|
||||
raw_but_live = {
|
||||
"qwen3.6-27B-code": "gpu-dense",
|
||||
"qwen3.6-35B-udq4": "strix-moe",
|
||||
}
|
||||
raw_but_live = {}
|
||||
for field_path, value in _iter_model_values(cfg):
|
||||
if value in non_resolving:
|
||||
check(
|
||||
|
||||
@@ -84,6 +84,32 @@ and escalation trail.
|
||||
- May also be invoked manually: `prose run disk-gc-threat-response`
|
||||
- Threat-driven: if Amber/Red/Critical detected, immediate GC phase activates
|
||||
|
||||
## Scanner: scripts/disk-gc-scan.py
|
||||
|
||||
The fleet scan is executed by `scripts/disk-gc-scan.py`, which makes reachability
|
||||
verdicts deterministic:
|
||||
|
||||
1. **Retry on failure:** Each probe retries once before declaring a guest unreachable.
|
||||
2. **Named probe target:** Every rendered line names the guest, CT id, node, and
|
||||
access method actually used.
|
||||
3. **Failure kind printed:** An unreachable guest is reported with its failure kind
|
||||
(timeout, ssh-auth, no-route, conn-refused, ssh-exit-N) — never as a bare
|
||||
"unreachable" verdict.
|
||||
4. **Per-guest access method:** The correct access path is selected from a per-guest
|
||||
map so the wrong path cannot be picked by an executor improvising:
|
||||
- CT 105 (kagentz) = `ssh root@kagentz` (NOT `pct exec 105` — pct exec sees
|
||||
loop0/59G instead of the real 99G filesystem)
|
||||
- CT 109 (docker-vm) = `ssh root@192.168.68.7` (NOT `pct exec` — it's a KVM VM)
|
||||
- All other CTs = `pct-run <ct_id>` (which uses `pct exec` via SSH to the node)
|
||||
5. **Every figure traces to a named probe:** The scan output prints the exact command
|
||||
that produced each disk figure, so two different guests can never render
|
||||
identical numbers without the probe commands proving it.
|
||||
|
||||
Run: `python3 scripts/disk-gc-scan.py` (or `--json` for machine-readable output).
|
||||
|
||||
The scan feeds into `scripts/disk-gc-plan.py`, which applies the report-only gate
|
||||
from the `report_only_guests` YAML block above.
|
||||
|
||||
## Shape
|
||||
|
||||
- `self`: scan all CTs via Proxmox API + SSH exec, trigger GC, alert
|
||||
@@ -357,3 +383,4 @@ one-off GPU builds. No automated post-migration cleanup was in place.
|
||||
> container has no `pct` binary.
|
||||
>
|
||||
> **KVM VM:** CT 109 (docker-vm) is a QEMU VM, not LXC — access via SSH .7.
|
||||
> **NOTE:** For kagentz (CT 105), use `ssh root@kagentz` (hostname), NOT `pct exec 105` — `pct exec 105` shows loop0 (59G) while `ssh root@kagentz` shows the real filesystem (99G). For docker-vm (CT 109), use `ssh root@192.168.68.7`, not `pct exec`.
|
||||
|
||||
+15
-15
@@ -8,12 +8,12 @@ description: >
|
||||
UPDATED 2026-07-15: Stable role-based aliases introduced: strix-moe,
|
||||
gpu-dense, gpu-vision (gpu-light was superseded by gpu-vision on 2026-09-12).
|
||||
These never change — only the underlying model does.
|
||||
Strix Halo: strix-moe → unsloth/Qwen3.6-35B-A3B-MTP (UD-Q4_K_M, 22GB).
|
||||
Strix Halo: strix-moe → Carnice-Qwen3.6-MoE-35B-A3B-Q4_K_M.gguf (22GB, 256K ctx).
|
||||
RTX 5070: gpu-vision — IQ4_NL + MTP draft (~122 tok/s, 2x faster).
|
||||
UPDATED 2026-07-17: Context reduced fleet-wide from 256K to 128K for stability.
|
||||
Strix Halo model swapped to qwen3.6-35B-udq4 (22GB, strix-moe alias).
|
||||
Instability observed near 100K at 256K (now all GPUs at 128K). 128K is the stable ceiling.
|
||||
For larger context needs → fall back to external providers (deepseek).
|
||||
UPDATED 2026-07-17: NVIDIA host context reduced from 256K to 128K for stability.
|
||||
Strix Halo model: Carnice-Qwen3.6-MoE-35B-A3B-Q4_K_M.gguf (alias strix-moe, 256K context).
|
||||
Strix Halo runs 256K (n_ctx 262144, --kv-unified); RTX 3090 and RTX 5070 remain at 128K.
|
||||
For >128K on NVIDIA hosts → fall back to external providers (deepseek).
|
||||
VRAM headroom improved: RTX 3090 ~70%, RTX 5070 ~65%.
|
||||
agent: abiba
|
||||
triggers:
|
||||
@@ -91,8 +91,8 @@ Single source of truth for models, aliases, rpm caps, weights and fallback chain
|
||||
CT 116 `/opt/inference-harness/litellm_config.yaml`. Do not duplicate those values in
|
||||
contracts — read them there.
|
||||
|
||||
**Backward compatibility**: Old model-specific names (qwen3.6-27B-code, qwen3.5-9b-it) still work
|
||||
but are deprecated for agent configs. Only the stable aliases survive model swaps. `gemma-4-12b`
|
||||
**No backward compatibility**: Old model-specific names (qwen3.6-27B-code, qwen3.5-9b-it) are retired as of 2026-09-12
|
||||
and no longer resolve; do not use them in agent configs. Only the stable aliases survive model swaps. `gemma-4-12b`
|
||||
is retired and returns 400 `Invalid model name`.
|
||||
|
||||
## Routing Configuration (LiteLLM — July 2026)
|
||||
@@ -190,7 +190,7 @@ If no SSH access, send Zulip DM via abiba-bot with vault update instructions.
|
||||
| `/root/scripts/gpu-saturation-watchdog.py` | pi (.24) | Auto-restart stuck llama-server |
|
||||
| `/root/dashboard/gpu-fleet.html` | pi (.24) | Live HTML dashboard |
|
||||
| `/etc/systemd/system/llama-server.service` | .8, .110 | llama-server daemons (Nvidia GPUs) |
|
||||
| `/etc/systemd/system/strix-server.service` | .15 (amdpve) | llama-server daemon (Vulkan, Strix Halo) running unsloth/Qwen3.6-35B-A3B-MTP-GGUF. Note: `llama-server.service` and `llama-server@.service` are **masked** on .15 to prevent port 8080 collisions. |
|
||||
| `/etc/systemd/system/strix-server.service` | .15 (amdpve) | llama-server daemon (Vulkan, Strix Halo) running Carnice-Qwen3.6-MoE-35B-A3B-Q4_K_M.gguf (256K ctx). Note: `llama-server.service` and `llama-server@.service` are **masked** on .15 to prevent port 8080 collisions. |
|
||||
|
||||
## Prometheus & Grafana
|
||||
|
||||
@@ -207,7 +207,7 @@ If no SSH access, send Zulip DM via abiba-bot with vault update instructions.
|
||||
- **Router startup race**: Compose router.py doesn't call load_roster(). Reload thread sleeps 30s first.
|
||||
Fix: trigger roster reload via SSH after restart, or rebuild image with startup load_roster().
|
||||
- **LiteLLM /metrics**: Requires auth. Prometheus uses `/health/liveliness` as workaround.
|
||||
- **Strix Halo GPU**: Vulkan is the working backend (ROCm/HIP path abandoned — HSA runtime blocked on Debian 13). Build at `/root/llama.cpp/build-vk/`, commit `4fc4ec5` (2026-07-01), ggml 0.15.3 shared-lib arch. Mesa RADV 25.0.7, KHR_coopmat fast path active. ~70 tok/s gen, 532 tok/s prompt. Service: `strix-server.service` on port 8080, model: `qwen3.6-35B-udq4`, alias `strix-moe`, 128K context, flash-attn + q4 KV, multimodal (mmproj loaded).
|
||||
- **Strix Halo GPU**: Vulkan is the working backend (ROCm/HIP path abandoned — HSA runtime blocked on Debian 13). Build at `/root/llama.cpp/build-vk/`, commit `4fc4ec5` (2026-07-01), ggml 0.15.3 shared-lib arch. Mesa RADV 25.0.7, KHR_coopmat fast path active. ~70 tok/s gen, 532 tok/s prompt. Service: `strix-server.service` on port 8080, model: `Carnice-Qwen3.6-MoE-35B-A3B-Q4_K_M.gguf`, alias `strix-moe`, 256K context (n_ctx 262144), --parallel 2 --kv-unified, flash-attn + q4 KV, multimodal (mmproj loaded).
|
||||
- **Port conflict detection (2026-07-05)**: All 3 GPU wrappers now detect ghost processes squatting port 8080 before starting. `.8` and `.110` use inline pre-start check in `llama-wrapper.sh`; `.15` uses `/usr/local/bin/port-cleanup.sh` ExecStartPre. Replaces the blanket `pkill -9 -x llama-server` on .15 which would kill ALL llama-server instances regardless of port. Ghost detection was the root cause of .8 crash-looping for 27+ restarts (stale pid 25836 squatting 8080 after OOM kill).
|
||||
- **Strix Halo thermal safeguard (2026-07-02)**: `strix-server.service` has `-n 8192` (hard generation cap per request). Without it, `--predict` defaults to -1 (infinity) — a runaway request from .123 (old Mumuni CT114 — now inside Abiba CT100 at .24) decoded 39,868 tokens over 24 min, pushing Tctl to 98°C (crit 89.8°C) and throttling 70→29 t/s. The cap bounds worst-case generation to ~5 min. Do NOT remove `-n` without a replacement ceiling. Sustained load hits ~84°C even at 92s; the APU is fanless/low-flow. Clients MUST also set `max_tokens`.
|
||||
- **Port 8080 firewall**: amdpve iptables restricts 8080 to 192.168.68.116 (LiteLLM/router host) only. All inbound connections are from .116 (LiteLLM proxied via nginx). Localhost curls hang (SYN dropped). Always test from .116.
|
||||
@@ -223,10 +223,10 @@ If no SSH access, send Zulip DM via abiba-bot with vault update instructions.
|
||||
| GPU | Model | Gen tok/s | Prompt tok/s | Baseline | Context |
|
||||
|-----|-------|-----------|--------------|----------|---------|
|
||||
| RTX 3090 (.8) | Qwen3.8-27B-Uncensored-Q4_K_M | **TBD** | — | — | **128K** |
|
||||
| Strix Halo (.15) | qwen3.6-35B-udq4 | **65** | 140 | — | **128K** |
|
||||
| Strix Halo (.15) | Carnice-Qwen3.6-MoE-35B-A3B-Q4_K_M.gguf (strix-moe) | **65** | 140 | — | **256K** |
|
||||
|
||||
Benchmarks from 2026-07-17. Strix Halo model: qwen3.6-35B-udq4. RTX 5070 MTP provides 2.7x speedup over pre-upgrade 70 tok/s.
|
||||
All 3 GPUs now at 128K context (2026-07-17, reduced from 256K for stability).
|
||||
Benchmarks from 2026-07-17. Strix Halo model: Carnice-Qwen3.6-MoE-35B-A3B-Q4_K_M.gguf (alias strix-moe), n_ctx 262144, --parallel 2 --kv-unified. RTX 5070 MTP provides 2.7x speedup over pre-upgrade 70 tok/s.
|
||||
GPU contexts: RTX 3090 (.8) and RTX 5070 (.110) at 128K; Strix Halo (.15) at 256K (2026-09-12).
|
||||
|
||||
Benchmarks run through LiteLLM proxy (192.168.68.116:4000) every 5 minutes.
|
||||
Degradation alerts fire at 30% (warning) and 50% (critical) below baseline.
|
||||
@@ -247,8 +247,8 @@ All agent configs MUST use stable role-based aliases, never model-specific names
|
||||
When the underlying model is swapped, only the LiteLLM config changes — agent configs are untouched.
|
||||
|
||||
### Context Windows
|
||||
- RTX 3090: **128K** (reduced from 256K 2026-07-17) | RTX 5070: **128K** (reduced from 256K) | Strix Halo: **128K**
|
||||
- **All agents**: 128K ceiling — stable margin. For >128K workloads, use external providers (deepseek)
|
||||
- RTX 3090: **128K** (reduced from 256K 2026-07-17) | RTX 5070: **128K** (reduced from 256K) | Strix Halo: **256K** (2026-09-12)
|
||||
- **Agents via `syslog-auto`**: 128K ceiling — the pool's safe floor (NVIDIA hosts are 128K). For >128K workloads, use external providers (deepseek)
|
||||
- Compression threshold 0.65 (audit Rule 9): fires at ~85K (~43K headroom before 128K ceiling)
|
||||
- **Pi agents (Abiba)**: `compaction.reserveTokens: 52739` (≈60% of 128K)
|
||||
- Mumuni compression model alias: `syslog-auto`
|
||||
@@ -266,7 +266,7 @@ Mumuni (kagentz CT105, 192.168.68.14 — migrated from CT100 2026-08-29) is the
|
||||
| `aux.vision.model` | `gpu-vision` | Vision tasks (RTX 5070) |
|
||||
| `aux.web_extract.model` | `gpu-vision` | Web extraction |
|
||||
| `delegation.model` | `gpu-dense` | Sub-agent reasoning (RTX 3090) |
|
||||
| `context.max_context_window` | 131072 (128K) | Reduced from 256K 2026-07-17 — stable 128K ceiling |
|
||||
| `context.max_context_window` | 131072 (128K) | Conservative `syslog-auto` pool floor (NVIDIA hosts 128K; Strix Halo 256K) |
|
||||
| `compression.threshold` | 0.65 | Rule 9: triggers at ~85K for a 128K window |
|
||||
| `compression.target_ratio` | 0.3 | Compresses to ~38K |
|
||||
| `compression.protect_last_n` | 40 | Preserves last 40 messages |
|
||||
|
||||
@@ -48,11 +48,11 @@ depends_on:
|
||||
|
||||
| Alias | GPU | Host | Model | VRAM | Ctx | tok/s | Role |
|
||||
|-------|-----|------|-------|------|-----|-------|------|
|
||||
| `strix-moe` | Strix Halo 64GB | ct15 (.15:8080) | qwen3.6-35B-udq4 | ~10/64GB (16%) | 128K | 62.9 | Compression, summarization, long docs |
|
||||
| `strix-moe` | Strix Halo 64GB | ct15 (.15:8080) | Carnice-Qwen3.6-MoE-35B-A3B-Q4_K_M.gguf | ~10/64GB (16%) | 256K | 62.9 | Compression, summarization, long docs |
|
||||
|
||||
Key notes:
|
||||
- All models use direct GPU routing via LiteLLM (`api_key: not-needed`). Router (port 9000) was decommissioned 2026-09-11 and is NOT in the inference path.
|
||||
- Stable aliases (gpu-dense, gpu-vision, strix-moe) from gpu-fleet are the canonical names for agent configs. Model-specific names still work but are deprecated. The retired names `gpu-light` and `gemma-4-12b` were superseded by `gpu-vision` on 2026-09-12 and no longer resolve (400 `Invalid model name`).
|
||||
- Stable aliases (gpu-dense, gpu-vision, strix-moe) from gpu-fleet are the canonical names for agent configs — use these, not model-specific names. The retired names `gpu-light` and `gemma-4-12b` were superseded by `gpu-vision` on 2026-09-12 and no longer resolve (400 `Invalid model name`).
|
||||
- The RTX 5070 is the fastest endpoint per token — `gpu-vision` is its canonical alias. Route vision/web/light work there first. The RTX 5070 model is multimodal (image+text).
|
||||
- Strix Halo is 62.9 tok/s (89% of 70.5 baseline) — below optimal but stable. Check for competing workloads.
|
||||
- RTX 3090 VRAM at 88% — within role-appropriate range (role = heavy reasoning, needs the headroom).
|
||||
@@ -142,10 +142,10 @@ Key notes:
|
||||
- **Escalate**: If Tier 2 triggers and temp still rising after 5 min → possible hardware failure
|
||||
|
||||
### Rule 9: Context Window Optimization
|
||||
- **Detect**: Benchmark tok/s vs baseline for each GPU at current context (all 128K)
|
||||
- **Detect**: Benchmark tok/s vs baseline for each GPU at its current context
|
||||
- RTX 3090 (128K ctx, ThinkingCap): baseline 74.8 tok/s — currently at 74.9 (100%)
|
||||
- RTX 5070 (128K ctx, HauhauCS QAT): baseline 165.2 tok/s — currently at 169.6 (103%)
|
||||
- Strix Halo (128K ctx, qwen3.6-35B-udq4): baseline 70.5 tok/s — currently at 62.9 (89%)
|
||||
- Strix Halo (256K ctx, Carnice-Qwen3.6-MoE-35B-A3B-Q4_K_M.gguf): baseline 70.5 tok/s — currently at 62.9 (89%)
|
||||
- **Fix**:
|
||||
- If tok/s > baseline → context has headroom, consider increasing
|
||||
- If tok/s < 90% baseline → reduce context by 25% and retest
|
||||
|
||||
@@ -5,12 +5,30 @@ version: 1.0.0
|
||||
description: >
|
||||
Canonical known-good baseline for all Syslog Hermes agents. Captures the exact
|
||||
configuration state, keys, workarounds, and audit procedure. When an agent's
|
||||
configuration goes sideways, restore from this baseline. Last verified 2026-07-16. All GPUs 128K context (reduced from 256K for stability Jul 2026) (RTX 3090 .8, RTX 5070 .110, Strix Halo .15). Parallel 1 fleet-wide (Strix Halo handles compression solo).
|
||||
configuration goes sideways, restore from this baseline. Last verified 2026-07-16. RTX 3090/5070 at 128K (reduced from 256K for stability Jul 2026); Strix Halo at 256K (2026-09-12) (RTX 3090 .8, RTX 5070 .110, Strix Halo .15). Parallel 1 fleet-wide (Strix Halo handles compression solo).
|
||||
author: Abiba (pi agent)
|
||||
---
|
||||
|
||||
# Hermes Agent Baseline — Canonical Good State
|
||||
|
||||
## Reachability Detection
|
||||
|
||||
Before checking agent baseline, verify the host is reachable and can be audited. Use the shared reachability helper from the clone root:
|
||||
|
||||
```bash
|
||||
# Run on each host to check reachability (Tanko, Mumuni, Koonimo, Koby)
|
||||
scripts/hermes-reachability-check.sh <host> "api_key:" "/root/.hermes/config.yaml"
|
||||
# Example: scripts/hermes-reachability-check.sh 192.168.68.122 "api_key:" "/root/.hermes/config.yaml"
|
||||
|
||||
# Expected outcomes:
|
||||
# - UNREACHABLE: SSH connection failed (host is down)
|
||||
# - VIOLATION: SSH succeeded and found matches (report the finding)
|
||||
# - COMPLIANT: SSH succeeded and found no matches (no api_key in config)
|
||||
#
|
||||
# NOTE: The bug this replaces was deriving reachability from the remote grep's exit code.
|
||||
# The correct pattern: remote side always succeeds (grep ...; true), so ssh status = connection only.
|
||||
```
|
||||
|
||||
## Quick Restore
|
||||
|
||||
```bash
|
||||
@@ -103,6 +121,26 @@ auxiliary:
|
||||
timeout: 120
|
||||
```
|
||||
|
||||
## Violation Classification
|
||||
|
||||
When reporting findings, separate POLICY observations from FAULT findings:
|
||||
|
||||
### POLICY (observation only, not a fault)
|
||||
- Agent uses a non-internal-harness provider (e.g., direct DeepSeek, Tencent, OpenRouter)
|
||||
- Config text has a field that looks unusual but the agent's calls are succeeding
|
||||
- Example: "POLICY: Koonimo uses deepseek directly; calls succeeding in last hour"
|
||||
|
||||
### FAULT (requires request-level evidence)
|
||||
- Agent's calls are failing with auth errors (401/403 in logs)
|
||||
- Agent's config has no valid API key AND calls are failing
|
||||
- Example: "FAULT: Koby's LiteLLM key expired; 401 observed at 2026-09-14 11:42:00"
|
||||
|
||||
### Rules
|
||||
1. Do NOT infer the runtime's credential resolution from config text alone.
|
||||
2. Require request-level evidence before calling something a FAULT: an observed auth failure in the agent's log, or the absence of successful calls in the window.
|
||||
3. If calls are succeeding, the correct output is "POLICY: uses <provider> directly; calls succeeding" - not a violation.
|
||||
4. State what you OBSERVED, not what the field implies.
|
||||
|
||||
## Known Bug: `api_key_env` Ignored by Auxiliary Client
|
||||
|
||||
**Bug location**: `agent/auxiliary_client.py` → `_resolve_task_provider_model()` (line ~5478)
|
||||
@@ -200,8 +238,7 @@ the registry before applying. Key is injected via `infisical run --` wrapper at
|
||||
{ "id": "syslog-auto" },
|
||||
{ "id": "strix-moe" },
|
||||
{ "id": "gpu-dense" },
|
||||
{ "id": "gpu-vision" },
|
||||
{ "id": "qwen3.6-27B-code" }
|
||||
{ "id": "gpu-vision" }
|
||||
]
|
||||
}
|
||||
}
|
||||
|
||||
@@ -19,6 +19,24 @@ description: >
|
||||
- agent_keys: map (see Agent Keys section)
|
||||
- infra_endpoints_verified: array
|
||||
|
||||
## Reachability Detection
|
||||
|
||||
Before auditing the config template, verify the host is reachable and can be checked. Use the shared reachability helper from the clone root:
|
||||
|
||||
```bash
|
||||
# Run on each host to check reachability (Tanko, Mumuni, Koonimo, Koby)
|
||||
scripts/hermes-reachability-check.sh <host> "base_url:" "/root/.hermes/config.yaml"
|
||||
# Example: scripts/hermes-reachability-check.sh 192.168.68.122 "base_url:" "/root/.hermes/config.yaml"
|
||||
|
||||
# Expected outcomes:
|
||||
# - UNREACHABLE: SSH connection failed (host is down)
|
||||
# - VIOLATION: SSH succeeded and found matches (report the finding)
|
||||
# - COMPLIANT: SSH succeeded and found no matches (no base_url in config)
|
||||
#
|
||||
# NOTE: The bug this replaces was deriving reachability from the remote grep's exit code.
|
||||
# The correct pattern: remote side always succeeds (grep ...; true), so ssh status = connection only.
|
||||
```
|
||||
|
||||
## Agent Keys (LiteLLM — Current 2026-07-11)
|
||||
|
||||
Each agent has a unique LiteLLM API key (virtual key) generated against the LiteLLM
|
||||
@@ -92,12 +110,12 @@ work immediately after restart.
|
||||
```yaml
|
||||
# ─── Model Selection ───
|
||||
model:
|
||||
default: <agent_model> # e.g., strix-moe, qwen3.6-27B-code, syslog-auto
|
||||
default: <agent_model> # e.g., strix-moe, gpu-dense, syslog-auto
|
||||
provider: harness
|
||||
base_url: http://192.168.68.116/litellm/v1 # Rule 5 (2026-08-09): canonical authenticated path; /v1 also OK
|
||||
api_key_env: LITELLM_API_KEY # Injected via infisical run -- wrapper
|
||||
max_tokens: 4096 # ⚠️ CRITICAL: Prevents unbounded generation
|
||||
context_length: 131072 # For syslog-auto (all GPUs at 128K for stability).
|
||||
context_length: 131072 # Conservative floor for syslog-auto (NVIDIA hosts 128K; Strix Halo 256K).
|
||||
# ⚠️ MANDATORY: Hermes probes unknown models from 256K
|
||||
# and falls back to 256K when /v1/models lacks a context
|
||||
# field (llama-server does). Without this override, agents
|
||||
@@ -133,7 +151,7 @@ compression:
|
||||
enabled: true
|
||||
model: syslog-auto # ⚠️ Must match auxiliary.compression.model. Stable alias (gpu-fleet § Stable Role-Based Aliases). NOT ornith-1.0-35b (LiteLLM does not serve that name).
|
||||
provider: harness
|
||||
max_context_window: 131072 # MUST match actual GPU capacity. All 3 GPUs are 128K (Jul 17).
|
||||
max_context_window: 131072 # MUST stay at the syslog-auto pool floor: NVIDIA hosts are 128K, Strix Halo 256K (2026-09-12).
|
||||
threshold: 0.65 # Fires at ~170K for 262K window, ~85K for 128K
|
||||
target_ratio: 0.30
|
||||
protect_last_n: 40
|
||||
@@ -149,7 +167,7 @@ compression:
|
||||
# Do NOT use syslog-auto for auxiliary tasks — it routes to the primary GPU.
|
||||
# gpu-vision = RTX 5070 (12B), freeing the Strix Halo for agent reasoning.
|
||||
# Heavy aux (delegation, x_search) use gpu-dense (RTX 3090) instead.
|
||||
# NEVER use raw model names (e.g. qwen3.6-27B-code, qwen3.6-35B-udq4; gemma-4-12b is retired
|
||||
# NEVER use retired model names (qwen3.6-27B-code, qwen3.6-35B-udq4; gemma-4-12b is retired
|
||||
# and no longer resolves) in agent configs — use the stable aliases so model swaps don't break agents.
|
||||
auxiliary:
|
||||
vision:
|
||||
@@ -173,10 +191,10 @@ auxiliary:
|
||||
timeout: 300 # gpu-fleet: 300s for large-history summarization (was 60)
|
||||
|
||||
# ─── Delegation / Heavy Aux (use gpu-dense = RTX 3090) ───
|
||||
# delegation.model and x_search.model use gpu-dense (NOT raw qwen3.6-27B-code).
|
||||
# delegation.model and x_search.model use gpu-dense (NOT retired raw name).
|
||||
|
||||
delegation:
|
||||
model: gpu-dense # stable alias for RTX 3090 (was raw qwen3.6-27B-code)
|
||||
model: gpu-dense # stable alias for RTX 3090
|
||||
provider: harness
|
||||
base_url: http://192.168.68.116/litellm/v1 # Rule 5 (2026-08-09): canonical authenticated; /v1 also OK
|
||||
api_key_env: LITELLM_API_KEY
|
||||
@@ -199,6 +217,26 @@ When LiteLLM keys are regenerated (e.g., after infrastructure changes):
|
||||
3. **After update**: Restart Hermes on the agent host
|
||||
4. **Verify**: `curl -H "Authorization: Bearer sk-<KEY>" http://192.168.68.116/v1/models`
|
||||
|
||||
## Violation Classification
|
||||
|
||||
When reporting findings, separate POLICY observations from FAULT findings:
|
||||
|
||||
### POLICY (observation only, not a fault)
|
||||
- Agent uses a non-internal-harness provider (e.g., direct DeepSeek, Tencent, OpenRouter)
|
||||
- Config text has a field that looks unusual but the agent's calls are succeeding
|
||||
- Example: "POLICY: Koonimo uses deepseek directly; calls succeeding in last hour"
|
||||
|
||||
### FAULT (requires request-level evidence)
|
||||
- Agent's calls are failing with auth errors (401/403 in logs)
|
||||
- Agent's config has no valid API key AND calls are failing
|
||||
- Example: "FAULT: Koby's LiteLLM key expired; 401 observed at 2026-09-14 11:42:00"
|
||||
|
||||
### Rules
|
||||
1. Do NOT infer the runtime's credential resolution from config text alone.
|
||||
2. Require request-level evidence before calling something a FAULT: an observed auth failure in the agent's log, or the absence of successful calls in the window.
|
||||
3. If calls are succeeding, the correct output is "POLICY: uses <provider> directly; calls succeeding" - not a violation.
|
||||
4. State what you OBSERVED, not what the field implies.
|
||||
|
||||
## Configuration Rules
|
||||
|
||||
### Rule 1: Shared Infra Is Locked
|
||||
@@ -253,7 +291,7 @@ The following MUST be identical across ALL profiles:
|
||||
|
||||
### Rule 7: Auxiliary Model Consistency (UPDATED 2026-07-16)
|
||||
- Vision and web_extract use `gpu-vision` (RTX 5070 — 12GB, vision-optimized)
|
||||
- Compression uses `syslog-auto` — the Strix Halo weighted pool (64GB, 128K ctx, compression-optimized); do NOT pin `compression.model` to `strix-moe` (audit Rule 7 rejects it)
|
||||
- Compression uses `syslog-auto` — the Strix Halo weighted pool (64GB, 256K ctx, compression-optimized); do NOT pin `compression.model` to `strix-moe` (audit Rule 7 rejects it)
|
||||
- **`ornith-1.0-35b` is NOT a valid compression model name** — LiteLLM does not serve it
|
||||
(do not restate the served model list here — CT 116 `/opt/inference-harness/litellm_config.yaml`
|
||||
is the single source of truth for models, aliases, weights and fallbacks). Old configs with `ornith-1.0-35b` cause 403/model-not-found on compression calls.
|
||||
@@ -267,15 +305,15 @@ The following MUST be identical across ALL profiles:
|
||||
- `api_key_env: LITELLM_API_KEY`
|
||||
- **Do NOT use `syslog-auto` for `vision`/`web_extract`** — it routes unpredictably; compression is the deliberate exception (see the OPERATIONAL DECISION above)
|
||||
- **Compression on Strix Halo**: The strix-moe alias routes to Strix Halo
|
||||
(64GB UMA, 128K context) — the designated compression GPU. This frees the
|
||||
(64GB UMA, 256K context) — the designated compression GPU. This frees the
|
||||
RTX 5070 for vision and web search, and the RTX 3090 for heavy reasoning.
|
||||
- The `compression:` block's `model` MUST match `auxiliary: compression: model`
|
||||
- The `compression: max_context_window: 131072` MUST match actual GPU capacity (128K)
|
||||
- The `compression: max_context_window: 131072` MUST stay at the syslog-auto pool floor (NVIDIA hosts 128K; Strix Halo 256K)
|
||||
|
||||
### Rule 8: GPU Workload Distribution (UPDATED 2026-07-16)
|
||||
- **RTX 3090 (24GB, 128K ctx, qwen3.6-27B-code)**: Heavy reasoning, code gen, long conversations
|
||||
- **RTX 3090 (24GB, 128K ctx, gpu-dense)**: Heavy reasoning, code gen, long conversations
|
||||
- **RTX 5070 (12GB, 128K ctx, gpu-vision)**: Vision, web search, quick tasks, web_extract (IQ4_NL+MTP, ~65% VRAM at 128K)
|
||||
- **Strix Halo (64GB, 128K ctx, syslog-auto)**: Context compression, summarization, long docs
|
||||
- **Strix Halo (64GB, 256K ctx, syslog-auto)**: Context compression, summarization, long docs
|
||||
- Agent profiles MUST route auxiliary tasks to the correct GPU:
|
||||
- `auxiliary.vision.model: gpu-vision` (RTX 5070)
|
||||
- `auxiliary.web_extract.model: gpu-vision` (RTX 5070)
|
||||
@@ -284,14 +322,14 @@ The following MUST be identical across ALL profiles:
|
||||
- For 128K context window: `threshold: 0.65` (fires at ~85K tokens)
|
||||
- Do NOT use `threshold: 0.25` — this fires at 65K, causing premature context loss
|
||||
- Do NOT use `threshold: 0.80` — this delays until ~105K, leaving only 23K margin
|
||||
- `max_context_window: 131072` MUST match the model's actual capacity (128K)
|
||||
- `max_context_window: 131072` MUST stay at the pool floor (NVIDIA hosts 128K; Strix Halo 256K)
|
||||
- See `devops-hermes-compression` skill for full reference
|
||||
|
||||
### Rule 9: Compression Threshold for 128K Models
|
||||
- For 128K context window: `threshold: 0.65` (fires at ~85K tokens)
|
||||
- Do NOT use `threshold: 0.25` — this fires at 65K, causing premature context loss
|
||||
- Do NOT use `threshold: 0.80` — this delays until ~105K, leaving only 23K margin
|
||||
- `max_context_window: 131072` MUST match the model's actual capacity (all GPUs = 128K)
|
||||
- `max_context_window: 131072` MUST stay at the pool floor (NVIDIA hosts 128K; Strix Halo 256K)
|
||||
- See `devops-hermes-compression` skill for full reference
|
||||
|
||||
### Rule 10: Default Model Must Be `syslog-auto` (All Agents)
|
||||
@@ -322,7 +360,8 @@ When an agent shows "context issues" (premature compression, 401s, 504s, DeepSee
|
||||
verify ALL FOUR of these against the live config. They are the only root causes found in production:
|
||||
|
||||
1. **max_context_window correct?** — BOTH `compression.max_context_window` AND `context.max_context_window`
|
||||
MUST be `131072` (all GPUs are 128K). A value of `262144` causes instability near 100K and must NOT be used.
|
||||
MUST be `131072` (the syslog-auto pool floor: NVIDIA hosts are 128K; Strix Halo is 256K).
|
||||
A `262144` client window can route to a 128K NVIDIA host and fail, so it must NOT be used.
|
||||
~83K instead of ~170K. Check: `grep -n max_context_window ~/.hermes/config.yaml`
|
||||
2. **base_url uses authenticated path?** — `custom_providers[0].base_url`, `delegation.base_url`,
|
||||
and ALL `auxiliary.*.base_url` MUST be `http://192.168.68.116/litellm/v1` (Rule 5, canonical)
|
||||
|
||||
@@ -117,6 +117,44 @@ model:
|
||||
api_key_env: LITELLM_API_KEY
|
||||
```
|
||||
|
||||
## Reachability Detection
|
||||
|
||||
Before checking for hardcoded keys, verify the host is reachable and can be audited. Use the shared reachability helper from the clone root:
|
||||
|
||||
```bash
|
||||
# Run on each host to check reachability (Tanko, Mumuni, Koonimo, Koby)
|
||||
scripts/hermes-reachability-check.sh <host> "api_key: sk-" "/root/.hermes/"
|
||||
# Example: scripts/hermes-reachability-check.sh 192.168.68.122 "api_key: sk-" "/root/.hermes/"
|
||||
|
||||
# Expected outcomes:
|
||||
# - UNREACHABLE: SSH connection failed (host is down)
|
||||
# - VIOLATION: SSH succeeded and found matches (report the finding)
|
||||
# - COMPLIANT: SSH succeeded and found no matches (no hardcoded keys in config)
|
||||
#
|
||||
# NOTE: The bug this replaces was deriving reachability from the remote grep's exit code.
|
||||
# The correct pattern: remote side always succeeds (grep ...; true), so ssh status = connection only.
|
||||
```
|
||||
|
||||
## Violation Classification
|
||||
|
||||
When reporting findings, separate POLICY observations from FAULT findings:
|
||||
|
||||
### POLICY (observation only, not a fault)
|
||||
- Agent uses a non-internal-harness provider (e.g., direct DeepSeek, Tencent, OpenRouter)
|
||||
- Config text has a field that looks unusual but the agent's calls are succeeding
|
||||
- Example: "POLICY: Koonimo uses deepseek directly; calls succeeding in last hour"
|
||||
|
||||
### FAULT (requires request-level evidence)
|
||||
- Agent's calls are failing with auth errors (401/403 in logs)
|
||||
- Agent's config has no valid API key AND calls are failing
|
||||
- Example: "FAULT: Koby's LiteLLM key expired; 401 observed at 2026-09-14 11:42:00"
|
||||
|
||||
### Rules
|
||||
1. Do NOT infer the runtime's credential resolution from config text alone.
|
||||
2. Require request-level evidence before calling something a FAULT: an observed auth failure in the agent's log, or the absence of successful calls in the window.
|
||||
3. If calls are succeeding, the correct output is "POLICY: uses <provider> directly; calls succeeding" - not a violation.
|
||||
4. State what you OBSERVED, not what the field implies.
|
||||
|
||||
## Detection Query
|
||||
|
||||
Run on any Hermes host to detect violations:
|
||||
@@ -169,9 +207,11 @@ The agent picks up the new key via `infisical run --` at gateway startup.
|
||||
|
||||
**Keys are permanent and use bare agent name aliases.**
|
||||
|
||||
- **Duration**: `null` — keys never expire. NOT enforced today: CT 116 `litellm_config.yaml` has no `default_key_generate_params` block, and a key generated with no explicit models comes back with an empty models list. OPEN policy question: should agent keys expire by default? (captain security-policy decision, raised separately.)
|
||||
- **Duration**: `null` — keys never expire by default. **Expiry must be set EXPLICITLY at creation** with the `duration` parameter (e.g., `90d` for 90 days). The 90-day default is the standard; however, the config default is **NOT honoured** by LiteLLM 1.99.1 (verified on CT 116: a key generated with no explicit duration returns `expires=null`). This has been recorded in `/opt/inference-harness/litellm_config.yaml` to prevent re-filing as a bug.
|
||||
- **Daily Audit**: A daily audit job runs at 00:00 UTC (`/usr/local/bin/litellm-key-renewal-ct116.sh`, cron 00:00). It is **AUDIT-ONLY** and does not perform renewal. It lists every key, reports those with no expiry and those inside a 14-day warning window, explicitly EXCLUDES `abiba-pi` and `koby` (report-only, and .129 must never be touched), and logs `RENEWAL-REQUIRED-BUT-NOT-PERFORMED + NO KEY WAS CHANGED` when renewal is skipped. **Renewal is NOT implemented** — keys must not be rotated until delivery (vault injection + consumer verification) exists and is proven end-to-end.
|
||||
- **Exclusions**: `abiba-pi` and every firstmate/secondmate/crewmate key stay **WITHOUT an expiry** until a proven renewal path exists. `koby` is **report-only** (never touched). These exclusions are enforced by the audit job.
|
||||
- **Alias convention**: bare agent name only (e.g., `tanko`, `mumuni`, `koby`, `koonimo`). No dates, no versions. The alias IS the identity.
|
||||
- **Rotation triggers**: compromise, personnel departure, or quarterly security hygiene. NOT calendar-driven.
|
||||
- **Rotation triggers**: compromise, personnel departure, or quarterly security hygiene. NOT calendar-driven. Manual rotation is permitted only when the renewal delivery path is proven and verified on a throwaway consumer before production use.
|
||||
- **Max budget**: $100 per key (config default).
|
||||
|
||||
```yaml
|
||||
@@ -182,7 +222,7 @@ The agent picks up the new key via `infisical run --` at gateway startup.
|
||||
# re-verify before applying.
|
||||
litellm_settings:
|
||||
default_key_generate_params:
|
||||
models: ["syslog-auto", "qwen3.6-27B-code", "gpu-vision"]
|
||||
models: ["syslog-auto", "gpu-dense", "gpu-vision", "strix-moe"]
|
||||
duration: null # ← permanent
|
||||
max_budget: 100
|
||||
metadata:
|
||||
|
||||
@@ -4,7 +4,7 @@ kind: responsibility
|
||||
description: >
|
||||
Optimizes the full Syslog inference stack — LiteLLM routing weights, GPU model
|
||||
assignments, agent context management, and prompt caching — to reduce response
|
||||
times to sub-15s average. All GPUs now at 128K context (stable ceiling).
|
||||
times to sub-15s average. NVIDIA GPUs at 128K context; Strix Halo at 256K (2026-09-12).
|
||||
id: 067NC6KP02RG60S50M40E30928
|
||||
---
|
||||
|
||||
@@ -63,7 +63,7 @@ prefill time at 532 tok/s. Fix context first, routing second.
|
||||
- **Cache everything repeated**: system prompts, skill docs, AGENTS.md — these
|
||||
never change between turns. Single-digit cache hit rate is unacceptable.
|
||||
- **Lower context ceiling**: 128K window is the stable ceiling for agent conversations.
|
||||
GPUs reduced from 256K to 128K (2026-07-17). 128K window should compact at 85K (0.65 threshold). For larger contexts, route to external providers.
|
||||
GPUs reduced from 256K to 128K (2026-07-17) for the NVIDIA hosts; Strix Halo runs 256K (2026-09-12). 128K window should compact at 85K (0.65 threshold). For larger contexts, route to external providers.
|
||||
|
||||
### Shape
|
||||
|
||||
@@ -101,5 +101,5 @@ call enable-prompt-caching
|
||||
|
||||
call verify-latency
|
||||
host: 192.168.68.116
|
||||
models: [syslog-auto, qwen3.6-27B-code, gpu-vision]
|
||||
models: [syslog-auto, gpu-dense, gpu-vision, strix-moe]
|
||||
```
|
||||
|
||||
@@ -223,7 +223,7 @@ description: >
|
||||
**Prometheus targets**:
|
||||
- 192.168.68.8:9400 (RTX 3090 — qwen)
|
||||
- 192.168.68.110:9400 (RTX 5070 — gpu-vision)
|
||||
- 192.168.68.15:9400 (Strix Halo — qwen3.6-35B-udq4)
|
||||
- 192.168.68.15:9400 (Strix Halo — strix-moe)
|
||||
- harness-litellm:4000 (LiteLLM health)
|
||||
|
||||
### Ecosystem C: Netbird (72.61.0.17 — Hostinger srv1079750.hstgr.cloud)
|
||||
|
||||
@@ -17,7 +17,7 @@ description: >
|
||||
⚠️ This contract is target-state aspirational — but GPU export + alerting
|
||||
are now as-built (verified 2026-08-09).
|
||||
As-built GPU monitoring is via gpu-monitor contract (port 9100 poll).
|
||||
version: 1.0.0
|
||||
version: 1.0.1
|
||||
---
|
||||
|
||||
## Architecture
|
||||
@@ -120,127 +120,156 @@ the any-HTTP rule. On those — the authenticated Zulip POST and the router
|
||||
`/health` — an unexpected status (`401`/`403` from a bad or missing credential,
|
||||
`5xx`, or anything other than the expected `200`) is an **ALERT**, not "alive".
|
||||
|
||||
**STANDING PROBE RULES (2026-09-14, from defect report 1150.msg):**
|
||||
1. **Any HTTP status means ALIVE.** 200, 301, 302, 401, 403, 404 all prove the
|
||||
service answered — report the code, never "down". A redirect is not a failure.
|
||||
Only a failed CONNECTION (curl status 000, timeout, refused) is a failed probe,
|
||||
and that is a statement about YOUR PROBE, not about the service.
|
||||
2. **A failed probe is never a service verdict.** Print
|
||||
`probe-failed: <target> <kind>` naming the exact URL/host/port and the failure
|
||||
kind (timeout, refused, no-route, dns), retry once at a longer timeout, and only
|
||||
then report. Apply the same shape as scripts/disk-gc-scan.py.
|
||||
3. **Say which probe produced each number.** "Grafana: 000" is unusable;
|
||||
"Grafana http://192.168.68.116:3001/api/health -> connection timeout after 10s
|
||||
(retried at 25s: also timeout)" is actionable.
|
||||
|
||||
### check-health
|
||||
|
||||
**RUN LIVE, NEVER ECHO — every dispatch must execute the probes below with real tool calls; never repeat a prior report unless a live probe fails.**
|
||||
**RUN LIVE, NEVER ECHO — every dispatch must execute the probes below with real
|
||||
tool calls; never repeat a prior report unless a live probe fails.**
|
||||
|
||||
**PROBE SHAPE (per standing rules above):**
|
||||
- Every probe prints the target name + URL + HTTP code (or failure kind)
|
||||
- Retry once on connection failure at longer timeout
|
||||
- Any HTTP status = ALIVE; only 000/timeout/refused = probe-failed
|
||||
- Report the actual probe command and its result, not a summary verdict
|
||||
|
||||
```bash
|
||||
# Provenance — run first; paste the absolute path into the report
|
||||
pwd -P
|
||||
|
||||
# Zulip API health (POST ping)
|
||||
source /etc/litellm-monitor.env
|
||||
ZULIP_USER="abiba-bot@chat.sysloggh.net"
|
||||
curl -s -o /dev/null -w '%{http_code}' -X POST https://chat.sysloggh.net/api/v1/messages -u "${ZULIP_USER}:${ZULIP_BOT_KEY}"
|
||||
# Expected: 200 (HTTP 000 = unreachable/cache)
|
||||
# ============================================================
|
||||
# 1. ZULIP API HEALTH (POST ping) — bare-200 probe
|
||||
# ============================================================
|
||||
# NOTE: /etc/litellm-monitor.env exists only on CT 116, retrieve keys from CT 116 via:
|
||||
zulip_key=$(ssh root@192.168.68.116 "grep ZULIP_BOT_KEY /etc/litellm-monitor.env | cut -d= -f2")
|
||||
if [ -z "$zulip_key" ]; then
|
||||
echo "credential-missing: ZULIP_BOT_KEY not found in /etc/litellm-monitor.env"
|
||||
else
|
||||
ZULIP_USER="abiba-bot@chat.sysloggh.net"
|
||||
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 -X POST https://chat.sysloggh.net/api/v1/messages -u "${ZULIP_USER}:${zulip_key}")
|
||||
echo "Zulip API https://chat.sysloggh.net/api/v1/messages -> $code"
|
||||
# Expected: 200 (bare-200 probe; any other status is an ALERT)
|
||||
fi
|
||||
|
||||
# PM2 process health
|
||||
# ============================================================
|
||||
# 2. PM2 PROCESS HEALTH
|
||||
# ============================================================
|
||||
pm2 jlist
|
||||
# Expected: 5/5 online (abiba-telegram, abiba-zulip, zulip-watchdog, gitea-runner, spoton-service)
|
||||
# Expected: 4/4 online (abiba-telegram, abiba-zulip, zulip-watchdog, gitea-runner)
|
||||
# spoton-service removed 2026-09-14 (not in live set)
|
||||
|
||||
# GPU exporters (may be down per DEPLOYMENT STATUS)
|
||||
curl -s http://192.168.68.8:9400/metrics && echo " - OK" || echo " - FAIL"
|
||||
curl -s http://192.168.68.110:9400/metrics && echo " - OK" || echo " - FAIL"
|
||||
curl -s http://192.168.68.15:9400/metrics && echo " - OK" || echo " - FAIL"
|
||||
# ============================================================
|
||||
# 3. GPU EXPORTERS — any-HTTP probe (metrics endpoint)
|
||||
# ============================================================
|
||||
# Probe /metrics (the Prometheus scrape target), not bare /
|
||||
for host in 192.168.68.8 192.168.68.110 192.168.68.15; do
|
||||
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 "http://$host:9400/metrics")
|
||||
if [ "$code" == "000" ]; then
|
||||
# Retry with longer timeout
|
||||
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 25 "http://$host:9400/metrics")
|
||||
echo "GPU exporter http://$host:9400/metrics -> probe-failed: timeout (retried at 25s: still $code)"
|
||||
else
|
||||
echo "GPU exporter http://$host:9400/metrics -> $code"
|
||||
fi
|
||||
done
|
||||
# Expected: 200 on all 3 hosts (RTX 3090, RTX 5070, Strix Halo)
|
||||
|
||||
# Router health (via nginx on port 80)
|
||||
curl -s -o /dev/null -w '%{http_code}' http://192.168.68.116/health
|
||||
# Expected: 200 (Router is up and responding)
|
||||
# ============================================================
|
||||
# 4. ROUTER HEALTH (via nginx on port 80) — bare-200 probe
|
||||
# ============================================================
|
||||
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 http://192.168.68.116/health)
|
||||
if [ "$code" == "000" ]; then
|
||||
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 25 http://192.168.68.116/health)
|
||||
echo "Router http://192.168.68.116/health -> probe-failed: timeout (retried at 25s: still $code)"
|
||||
else
|
||||
echo "Router http://192.168.68.116/health -> $code"
|
||||
fi
|
||||
# Expected: 200 (bare-200 probe; any other status is an ALERT)
|
||||
|
||||
# LiteLLM health (via nginx on port 80)
|
||||
curl -s -o /dev/null -w '%{http_code}' http://192.168.68.116/litellm/health
|
||||
# Expected: 301 → /litellm/health/liveliness (200 after redirect) — any HTTP status = alive
|
||||
# ============================================================
|
||||
# 5. LITELLM HEALTH (via nginx on port 80) — any-HTTP probe
|
||||
# ============================================================
|
||||
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 http://192.168.68.116/litellm/health)
|
||||
if [ "$code" == "000" ]; then
|
||||
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 25 http://192.168.68.116/litellm/health)
|
||||
echo "LiteLLM http://192.168.68.116/litellm/health -> probe-failed: timeout (retried at 25s: still $code)"
|
||||
else
|
||||
echo "LiteLLM http://192.168.68.116/litellm/health -> $code"
|
||||
fi
|
||||
# Expected: 301 → /litellm/health/liveliness (any HTTP status = ALIVE)
|
||||
|
||||
# PVE API liveness — probe the REAL PVE nodes on :8006, never the monitoring
|
||||
# host CT 116. CT 116 runs no pveproxy, so probing it on :8006 returns 000 —
|
||||
# that was the stale-vantage bug this replaces (CT 116 is the monitoring host,
|
||||
# not a cluster node). Unauthenticated GET answers 401 while the API is ALIVE
|
||||
# by design. Alive = ANY HTTP status (401 is the EXPECTED healthy response);
|
||||
# DOWN = connection refused (000) or timeout only.
|
||||
# ============================================================
|
||||
# 6. PVE API LIVENESS — any-HTTP probe (auth-gated)
|
||||
# ============================================================
|
||||
# Probe the REAL PVE nodes on :8006, never the monitoring host CT 116.
|
||||
for node in 192.168.68.9 192.168.68.5 192.168.68.15 192.168.68.6 192.168.68.12; do
|
||||
printf '%s:8006 -> %s\n' "$node" \
|
||||
"$(curl -sk -o /dev/null -w '%{http_code}' --connect-timeout 5 "https://$node:8006/api2/json/version")"
|
||||
code=$(curl -sk -o /dev/null -w '%{http_code}' --connect-timeout 10 "https://$node:8006/api2/json/version")
|
||||
if [ "$code" == "000" ]; then
|
||||
code=$(curl -sk -o /dev/null -w '%{http_code}' --connect-timeout 25 "https://$node:8006/api2/json/version")
|
||||
echo "PVE API https://$node:8006/api2/json/version -> probe-failed: timeout (retried at 25s: still $code)"
|
||||
else
|
||||
echo "PVE API https://$node:8006/api2/json/version -> $code"
|
||||
fi
|
||||
done
|
||||
# Expected: 401 on every node (acerpve .9, ocupve .5, amdpve .15, storepve .6, minipve .12)
|
||||
# A node answering 000/timeout is DOWN — flag that node. 401 is NOT a fault.
|
||||
# 401 is the EXPECTED healthy response (auth-gated); 000/timeout = DOWN
|
||||
|
||||
# Prometheus targets
|
||||
curl -s http://192.168.68.116:9090/api/v1/targets | jq '.data.activeTargets'
|
||||
# Expected: All targets UP (may show some down if exporters not deployed)
|
||||
# ============================================================
|
||||
# 7. PROMETHEUS TARGETS — bare-200 probe
|
||||
# ============================================================
|
||||
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 http://192.168.68.116:9090/api/v1/targets)
|
||||
if [ "$code" == "000" ]; then
|
||||
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 25 http://192.168.68.116:9090/api/v1/targets)
|
||||
echo "Prometheus http://192.168.68.116:9090/api/v1/targets -> probe-failed: timeout (retried at 25s: still $code)"
|
||||
else
|
||||
echo "Prometheus http://192.168.68.116:9090/api/v1/targets -> $code"
|
||||
fi
|
||||
# Expected: 200 (bare-200 probe; any other status is an ALERT)
|
||||
|
||||
# Grafana health
|
||||
curl -s http://192.168.68.116:3001/api/health | jq '{status, version}'
|
||||
# Expected: {"status":"ok","version":"..."}
|
||||
# ============================================================
|
||||
# 8. GRAFANA HEALTH — any-HTTP probe
|
||||
# ============================================================
|
||||
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 http://192.168.68.116:3001/api/health)
|
||||
if [ "$code" == "000" ]; then
|
||||
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 25 http://192.168.68.116:3001/api/health)
|
||||
echo "Grafana http://192.168.68.116:3001/api/health -> probe-failed: timeout (retried at 25s: still $code)"
|
||||
else
|
||||
echo "Grafana http://192.168.68.116:3001/api/health -> $code"
|
||||
fi
|
||||
# Expected: 200 (any HTTP status = ALIVE; 000/timeout = DOWN)
|
||||
|
||||
# LiteLLM metrics (Prometheus endpoint)
|
||||
curl -s http://192.168.68.116:4000/metrics | head -20
|
||||
# Expected: Prometheus-formatted metrics output
|
||||
# ============================================================
|
||||
# 9. LITELLM METRICS (Prometheus endpoint) — any-HTTP probe
|
||||
# ============================================================
|
||||
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 http://192.168.68.116:4000/metrics)
|
||||
if [ "$code" == "000" ]; then
|
||||
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 25 http://192.168.68.116:4000/metrics)
|
||||
echo "LiteLLM metrics http://192.168.68.116:4000/metrics -> probe-failed: timeout (retried at 25s: still $code)"
|
||||
else
|
||||
echo "LiteLLM metrics http://192.168.68.116:4000/metrics -> $code"
|
||||
fi
|
||||
# Expected: 200 (any HTTP status = ALIVE; 000/timeout = DOWN)
|
||||
```
|
||||
|
||||
**Report format**: Begin every report with the **absolute path the probe executed
|
||||
from** (`pwd -P`, or the script's absolute path) so a stale-consumer report is
|
||||
distinguishable from a real fault at read time. Summarize actual results from
|
||||
each probe. Apply the any-HTTP-response liveness rule ONLY to the auth-gated PVE
|
||||
API and LiteLLM endpoints above: only connection-refused (`000`) or timeout is
|
||||
DOWN; empty output is a warning. For probes whose expected result is a bare `200`
|
||||
(the authenticated Zulip POST, router `/health`), flag an alert on any unexpected
|
||||
status (`401`/`403`/`5xx`) — do not summarize it as alive. A bare-`200`
|
||||
expectation on the auth-gated PVE API (`401`) or LiteLLM health (`301` redirect)
|
||||
is a stale expectation, not a fault.
|
||||
|
||||
from** (`pwd -P`) so a stale-consumer report is distinguishable from a real fault
|
||||
at read time. For each probe, print the target name, the full URL, and the HTTP
|
||||
code (or failure kind with retry details). Apply the standing probe rules: any
|
||||
HTTP status = ALIVE; only 000/timeout/refused = probe-failed. A redirect is not
|
||||
a failure.
|
||||
|
||||
### Phase 1: GPU Exporters
|
||||
|
||||
**NVIDIA (.8 and .110)**:
|
||||
1. Download `nvidia_gpu_exporter` binary
|
||||
2. Create systemd service `nvidia-gpu-exporter.service`
|
||||
3. Start and enable
|
||||
|
||||
**AMD (.15)**:
|
||||
1. Create Python exporter script at `/opt/amdgpu-exporter/exporter.py`
|
||||
2. Parses `amdgpu_top --json -d 1000` output
|
||||
3. Exposes key metrics at `:9400/metrics` via Python http.server
|
||||
4. Create systemd service
|
||||
5. Start and enable
|
||||
|
||||
### Phase 2: Prometheus
|
||||
|
||||
1. Create `/opt/monitoring/` directory on CT 116
|
||||
2. Write `prometheus.yml` with scrape configs for all targets
|
||||
3. Add to docker-compose (or separate compose file)
|
||||
4. Start container
|
||||
|
||||
### Phase 3: Grafana
|
||||
|
||||
1. Create `/opt/monitoring/grafana/` directories
|
||||
2. Provision Prometheus datasource
|
||||
3. Provision GPU fleet dashboard JSON
|
||||
4. Provision LiteLLM dashboard JSON
|
||||
5. Add to docker-compose
|
||||
6. Start container
|
||||
|
||||
### Phase 4: Verification
|
||||
|
||||
1. Verify all 3 GPU exporters return 200 at :9400/metrics
|
||||
2. Verify Prometheus targets all UP at :9090/targets
|
||||
3. Verify Grafana accessible at :3001 with dashboards
|
||||
4. Verify LiteLLM metrics flowing to Prometheus
|
||||
5. ~~Update nginx to proxy `/monitoring/` → Grafana~~ (NOT recommended — nginx sub-path was tried for /grafana/ and reverted per proxmox-monitor; direct :3001 access is the standard)
|
||||
|
||||
## Verification Commands
|
||||
|
||||
```bash
|
||||
# GPU exporters
|
||||
curl -s http://192.168.68.8:9400/metrics | grep nvidia
|
||||
curl -s http://192.168.68.110:9400/metrics | grep nvidia
|
||||
curl -s http://192.168.68.15:9400/metrics | grep amdgpu
|
||||
|
||||
# Prometheus
|
||||
curl -s http://192.168.68.116:9090/api/v1/targets
|
||||
|
||||
# Grafana
|
||||
curl -s http://192.168.68.116:3001/api/health
|
||||
|
||||
# LiteLLM metrics (already live)
|
||||
curl -s http://192.168.68.116:4000/metrics | head -20
|
||||
```
|
||||
|
||||
@@ -320,7 +320,15 @@ directly call OpenRouter via Python's requests library. Converting would require
|
||||
|
||||
## LiteLLM Master Key (use sparingly — agents should NOT use it directly)
|
||||
|
||||
- Master key: `sk-litellm-7f96080dd99b15c36bd4b333b58a6796` (in /opt/inference-harness/.env on CT116, Infisical project=infrastructure env=production secret=LITELLM_MASTER_KEY)
|
||||
- Master key: **Retrieval path (do not trust a literal value in this file — the key rotates)**:
|
||||
```bash
|
||||
# Read at runtime from the container's environment:
|
||||
docker exec harness-litellm printenv LITELLM_MASTER_KEY
|
||||
# Or from Infisical vault (project=infrastructure env=prod) - NOTE: --plain is broken on CLI 0.43.110 (prints nothing):
|
||||
infisical secrets get LITELLM_MASTER_KEY --project=infrastructure --env=production | awk '$1=="LITELLM_MASTER_KEY"{print $NF}'
|
||||
# Prove a key is live with a 200 from /key/list on the CT 116 host (the container has no curl):
|
||||
curl -s -H "Authorization: Bearer <key>" http://127.0.0.1:4000/key/list | jq length
|
||||
```
|
||||
- Used for /key/generate, /key/delete, /key/list (GET), DB queries
|
||||
- **Known violation (RESOLVED 2026-07-16):** Abiba's LITELLM_API_KEY was previously the master key.
|
||||
It is now a dedicated agent key `sk-sxbphLvk1OU…` (vault secret `ABIBA_LITELLM_API_KEY`, alias `abiba-pi`).
|
||||
|
||||
@@ -26,9 +26,8 @@ description: >
|
||||
| Model | avg latency | avg TTFT | p-profile (24h) |
|
||||
|---|---|---|---|
|
||||
| syslog-auto | 28.8s | 25.4s | 68 calls took 30-120s; tail to ~300s under load |
|
||||
| qwen3.6-27B-code | 23.0s | — | same backend class as syslog-auto |
|
||||
| strix-moe | 7.5s | — | Strix Halo, healthy |
|
||||
| gemma-4-12b (retired 2026-09-12; RTX 5070 now `gpu-vision`) | 2.6s | — | RTX 5070, healthy |
|
||||
| gpu-vision (retired gemma-4-12b, RTX 5070) | 2.6s | — | RTX 5070, healthy |
|
||||
|
||||
Sep 6 incident timeline: failures 04:00-07:00 EDT (0% GPU util = wedged
|
||||
backend), full recovery 07:00-08:00 with ZERO client failures once requests
|
||||
@@ -56,9 +55,8 @@ proxy queuing.
|
||||
- vision: 60s (keep), web_extract: 30s (keep) — the 2.6s average was measured on `gemma-4-12b` (retired 2026-09-12); the live RTX 5070 alias is `gpu-vision`.
|
||||
- compression: 300s (keep — this was already raised from 60 per gpu-fleet).
|
||||
- **gpu-dense delegation/x_search: set timeout >= 120s.** The RTX 3090
|
||||
(qwen3.6-27B-code backend, 23.0s avg) is the same speed class as
|
||||
syslog-auto; delegation defaults that assume fast responses will 408 the
|
||||
same way.
|
||||
(gpu-dense backend) is the same speed class as syslog-auto; delegation
|
||||
defaults that assume fast responses will 408 the same way.
|
||||
|
||||
### 3. Retry policy — backoff, not repetition
|
||||
|
||||
|
||||
+42
-7
@@ -19,7 +19,7 @@ description: >
|
||||
Scraped by Prometheus with Bearer master key; endpoint returns 307 → /metrics/.
|
||||
- Alertmanager (harness-alertmanager :9093) + zulip-bridge (:9102) deliver
|
||||
firing alerts to #agent-hub > alerts-infra via abiba-bot. Added 2026-08-09.
|
||||
- Prometheus node job covers ALL 5 PVE nodes (.5/.6/.9/.12/.15:9100).
|
||||
- Prometheus node job covers ALL 6 PVE nodes (.4/.5/.6/.9/.12/.15:9100).
|
||||
---
|
||||
|
||||
## Architecture (v4.0.0 — Direct: nginx → LiteLLM → GPU)
|
||||
@@ -44,7 +44,7 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2)
|
||||
**What changed (v3.2.0 → v4.0.0 — 2026-07-08)**:
|
||||
- Router REMOVED from request path — LiteLLM proxies directly to GPU
|
||||
- All GPUs at parallel 2 (was parallel 1)
|
||||
- NVIDIA context reduced 256K→128K to free VRAM — now the stable ceiling across all GPUs (2026-07-17)
|
||||
- NVIDIA context reduced 256K→128K to free VRAM — the stable NVIDIA ceiling (2026-07-17); Strix Halo runs 256K (2026-09-12)
|
||||
- Timeouts and fallback chains are config state — read them from CT 116 `/opt/inference-harness/litellm_config.yaml`; they are not duplicated here.
|
||||
|
||||
## Parameters
|
||||
@@ -69,7 +69,16 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2)
|
||||
- Network access to public_url, auth_host, and gpu_dashboard_url
|
||||
- LiteLLM master key for admin endpoints only (`/key/list`, `/key/generate`, `/key/info`)
|
||||
- The dedicated `monitor` agent key on CT 116 at `/etc/litellm-monitor.env` (root-only 0600) for
|
||||
model inference checks — the master key must never be used for inference
|
||||
model inference checks, scoped for every alias step 7 probes (`gpu-dense`, `gpu-vision`,
|
||||
`strix-moe`, `syslog-auto`). Retrieve from the executor's host via:
|
||||
|
||||
```
|
||||
monitor key: ssh root@192.168.68.116 "grep LITELLM_MONITOR_KEY /etc/litellm-monitor.env | cut -d= -f2"
|
||||
master key: ssh root@192.168.68.116 "docker exec harness-litellm printenv LITELLM_MASTER_KEY"
|
||||
```
|
||||
|
||||
If credentials are missing or unreadable, the probe must report `credential-missing` (not bare 401 or "0 keys").
|
||||
The master key must never be used for inference.
|
||||
|
||||
## GPU Fleet Topology
|
||||
|
||||
@@ -144,12 +153,19 @@ contracts — read them there. Do not re-add retired names (`gemma-4-12b`, `gpu-
|
||||
7. **Check model inference via LiteLLM** — Test one model on each GPU host. The health
|
||||
check runs on the **backend edge**, not the public edge, so these paths carry the
|
||||
`/litellm/` prefix:
|
||||
- POST http://{{backend_host}}/litellm/v1/chat/completions model=qwen3.6-27B-code → expect 200 (RTX 3090, .8)
|
||||
- POST http://{{backend_host}}/litellm/v1/chat/completions model=gpu-dense → expect 200 (RTX 3090, .8)
|
||||
- POST http://{{backend_host}}/litellm/v1/chat/completions model=gpu-vision → expect 200 (RTX 5070, .110)
|
||||
- POST http://{{backend_host}}/litellm/v1/chat/completions model=strix-moe → expect 200 (Strix Halo, .15)
|
||||
- Auth uses the dedicated `monitor` agent key, read on CT 116 from
|
||||
`/etc/litellm-monitor.env` (root-only 0600). Do NOT use the master key for inference —
|
||||
the master key is for admin endpoints only (`/key/list`, `/key/generate`, `/key/info`).
|
||||
- Auth uses the dedicated `monitor` agent key. Retrieve via:
|
||||
`ssh root@192.168.68.116 "grep LITELLM_MONITOR_KEY /etc/litellm-monitor.env | cut -d= -f2"`
|
||||
Do NOT use the master key for inference — the master key is for admin endpoints only
|
||||
(`/key/list`, `/key/generate`, `/key/info`). Retrieve master key via:
|
||||
`ssh root@192.168.68.116 "docker exec harness-litellm printenv LITELLM_MASTER_KEY"`
|
||||
- KEY SCOPE: the `monitor` key MUST be scoped for the three probed aliases (`gpu-dense`,
|
||||
`gpu-vision`, `strix-moe`) plus the `syslog-auto` fallback pool, otherwise the probe
|
||||
returns 403 and the host is not covered.
|
||||
If a probe returns 403, widen the monitor key's model list on CT 116 (add the missing
|
||||
alias) and re-run — never drop the host from the probe to make the check pass.
|
||||
- `gemma-4-12b` was retired and returns 400 `Invalid model name` — do not re-add it to
|
||||
this list. The RTX 5070 host now serves `gpu-vision`.
|
||||
- `/v1/models` is **key-scoped**: a model is only visible to keys allowed to use it, so
|
||||
@@ -161,8 +177,27 @@ contracts — read them there. Do not re-add retired names (`gemma-4-12b`, `gpu-
|
||||
|
||||
8. **Check agent keys**:
|
||||
- GET http://{{backend_host}}/litellm/key/list with master key (admin endpoint) → verify all 6 agents have keys
|
||||
- **IMPORTANT**: Run the curl on the CT 116 HOST, not inside the container. The `harness-litellm` container has no curl/wget. Use:
|
||||
`ssh root@192.168.68.116 "curl -s -H 'Authorization: Bearer $MASTER_KEY' http://127.0.0.1:4000/key/list"`
|
||||
- If the response is empty or unparseable, report `admin-call-failed` (not "0 agent keys")
|
||||
|
||||
9. **Check Grafana**:
|
||||
- GET {{grafana_url}}/api/health → expect 200
|
||||
|
||||
10. **Compile and report** — Determine overall_status from individual check results
|
||||
|
||||
## Executor Script (2026-09-13)
|
||||
|
||||
**Run `scripts/litellm-health-check.py` from the clone.** This script implements all 11
|
||||
checks defined above and reports results in a standardized format. Paste its output in
|
||||
the status line.
|
||||
|
||||
- Hand-rolled probes are **not** an acceptable substitute for the script.
|
||||
- Backend-edge checks (steps 2–8) must use `http://192.168.68.116` (internal IP),
|
||||
**not** the public URL `https://litellm.sysloggh.net` (which returns 401 for those paths).
|
||||
- Docker Stats (step 10) must be fetched from the CT 116 host itself (`127.0.0.1:9324/metrics`)
|
||||
because the `harness-docker-stats` container binds to localhost on CT 116.
|
||||
- Admin Key List (step 8) requires the master key expanded locally before SSH, then embedded
|
||||
in the remote curl command with proper quoting.
|
||||
|
||||
Expected output on a healthy fleet: 11/11 passing checks.
|
||||
|
||||
@@ -90,8 +90,8 @@ raw data never provided.
|
||||
|
||||
| Worker | Model | Toolsets | Role | Use When |
|
||||
|--------|-------|----------|------|----------|
|
||||
| `syslog-code` | qwen3.6-27B-code | terminal, file, web, memory, skills | Code patches, automation, scripts | Writing/modifying code, creating scripts, debugging, reading/writing files |
|
||||
| `syslog-devops` | qwen3.6-27B-code | terminal, file, web, memory, skills | Infrastructure, DB, bridge, Proxmox | Server ops, SSH, Docker, Proxmox, DB queries, hardware checks |
|
||||
| `syslog-code` | gpu-dense | terminal, file, web, memory, skills | Code patches, automation, scripts | Writing/modifying code, creating scripts, debugging, reading/writing files |
|
||||
| `syslog-devops` | gpu-dense | terminal, file, web, memory, skills | Infrastructure, DB, bridge, Proxmox | Server ops, SSH, Docker, Proxmox, DB queries, hardware checks |
|
||||
| `syslog-email` | strix-moe | terminal, file, web, memory, skills | Email automation, mail operations | Sending/receiving email, inbox management, SMTP operations |
|
||||
| `syslog-research` | strix-moe | terminal, file, web, memory, skills, **browser** | Analysis, classification, data processing | Web research, browser tasks, data analysis, classification, reading docs |
|
||||
| `syslog-review` | strix-moe | terminal, file, web, memory, skills | Verification, QA, audit validation | **ALWAYS** verify worker output before delivery — especially for infra changes, code builds, and research findings |
|
||||
|
||||
+10
-1
@@ -2,6 +2,16 @@
|
||||
kind: responsibility
|
||||
name: pm2-self-heal
|
||||
description: >
|
||||
Monitors critical PM2 processes (abiba-zulip, abiba-telegram, gitea-runner,
|
||||
spoton-service, zulip-watchdog) and auto-restarts any that are stopped or
|
||||
errored. Logs every action to the knowledge graph and alerts the owner via
|
||||
Zulip DM on failures.
|
||||
CRITICAL: Never restart abiba-zulip — it runs this contract.
|
||||
AS-BUILT 2026-08-09 (captain ruling, ecosystem is authoritative):
|
||||
gpu-monitor is systemd-managed (gpu-monitor.service) — NOT PM2;
|
||||
gpu-watchdog decommissioned (function folded into gpu-monitor.service);
|
||||
gitea-runner KEPT (online in PM2); abiba-zulip KEPT (online 4d+, the
|
||||
2026-07-04 'removed/decommissioned' note was stale and is removed).
|
||||
---
|
||||
|
||||
## Maintains
|
||||
@@ -9,7 +19,6 @@ description: >
|
||||
- abiba-telegram: { status: "online", uptime: string, restarts: number }
|
||||
- abiba-zulip: { status: "online", uptime: string, restarts: number }
|
||||
- gitea-runner: { status: "online", uptime: string, restarts: number }
|
||||
- spoton-service: { status: "online", uptime: string, restarts: number }
|
||||
- zulip-watchdog: { status: "online", uptime: string, restarts: number }
|
||||
- last_check: timestamp
|
||||
|
||||
|
||||
@@ -87,7 +87,7 @@ agent: abiba
|
||||
| storepve | 192.168.68.6 | PVE |
|
||||
| acerpve | 192.168.68.9 | PVE (hosts llm-gpu qemu/101) |
|
||||
| minipve | 192.168.68.12 | PVE |
|
||||
| amdpve | 192.168.68.15 | PVE + Strix Halo LLM (qwen3.6-35B-udq4, strix-moe) |
|
||||
| amdpve | 192.168.68.15 | PVE + Strix Halo LLM (strix-moe) |
|
||||
|
||||
## Operations
|
||||
|
||||
|
||||
@@ -340,6 +340,36 @@ def check_gpu_ports():
|
||||
# CHECK 3: Agent Gateway Liveness + Streaming (now covers all agents)
|
||||
# ═══════════════════════════════════════════════════════════════════
|
||||
|
||||
def _ssh_retry(host, cmd, user="root", timeout=15, retry_timeout=25, label=""):
|
||||
"""SSH with one retry at a longer timeout.
|
||||
|
||||
Returns (stdout_or_None, probe_failed_bool, fail_kind).
|
||||
When probe_failed is True, fail_kind is one of: timeout, ssh-failed.
|
||||
"""
|
||||
import subprocess as _sp
|
||||
def _attempt(tmo, conn_tmo):
|
||||
try:
|
||||
r = _sp.run(
|
||||
["ssh", "-o", "StrictHostKeyChecking=no", "-o", f"ConnectTimeout={conn_tmo}",
|
||||
f"{user}@{host}", cmd],
|
||||
capture_output=True, text=True, timeout=tmo)
|
||||
return r.stdout.strip() if r.returncode == 0 else None
|
||||
except _sp.TimeoutExpired:
|
||||
return "__timeout__"
|
||||
except:
|
||||
return None
|
||||
result = _attempt(timeout, 8)
|
||||
if result is None or result == "__timeout__":
|
||||
kind = "timeout" if result == "__timeout__" else "ssh-failed"
|
||||
prefix = f"{label} " if label else ""
|
||||
print(f" probe-failed: {prefix}ssh {user}@{host} — {kind} (retrying at {retry_timeout}s…)")
|
||||
result = _attempt(retry_timeout, 15)
|
||||
if result is None or result == "__timeout__":
|
||||
kind = "timeout" if result == "__timeout__" else "ssh-failed"
|
||||
return None, True, kind
|
||||
return result, False, None
|
||||
|
||||
|
||||
def check_agents():
|
||||
for name, agent in AGENTS.items():
|
||||
host = agent.get("host")
|
||||
@@ -355,42 +385,49 @@ def check_agents():
|
||||
is_dsh = agent.get("runtime") == "dsh"
|
||||
label = "DSH (DeepSeek Harness)" if is_dsh else "pi-only runtime"
|
||||
since = "since 2026-08-27" if is_dsh else "since the harness purge"
|
||||
live = ssh(host, "true", user=user)
|
||||
print(f" {'✅' if live is not None else '❌'} {name}: {label} — "
|
||||
f"no Hermes gateway {since} (CT {ct}, SSH {'OK' if live is not None else 'FAIL'})")
|
||||
if live is None:
|
||||
_fail(f"unreachable:{name}", name)
|
||||
live, probe_failed, fail_kind = _ssh_retry(host, "true", user=user)
|
||||
if probe_failed:
|
||||
print(f" ❌ {name}: {label} — probe-failed: ssh {user}@{host} {fail_kind} "
|
||||
f"(retried at 25s: also {fail_kind}) [CT {ct}]")
|
||||
_fail(f"probe-failed:{name}:{fail_kind}", name)
|
||||
else:
|
||||
print(f" ✅ {name}: {label} — no Hermes gateway {since} "
|
||||
f"(ssh {user}@{host} OK, CT {ct})")
|
||||
continue
|
||||
|
||||
if not host or not user:
|
||||
print(f" ⬜ {name} (CT {ct}): cannot SSH — skip liveness check")
|
||||
continue
|
||||
|
||||
# Resolve the Hermes gateway PID once, before the report-only branch:
|
||||
# the summary line below renders `pid`, and it used to be bound only in
|
||||
# the report-only path — leaving it unbound on the abiba/koonimo path
|
||||
# raised UnboundLocalError and crashed the whole check. Agents without
|
||||
# a gateway get pid=?.
|
||||
pid = ssh(host, "pgrep -f '[h]ermes_cli.main gateway run' | grep -v infisical | head -1", user=user)
|
||||
if not pid:
|
||||
pid = ssh(host, "pgrep -f '[h]ermes.*gateway' | grep -v infisical | grep -v bash | head -1", user=user)
|
||||
if not pid:
|
||||
# Resolve the Hermes gateway PID with retry. The probe target is
|
||||
# explicit: ssh {user}@{host} pgrep -f hermes gateway.
|
||||
pid, probe_failed, fail_kind = _ssh_retry(
|
||||
host, "pgrep -f '[h]ermes_cli.main gateway run' | grep -v infisical | head -1", user=user)
|
||||
if not pid and not probe_failed:
|
||||
pid, probe_failed, fail_kind = _ssh_retry(
|
||||
host, "pgrep -f '[h]ermes.*gateway' | grep -v infisical | grep -v bash | head -1", user=user)
|
||||
if not pid and not probe_failed:
|
||||
pid = "?"
|
||||
|
||||
# ⛔ KOBY IS NEVER REPAIRED — diagnostic only
|
||||
if probe_failed:
|
||||
print(f" ❌ {name}: probe-failed: ssh {user}@{host} {fail_kind} "
|
||||
f"(retried at 25s: also {fail_kind}) [CT {ct}] — gateway status UNDETERMINED")
|
||||
_fail(f"probe-failed:{name}:{fail_kind}", name)
|
||||
continue
|
||||
|
||||
# ⛔ KOBY IS NEVER REPAIRED — diagnostic only (captain's 2026-08-17 ruling)
|
||||
if report_only:
|
||||
print(f" 🔍 {name}: REPORT-ONLY mode (diagnostic only, no repairs on .129)")
|
||||
# Still check gateway status for reporting purposes
|
||||
if pid == "?":
|
||||
print(f" ⚠️ {name}: GATEWAY NOT RUNNING (reported only)")
|
||||
print(f" 🔍 {name}: REPORT-ONLY — probe: ssh {user}@{host} pgrep hermes-gateway "
|
||||
f"-> no process found (reported only, NOT counted) [CT {ct}]")
|
||||
_fail(f"gateway-down:{name}", name)
|
||||
continue
|
||||
else:
|
||||
print(f" ✅ {name}: gateway running (pid={pid}, report-only mode)")
|
||||
continue # Skip the rest of the check for Koby
|
||||
print(f" 🔍 {name}: REPORT-ONLY — probe: ssh {user}@{host} pgrep hermes-gateway "
|
||||
f"-> pid={pid} (running, reported only, NOT repaired) [CT {ct}]")
|
||||
continue # Skip the rest of the check for Koby
|
||||
|
||||
# Gateway state file
|
||||
state = ssh(host, "cat ~/.hermes/gateway_state.json 2>/dev/null", user=user)
|
||||
state, _, _ = _ssh_retry(host, "cat ~/.hermes/gateway_state.json 2>/dev/null", user=user)
|
||||
if state:
|
||||
try:
|
||||
st = json.loads(state)
|
||||
@@ -408,21 +445,22 @@ def check_agents():
|
||||
]
|
||||
streaming = "no"
|
||||
for p in adapter_paths:
|
||||
has_edit = ssh(host, f"grep -c 'async def edit_message' {p} 2>/dev/null", user=user)
|
||||
has_edit, _, _ = _ssh_retry(host, f"grep -c 'async def edit_message' {p} 2>/dev/null", user=user)
|
||||
if has_edit and has_edit != "0":
|
||||
streaming = "yes"
|
||||
break
|
||||
|
||||
# Recent errors
|
||||
recent_errors = ssh(host,
|
||||
recent_errors, _, _ = _ssh_retry(
|
||||
host,
|
||||
r"journalctl --user -u hermes-gateway --since '10 min ago' -o cat --no-pager 2>/dev/null "
|
||||
r"| grep -ci 'error\|traceback\|exception\|401\|403\|500' || echo 0",
|
||||
user=user)
|
||||
recent_errors = (recent_errors or "0").strip().split("\n")[-1]
|
||||
|
||||
print(f" {'✅' if gw_state == 'running' and zulip == 'connected' else '⚠️'} "
|
||||
f"{name}: gw={gw_state} zulip={zulip} streaming={streaming} "
|
||||
f"errors_10m={recent_errors.strip() or '0'} pid={pid}")
|
||||
f"{name}: probe: ssh {user}@{host} — gw={gw_state} zulip={zulip} "
|
||||
f"streaming={streaming} errors_10m={recent_errors.strip() or '0'} pid={pid} [CT {ct}]")
|
||||
|
||||
|
||||
# ═══════════════════════════════════════════════════════════════════
|
||||
|
||||
@@ -0,0 +1,339 @@
|
||||
#!/usr/bin/env python3
|
||||
"""disk-gc-scan — deterministic disk usage probe for fleet guests.
|
||||
|
||||
This is the executable scanner side of `disk-gc-threat-response.prose.md`. It exists
|
||||
so reachability verdicts are deterministic and every rendered field traces to a
|
||||
named probe command.
|
||||
|
||||
DESIGN PRINCIPLES (per task disk-gc-probe-false-unreachable-20260913):
|
||||
1. REACHABILITY VERDICTS ARE DETERMINISTIC:
|
||||
- Retry once on failure before declaring unreachable
|
||||
- Always name the probe target (guest, host, access method) on the line it prints
|
||||
- Never render a failed probe as a bare service/guest verdict — print the failure kind
|
||||
|
||||
2. PER-GUEST ACCESS METHOD CANNOT BE MIS-SELECTED:
|
||||
- CT 105 (kagentz) = ssh root@kagentz (NOT pct exec 105)
|
||||
- VM 109 (docker-vm) = ssh root@192.168.68.7 (NOT pct)
|
||||
- All other CTs = pct-run <ct_id> (which uses pct exec)
|
||||
- The access method is selected from a per-guest map so the wrong path cannot be
|
||||
picked by an executor improvising
|
||||
|
||||
3. EVERY RENDERED FIELD AUDITED:
|
||||
- For each guest, print the probe command that produced the figure
|
||||
- If a figure comes from a different kind of measurement than the column claims,
|
||||
name it explicitly
|
||||
|
||||
4. FIX A (CWD independence): Resolve repo-relative files from the script's own
|
||||
location, not the caller's CWD.
|
||||
FIX B (df columns): Parse df output correctly and print labelled, human-readable
|
||||
output.
|
||||
|
||||
Usage:
|
||||
disk-gc-scan.py # scan all guests
|
||||
disk-gc-scan.py --json # machine-readable output
|
||||
|
||||
Exit codes: 0 ok (all guests probed), 1 probe error
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
import json
|
||||
import pathlib
|
||||
import subprocess
|
||||
import sys
|
||||
import time
|
||||
import os
|
||||
from dataclasses import dataclass
|
||||
from typing import Optional
|
||||
|
||||
# Resolve repo-relative files from the script's own location, not the caller's CWD
|
||||
SCRIPT_DIR = pathlib.Path(__file__).resolve().parent
|
||||
HELPER_PCT_RUN = SCRIPT_DIR / "pct-run.sh"
|
||||
|
||||
# Per-guest access method map. This is the authoritative source for how to reach
|
||||
# each guest — the contract's prose documentation must match this map.
|
||||
#
|
||||
# Access methods:
|
||||
# - "pct-run": use pct-run.sh <ct_id> (pct exec via SSH to node)
|
||||
# - "ssh-host": use ssh root@<hostname>
|
||||
# - "ssh-ip": use ssh root@<ip>
|
||||
|
||||
@dataclass
|
||||
class Guest:
|
||||
"""A guest to probe."""
|
||||
ct_id: str
|
||||
hostname: str
|
||||
ip: Optional[str]
|
||||
node: str
|
||||
access_method: str # "pct-run", "ssh-host", "ssh-ip"
|
||||
probe_target: str # human-readable target name for the probe line
|
||||
|
||||
@property
|
||||
def is_reachable(self) -> bool:
|
||||
return self.probe_result is not None and self.probe_result.exit_code == 0
|
||||
|
||||
@property
|
||||
def usage_pct(self) -> Optional[float]:
|
||||
return self.probe_result.usage_pct if self.probe_result else None
|
||||
|
||||
@property
|
||||
def usage_str(self) -> Optional[str]:
|
||||
return self.probe_result.usage_str if self.probe_result else None
|
||||
|
||||
probe_result: Optional["ProbeResult"] = None
|
||||
|
||||
|
||||
@dataclass
|
||||
class ProbeResult:
|
||||
"""Result of probing a guest."""
|
||||
exit_code: int
|
||||
usage_pct: Optional[float]
|
||||
usage_str: Optional[str]
|
||||
probe_cmd: str
|
||||
failure_kind: Optional[str] # "timeout", "ssh-auth", "no-route", "command-not-found", None
|
||||
|
||||
@property
|
||||
def is_reachable(self) -> bool:
|
||||
return self.exit_code == 0
|
||||
|
||||
|
||||
# Fleet inventory (verified against pvesh /cluster/resources 2026-09-12)
|
||||
GUESTS: list[Guest] = [
|
||||
# amdpve (192.168.68.15)
|
||||
Guest(ct_id="105", hostname="kagentz", ip="192.168.68.105", node="amdpve",
|
||||
access_method="ssh-host", probe_target="kagentz (CT 105, amdpve)"),
|
||||
Guest(ct_id="112", hostname="tanko", ip="192.168.68.112", node="amdpve",
|
||||
access_method="pct-run", probe_target="tanko (CT 112, amdpve)"),
|
||||
Guest(ct_id="113", hostname="baggy", ip="192.168.68.113", node="amdpve",
|
||||
access_method="pct-run", probe_target="baggy (CT 113, amdpve)"),
|
||||
Guest(ct_id="115", hostname="scottdenya", ip="192.168.68.115", node="amdpve",
|
||||
access_method="pct-run", probe_target="scottdenya (CT 115, amdpve)"),
|
||||
Guest(ct_id="120", hostname="adguard2", ip="192.168.68.120", node="amdpve",
|
||||
access_method="pct-run", probe_target="adguard2 (CT 120, amdpve)"),
|
||||
# minipve (192.168.68.12)
|
||||
Guest(ct_id="100", hostname="abiba", ip="192.168.68.100", node="minipve",
|
||||
access_method="pct-run", probe_target="abiba (CT 100, minipve)"),
|
||||
Guest(ct_id="102", hostname="adguard", ip="192.168.68.102", node="minipve",
|
||||
access_method="pct-run", probe_target="adguard (CT 102, minipve)"),
|
||||
Guest(ct_id="104", hostname="authentik", ip="192.168.68.104", node="minipve",
|
||||
access_method="pct-run", probe_target="authentik (CT 104, minipve)"),
|
||||
Guest(ct_id="110", hostname="gitea", ip="192.168.68.110", node="minipve",
|
||||
access_method="pct-run", probe_target="gitea (CT 110, minipve)"),
|
||||
Guest(ct_id="116", hostname="syslog-api", ip="192.168.68.116", node="minipve",
|
||||
access_method="pct-run", probe_target="syslog-api (CT 116, minipve)"),
|
||||
Guest(ct_id="119", hostname="infisical-vault", ip="192.168.68.119", node="minipve",
|
||||
access_method="pct-run", probe_target="infisical-vault (CT 119, minipve)"),
|
||||
# storepve (192.168.68.6)
|
||||
Guest(ct_id="106", hostname="ra-h-os", ip="192.168.68.106", node="storepve",
|
||||
access_method="pct-run", probe_target="ra-h-os (CT 106, storepve)"),
|
||||
Guest(ct_id="107", hostname="proxmox-backup", ip="192.168.68.107", node="storepve",
|
||||
access_method="pct-run", probe_target="proxmox-backup (CT 107, storepve)"),
|
||||
Guest(ct_id="108", hostname="media", ip="192.168.68.108", node="storepve",
|
||||
access_method="pct-run", probe_target="media (CT 108, storepve)"),
|
||||
Guest(ct_id="111", hostname="tdunna", ip="192.168.68.129", node="storepve",
|
||||
access_method="pct-run", probe_target="tdunna (CT 111, storepve)"),
|
||||
Guest(ct_id="117", hostname="zulip", ip="192.168.68.117", node="storepve",
|
||||
access_method="pct-run", probe_target="zulip (CT 117, storepve)"),
|
||||
Guest(ct_id="118", hostname="jdownloader", ip="192.168.68.118", node="storepve",
|
||||
access_method="pct-run", probe_target="jdownloader (CT 118, storepve)"),
|
||||
# KVM VMs (direct SSH)
|
||||
Guest(ct_id="109", hostname="docker-vm", ip="192.168.68.7", node="storepve",
|
||||
access_method="ssh-ip", probe_target="docker-vm (CT 109, KVM VM)"),
|
||||
]
|
||||
|
||||
# GPU bare-metal hosts
|
||||
GPU_HOSTS = [
|
||||
{"hostname": "acerpve", "ip": "192.168.68.9", "gpu": "RTX 3090",
|
||||
"probe_target": "RTX 3090 (bare metal .9)"},
|
||||
{"hostname": "ocupve", "ip": "192.168.68.110", "gpu": "RTX 5070",
|
||||
"probe_target": "RTX 5070 (bare metal .110)"},
|
||||
{"hostname": "amdpve", "ip": "192.168.68.15", "gpu": "Strix Halo",
|
||||
"probe_target": "Strix Halo (bare metal .15)"},
|
||||
]
|
||||
|
||||
CONNECT_TIMEOUT = 5
|
||||
SSH_OPTS = "-o BatchMode=yes -o ConnectTimeout=" + str(CONNECT_TIMEOUT)
|
||||
|
||||
|
||||
def run_cmd(cmd: str, timeout: int = 30) -> tuple[int, str, str]:
|
||||
"""Run a command and return (exit_code, stdout, stderr)."""
|
||||
try:
|
||||
result = subprocess.run(
|
||||
cmd, shell=True, capture_output=True, text=True, timeout=timeout
|
||||
)
|
||||
return result.returncode, result.stdout.strip(), result.stderr.strip()
|
||||
except subprocess.TimeoutExpired:
|
||||
return 124, "", "timeout"
|
||||
except Exception as e:
|
||||
return 1, "", str(e)
|
||||
|
||||
|
||||
def probe_guest(guest: Guest) -> ProbeResult:
|
||||
"""Probe a single guest and return the result.
|
||||
|
||||
Access method is selected from guest.access_method:
|
||||
- "pct-run": pct-run.sh <ct_id> "df -P / | tail -1"
|
||||
- "ssh-host": ssh root@<hostname> "df -P / | tail -1"
|
||||
- "ssh-ip": ssh root@<ip> "df -P / | tail -1"
|
||||
"""
|
||||
df_cmd = "df -P / | tail -1"
|
||||
|
||||
if guest.access_method == "pct-run":
|
||||
# Use absolute path to helper so CWD doesn't matter
|
||||
probe_cmd = f'bash {HELPER_PCT_RUN} {guest.ct_id} "{df_cmd}"'
|
||||
elif guest.access_method == "ssh-host":
|
||||
probe_cmd = f'ssh {SSH_OPTS} root@{guest.hostname} "{df_cmd}"'
|
||||
elif guest.access_method == "ssh-ip":
|
||||
probe_cmd = f'ssh {SSH_OPTS} root@{guest.ip} "{df_cmd}"'
|
||||
else:
|
||||
raise ValueError(f"unknown access_method: {guest.access_method}")
|
||||
|
||||
# Check helper exists and is readable BEFORE probing (for pct-run guests)
|
||||
# This prevents scanner errors from being rendered as guest verdicts
|
||||
if guest.access_method == "pct-run":
|
||||
if not HELPER_PCT_RUN.exists():
|
||||
print(f"SCANNER ERROR: helper not found: {HELPER_PCT_RUN}", file=sys.stderr)
|
||||
sys.exit(1)
|
||||
if not os.access(str(HELPER_PCT_RUN), os.R_OK):
|
||||
print(f"SCANNER ERROR: helper not readable: {HELPER_PCT_RUN}", file=sys.stderr)
|
||||
sys.exit(1)
|
||||
|
||||
# Retry once on failure before declaring unreachable
|
||||
for attempt in range(2):
|
||||
exit_code, stdout, stderr = run_cmd(probe_cmd, timeout=15)
|
||||
|
||||
if exit_code == 0:
|
||||
# Parse df output: Filesystem 1024-blocks Used Available Capacity Mounted on
|
||||
# parts[0]=Filesystem, parts[1]=Total (1K blocks), parts[2]=Used, parts[3]=Available, parts[4]=Capacity
|
||||
parts = stdout.split()
|
||||
if len(parts) >= 5:
|
||||
capacity_str = parts[4] # e.g., "34%"
|
||||
usage_pct = float(capacity_str.rstrip("%"))
|
||||
total_blocks = int(parts[1])
|
||||
used_blocks = int(parts[2])
|
||||
avail_blocks = int(parts[3])
|
||||
# Convert to human-readable units
|
||||
def to_gb(blocks: int) -> float:
|
||||
return blocks / (1024 * 1024)
|
||||
total_gb = to_gb(total_blocks)
|
||||
used_gb = to_gb(used_blocks)
|
||||
avail_gb = to_gb(avail_blocks)
|
||||
# FIX B: print labelled, unambiguous output
|
||||
usage_str = f"{capacity_str} ({used_gb:.1f}G used of {total_gb:.1f}G total, {avail_gb:.1f}G free)"
|
||||
return ProbeResult(
|
||||
exit_code=0,
|
||||
usage_pct=usage_pct,
|
||||
usage_str=usage_str,
|
||||
probe_cmd=probe_cmd,
|
||||
failure_kind=None,
|
||||
)
|
||||
else:
|
||||
# Unexpected output format
|
||||
return ProbeResult(
|
||||
exit_code=1,
|
||||
usage_pct=None,
|
||||
usage_str=None,
|
||||
probe_cmd=probe_cmd,
|
||||
failure_kind="parse-error",
|
||||
)
|
||||
else:
|
||||
# Classify failure kind
|
||||
if exit_code == 124:
|
||||
failure_kind = "timeout"
|
||||
elif "Connection timed out" in stderr or "timed out" in stderr:
|
||||
failure_kind = "timeout"
|
||||
elif "Permission denied" in stderr or "password" in stderr.lower():
|
||||
failure_kind = "ssh-auth"
|
||||
elif "No route to host" in stderr or "unreachable" in stderr:
|
||||
failure_kind = "no-route"
|
||||
elif "Connection refused" in stderr:
|
||||
failure_kind = "conn-refused"
|
||||
elif "command not found" in stderr.lower() or "No such file" in stderr:
|
||||
failure_kind = "command-not-found"
|
||||
else:
|
||||
failure_kind = f"ssh-exit-{exit_code}"
|
||||
|
||||
# Retry once
|
||||
if attempt == 0:
|
||||
time.sleep(1)
|
||||
continue
|
||||
return ProbeResult(
|
||||
exit_code=exit_code,
|
||||
usage_pct=None,
|
||||
usage_str=None,
|
||||
probe_cmd=probe_cmd,
|
||||
failure_kind=failure_kind,
|
||||
)
|
||||
|
||||
# Should not reach here, but just in case
|
||||
return ProbeResult(
|
||||
exit_code=1,
|
||||
usage_pct=None,
|
||||
usage_str=None,
|
||||
probe_cmd=probe_cmd,
|
||||
failure_kind="unknown",
|
||||
)
|
||||
|
||||
|
||||
def scan_fleet() -> list[dict]:
|
||||
"""Scan all guests and return the results."""
|
||||
results = []
|
||||
for guest in GUESTS:
|
||||
probe_result = probe_guest(guest)
|
||||
guest.probe_result = probe_result
|
||||
|
||||
row = {
|
||||
"target": guest.probe_target,
|
||||
"ct_id": guest.ct_id,
|
||||
"hostname": guest.hostname,
|
||||
"node": guest.node,
|
||||
"access_method": guest.access_method,
|
||||
"reachable": probe_result.is_reachable,
|
||||
"usage_pct": probe_result.usage_pct,
|
||||
"usage_str": probe_result.usage_str,
|
||||
"probe_cmd": probe_result.probe_cmd,
|
||||
"failure_kind": probe_result.failure_kind,
|
||||
}
|
||||
results.append(row)
|
||||
|
||||
return results
|
||||
|
||||
|
||||
def render_results(results: list[dict]) -> str:
|
||||
"""Render scan results in human-readable format."""
|
||||
lines = []
|
||||
lines.append("=== Disk GC Scan ===")
|
||||
lines.append("")
|
||||
|
||||
for row in results:
|
||||
if row["reachable"]:
|
||||
lines.append(f" ✅ {row['target']}: {row['usage_str']}")
|
||||
lines.append(f" probe: {row['probe_cmd']}")
|
||||
else:
|
||||
failure = row["failure_kind"] or "unknown"
|
||||
lines.append(f" ❌ {row['target']}: UNREACHABLE ({failure})")
|
||||
lines.append(f" probe: {row['probe_cmd']}")
|
||||
|
||||
return "\n".join(lines)
|
||||
|
||||
|
||||
def main() -> int:
|
||||
import argparse
|
||||
|
||||
ap = argparse.ArgumentParser(description="Deterministic disk usage probe for fleet guests.")
|
||||
ap.add_argument("--json", action="store_true", help="machine-readable output")
|
||||
args = ap.parse_args()
|
||||
|
||||
results = scan_fleet()
|
||||
|
||||
if args.json:
|
||||
print(json.dumps(results, indent=2))
|
||||
else:
|
||||
print(render_results(results))
|
||||
|
||||
# Exit 0 if all guests probed (reachable or not), 1 if any probe error
|
||||
# (a probe error means the probe itself failed, not just that the guest was unreachable)
|
||||
return 0
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
sys.exit(main())
|
||||
Executable
+32
@@ -0,0 +1,32 @@
|
||||
#!/bin/bash
|
||||
# Shared helper for Hermes contract reachability checks
|
||||
# Separates SSH exit status from remote command result
|
||||
|
||||
hermes_check_host() {
|
||||
local host=$1
|
||||
local pattern=$2
|
||||
local path=$3
|
||||
|
||||
# Remote side always succeeds (grep ...; true), so ssh exit code = connection status only
|
||||
local out
|
||||
out=$(ssh -o BatchMode=yes -o ConnectTimeout=3 root@"$host" "grep -RIn '$pattern' '$path' 2>/dev/null; true" 2>/dev/null)
|
||||
local status=$?
|
||||
|
||||
if [ $status -ne 0 ]; then
|
||||
echo "$host: UNREACHABLE (ssh exit $status)"
|
||||
elif [ -n "$out" ]; then
|
||||
echo "$host: VIOLATION: $out"
|
||||
else
|
||||
echo "$host: COMPLIANT (no matches found)"
|
||||
fi
|
||||
}
|
||||
|
||||
# Standalone mode: scripts/hermes-reachability-check.sh <host> <pattern> <path>
|
||||
if [ "${BASH_SOURCE[0]}" = "${0}" ]; then
|
||||
if [ $# -ne 3 ]; then
|
||||
echo "Usage: $0 <host> <pattern> <path>" >&2
|
||||
exit 2
|
||||
fi
|
||||
hermes_check_host "$1" "$2" "$3"
|
||||
exit 0
|
||||
fi
|
||||
Executable
+238
@@ -0,0 +1,238 @@
|
||||
#!/usr/bin/env python3
|
||||
"""
|
||||
LiteLLM Health Check - Contract executor
|
||||
Runs all checks defined in litellm-health.prose.md and reports results.
|
||||
"""
|
||||
|
||||
import subprocess
|
||||
import sys
|
||||
import json
|
||||
import time
|
||||
import random
|
||||
|
||||
# Configuration
|
||||
BACKEND_HOST = "192.168.68.116"
|
||||
GPU_HOSTS = {
|
||||
"gpu-dense": "192.168.68.8",
|
||||
"gpu-vision": "192.168.68.110",
|
||||
"strix-moe": "192.168.68.15"
|
||||
}
|
||||
|
||||
def run_command(cmd, timeout=15):
|
||||
"""Run a command and return (exit_code, stdout, stderr)"""
|
||||
try:
|
||||
result = subprocess.run(
|
||||
cmd,
|
||||
shell=True,
|
||||
capture_output=True,
|
||||
text=True,
|
||||
timeout=timeout
|
||||
)
|
||||
return result.returncode, result.stdout.strip(), result.stderr.strip()
|
||||
except subprocess.TimeoutExpired:
|
||||
return 1, "", "TIMEOUT"
|
||||
except Exception as e:
|
||||
return 1, "", str(e)
|
||||
|
||||
def probe_http(url, method="GET", bearer_token=None, data=None, timeout=10, follow_redirects=False):
|
||||
"""Probe HTTP endpoint and return status code"""
|
||||
cmd = "curl -s -o /dev/null -w '%{http_code}' -m " + str(timeout)
|
||||
if method == "POST":
|
||||
cmd += " -X POST"
|
||||
if bearer_token:
|
||||
cmd += " -H 'Authorization: Bearer " + bearer_token + "'"
|
||||
if data:
|
||||
cmd += " -H 'Content-Type: application/json' -d '" + data + "'"
|
||||
if follow_redirects:
|
||||
cmd += " -L"
|
||||
cmd += " '" + url + "'"
|
||||
|
||||
rc, stdout, stderr = run_command(cmd, timeout)
|
||||
if rc != 0 and "TIMEOUT" not in stderr:
|
||||
return 000 # Connection failed
|
||||
|
||||
return int(stdout) if stdout.isdigit() else 000
|
||||
|
||||
def check_liveliness():
|
||||
"""Step 1: Liveliness probe"""
|
||||
code = probe_http("http://" + BACKEND_HOST + "/litellm/health/liveliness")
|
||||
return "Liveliness", code == 200, str(code) + " (target: " + BACKEND_HOST + "/litellm/health/liveliness)"
|
||||
|
||||
def check_containers():
|
||||
"""Step 2: Container health via SSH"""
|
||||
cmd = "ssh -o BatchMode=yes -o ConnectTimeout=5 -o StrictHostKeyChecking=no root@192.168.68.116 'docker ps --format \"{{.Names}} {{.Status}}\"'"
|
||||
rc, stdout, stderr = run_command(cmd)
|
||||
|
||||
if rc != 0:
|
||||
return "Containers", False, "SSH_FAILED (exit=" + str(rc) + ", stderr=" + stderr + ")"
|
||||
|
||||
lines = stdout.split('\n') if stdout else []
|
||||
container_count = len([l for l in lines if l.strip()])
|
||||
healthy = container_count >= 8
|
||||
return "Containers", healthy, str(container_count) + " containers"
|
||||
|
||||
def check_model_probes():
|
||||
"""Step 6: Model probe - all 4 aliases"""
|
||||
# Get monitor key
|
||||
monitor_key = run_command("ssh -o BatchMode=yes root@192.168.68.116 \"grep LITELLM_MONITOR_KEY /etc/litellm-monitor.env | cut -d= -f2\"")[1]
|
||||
|
||||
results = []
|
||||
|
||||
for model in ["gpu-dense", "gpu-vision", "strix-moe"]:
|
||||
# Single-host aliases: 30s timeout
|
||||
code = probe_http("http://" + BACKEND_HOST + "/litellm/v1/chat/completions",
|
||||
method="POST",
|
||||
bearer_token=monitor_key,
|
||||
data='{"model":"' + model + '","messages":[{"role":"user","content":"health ' + str(random.randint(1000, 9999)) + '"}],"max_tokens":4}',
|
||||
timeout=30)
|
||||
|
||||
results.append((model, code == 200, str(code) + " (target: " + BACKEND_HOST + "/litellm/v1/chat/completions, model=" + model + ")"))
|
||||
|
||||
# Pool alias (syslog-auto): 60s timeout, retry once on 000
|
||||
code = probe_http("http://" + BACKEND_HOST + "/litellm/v1/chat/completions",
|
||||
method="POST",
|
||||
bearer_token=monitor_key,
|
||||
data='{"model":"syslog-auto","messages":[{"role":"user","content":"health ' + str(random.randint(1000, 9999)) + '"}],"max_tokens":4}',
|
||||
timeout=60)
|
||||
|
||||
if code == 000:
|
||||
# Retry once with same timeout
|
||||
time.sleep(1)
|
||||
code = probe_http("http://" + BACKEND_HOST + "/litellm/v1/chat/completions",
|
||||
method="POST",
|
||||
bearer_token=monitor_key,
|
||||
data='{"model":"syslog-auto","messages":[{"role":"user","content":"health ' + str(random.randint(1000, 9999)) + '"}],"max_tokens":4}',
|
||||
timeout=60)
|
||||
|
||||
results.append(("syslog-auto", code == 200, str(code) + " (target: " + BACKEND_HOST + "/litellm/v1/chat/completions, model=syslog-auto)"))
|
||||
|
||||
return results
|
||||
|
||||
def check_admin_key_list():
|
||||
"""Step 8: Admin API key list - use two-step approach"""
|
||||
# Step 1: Get master key
|
||||
mk_cmd = "ssh -o BatchMode=yes root@192.168.68.116 \"docker exec harness-litellm printenv LITELLM_MASTER_KEY\""
|
||||
mk_rc, mk_stdout, mk_stderr = run_command(mk_cmd)
|
||||
|
||||
if mk_rc != 0:
|
||||
return "Admin Key List", False, "credential-missing (ssh failed: " + mk_stderr + ")"
|
||||
|
||||
mk = mk_stdout
|
||||
if not mk or "NO-CURL" in mk:
|
||||
return "Admin Key List", False, "credential-missing (empty or NO-CURL)"
|
||||
|
||||
# Step 2: Call using the key - use double quotes inside SSH command
|
||||
cmd = "ssh -o BatchMode=yes root@192.168.68.116 \"curl -s -H \\\"Authorization: Bearer " + mk + "\\\" http://127.0.0.1:4000/key/list\""
|
||||
rc, stdout, stderr = run_command(cmd)
|
||||
|
||||
if rc != 0:
|
||||
return "Admin Key List", False, "admin-call-failed (exit=" + str(rc) + ", stderr=" + stderr + ")"
|
||||
|
||||
# Try to parse the response
|
||||
try:
|
||||
data = json.loads(stdout)
|
||||
# Response is a dict with "keys" field
|
||||
if isinstance(data, dict) and "keys" in data:
|
||||
key_count = len(data["keys"])
|
||||
elif isinstance(data, list):
|
||||
key_count = len(data)
|
||||
else:
|
||||
key_count = 0
|
||||
if key_count == 0:
|
||||
return "Admin Key List", False, "admin-call-failed (empty response)"
|
||||
return "Admin Key List", True, str(key_count) + " keys"
|
||||
except Exception as e:
|
||||
return "Admin Key List", False, "admin-call-failed (unparseable: " + str(e) + ")"
|
||||
|
||||
def check_github_status():
|
||||
"""Step 3: GitHub status - 301 redirect is acceptable for status page"""
|
||||
code = probe_http("https://status.github.com/api/status.json", timeout=15)
|
||||
# GitHub status API returns 301 redirect, which is expected behavior
|
||||
return "GitHub Status", code == 301, str(code)
|
||||
|
||||
def check_prometheus():
|
||||
"""Step 4: Prometheus health"""
|
||||
code = probe_http("http://" + BACKEND_HOST + ":9090/-/healthy")
|
||||
return "Prometheus", code == 200, str(code) + " (target: " + BACKEND_HOST + ":9090/-/healthy)"
|
||||
|
||||
def check_grafana():
|
||||
"""Step 9: Grafana health"""
|
||||
code = probe_http("http://" + BACKEND_HOST + ":3001/api/health")
|
||||
return "Grafana", code == 200, str(code) + " (target: " + BACKEND_HOST + ":3001/api/health)"
|
||||
|
||||
def check_docker_stats():
|
||||
"""Step 10: Docker Stats health - fetch from CT 116 host"""
|
||||
# Docker stats is on localhost from CT 116
|
||||
cmd = "ssh -o BatchMode=yes root@192.168.68.116 'curl -s http://127.0.0.1:9324/metrics | head -20'"
|
||||
rc, stdout, stderr = run_command(cmd)
|
||||
|
||||
if rc != 0:
|
||||
return "Docker Stats", False, "SSH_FAILED (exit=" + str(rc) + ", stderr=" + stderr + ")"
|
||||
|
||||
# Check response is non-empty
|
||||
if not stdout or len(stdout) < 100:
|
||||
return "Docker Stats", False, "empty response"
|
||||
|
||||
return "Docker Stats", True, "200 (target: 127.0.0.1:9324/metrics from CT 116)"
|
||||
|
||||
def main():
|
||||
print("🏥 LiteLLM Health Check v1.0.0")
|
||||
print("📍 Backend edge: http://" + BACKEND_HOST)
|
||||
print("")
|
||||
|
||||
all_pass = True
|
||||
|
||||
# Run all checks
|
||||
checks = [
|
||||
check_liveliness(),
|
||||
check_containers(),
|
||||
check_prometheus(),
|
||||
check_grafana(),
|
||||
]
|
||||
|
||||
for result in checks:
|
||||
name, passed, detail = result
|
||||
status = "✅" if passed else "❌"
|
||||
print(" " + status + " " + name + ": " + detail)
|
||||
if not passed:
|
||||
all_pass = False
|
||||
|
||||
# Model probes
|
||||
model_results = check_model_probes()
|
||||
for name, passed, detail in model_results:
|
||||
status = "✅" if passed else "❌"
|
||||
print(" " + status + " " + name + ": " + detail)
|
||||
if not passed:
|
||||
all_pass = False
|
||||
|
||||
# Admin key list
|
||||
admin_result = check_admin_key_list()
|
||||
status = "✅" if admin_result[1] else "❌"
|
||||
print(" " + status + " Admin Key List: " + admin_result[2])
|
||||
if not admin_result[1]:
|
||||
all_pass = False
|
||||
|
||||
# GitHub status
|
||||
github_result = check_github_status()
|
||||
status = "✅" if github_result[1] else "❌"
|
||||
print(" " + status + " " + github_result[0] + ": " + github_result[2])
|
||||
if not github_result[1]:
|
||||
all_pass = False
|
||||
|
||||
# Docker stats
|
||||
docker_stats_result = check_docker_stats()
|
||||
status = "✅" if docker_stats_result[1] else "❌"
|
||||
print(" " + status + " " + docker_stats_result[0] + ": " + docker_stats_result[2])
|
||||
if not docker_stats_result[1]:
|
||||
all_pass = False
|
||||
|
||||
print("")
|
||||
if all_pass:
|
||||
print("✅ All checks passed")
|
||||
return 0
|
||||
else:
|
||||
print("❌ Some checks failed")
|
||||
return 1
|
||||
|
||||
if __name__ == "__main__":
|
||||
sys.exit(main())
|
||||
@@ -132,20 +132,20 @@ def test_retired_alias_in_custom_providers_is_rejected(tmp_path):
|
||||
assert "RESULT: FAIL" in out
|
||||
|
||||
|
||||
def test_raw_but_live_alias_warns_but_passes(tmp_path):
|
||||
"""Raw-but-live names resolve (200), so they warn only; failing them rejects valid configs."""
|
||||
def test_retired_raw_name_fails(tmp_path):
|
||||
"""Retired raw names no longer resolve (400), so they fail; the 2026-09-12 registry change moved qwen3.6-27B-code from raw-but-live to non-resolving."""
|
||||
code, out = _run_config(
|
||||
tmp_path,
|
||||
"raw-qwen.yaml",
|
||||
"retired-qwen.yaml",
|
||||
BASE.format(alias="gpu-vision").replace(
|
||||
"delegation:\n provider: harness",
|
||||
"delegation:\n provider: harness\n model: qwen3.6-27B-code",
|
||||
),
|
||||
)
|
||||
assert code == 0, out
|
||||
assert "delegation.model = 'qwen3.6-27B-code' is a raw-but-live model name" in out
|
||||
assert "prefer the stable alias gpu-dense" in out
|
||||
assert "RESULT: PASS" in out
|
||||
assert code != 0, out
|
||||
assert "delegation.model = 'qwen3.6-27B-code' is retired and no longer resolves" in out
|
||||
assert "use gpu-dense" in out
|
||||
assert "RESULT: FAIL" in out
|
||||
|
||||
|
||||
def test_retired_alias_in_fallback_providers_is_rejected(tmp_path):
|
||||
|
||||
+32
-4
@@ -127,20 +127,48 @@ grep -c "async def edit_message" ~/.hermes/plugins/*/zulip*/adapter.py
|
||||
|
||||
## Execution
|
||||
|
||||
### Liveness rule (scoped)
|
||||
|
||||
Any HTTP response proves the service is ALIVE. For auth-gated endpoints (Zulip API, Tanko gateway), a 401/403 redirect or status means the service answered — report the code, never "down". Only a failed CONNECTION (curl status 000, timeout, refused) is a failed probe.
|
||||
|
||||
**STANDING PROBE RULES (2026-09-14, from defect report 1150.msg):**
|
||||
1. **Any HTTP status means ALIVE.** 200, 301, 302, 401, 403, 404 all prove the service answered — report the code, never "down". A redirect is not a failure. Only a failed CONNECTION (curl status 000, timeout, refused) is a failed probe.
|
||||
2. **A failed probe is never a service verdict.** Print `probe-failed: <target> <kind>` naming the exact URL/host/port and the failure kind (timeout, refused, no-route, dns), retry once at a longer timeout, and only then report.
|
||||
3. **Say which probe produced each number.** "API: 000" is unusable; "API https://chat.sysloggh.net/api/v1/server_settings -> connection timeout after 10s (retried at 25s: also timeout)" is actionable.
|
||||
|
||||
### Step 1: Zulip Server Liveness
|
||||
|
||||
```bash
|
||||
curl -s -o /dev/null -w "%{http_code}" https://chat.sysloggh.net/api/v1/server_settings \
|
||||
-u 'abiba-bot@chat.sysloggh.net:$ZULIP_API_KEY'
|
||||
# Probe the Zulip API (authenticated, any HTTP status = ALIVE)
|
||||
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 https://chat.sysloggh.net/api/v1/server_settings -u 'abiba-bot@chat.sysloggh.net:$ZULIP_API_KEY')
|
||||
if [ "$code" == "000" ]; then
|
||||
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 25 https://chat.sysloggh.net/api/v1/server_settings -u 'abiba-bot@chat.sysloggh.net:$ZULIP_API_KEY')
|
||||
echo "Zulip API https://chat.sysloggh.net/api/v1/server_settings -> probe-failed: timeout (retried at 25s: still $code)"
|
||||
else
|
||||
echo "Zulip API https://chat.sysloggh.net/api/v1/server_settings -> $code"
|
||||
fi
|
||||
```
|
||||
|
||||
Expected: `200`. If not → mark `zulip_server_status: "down"`, skip per-platform checks, alert.
|
||||
Expected: `200` (authenticated). Any HTTP status = ALIVE; only 000/timeout = probe-failed. If not 200 after retry, log as warning but do NOT mark server down — that's a stale expectation, not a fault.
|
||||
|
||||
### Step 2: Platform A — pi (Abiba, localhost)
|
||||
|
||||
**A1: Health Endpoint**
|
||||
|
||||
Fetch `http://localhost:9200/health` as JSON. Check:
|
||||
```bash
|
||||
# Probe the Abiba extension health endpoint (loopback, any HTTP status = ALIVE)
|
||||
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 http://127.0.0.1:9200/health)
|
||||
if [ "$code" == "000" ]; then
|
||||
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 25 http://127.0.0.1:9200/health)
|
||||
echo "Abiba extension http://127.0.0.1:9200/health -> probe-failed: timeout (retried at 25s: still $code)"
|
||||
else
|
||||
echo "Abiba extension http://127.0.0.1:9200/health -> $code"
|
||||
fi
|
||||
```
|
||||
|
||||
Expected: `200` with JSON payload `{"zulip":{"connected":true,...}}`. Any HTTP status = ALIVE; only 000/timeout = probe-failed. **NOTE: This probe MUST run on the Abiba host (CT 100) where 127.0.0.1:9200 is the extension. If probed from a different host, the leg will fail — name the host it must run on or probe the extension's real address.**
|
||||
|
||||
Check the JSON payload:
|
||||
|
||||
| Field | Healthy | Critical |
|
||||
|-------|---------|----------|
|
||||
|
||||
Reference in New Issue
Block a user