Compare commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
05366bd58d | ||
|
|
e648b5ac0e | ||
|
|
38d7e8b064 | ||
|
|
e94fadfa60 | ||
|
|
ba76f2c7d3 | ||
|
|
7906b2d52d | ||
|
|
c81cf5b6f0 | ||
|
|
ef168d9690 | ||
|
|
edcf465831 | ||
|
|
89651cf37c | ||
|
|
6f40a3be60 | ||
|
|
d9368467ff | ||
|
|
e38598eea4 | ||
|
|
876b011359 | ||
|
|
d23cce89e1 | ||
|
|
9ada2b23c7 | ||
|
|
bd0065bb31 | ||
|
|
f99f7e1e34 | ||
|
|
8210fd905c | ||
|
|
b3644f0292 | ||
|
|
672bf8a912 | ||
|
|
6e612ce37b | ||
|
|
357f808a82 | ||
|
|
6272d29978 | ||
|
|
4bd6132cf5 | ||
|
|
6d65cba064 | ||
|
|
de1428b4ae | ||
|
|
4ea2d0309f | ||
|
|
7b8cc5f9ac | ||
|
|
d9e06863d8 | ||
|
|
a2edc2f56f | ||
|
|
78b501798f | ||
|
|
1f1b47f59d | ||
|
|
3d55764799 | ||
|
|
ce48070f21 | ||
|
|
9e581ab203 | ||
|
|
9cac3589cf | ||
|
|
221f9f79f3 | ||
|
|
1dc040251d | ||
|
|
21f7b6171c | ||
|
|
bb17c2120f | ||
|
|
dc42ecc235 | ||
|
|
baaac9d7c6 | ||
|
|
2f65c38213 | ||
|
|
d9eb18c024 | ||
|
|
3f07b9bbcc | ||
|
|
1a598d0fcb | ||
|
|
e3752af162 | ||
|
|
22d2c3acac | ||
|
|
5f582e2c9c | ||
|
|
b1b3b4c010 | ||
|
|
5d9b9847bc | ||
|
|
d29da3cc69 | ||
|
|
2ba1016ca8 | ||
|
|
9c8637bcaf | ||
|
|
aec62f7e77 | ||
|
|
b326a8944a | ||
|
|
7730cc7c03 | ||
|
|
af27530edc | ||
|
|
862356bcac | ||
|
|
767bd22d9c | ||
|
|
e97145c88f | ||
|
|
55f1208eb8 | ||
|
|
b9adf353ee | ||
|
|
aeb66ea22d | ||
|
|
288f74cf84 | ||
|
|
c66671dbee | ||
|
|
b80d3142aa | ||
|
|
266fa1f835 | ||
|
|
85ea1f4f3d | ||
|
|
b2a259fa23 |
@@ -43,17 +43,21 @@ jobs:
|
||||
echo "=== Prose Contract Frontmatter Validation ==="
|
||||
FAILED=0
|
||||
for f in $(find . -name "*.prose.md" -not -path "./.git/*" -not -path "./runs/*"); do
|
||||
# NOTE: use herestrings, not `echo "$FM" | grep ...`. Under the runner's
|
||||
# `-e -o pipefail`, `grep -q` exits on first match and can SIGPIPE the
|
||||
# producer, making the pipeline report non-zero and raising a false
|
||||
# "Missing name/description" whose file set varies run to run.
|
||||
FM=$(sed -n '/^---$/,/^---$/p' "$f" | sed '1d;$d')
|
||||
[ -z "$FM" ] && { echo " ❌ $f: No YAML frontmatter"; FAILED=$((FAILED+1)); continue; }
|
||||
|
||||
KIND=$(echo "$FM" | grep '^kind:' | awk '{print $2}')
|
||||
KIND=$(grep '^kind:' <<< "$FM" | awk '{print $2}')
|
||||
case "$KIND" in
|
||||
function|responsibility|gateway|pattern|test|template|architecture|enforcement) echo " ✅ $f: kind=$KIND" ;;
|
||||
*) echo " ❌ $f: Invalid kind='$KIND'"; FAILED=$((FAILED+1)) ;;
|
||||
esac
|
||||
|
||||
echo "$FM" | grep -q '^name:' || { echo " ❌ $f: Missing name"; FAILED=$((FAILED+1)); }
|
||||
echo "$FM" | grep -q '^description:' || { echo " ❌ $f: Missing description"; FAILED=$((FAILED+1)); }
|
||||
grep -q '^name:' <<< "$FM" || { echo " ❌ $f: Missing name"; FAILED=$((FAILED+1)); }
|
||||
grep -q '^description:' <<< "$FM" || { echo " ❌ $f: Missing description"; FAILED=$((FAILED+1)); }
|
||||
done
|
||||
[ $FAILED -gt 0 ] && { echo "❌ FRONTMATTER FAILED ($FAILED error(s))"; exit 1; }
|
||||
echo "✅ Frontmatter validation passed"
|
||||
|
||||
@@ -55,7 +55,7 @@ Two incidents taught us this:
|
||||
### Stage 3 — AI Review
|
||||
- Diff is sent to `syslog-auto` model via LiteLLM
|
||||
- Review checks against infrastructure-control ground truth:
|
||||
- CT IDs match PVE cluster (100-117, no 122/123)
|
||||
- CT IDs match the PVE cluster inventory in `infrastructure-control.prose.md` Appendix B (100-120 with gaps; no 122/123)
|
||||
- Grafana is direct LAN :3001, NOT behind nginx
|
||||
- Zulip is CT 117 on storepve (bridge IP .19)
|
||||
- Strix Halo :8080 is firewalled to .116 only
|
||||
|
||||
@@ -88,7 +88,7 @@ prose run memory-audit-maintenance memory_threshold=90 verify_configs=true
|
||||
prose run hermes-config-template agent_name=syslog-devops default_model=claude-sonnet-4
|
||||
|
||||
# Configure an agent with a different auxiliary model
|
||||
prose run hermes-config-template agent_name=syslog-code default_model=qwen3.6-27B-code auxiliary_model=gemma-4-12b
|
||||
prose run hermes-config-template agent_name=syslog-code default_model=gpu-dense auxiliary_model=gpu-vision
|
||||
```
|
||||
|
||||
### Option B: Manual Execution
|
||||
@@ -116,7 +116,7 @@ Run on trigger or schedule. Maintain persistent world-model state across runs.
|
||||
| `zulip-health` | Zulip | Checks Zulip connectivity, message flow, and bot responsiveness. |
|
||||
| `zulip-mention-reliability` | Zulip | Diagnoses and fixes @mention detection issues in Zulip. |
|
||||
| `zulip-approval-fix` | Zulip | Fixes broken /approve and /deny slash commands for Hermes agents. |
|
||||
| `litellm-self-heal` | LiteLLM | Consolidated health check + self-healing for the full nginx → LiteLLM → GPU chain. Verifies 8 containers, 3 GPUs, model inference, and agent keys. Applies 9 remediation rules. (litellm-health merged into this contract 2026-07-09.) |
|
||||
| `litellm-self-heal` | LiteLLM | Applies remediation rules for LiteLLM stack failures detected by `litellm-health` (full nginx → LiteLLM → GPU chain). 9 remediation rules. |
|
||||
| `gpu-fleet` | GPU | Manages the GPU inference fleet: model deployment, registration, health checks, LiteLLM sync. |
|
||||
| `gpu-monitor` | GPU | Comprehensive GPU fleet monitor — polls sidecars, router, LiteLLM every 15s, renders SSE dashboard. |
|
||||
| `proxmox-monitor` | Infra | Proxmox cluster + Docker monitoring via the existing Grafana/Prometheus stack on CT 116. |
|
||||
@@ -149,7 +149,7 @@ Called on-demand as single-render tools.
|
||||
| Contract | Description |
|
||||
|---|---|
|
||||
| `litellm-api-keys` | Manages LiteLLM API keys for agent identity. Create, rotate, verify, and list agent keys. References gpu-fleet for current key inventory. |
|
||||
| `litellm-health` | ⚠️ **DEPRECATED** — consolidated into `litellm-self-heal` (2026-07-09). Retained for reference only. |
|
||||
| `litellm-health` | LiteLLM health check: public vs backend surfaces, CT 116 containers, GPU fleet, one model per GPU host, and agent keys. Owner of the probes; `litellm-self-heal` owns remediation. |
|
||||
| `infrastructure-monitoring` | Target-state for Prometheus + GPU exporters + Grafana. Core stack deployed, GPU exporters NOT live. |
|
||||
| `stirling-pdf-agent-access` | Documents the Stirling-PDF API access pattern for agents — global API key, 12 operations, curl examples. Agents use the `stirling-pdf-api` shared skill for templates. |
|
||||
| `hello-world` | Minimal test contract — verifies the OpenProse execution pipeline works. |
|
||||
|
||||
@@ -34,7 +34,7 @@ verification, and DM loopback testing.
|
||||
| @all-bots user ID | 20 | ✅ (config, verified by API at runtime) |
|
||||
| PM2 process name | abiba-zulip | ✅ |
|
||||
| Provider | syslog-harness (http://192.168.68.116/v1) | ✅ |
|
||||
| Default model | deepseek-v4-pro | ✅ (settings.json) |
|
||||
| Default model | syslog-auto | ✅ (settings.json) |
|
||||
|
||||
## Architecture
|
||||
|
||||
|
||||
+85
-18
@@ -36,6 +36,59 @@ def warn(rule, message):
|
||||
WARNINGS.append(f"[{rule}] {message}")
|
||||
|
||||
|
||||
# Derivation rule: a model name is any scalar under a mapping key named `model` or
|
||||
# `model_name`, at any depth. The top-level `model:` SECTION is the exception where the model
|
||||
# name lives under `default`/`model`/`model_name` inside that section, so it is descended
|
||||
# specially. The only other exception is key `models` (litellm key-generation params carry a
|
||||
# list of model names). EXTEND THE ALLOWLIST for a new exception; do NOT add another field by
|
||||
# hand.
|
||||
MODEL_KEYS = ("model", "model_name")
|
||||
MODEL_SECTION_KEYS = ("default", "model", "model_name")
|
||||
MODEL_LIST_KEYS = ("models",)
|
||||
|
||||
|
||||
def _iter_model_values(node, path=""):
|
||||
"""Yield (path, value) for every model-name-bearing scalar in a config."""
|
||||
if isinstance(node, dict):
|
||||
for key, value in node.items():
|
||||
child = f"{path}.{key}" if path else key
|
||||
if key in MODEL_KEYS:
|
||||
if isinstance(value, dict):
|
||||
for subkey in MODEL_SECTION_KEYS:
|
||||
subvalue = value.get(subkey)
|
||||
if isinstance(subvalue, str):
|
||||
yield (f"{child}.{subkey}", subvalue)
|
||||
for subkey, subvalue in value.items():
|
||||
if isinstance(subvalue, (dict, list)):
|
||||
yield from _iter_model_values(subvalue, f"{child}.{subkey}")
|
||||
elif isinstance(value, list):
|
||||
yield from _iter_model_values(value, child)
|
||||
else:
|
||||
yield (child, value)
|
||||
elif key in MODEL_LIST_KEYS:
|
||||
yield from _iter_model_list(value, child)
|
||||
elif isinstance(value, (dict, list)):
|
||||
yield from _iter_model_values(value, child)
|
||||
elif isinstance(node, list):
|
||||
for i, item in enumerate(node):
|
||||
yield from _iter_model_values(item, f"{path}[{i}]")
|
||||
|
||||
|
||||
def _iter_model_list(node, path):
|
||||
"""Yield scalars under an allowlisted `models` key (list of names or list of dicts)."""
|
||||
if isinstance(node, list):
|
||||
for i, item in enumerate(node):
|
||||
yield from _iter_model_list(item, f"{path}[{i}]")
|
||||
elif isinstance(node, dict):
|
||||
for key, value in node.items():
|
||||
if key in MODEL_KEYS and isinstance(value, str):
|
||||
yield (f"{path}.{key}", value)
|
||||
elif isinstance(value, (dict, list)):
|
||||
yield from _iter_model_list(value, f"{path}.{key}")
|
||||
else:
|
||||
yield (path, node)
|
||||
|
||||
|
||||
def audit(path):
|
||||
with open(path) as f:
|
||||
cfg = yaml.safe_load(f)
|
||||
@@ -89,15 +142,16 @@ def audit(path):
|
||||
)
|
||||
|
||||
# --- Rule 8: GPU Workload Distribution ---
|
||||
# gpu-light (and gemma-4-12b) were retired 2026-09-12; the RTX 5070 stable alias is gpu-vision.
|
||||
check(
|
||||
aux.get("vision", {}).get("model") == "gpu-light",
|
||||
aux.get("vision", {}).get("model") == "gpu-vision",
|
||||
"Rule 8",
|
||||
f"auxiliary.vision.model must be gpu-light (got {aux.get('vision', {}).get('model')!r}) — RTX 5070 stable alias",
|
||||
f"auxiliary.vision.model must be gpu-vision (got {aux.get('vision', {}).get('model')!r}) — RTX 5070 stable alias",
|
||||
)
|
||||
check(
|
||||
aux.get("web_extract", {}).get("model") == "gpu-light",
|
||||
aux.get("web_extract", {}).get("model") == "gpu-vision",
|
||||
"Rule 8",
|
||||
f"auxiliary.web_extract.model must be gpu-light (got {aux.get('web_extract', {}).get('model')!r}) — RTX 5070 stable alias",
|
||||
f"auxiliary.web_extract.model must be gpu-vision (got {aux.get('web_extract', {}).get('model')!r}) — RTX 5070 stable alias",
|
||||
)
|
||||
|
||||
# --- Rule 9: Compression Threshold ---
|
||||
@@ -109,7 +163,7 @@ def audit(path):
|
||||
check(
|
||||
comp.get("max_context_window") == 131072,
|
||||
"Rule 9",
|
||||
f"compression.max_context_window must be 131072 (got {comp.get('max_context_window')!r}) — matches 128K GPU capacity",
|
||||
f"compression.max_context_window must be 131072 (got {comp.get('max_context_window')!r}) — syslog-auto pool floor (NVIDIA hosts 128K; Strix Halo 256K)",
|
||||
)
|
||||
|
||||
# --- Rule 10: Default Model Must Be syslog-auto ---
|
||||
@@ -178,21 +232,34 @@ def audit(path):
|
||||
f"custom_providers[0].base_url must end with /v1 (got {cp.get('base_url')!r})",
|
||||
)
|
||||
|
||||
# --- No raw model names (Rule 7/8 spirit) ---
|
||||
raw_names = {"gemma-4-12b", "qwen3.6-27B-code", "qwen3.6-35B-udq4", "ornith-1.0-35b"}
|
||||
for section_path, section_dict in [
|
||||
("model", model), ("compression", comp),
|
||||
("auxiliary.vision", aux.get("vision", {})),
|
||||
("auxiliary.web_extract", aux.get("web_extract", {})),
|
||||
("auxiliary.compression", aux.get("compression", {})),
|
||||
("delegation", deleg),
|
||||
]:
|
||||
m = section_dict.get("model", "")
|
||||
if m in raw_names:
|
||||
# --- Retired/raw model names (Rule 7/8 spirit) ---
|
||||
# The audit's job is to catch configs that are BROKEN, not to enforce a style preference.
|
||||
# NON-RESOLVING names (removed 2026-09-12, verified 400/403 via live LiteLLM) must hard-FAIL:
|
||||
# gpu-light -> gpu-vision ; gemma-4-12b -> gpu-vision
|
||||
# crew-auto -> syslog-auto (its 64K cap is retired; no cap in force) ; ornith-1.0-35b -> strix-moe
|
||||
# RESOLVING names (verified 200) are discouraged but working, so they only WARN:
|
||||
# qwen3.6-27B-code -> gpu-dense ; qwen3.6-35B-udq4 -> strix-moe
|
||||
# Failing a working alias would reject valid configs - the exact defect this change fixes.
|
||||
non_resolving = {
|
||||
"gpu-light": "gpu-vision",
|
||||
"gemma-4-12b": "gpu-vision",
|
||||
"crew-auto": "syslog-auto (its 64K cap is retired; no cap in force)",
|
||||
"ornith-1.0-35b": "strix-moe",
|
||||
"qwen3.6-27B-code": "gpu-dense",
|
||||
"qwen3.6-35B-udq4": "strix-moe",
|
||||
}
|
||||
raw_but_live = {}
|
||||
for field_path, value in _iter_model_values(cfg):
|
||||
if value in non_resolving:
|
||||
check(
|
||||
False,
|
||||
"Rule 7/8",
|
||||
f"{field_path} = {value!r} is retired and no longer resolves (2026-09-12) — use {non_resolving[value]}",
|
||||
)
|
||||
elif value in raw_but_live:
|
||||
warn(
|
||||
"Rule 7/8",
|
||||
f"{section_path}.model = {m!r} — raw model name, use stable alias instead "
|
||||
f"(gpu-light, gpu-dense, strix-moe, syslog-auto)",
|
||||
f"{field_path} = {value!r} is a raw-but-live model name — prefer the stable alias {raw_but_live[value]}",
|
||||
)
|
||||
|
||||
# --- Report ---
|
||||
|
||||
@@ -1,5 +1,5 @@
|
||||
registry_version: 0.1.0
|
||||
last_updated: '2026-07-13T00:00:00Z'
|
||||
last_updated: '2026-09-11T00:00:00Z'
|
||||
updated_by: mumuni
|
||||
categories:
|
||||
- compliance
|
||||
@@ -628,7 +628,7 @@ contracts:
|
||||
sensitivity: high
|
||||
status: active
|
||||
owner: abiba
|
||||
version: 3.1.0
|
||||
version: 3.3.0
|
||||
trigger:
|
||||
type: scheduled
|
||||
cadence: '*/15 * * * *'
|
||||
@@ -693,8 +693,8 @@ contracts:
|
||||
version: 1.0.0
|
||||
trigger:
|
||||
type: scheduled
|
||||
cadence: '*/10 * * * *'
|
||||
description: "Every 10 minutes \u2014 LiteLLM proxy health"
|
||||
cadence: '5 3,7,11,15,19,23 * * *'
|
||||
description: "4-hourly staggered dispatch via fm-send (run contract litellm-health)"
|
||||
cron_job_id: null
|
||||
execution:
|
||||
agent: abiba
|
||||
@@ -750,7 +750,7 @@ contracts:
|
||||
sensitivity: critical
|
||||
status: active
|
||||
owner: abiba
|
||||
version: 1.0.0
|
||||
version: 1.1.0
|
||||
trigger:
|
||||
type: event_driven
|
||||
description: Triggered by relay message from litellm-health or infrastructure-monitoring
|
||||
@@ -1365,7 +1365,7 @@ contracts:
|
||||
sensitivity: high
|
||||
status: active
|
||||
owner: ops
|
||||
version: 1.0.0
|
||||
version: 1.1.0
|
||||
trigger:
|
||||
type: scheduled
|
||||
cadence: 0 2 * * 0
|
||||
|
||||
@@ -2,6 +2,10 @@
|
||||
|
||||
Generated: 2026-07-13 20:59:18 ET
|
||||
|
||||
> **Point-in-time snapshot.** Schedules and cadences are authoritative in
|
||||
> `contract-registry.yaml`; any schedule quoted below may be stale. Do not use
|
||||
> this file as the source of truth for a contract's trigger.
|
||||
|
||||
---
|
||||
|
||||
## hermes-key-enforcement
|
||||
@@ -442,7 +446,7 @@ IMPORTANT: If the contract file does not exist in prose-contracts/main, report f
|
||||
|
||||
## litellm-health
|
||||
|
||||
**Category:** monitoring | **Domain:** litellm | **Owner:** abiba | **Schedule:** */10 * * * *
|
||||
**Category:** monitoring | **Domain:** litellm | **Owner:** abiba | **Schedule:** see contract-registry.yaml (authoritative)
|
||||
|
||||
```
|
||||
Contract Enforcement: litellm-health
|
||||
@@ -450,7 +454,7 @@ Contract Enforcement: litellm-health
|
||||
Category: monitoring
|
||||
Domain: litellm
|
||||
Owner: abiba
|
||||
Schedule: Every 10 minutes — LiteLLM proxy health
|
||||
Schedule: see contract-registry.yaml (authoritative)
|
||||
|
||||
This is a monitoring contract. Execute the monitoring checks defined in the contract. Report any deviations from expected state.
|
||||
|
||||
|
||||
@@ -1,11 +1,14 @@
|
||||
---
|
||||
report_only_agents:
|
||||
- koby # ⛔ KOBY IS NEVER REPAIRED (Rule 17, 2026-08-17) — detect + report, never fix on .129
|
||||
# ⛔ The guest/host-keyed report-only gate the GC executor MUST honour lives in the body
|
||||
# "Hard gate" YAML block below — that block is authoritative and is the only copy.
|
||||
kind: responsibility
|
||||
name: disk-gc-threat-response
|
||||
description: >
|
||||
Recurring disk health scan, garbage collection, and threat response
|
||||
across 15 Proxmox CTs + 3 GPU bare-metal hosts. Triggered by incident
|
||||
across 20 Proxmox guests (17 LXC + 3 QEMU VMs) + 3 GPU bare-metal hosts
|
||||
(fleet verified against `pvesh get /cluster/resources` 2026-09-12). Triggered by incident
|
||||
2026-07-04 where CT 105 (kagentz) hit 87% disk (49G/59G) from
|
||||
Docker image bloat — 5 dangling images, 15 build cache layers.
|
||||
Recovered 35.67GB. Second incident 2026-07-09: amdpve (.15) Docker
|
||||
@@ -25,7 +28,7 @@ logged within 5 minutes of discovery.
|
||||
|
||||
## Scope
|
||||
|
||||
All 15 CTs via `pct-run` + 3 GPU bare-metal hosts via direct SSH.
|
||||
All 20 Proxmox guests (17 LXC via `pct-run` + 3 QEMU VMs via direct SSH) + 3 GPU bare-metal hosts via direct SSH.
|
||||
Docker hosts get special attention:
|
||||
|
||||
| Host | CT | Disk Risk | GC Strategy |
|
||||
@@ -99,35 +102,67 @@ and escalation trail.
|
||||
|
||||
## Execution
|
||||
|
||||
### Hard gate: report-only guests (READ THIS BEFORE RUNNING GC)
|
||||
|
||||
**CT 111 / hostname `tdunna` / 192.168.68.129 is DETECT-AND-REPORT-ONLY.** It belongs to Theo.
|
||||
The captain ruled 2026-08-17 and re-confirmed 2026-09-10 that Theo handles CT 111 himself.
|
||||
At **every** threat level — AMBER, RED, or CRITICAL — the executor must:
|
||||
|
||||
- push the threat row and alert the owner, and
|
||||
- **never** call `gc-executor`, and **never** run any GC command against that guest: no
|
||||
`apt-get clean/autoremove`, no `journalctl --vacuum-*`, no `find /var/log -delete`, no
|
||||
`/tmp`/`/var/tmp` deletion, no snap removal, no `docker system prune`.
|
||||
|
||||
This gate is keyed on **guest id / hostname / IP**, not on an agent name. The frontmatter
|
||||
`report_only_agents` marker (e.g. `koby`) names an AGENT while the scan unit is a GUEST, so an
|
||||
agent-name marker can silently miss the guest it lives on — it must never be the only gate.
|
||||
|
||||
**The authoritative machine-readable exclusion list is the YAML block below.** The executor
|
||||
reads it at run time; `scripts/disk-gc-plan.py` turns a fleet scan into the action plan using it.
|
||||
Extend the list here, never by hand-maintaining a second copy. The Execution loop below MUST
|
||||
call that planner and MUST NOT reimplement the gate.
|
||||
|
||||
```yaml
|
||||
# disk-gc report-only guests — authoritative. Keyed on guest/host, not agent.
|
||||
report_only_guests:
|
||||
- guest: 111
|
||||
hostname: tdunna
|
||||
ip: 192.168.68.129
|
||||
node: storepve
|
||||
reason: "Theo's box — captain ruling 2026-08-17, re-confirmed 2026-09-10"
|
||||
```
|
||||
|
||||
### Loop
|
||||
|
||||
```prose
|
||||
let fleet = call disk-scanner
|
||||
scope: all
|
||||
|
||||
let threats = []
|
||||
for ct in fleet:
|
||||
if ct.usage_pct >= 95:
|
||||
push threats { ct: ct.id, level: "CRITICAL", pct: ct.usage_pct }
|
||||
else if ct.usage_pct >= 85:
|
||||
push threats { ct: ct.id, level: "RED", pct: ct.usage_pct }
|
||||
else if ct.usage_pct >= 75:
|
||||
push threats { ct: ct.id, level: "AMBER", pct: ct.usage_pct }
|
||||
-- The report-only gate is IMPLEMENTED IN scripts/disk-gc-plan.py and MUST NOT be
|
||||
-- reimplemented here. That planner reads the contract's `report_only_guests` YAML block
|
||||
-- and matches on guest id OR hostname OR IP, so the tested gate is the executed gate.
|
||||
let plan = call disk-gc-plan
|
||||
fleet: fleet
|
||||
|
||||
-- sort by severity descending
|
||||
sort threats by pct desc
|
||||
|
||||
for threat in threats:
|
||||
for row in plan:
|
||||
if row.action == "report-only":
|
||||
-- Excluded guest: alert only. No gc-executor call is constructed for it, at any level.
|
||||
call alerter
|
||||
threat: row
|
||||
result: { action: "report-only", reason: row.reason }
|
||||
else:
|
||||
let result = call gc-executor
|
||||
ct: threat.ct
|
||||
level: threat.level
|
||||
strategy: lookup-gc-strategy(threat.ct)
|
||||
ct: row.target
|
||||
level: row.level
|
||||
strategy: lookup-gc-strategy(row.target)
|
||||
|
||||
call alerter
|
||||
threat: threat
|
||||
threat: row
|
||||
result: result
|
||||
|
||||
call summary-reporter
|
||||
fleet: fleet
|
||||
threats: threats
|
||||
plan: plan
|
||||
```
|
||||
|
||||
## GC Strategies by Host Type
|
||||
@@ -239,7 +274,9 @@ dangling images and orphaned build cache. No automated GC was in place.
|
||||
## Incident Log: 2026-07-09 — amdpve docker bloat
|
||||
|
||||
### Discovery
|
||||
Scheduled fleet disk scan via `pct-run` across all 15 CTs + 3 GPU bare-metal hosts.
|
||||
Scheduled fleet disk scan across all 20 Proxmox guests (17 LXC via `pct-run`, 3 QEMU VMs via direct SSH) + 3 GPU bare-metal hosts.
|
||||
> **Report-only gate applies to this scan:** CT 111 (`tdunna`, 192.168.68.129) is alerted but never
|
||||
> garbage-collected at any level.
|
||||
amdpve (.15) flagged at 78% (AMBER threshold: 75%).
|
||||
|
||||
### Diagnosis
|
||||
@@ -272,25 +309,35 @@ one-off GPU builds. No automated post-migration cleanup was in place.
|
||||
- Contract now scans GPU bare-metal hosts alongside CTs
|
||||
- Access via `pct-run` script for all CTs (no hardcoded IPs)
|
||||
|
||||
## Access Matrix (documented 2026-07-09)
|
||||
## Access Matrix (verified against `pvesh get /cluster/resources` 2026-09-12)
|
||||
|
||||
### CT Access (via pct-run)
|
||||
| CT | Name | Node | Status |
|
||||
|----|------|------|--------|
|
||||
| 100 | abiba | minipve | local |
|
||||
| 102 | adguard | minipve | ✅ reachable |
|
||||
| 104 | authentik | minipve | ✅ reachable |
|
||||
| 105 | kagentz | minipve | ✅ reachable |
|
||||
| 106 | ra-h-os | storepve | ✅ reachable |
|
||||
| 107 | pbs | storepve | ✅ reachable |
|
||||
| 108 | media | storepve | ✅ reachable |
|
||||
| 110 | gitea | minipve | ✅ reachable |
|
||||
| 111 | tdunna | amdpve | ✅ reachable |
|
||||
| 112 | tanko | amdpve | ✅ reachable |
|
||||
| 113 | baggy | amdpve | ✅ reachable |
|
||||
| 115 | scottdenya | amdpve | ✅ reachable |
|
||||
| 116 | syslog-api | minipve | ✅ reachable |
|
||||
| 117 | zulip | storepve | ✅ reachable |
|
||||
### Guest Access (via `pct-run` — CT id only, node resolved by `scripts/pct-run.sh`)
|
||||
| Guest | Name | Node | Type | Status |
|
||||
|------|------|------|------|--------|
|
||||
| 100 | abiba | minipve | lxc | ✅ reachable (probed via pct-run like any other guest; no local shortcut) |
|
||||
| 102 | adguard | minipve | lxc | ✅ reachable |
|
||||
| 104 | authentik | minipve | lxc | ✅ reachable |
|
||||
| 105 | kagentz | **amdpve** | lxc | ✅ reachable (was documented as minipve — corrected) |
|
||||
| 106 | ra-h-os | storepve | lxc | ✅ reachable |
|
||||
| 107 | pbs | storepve | lxc | ✅ reachable |
|
||||
| 108 | media | storepve | lxc | ✅ reachable |
|
||||
| 110 | gitea | minipve | lxc | ✅ reachable |
|
||||
| 111 | tdunna | **storepve** | lxc | ⛔ **REPORT-ONLY** (192.168.68.129, Theo's box — no GC at any level) |
|
||||
| 112 | tanko | amdpve | lxc | ✅ reachable |
|
||||
| 113 | baggy | amdpve | lxc | ✅ reachable |
|
||||
| 115 | scottdenya | amdpve | lxc | ✅ reachable |
|
||||
| 116 | syslog-api | minipve | lxc | ✅ reachable |
|
||||
| 117 | zulip | storepve | lxc | ✅ reachable |
|
||||
| 118 | jdownloader | storepve | lxc | ✅ reachable |
|
||||
| 119 | infisical-vault | minipve | lxc | ✅ reachable |
|
||||
| 120 | adguard2 | amdpve | lxc | ✅ reachable |
|
||||
|
||||
### QEMU VMs (via direct SSH)
|
||||
| VM | Name | Node | IP | Status |
|
||||
|----|------|------|-----|--------|
|
||||
| 101 | llm-gpu (workload now bare metal .8) | acerpve | — | ✅ reachable |
|
||||
| 103 | ocu-llm (workload now bare metal .110) | ocupve | — | ✅ reachable |
|
||||
| 109 | docker-vm | storepve | 192.168.68.7 | ✅ reachable |
|
||||
|
||||
### GPU Bare Metal (via direct SSH)
|
||||
| Host | IP | GPU | Status |
|
||||
@@ -299,9 +346,15 @@ one-off GPU builds. No automated post-migration cleanup was in place.
|
||||
| ocu-llm | 192.168.68.110 | RTX 5070 | ✅ reachable |
|
||||
| amdpve | 192.168.68.15 | Strix Halo | ✅ reachable |
|
||||
|
||||
### KVM VM (via direct SSH)
|
||||
| Host | IP | Role | Status |
|
||||
|------|-----|------|--------|
|
||||
| docker-vm | 192.168.68.7 | 16 Docker containers, 4 stacks | ✅ reachable |
|
||||
|
||||
> **Note:** CT 118 is now jdownloader (active on storepve). CT 119 (infisical-vault) added on minipve.\n> **Migrated:** CT 101 → .8, CT 103 → .110 (bare metal GPU).\n> **KVM VM:** CT 109 (docker-vm) is a KVM VM, not LXC — access via SSH .7.
|
||||
> **Fleet count:** 20 Proxmox guests (17 LXC + 3 QEMU VMs) + 3 GPU bare-metal hosts. Corrected
|
||||
> 2026-09-12: CT 105 → amdpve, CT 111 → storepve, and guests 118/119/120 were missing.
|
||||
>
|
||||
> **CT 100 probe gap (folded in):** CT 100 previously reported "unreachable (not reported)" every
|
||||
> run. Root cause is the same stale access layer: `pct-run` resolves the guest's node from its map,
|
||||
> and the map/contract must reflect `pvesh /cluster/resources`. Verified working from inside CT 100:
|
||||
> `scripts/pct-run.sh 100 "df -P / | tail -1"` → `23% /`. Probe CT 100 through `pct-run` like any
|
||||
> other guest — never through a local-only path, since the scanner itself runs inside CT 100 and a
|
||||
> container has no `pct` binary.
|
||||
>
|
||||
> **KVM VM:** CT 109 (docker-vm) is a QEMU VM, not LXC — access via SSH .7.
|
||||
> **NOTE:** For kagentz (CT 105), use `ssh root@kagentz` (hostname), NOT `pct exec 105` — `pct exec 105` shows loop0 (59G) while `ssh root@kagentz` shows the real filesystem (99G). For docker-vm (CT 109), use `ssh root@192.168.68.7`, not `pct exec`.
|
||||
|
||||
@@ -13,7 +13,7 @@ the description:
|
||||
|
||||
1. **What system does this contract touch?** Name the hosts, CTs, containers,
|
||||
and services explicitly. "The inference fleet" is vague. "GPU .8 (RTX 3090,
|
||||
qwen), .110 (RTX 5070, gemma), .15 (Strix Halo, strix-moe), and LiteLLM on CT
|
||||
qwen), .110 (RTX 5070, gpu-vision), .15 (Strix Halo, strix-moe), and LiteLLM on CT
|
||||
116" is specific.
|
||||
|
||||
2. **Who runs this contract, and when?** State the agent, the trigger (cron,
|
||||
|
||||
@@ -261,6 +261,10 @@ repaired.
|
||||
not exist` on .15. `agent-health-check.py` now carries the live-verified
|
||||
`storepve` mapping (the script is not the topology source of truth); the
|
||||
CRITICAL contract itself needs an authorized correction.
|
||||
**✅ Resolved 2026-09-12:** the topology was corrected in its owner,
|
||||
`infrastructure-control.prose.md` (CT 111 → storepve, CT 105 → amdpve), and
|
||||
`scripts/pct-run.sh` now matches. This snapshot is left as observed; treat
|
||||
those owner documents as authoritative.
|
||||
2. **Strix Halo `:8080` firewall claim is stale.** `prose-ai-review.sh`
|
||||
ground-truth rule #4 and `gpu-monitor.prose.md` say `:8080` is firewalled to
|
||||
`.116` only and `.24` cannot probe it. Live on .15:
|
||||
|
||||
+46
-74
@@ -6,13 +6,14 @@ description: >
|
||||
registration, health checks, LiteLLM sync, agent key management, GPU
|
||||
saturation watchdog, Prometheus/Grafana monitoring, and self-healing.
|
||||
UPDATED 2026-07-15: Stable role-based aliases introduced: strix-moe,
|
||||
gpu-dense, gpu-light. These never change — only the underlying model does.
|
||||
Strix Halo: strix-moe → unsloth/Qwen3.6-35B-A3B-MTP (UD-Q4_K_M, 22GB).
|
||||
RTX 5070: gemma-4-12b Q4_K_M → IQ4_NL + MTP draft (122 tok/s, 2x faster).
|
||||
UPDATED 2026-07-17: Context reduced fleet-wide from 256K to 128K for stability.
|
||||
Strix Halo model swapped to qwen3.6-35B-udq4 (22GB, strix-moe alias).
|
||||
Instability observed near 100K at 256K (now all GPUs at 128K). 128K is the stable ceiling.
|
||||
For larger context needs → fall back to external providers (deepseek).
|
||||
gpu-dense, gpu-vision (gpu-light was superseded by gpu-vision on 2026-09-12).
|
||||
These never change — only the underlying model does.
|
||||
Strix Halo: strix-moe → Carnice-Qwen3.6-MoE-35B-A3B-Q4_K_M.gguf (22GB, 256K ctx).
|
||||
RTX 5070: gpu-vision — IQ4_NL + MTP draft (~122 tok/s, 2x faster).
|
||||
UPDATED 2026-07-17: NVIDIA host context reduced from 256K to 128K for stability.
|
||||
Strix Halo model: Carnice-Qwen3.6-MoE-35B-A3B-Q4_K_M.gguf (alias strix-moe, 256K context).
|
||||
Strix Halo runs 256K (n_ctx 262144, --kv-unified); RTX 3090 and RTX 5070 remain at 128K.
|
||||
For >128K on NVIDIA hosts → fall back to external providers (deepseek).
|
||||
VRAM headroom improved: RTX 3090 ~70%, RTX 5070 ~65%.
|
||||
agent: abiba
|
||||
triggers:
|
||||
@@ -24,7 +25,7 @@ triggers:
|
||||
|
||||
## Maintains
|
||||
|
||||
- gpu_roster: { models: map, hosts: map } — Single source of truth for all GPU models
|
||||
- gpu_roster: { models: map, hosts: map } — GPU host/model roster; the authoritative alias/weight/fallback registry is CT 116 `litellm_config.yaml`
|
||||
- router: { status: "healthy", roster_loaded: bool, models: array }
|
||||
- litellm: { status: "healthy", keys: array, models: array }
|
||||
- agent_keys: { agent: api_key } — All agent API keys registered in LiteLLM DB
|
||||
@@ -79,54 +80,28 @@ triggers:
|
||||
Agent configs, cron jobs, and workflows MUST use these aliases, never model-specific names.
|
||||
When a model is swapped on a GPU, ONLY the infrastructure layer changes — agent configs are untouched.
|
||||
|
||||
| Alias | GPU | Current Model | Will Route To |
|
||||
|-------|-----|---------------|---------------|
|
||||
| `strix-moe` | Strix Halo (.15) | qwen3.6-35B-udq4 | Whatever runs on Strix Halo |
|
||||
| `gpu-dense` | RTX 3090 (.8) | Qwen3.8-27B-Uncensored-Q4_K_M | Whatever runs on RTX 3090 |
|
||||
| `gpu-light` | RTX 5070 (.110) | gemma-4-12b | Whatever runs on RTX 5070 |
|
||||
| Alias | Serves | Where | Kind |
|
||||
|-------|--------|-------|------|
|
||||
| `gpu-dense` | heavy reasoning | RTX 3090 (192.168.68.8) | direct alias |
|
||||
| `gpu-vision` | vision / web extract / light tasks | RTX 5070 (192.168.68.110) | direct alias AND `syslog-auto` pool member |
|
||||
| `strix-moe` | compression (MoE) | Strix Halo (192.168.68.15) | direct alias |
|
||||
| `syslog-auto` | balanced default | weighted pool across the three GPU hosts | pool router |
|
||||
|
||||
**Backward compatibility**: Old model-specific names (qwen3.6-27B-code, gemma-4-12b, qwen3.5-9b-it) still work
|
||||
but are deprecated for agent configs. Only the stable aliases survive model swaps.
|
||||
Single source of truth for models, aliases, rpm caps, weights and fallback chains:
|
||||
CT 116 `/opt/inference-harness/litellm_config.yaml`. Do not duplicate those values in
|
||||
contracts — read them there.
|
||||
|
||||
## Current Model Assignments (2026-07-15)
|
||||
|
||||
| Model | GPU | Host | VRAM | Ctx | KV Cache | Parallel | Batch/Ubatch | Status |
|
||||
|-------|-----|------|------|-----|----------|----------|-------------|--------|
|
||||
**No backward compatibility**: Old model-specific names (qwen3.6-27B-code, qwen3.5-9b-it) are retired as of 2026-09-12
|
||||
and no longer resolve; do not use them in agent configs. Only the stable aliases survive model swaps. `gemma-4-12b`
|
||||
is retired and returns 400 `Invalid model name`.
|
||||
|
||||
## Routing Configuration (LiteLLM — July 2026)
|
||||
|
||||
### syslog-auto Weighted Pool (Direct GPU — bypasses router)
|
||||
|
||||
| Model | GPU | Weight | RPM Cap | Timeout |
|
||||
|-------|-----|--------|---------|---------|
|
||||
| Qwen3.8-27B-Uncensored-Q4_K_M | RTX 3090 (.8:8080) | **0.55** | 500 | **300s** |
|
||||
Model, alias, rpm/weight and fallback values are owned by CT 116
|
||||
`/opt/inference-harness/litellm_config.yaml` (see § Stable Role-Based Aliases above).
|
||||
|
||||
Note: All syslog-auto entries route directly to GPUs with `api_key: not-needed`. The router (port 9000) was decommissioned 2026-09-11 and is NOT in the inference path.
|
||||
|
||||
### Direct Model Endpoints
|
||||
|
||||
| Model | RPM Cap | Notes |
|
||||
|-------|---------|-------|
|
||||
|
||||
### Stable Aliases (for agent configs — never change)
|
||||
|
||||
| Alias | RPM Cap | Routes To | Purpose |
|
||||
|-------|---------|-----------|---------|
|
||||
| `strix-moe` | 40 | Strix Halo | Compression tasks (MoE models) |
|
||||
| `gpu-dense` | 500 | RTX 3090 | Heavy reasoning |
|
||||
| `gpu-light` | 500 | RTX 5070 | Vision, web extract, light tasks |
|
||||
|
||||
### Fallback Chains
|
||||
- gemma → qwen
|
||||
- qwen → gemma
|
||||
- strix-moe → qwen → gemma
|
||||
- syslog-auto → qwen → gemma → qwen3.6-35B-udq4
|
||||
|
||||
### Why Strix Halo RPM Is Capped
|
||||
- Direct (strix-moe): 40 RPM (tight) — Strix Halo is shared with compression tasks
|
||||
- Via syslog-auto: 60 RPM (moderate) — prevents flooding when multiple agents use syslog-auto simultaneously
|
||||
- Combined max: ~100 RPM across both paths — Strix Halo can sustain this at 80°C
|
||||
|
||||
## Operations
|
||||
|
||||
### add-model
|
||||
@@ -176,9 +151,7 @@ Show full fleet status: GPUs, models, VRAM, context windows, parallel slots, act
|
||||
2. Check llama-server processes: `ps aux | grep llama-server` on all 3 hosts
|
||||
3. Check LiteLLM: `curl http://192.168.68.116/health` (expect "I'm alive!")
|
||||
4. Check LiteLLM models: `curl -H "Authorization: Bearer $MASTER_KEY" http://192.168.68.116/v1/models`
|
||||
5. Check LiteLLM timeouts: `grep -n 'timeout:' /opt/inference-harness/litellm_config.yaml`
|
||||
- Qwen3.5-9B: 120s, qwen3.6-27B-code: 300s, Carnice-Qwen3.6-MoE-35B-A3B/strix-moe: 300s (strix-moe alias retained, legacy name qwen3.6-35B-udq4 deprecated)
|
||||
- global request_timeout: 300s, nginx proxy_read_timeout: 600s
|
||||
5. Check LiteLLM timeouts: `grep -n 'timeout:' /opt/inference-harness/litellm_config.yaml` — read the live values from the authority config; do not assert them from this contract.
|
||||
6. Check AMD metrics: `curl http://192.168.68.15:9400/metrics` (Radeon 8060S, util%, VRAM, temp, power)
|
||||
7. Check port conflicts: verify only one llama-server on :8080 per host
|
||||
8. Verify agent keys: 9 keys in LiteLLM DB (`GET /key/list`)
|
||||
@@ -217,7 +190,7 @@ If no SSH access, send Zulip DM via abiba-bot with vault update instructions.
|
||||
| `/root/scripts/gpu-saturation-watchdog.py` | pi (.24) | Auto-restart stuck llama-server |
|
||||
| `/root/dashboard/gpu-fleet.html` | pi (.24) | Live HTML dashboard |
|
||||
| `/etc/systemd/system/llama-server.service` | .8, .110 | llama-server daemons (Nvidia GPUs) |
|
||||
| `/etc/systemd/system/strix-server.service` | .15 (amdpve) | llama-server daemon (Vulkan, Strix Halo) running unsloth/Qwen3.6-35B-A3B-MTP-GGUF. Note: `llama-server.service` and `llama-server@.service` are **masked** on .15 to prevent port 8080 collisions. |
|
||||
| `/etc/systemd/system/strix-server.service` | .15 (amdpve) | llama-server daemon (Vulkan, Strix Halo) running Carnice-Qwen3.6-MoE-35B-A3B-Q4_K_M.gguf (256K ctx). Note: `llama-server.service` and `llama-server@.service` are **masked** on .15 to prevent port 8080 collisions. |
|
||||
|
||||
## Prometheus & Grafana
|
||||
|
||||
@@ -234,7 +207,7 @@ If no SSH access, send Zulip DM via abiba-bot with vault update instructions.
|
||||
- **Router startup race**: Compose router.py doesn't call load_roster(). Reload thread sleeps 30s first.
|
||||
Fix: trigger roster reload via SSH after restart, or rebuild image with startup load_roster().
|
||||
- **LiteLLM /metrics**: Requires auth. Prometheus uses `/health/liveliness` as workaround.
|
||||
- **Strix Halo GPU**: Vulkan is the working backend (ROCm/HIP path abandoned — HSA runtime blocked on Debian 13). Build at `/root/llama.cpp/build-vk/`, commit `4fc4ec5` (2026-07-01), ggml 0.15.3 shared-lib arch. Mesa RADV 25.0.7, KHR_coopmat fast path active. ~70 tok/s gen, 532 tok/s prompt. Service: `strix-server.service` on port 8080, model: `qwen3.6-35B-udq4`, alias `strix-moe`, 128K context, flash-attn + q4 KV, multimodal (mmproj loaded).
|
||||
- **Strix Halo GPU**: Vulkan is the working backend (ROCm/HIP path abandoned — HSA runtime blocked on Debian 13). Build at `/root/llama.cpp/build-vk/`, commit `4fc4ec5` (2026-07-01), ggml 0.15.3 shared-lib arch. Mesa RADV 25.0.7, KHR_coopmat fast path active. ~70 tok/s gen, 532 tok/s prompt. Service: `strix-server.service` on port 8080, model: `Carnice-Qwen3.6-MoE-35B-A3B-Q4_K_M.gguf`, alias `strix-moe`, 256K context (n_ctx 262144), --parallel 2 --kv-unified, flash-attn + q4 KV, multimodal (mmproj loaded).
|
||||
- **Port conflict detection (2026-07-05)**: All 3 GPU wrappers now detect ghost processes squatting port 8080 before starting. `.8` and `.110` use inline pre-start check in `llama-wrapper.sh`; `.15` uses `/usr/local/bin/port-cleanup.sh` ExecStartPre. Replaces the blanket `pkill -9 -x llama-server` on .15 which would kill ALL llama-server instances regardless of port. Ghost detection was the root cause of .8 crash-looping for 27+ restarts (stale pid 25836 squatting 8080 after OOM kill).
|
||||
- **Strix Halo thermal safeguard (2026-07-02)**: `strix-server.service` has `-n 8192` (hard generation cap per request). Without it, `--predict` defaults to -1 (infinity) — a runaway request from .123 (old Mumuni CT114 — now inside Abiba CT100 at .24) decoded 39,868 tokens over 24 min, pushing Tctl to 98°C (crit 89.8°C) and throttling 70→29 t/s. The cap bounds worst-case generation to ~5 min. Do NOT remove `-n` without a replacement ceiling. Sustained load hits ~84°C even at 92s; the APU is fanless/low-flow. Clients MUST also set `max_tokens`.
|
||||
- **Port 8080 firewall**: amdpve iptables restricts 8080 to 192.168.68.116 (LiteLLM/router host) only. All inbound connections are from .116 (LiteLLM proxied via nginx). Localhost curls hang (SYN dropped). Always test from .116.
|
||||
@@ -243,16 +216,17 @@ If no SSH access, send Zulip DM via abiba-bot with vault update instructions.
|
||||
- **Alert migration**: All alerts now go to `#agent-hub` topics (`alerts-gpu`, `alerts-pm2`, `alerts-infra`) instead of DMs. Cross-agent visibility enabled.
|
||||
- **tok/s benchmarks**: Measured every 5 min via LiteLLM proxy. Baselines tracked with 30%/50% degradation thresholds.
|
||||
- **NetBird 502**: Tanko routes through NetBird for litellm.sysloggh.net. Use direct IP if NetBird down.
|
||||
- **Alias-retirement sweep (2026-09-12)**: `gemma-4-12b`, `gpu-light` and `crew-auto` are retired and replaced by `gpu-vision` / no cap respectively. The agent-facing templates (`hermes-config-template.prose.md`, `hermes-agent-baseline.prose.md`), `litellm-api-keys.prose.md`, `gpu-self-heal.prose.md`, `hermes-key-enforcement.prose.md`, `inference-optimization.prose.md`, `litellm-client-timeouts.prose.md` and the executable `audit-hermes-config.py` were all updated to the live canonical alias in the same change. **koby's config on .129 still names `gpu-light` (and `gemma-4-E4B`); .129 is report-only, so that is recorded for its owner and NOT edited here.**
|
||||
|
||||
## GPU Inference Benchmarks (Current)
|
||||
|
||||
| GPU | Model | Gen tok/s | Prompt tok/s | Baseline | Context |
|
||||
|-----|-------|-----------|--------------|----------|---------|
|
||||
| RTX 3090 (.8) | Qwen3.8-27B-Uncensored-Q4_K_M | **TBD** | — | — | **128K** |
|
||||
| Strix Halo (.15) | qwen3.6-35B-udq4 | **65** | 140 | — | **128K** |
|
||||
| Strix Halo (.15) | Carnice-Qwen3.6-MoE-35B-A3B-Q4_K_M.gguf (strix-moe) | **65** | 140 | — | **256K** |
|
||||
|
||||
Benchmarks from 2026-07-17. Strix Halo model: qwen3.6-35B-udq4. RTX 5070 MTP provides 2.7x speedup over pre-upgrade 70 tok/s.
|
||||
All 3 GPUs now at 128K context (2026-07-17, reduced from 256K for stability).
|
||||
Benchmarks from 2026-07-17. Strix Halo model: Carnice-Qwen3.6-MoE-35B-A3B-Q4_K_M.gguf (alias strix-moe), n_ctx 262144, --parallel 2 --kv-unified. RTX 5070 MTP provides 2.7x speedup over pre-upgrade 70 tok/s.
|
||||
GPU contexts: RTX 3090 (.8) and RTX 5070 (.110) at 128K; Strix Halo (.15) at 256K (2026-09-12).
|
||||
|
||||
Benchmarks run through LiteLLM proxy (192.168.68.116:4000) every 5 minutes.
|
||||
Degradation alerts fire at 30% (warning) and 50% (critical) below baseline.
|
||||
@@ -265,42 +239,40 @@ History stored at `/root/data/toks-history.json` with 7-day rolling window.
|
||||
### Stable Aliases — CRITICAL
|
||||
|
||||
All agent configs MUST use stable role-based aliases, never model-specific names:
|
||||
- `compression.model: strix-moe` (NOT `qwen3.6-35B-udq4`)
|
||||
- `auxiliary.vision.model: gpu-light` (NOT `gemma-4-12b`)
|
||||
- `delegation.model: gpu-dense` (NOT `qwen3.6-27B-code`)
|
||||
- `auxiliary.web_extract.model: gpu-light`
|
||||
- `compression.model: syslog-auto`
|
||||
- `auxiliary.vision.model: gpu-vision`
|
||||
- `delegation.model: gpu-dense`
|
||||
- `auxiliary.web_extract.model: gpu-vision`
|
||||
|
||||
When the underlying model is swapped, only the LiteLLM config changes — agent configs are untouched.
|
||||
|
||||
### Context Windows
|
||||
- RTX 3090: **128K** (reduced from 256K 2026-07-17) | RTX 5070: **128K** (reduced from 256K) | Strix Halo: **128K**
|
||||
- **All agents**: 128K ceiling — stable margin. For >128K workloads, use external providers (deepseek)
|
||||
- Compression threshold 0.60: fires at ~77K (~51K headroom before 128K ceiling)
|
||||
- RTX 3090: **128K** (reduced from 256K 2026-07-17) | RTX 5070: **128K** (reduced from 256K) | Strix Halo: **256K** (2026-09-12)
|
||||
- **Agents via `syslog-auto`**: 128K ceiling — the pool's safe floor (NVIDIA hosts are 128K). For >128K workloads, use external providers (deepseek)
|
||||
- Compression threshold 0.65 (audit Rule 9): fires at ~85K (~43K headroom before 128K ceiling)
|
||||
- **Pi agents (Abiba)**: `compaction.reserveTokens: 52739` (≈60% of 128K)
|
||||
- Mumuni compression model alias: `strix-moe` with 300s timeout
|
||||
- Mumuni compression model alias: `syslog-auto`
|
||||
|
||||
### Mumuni Agent Profile
|
||||
|
||||
Mumuni (kagentz CT105, 192.168.68.14 — migrated from CT100 2026-08-29) is the primary business assistant. This profile is the reference for all agent configs:
|
||||
Mumuni (kagentz CT105, 192.168.68.14 — migrated from CT100 2026-08-29) is the primary business assistant. This profile is the reference for all agent configs. The compression values below are the current required values per template Rules 7/9 and `audit-hermes-config.py`; whether Mumuni's LIVE config currently complies is a separate operational question.
|
||||
|
||||
| Setting | Value | Notes |
|
||||
|---------|-------|-------|
|
||||
| `model.default` | `syslog-auto` | Weighted pool (55% qwen, 30% strix, 15% gemma) |
|
||||
| `model.default` | `syslog-auto` | Balanced default (pool router) |
|
||||
| `model.provider` | `custom:litellm` | LiteLLM on CT116 |
|
||||
| `compression.model` | `strix-moe` | Stable alias — survives model swaps |
|
||||
| `aux.compression.model` | `strix-moe` | Compression auxiliary model |
|
||||
| `aux.vision.model` | `gpu-light` | Vision tasks (RTX 5070) |
|
||||
| `aux.web_extract.model` | `gpu-light` | Web extraction |
|
||||
| `compression.model` | `syslog-auto` | Rule 7: auto-routing, prevents Strix Halo overload |
|
||||
| `aux.compression.model` | `syslog-auto` | Must match `compression.model` (Rule 7) |
|
||||
| `aux.vision.model` | `gpu-vision` | Vision tasks (RTX 5070) |
|
||||
| `aux.web_extract.model` | `gpu-vision` | Web extraction |
|
||||
| `delegation.model` | `gpu-dense` | Sub-agent reasoning (RTX 3090) |
|
||||
| `context.max_context_window` | 131072 (128K) | Reduced from 256K 2026-07-17 — stable 128K ceiling |
|
||||
| `compression.threshold` | 0.60 | Triggers at ~77K (~60% of 128K) — optimized for 128K context |
|
||||
| `context.max_context_window` | 131072 (128K) | Conservative `syslog-auto` pool floor (NVIDIA hosts 128K; Strix Halo 256K) |
|
||||
| `compression.threshold` | 0.65 | Rule 9: triggers at ~85K for a 128K window |
|
||||
| `compression.target_ratio` | 0.3 | Compresses to ~38K |
|
||||
| `compression.protect_last_n` | 40 | Preserves last 40 messages |
|
||||
| `memory.memory_char_limit` | 800 | Brief memory entries |
|
||||
| `personalities` | `creative` | Creative assistant personality |
|
||||
| Platforms | cli, homeassistant, signal, telegram, zulip | All Hermes platforms |
|
||||
| Main model timeout | 300s | LiteLLM global timeout |
|
||||
| Compression model timeout | 300s | strix-moe timeout increased from 120s |
|
||||
|
||||
### Agent Update Status (2026-07-15)
|
||||
|
||||
|
||||
@@ -27,8 +27,8 @@ agent: abiba
|
||||
┌──────┐ ┌──────┐ ┌────────┐
|
||||
│.8:8080│ │.110 │ │.116:80 │
|
||||
│RTX3090│ │:8080 │ │nginx │
|
||||
│gemma │ │RTX5070│ │router │
|
||||
└──────┘ │qwen27B│ │LiteLLM │
|
||||
│qwen │ │RTX5070│ │router │
|
||||
└──────┘ │vision │ │LiteLLM │
|
||||
└──────┘ │dashboard│
|
||||
└────────┘
|
||||
```
|
||||
|
||||
+13
-12
@@ -12,7 +12,7 @@ description: >
|
||||
Router (port 9000) DECOMMISSIONED 2026-09-11; references replaced with direct GPU routing.
|
||||
Benchmark baselines refreshed to live values.
|
||||
Prometheus exporters removed — not deployed; fall back to direct sidecar probes.
|
||||
Stable role-based aliases (strix-moe, gpu-dense, gpu-light) from gpu-fleet.
|
||||
Stable role-based aliases (strix-moe, gpu-dense, gpu-vision) from gpu-fleet.
|
||||
agent: abiba
|
||||
depends_on:
|
||||
- gpu-monitor.prose.md (live data source on .24:9100)
|
||||
@@ -48,12 +48,12 @@ depends_on:
|
||||
|
||||
| Alias | GPU | Host | Model | VRAM | Ctx | tok/s | Role |
|
||||
|-------|-----|------|-------|------|-----|-------|------|
|
||||
| `strix-moe` | Strix Halo 64GB | ct15 (.15:8080) | qwen3.6-35B-udq4 | ~10/64GB (16%) | 128K | 62.9 | Compression, summarization, long docs |
|
||||
| `strix-moe` | Strix Halo 64GB | ct15 (.15:8080) | Carnice-Qwen3.6-MoE-35B-A3B-Q4_K_M.gguf | ~10/64GB (16%) | 256K | 62.9 | Compression, summarization, long docs |
|
||||
|
||||
Key notes:
|
||||
- All models use direct GPU routing via LiteLLM (`api_key: not-needed`). Router (port 9000) was decommissioned 2026-09-11 and is NOT in the inference path.
|
||||
- Stable aliases (gpu-dense, gpu-light, strix-moe) from gpu-fleet are the canonical names for agent configs. Model-specific names still work but are deprecated.
|
||||
- RTX 5070 tok/s is ~145 for Qwen3.5-9B — gpu-light is the fastest endpoint. Route vision/web/light work there first. NOTE: Qwen3.5-9B is multimodal (image+text), NOT text-only like gemma-4-12b was.
|
||||
- Stable aliases (gpu-dense, gpu-vision, strix-moe) from gpu-fleet are the canonical names for agent configs — use these, not model-specific names. The retired names `gpu-light` and `gemma-4-12b` were superseded by `gpu-vision` on 2026-09-12 and no longer resolve (400 `Invalid model name`).
|
||||
- The RTX 5070 is the fastest endpoint per token — `gpu-vision` is its canonical alias. Route vision/web/light work there first. The RTX 5070 model is multimodal (image+text).
|
||||
- Strix Halo is 62.9 tok/s (89% of 70.5 baseline) — below optimal but stable. Check for competing workloads.
|
||||
- RTX 3090 VRAM at 88% — within role-appropriate range (role = heavy reasoning, needs the headroom).
|
||||
- RTX 5070 VRAM at 83% — role-appropriate for vision/web (smaller batch sizes).
|
||||
@@ -64,7 +64,7 @@ Key notes:
|
||||
- **Detect**: Any GPU temp >85°C sustained for 2+ consecutive polls
|
||||
- **Fix**:
|
||||
1. Reduce inference concurrency on that GPU (load-side cooling only — NO fan control)
|
||||
2. Redirect new requests to cooler GPUs via LiteLLM fallback chains (gemma → qwen, qwen → gemma)
|
||||
2. Redirect new requests to cooler GPUs via LiteLLM fallback chains (gpu-vision → gpu-dense, gpu-dense → gpu-vision)
|
||||
3. If all GPUs hot, alert about cooling infrastructure
|
||||
- **Verify**: Temp drops below 80°C within 5 minutes
|
||||
- **Escalate after**: 3 verification failures → Zulip alert
|
||||
@@ -142,10 +142,10 @@ Key notes:
|
||||
- **Escalate**: If Tier 2 triggers and temp still rising after 5 min → possible hardware failure
|
||||
|
||||
### Rule 9: Context Window Optimization
|
||||
- **Detect**: Benchmark tok/s vs baseline for each GPU at current context (all 128K)
|
||||
- **Detect**: Benchmark tok/s vs baseline for each GPU at its current context
|
||||
- RTX 3090 (128K ctx, ThinkingCap): baseline 74.8 tok/s — currently at 74.9 (100%)
|
||||
- RTX 5070 (128K ctx, HauhauCS QAT): baseline 165.2 tok/s — currently at 169.6 (103%)
|
||||
- Strix Halo (128K ctx, qwen3.6-35B-udq4): baseline 70.5 tok/s — currently at 62.9 (89%)
|
||||
- Strix Halo (256K ctx, Carnice-Qwen3.6-MoE-35B-A3B-Q4_K_M.gguf): baseline 70.5 tok/s — currently at 62.9 (89%)
|
||||
- **Fix**:
|
||||
- If tok/s > baseline → context has headroom, consider increasing
|
||||
- If tok/s < 90% baseline → reduce context by 25% and retest
|
||||
@@ -157,13 +157,14 @@ Key notes:
|
||||
### Rule 10: Workload Distribution Optimization (updated 2026-07-18)
|
||||
- **Detect**: GPU roles misaligned with hardware capabilities
|
||||
- **Target distribution**:
|
||||
- RTX 3090 (gpu-dense, 24GB, 74.9 tok/s) → Heavy reasoning, code gen, long conversations (slowest per-token but largest context capacity). Weight: 0.55 (LiteLLM).
|
||||
- RTX 5070 (gpu-light, 12GB, ~145 tok/s) → Vision (image+text), web search, lightweight tasks. Weight: 0.15 (LiteLLM).
|
||||
- Strix Halo (strix-moe, 64GB, 62.9 tok/s) → Context compression, summarization, long docs (MoE model). Weight: 0.30 (LiteLLM).
|
||||
- RTX 3090 (gpu-dense, 24GB, 74.9 tok/s) → Heavy reasoning, code gen, long conversations (slowest per-token but largest context capacity).
|
||||
- RTX 5070 (gpu-vision, 12GB, ~145 tok/s) → Vision (image+text), web search, lightweight tasks.
|
||||
- Strix Halo (strix-moe, 64GB, 62.9 tok/s) → Context compression, summarization, long docs (MoE model).
|
||||
- **Note**: RTX 5070 is the fastest endpoint per token. Route high-volume, low-complexity work there first.
|
||||
- **Weights are not restated here** — the live `syslog-auto` pool weights and rpm caps live in CT 116 `/opt/inference-harness/litellm_config.yaml`, the single source of truth.
|
||||
- **Fix**:
|
||||
- Alert if any GPU is handling workload outside its designated role
|
||||
- Recommend agent alias updates to match workload to GPU role (use stable aliases: gpu-dense, gpu-light, strix-moe)
|
||||
- Recommend agent alias updates to match workload to GPU role (use stable aliases: gpu-dense, gpu-vision, strix-moe)
|
||||
- Track per-GPU request distribution via LiteLLM spend logs
|
||||
- **Verify**: Each GPU's request pattern matches its designated role within 24h
|
||||
- **Escalate**: If role mismatch persists >48h → agent alias audit needed
|
||||
@@ -311,6 +312,6 @@ Pushed to `SyslogSolution/health-logs/gpu/{run_id}.json` — versioned, searchab
|
||||
If monitor response > 1MB, log a warning and skip the cycle rather than crashing.
|
||||
|
||||
### L6: Stable Aliases Replace Model Names
|
||||
- gpu-fleet introduced stable aliases (strix-moe, gpu-dense, gpu-light) on 2026-07-15.
|
||||
- gpu-fleet introduced stable aliases (strix-moe, gpu-dense, gpu-light) on 2026-07-15; `gpu-light` was superseded by `gpu-vision` on 2026-09-12.
|
||||
- Self-heal must use aliases for reporting and alerting, not model-specific names.
|
||||
- **Rule**: All alert messages and KG nodes use the stable alias as the GPU identifier.
|
||||
|
||||
@@ -5,7 +5,7 @@ version: 1.0.0
|
||||
description: >
|
||||
Canonical known-good baseline for all Syslog Hermes agents. Captures the exact
|
||||
configuration state, keys, workarounds, and audit procedure. When an agent's
|
||||
configuration goes sideways, restore from this baseline. Last verified 2026-07-16. All GPUs 128K context (reduced from 256K for stability Jul 2026) (RTX 3090 .8, RTX 5070 .110, Strix Halo .15). Parallel 1 fleet-wide (Strix Halo handles compression solo).
|
||||
configuration goes sideways, restore from this baseline. Last verified 2026-07-16. RTX 3090/5070 at 128K (reduced from 256K for stability Jul 2026); Strix Halo at 256K (2026-09-12) (RTX 3090 .8, RTX 5070 .110, Strix Halo .15). Parallel 1 fleet-wide (Strix Halo handles compression solo).
|
||||
author: Abiba (pi agent)
|
||||
---
|
||||
|
||||
@@ -24,11 +24,12 @@ done
|
||||
|
||||
| Agent | CT | Node | IP | LiteLLM Alias | Key Source | Platform |
|
||||
|-------|-----|------|-----|---------------|------------|----------|
|
||||
| Koby | 111 | amdpve | .129 | `koby` | Infisical vault | **Hermes** |
|
||||
| Koby | 111 | storepve | .129 | `koby` | Infisical vault | **Hermes** |
|
||||
| Koonimo | 113 | amdpve | .114 | `koonimo` | Infisical vault | Hermes |
|
||||
| Shumba | — | 192.168.68.119 | N/A | N/A (DeepSeek) | Hermes (RETIRED — CT119 now Infisical vault) |
|
||||
|
||||
> **Note**: CT hostnames (tdunna→CT111, baggy→CT113) differ from agent identities (koby, koonimo).
|
||||
> CT 111 (tdunna, 192.168.68.129, storepve) is report-only — Theo's box; alert only, never garbage-collect.
|
||||
|
||||
Access: `pct-run <CT_ID> <command>` — no IPs needed. GPU hosts (.8, .110, .15) use SSH.
|
||||
Keys are stored in Infisical vault (project=agents, env=production) and injected at
|
||||
@@ -80,7 +81,7 @@ custom_providers:
|
||||
auxiliary:
|
||||
vision:
|
||||
provider: harness
|
||||
model: gemma-4-12b # or syslog-auto
|
||||
model: gpu-vision # RTX 5070 stable alias (Rule 8; do not use syslog-auto for aux)
|
||||
base_url: http://192.168.68.116/litellm/v1
|
||||
api_key_env: LITELLM_API_KEY
|
||||
api_key: <value from: infisical secrets get LITELLM_API_KEY --project=agents --env=production> # ← MANDATORY workaround
|
||||
@@ -95,7 +96,7 @@ auxiliary:
|
||||
threshold: 0.65
|
||||
target_ratio: 0.3
|
||||
provider: harness
|
||||
model: syslog-auto # or gemma-4-12b
|
||||
model: syslog-auto # Rule 7: compression must be syslog-auto
|
||||
base_url: http://192.168.68.116/litellm/v1
|
||||
api_key_env: LITELLM_API_KEY
|
||||
api_key: <value from: infisical secrets get LITELLM_API_KEY --project=agents --env=production> # ← MANDATORY workaround
|
||||
@@ -185,7 +186,9 @@ Abiba (CT100) runs pi via PM2 with the Zulip extension.
|
||||
Config files: `~/.pi/agent/models.json`, `~/.pi/agent/settings.json`.
|
||||
|
||||
**models.json** — Must only list models authorized for the agent's LiteLLM key.
|
||||
Key is injected via `infisical run --` wrapper at PM2 startup:
|
||||
`/v1/models` is key-scoped and the live registry is CT 116
|
||||
`/opt/inference-harness/litellm_config.yaml`; treat the list below as a snapshot and re-read
|
||||
the registry before applying. Key is injected via `infisical run --` wrapper at PM2 startup:
|
||||
```json
|
||||
{
|
||||
"providers": {
|
||||
@@ -197,9 +200,7 @@ Key is injected via `infisical run --` wrapper at PM2 startup:
|
||||
{ "id": "syslog-auto" },
|
||||
{ "id": "strix-moe" },
|
||||
{ "id": "gpu-dense" },
|
||||
{ "id": "gpu-light" },
|
||||
{ "id": "qwen3.6-27B-code" },
|
||||
{ "id": "gemma-4-12b" }
|
||||
{ "id": "gpu-vision" }
|
||||
]
|
||||
}
|
||||
}
|
||||
|
||||
@@ -92,17 +92,17 @@ work immediately after restart.
|
||||
```yaml
|
||||
# ─── Model Selection ───
|
||||
model:
|
||||
default: <agent_model> # e.g., strix-moe, qwen3.6-27B-code, syslog-auto
|
||||
default: <agent_model> # e.g., strix-moe, gpu-dense, syslog-auto
|
||||
provider: harness
|
||||
base_url: http://192.168.68.116/litellm/v1 # Rule 5 (2026-08-09): canonical authenticated path; /v1 also OK
|
||||
api_key_env: LITELLM_API_KEY # Injected via infisical run -- wrapper
|
||||
max_tokens: 4096 # ⚠️ CRITICAL: Prevents unbounded generation
|
||||
context_length: 131072 # For syslog-auto (all GPUs at 128K for stability).
|
||||
context_length: 131072 # Conservative floor for syslog-auto (NVIDIA hosts 128K; Strix Halo 256K).
|
||||
# ⚠️ MANDATORY: Hermes probes unknown models from 256K
|
||||
# and falls back to 256K when /v1/models lacks a context
|
||||
# field (llama-server does). Without this override, agents
|
||||
# silently run syslog-auto at 256K (verified 2026-08-09).
|
||||
# Set 65536 if using gemma-4-12b directly (tight VRAM).
|
||||
# Set 65536 if pinning a single model directly (tight VRAM).
|
||||
|
||||
fallback_providers:
|
||||
provider: deepseek
|
||||
@@ -133,7 +133,7 @@ compression:
|
||||
enabled: true
|
||||
model: syslog-auto # ⚠️ Must match auxiliary.compression.model. Stable alias (gpu-fleet § Stable Role-Based Aliases). NOT ornith-1.0-35b (LiteLLM does not serve that name).
|
||||
provider: harness
|
||||
max_context_window: 131072 # MUST match actual GPU capacity. All 3 GPUs are 128K (Jul 17).
|
||||
max_context_window: 131072 # MUST stay at the syslog-auto pool floor: NVIDIA hosts are 128K, Strix Halo 256K (2026-09-12).
|
||||
threshold: 0.65 # Fires at ~170K for 262K window, ~85K for 128K
|
||||
target_ratio: 0.30
|
||||
protect_last_n: 40
|
||||
@@ -143,25 +143,25 @@ compression:
|
||||
|
||||
# ─── Auxiliary Tasks (CONSISTENCY RULE) ───
|
||||
# All auxiliary services MUST use identical model, base_url, and api_key_env:
|
||||
# model: gpu-light # stable alias (NOT raw "gemma-4-12b")
|
||||
# model: gpu-vision # stable alias (NOT a raw model name)
|
||||
# base_url: http://192.168.68.116/litellm/v1 # Rule 5 (2026-08-09): canonical authenticated; /v1 also OK
|
||||
# api_key_env: LITELLM_API_KEY
|
||||
# Do NOT use syslog-auto for auxiliary tasks — it routes to the primary GPU.
|
||||
# gpu-light = RTX 5070 (12B), freeing the Strix Halo for agent reasoning.
|
||||
# gpu-vision = RTX 5070 (12B), freeing the Strix Halo for agent reasoning.
|
||||
# Heavy aux (delegation, x_search) use gpu-dense (RTX 3090) instead.
|
||||
# NEVER use raw model names (gemma-4-12b, qwen3.6-27B-code, qwen3.6-35B-udq4)
|
||||
# in agent configs — use the stable aliases so model swaps don't break agents.
|
||||
# NEVER use retired model names (qwen3.6-27B-code, qwen3.6-35B-udq4; gemma-4-12b is retired
|
||||
# and no longer resolves) in agent configs — use the stable aliases so model swaps don't break agents.
|
||||
auxiliary:
|
||||
vision:
|
||||
provider: harness
|
||||
model: gpu-light # stable alias for RTX 5070 (was raw gemma-4-12b)
|
||||
model: gpu-vision # stable alias for RTX 5070
|
||||
base_url: http://192.168.68.116/litellm/v1 # Rule 5 (2026-08-09): canonical authenticated; /v1 also OK
|
||||
api_key_env: LITELLM_API_KEY
|
||||
timeout: 60
|
||||
download_timeout: 30
|
||||
web_extract:
|
||||
provider: harness
|
||||
model: gpu-light # stable alias for RTX 5070
|
||||
model: gpu-vision # stable alias for RTX 5070
|
||||
base_url: http://192.168.68.116/litellm/v1 # Rule 5 (2026-08-09): canonical authenticated; /v1 also OK
|
||||
api_key_env: LITELLM_API_KEY
|
||||
timeout: 30
|
||||
@@ -173,10 +173,10 @@ auxiliary:
|
||||
timeout: 300 # gpu-fleet: 300s for large-history summarization (was 60)
|
||||
|
||||
# ─── Delegation / Heavy Aux (use gpu-dense = RTX 3090) ───
|
||||
# delegation.model and x_search.model use gpu-dense (NOT raw qwen3.6-27B-code).
|
||||
# delegation.model and x_search.model use gpu-dense (NOT retired raw name).
|
||||
|
||||
delegation:
|
||||
model: gpu-dense # stable alias for RTX 3090 (was raw qwen3.6-27B-code)
|
||||
model: gpu-dense # stable alias for RTX 3090
|
||||
provider: harness
|
||||
base_url: http://192.168.68.116/litellm/v1 # Rule 5 (2026-08-09): canonical authenticated; /v1 also OK
|
||||
api_key_env: LITELLM_API_KEY
|
||||
@@ -252,10 +252,11 @@ The following MUST be identical across ALL profiles:
|
||||
- For agents needing longer outputs: raise to 8192, but never omit
|
||||
|
||||
### Rule 7: Auxiliary Model Consistency (UPDATED 2026-07-16)
|
||||
- Vision and web_extract use `gemma-4-12b` (RTX 5070 — 12GB, vision-optimized)
|
||||
- Compression uses `strix-moe` (stable alias for Strix Halo — 64GB, 128K ctx, compression-optimized)
|
||||
- **`strix-moe` is the only valid compression model name** — LiteLLM does NOT serve `ornith-1.0-35b`
|
||||
(it serves `strix-moe`, `qwen3.6-35B-udq4`, `gpu-dense`, `gpu-light`, `syslog-auto`, `gemma-4-12b`, `qwen3.6-27B-code`). Old configs with `ornith-1.0-35b` cause 403/model-not-found on compression calls.
|
||||
- Vision and web_extract use `gpu-vision` (RTX 5070 — 12GB, vision-optimized)
|
||||
- Compression uses `syslog-auto` — the Strix Halo weighted pool (64GB, 256K ctx, compression-optimized); do NOT pin `compression.model` to `strix-moe` (audit Rule 7 rejects it)
|
||||
- **`ornith-1.0-35b` is NOT a valid compression model name** — LiteLLM does not serve it
|
||||
(do not restate the served model list here — CT 116 `/opt/inference-harness/litellm_config.yaml`
|
||||
is the single source of truth for models, aliases, weights and fallbacks). Old configs with `ornith-1.0-35b` cause 403/model-not-found on compression calls.
|
||||
- **OPERATIONAL DECISION (2026-07-23): Use `syslog-auto` for compression across all agents.**
|
||||
The `syslog-auto` alias routes to the Strix Halo, but uses the weighted pool instead of pinning
|
||||
to `strix-moe` directly. This prevents sustained Strix Halo thermal load because the pool can
|
||||
@@ -264,41 +265,41 @@ The following MUST be identical across ALL profiles:
|
||||
- All auxiliary services MUST use identical routing:
|
||||
- `base_url: http://192.168.68.116/litellm/v1` (Rule 5, 2026-08-09: canonical authenticated; `/v1` also OK)
|
||||
- `api_key_env: LITELLM_API_KEY`
|
||||
- **Do NOT use `syslog-auto`** for auxiliary tasks — it routes unpredictably
|
||||
- **Do NOT use `syslog-auto` for `vision`/`web_extract`** — it routes unpredictably; compression is the deliberate exception (see the OPERATIONAL DECISION above)
|
||||
- **Compression on Strix Halo**: The strix-moe alias routes to Strix Halo
|
||||
(64GB UMA, 128K context) — the designated compression GPU. This frees the
|
||||
(64GB UMA, 256K context) — the designated compression GPU. This frees the
|
||||
RTX 5070 for vision and web search, and the RTX 3090 for heavy reasoning.
|
||||
- The `compression:` block's `model` MUST match `auxiliary: compression: model`
|
||||
- The `compression: max_context_window: 131072` MUST match actual GPU capacity (128K)
|
||||
- The `compression: max_context_window: 131072` MUST stay at the syslog-auto pool floor (NVIDIA hosts 128K; Strix Halo 256K)
|
||||
|
||||
### Rule 8: GPU Workload Distribution (UPDATED 2026-07-16)
|
||||
- **RTX 3090 (24GB, 128K ctx, qwen3.6-27B-code)**: Heavy reasoning, code gen, long conversations
|
||||
- **RTX 5070 (12GB, 128K ctx, gemma-4-12b)**: Vision, web search, quick tasks, web_extract (IQ4_NL+MTP, ~65% VRAM at 128K)
|
||||
- **Strix Halo (64GB, 128K ctx, syslog-auto)**: Context compression, summarization, long docs
|
||||
- **RTX 3090 (24GB, 128K ctx, gpu-dense)**: Heavy reasoning, code gen, long conversations
|
||||
- **RTX 5070 (12GB, 128K ctx, gpu-vision)**: Vision, web search, quick tasks, web_extract (IQ4_NL+MTP, ~65% VRAM at 128K)
|
||||
- **Strix Halo (64GB, 256K ctx, syslog-auto)**: Context compression, summarization, long docs
|
||||
- Agent profiles MUST route auxiliary tasks to the correct GPU:
|
||||
- `auxiliary.vision.model: gemma-4-12b` (RTX 5070)
|
||||
- `auxiliary.web_extract.model: gemma-4-12b` (RTX 5070)
|
||||
- `auxiliary.vision.model: gpu-vision` (RTX 5070)
|
||||
- `auxiliary.web_extract.model: gpu-vision` (RTX 5070)
|
||||
- `auxiliary.compression.model: syslog-auto` (Strix Halo)
|
||||
- Default model (`model.default`) and custom_provider remain `syslog-auto` for auto-routing
|
||||
- For 128K context window: `threshold: 0.65` (fires at ~85K tokens)
|
||||
- Do NOT use `threshold: 0.25` — this fires at 65K, causing premature context loss
|
||||
- Do NOT use `threshold: 0.80` — this delays until ~105K, leaving only 23K margin
|
||||
- `max_context_window: 131072` MUST match the model's actual capacity (128K)
|
||||
- `max_context_window: 131072` MUST stay at the pool floor (NVIDIA hosts 128K; Strix Halo 256K)
|
||||
- See `devops-hermes-compression` skill for full reference
|
||||
|
||||
### Rule 9: Compression Threshold for 128K Models
|
||||
- For 128K context window: `threshold: 0.65` (fires at ~85K tokens)
|
||||
- Do NOT use `threshold: 0.25` — this fires at 65K, causing premature context loss
|
||||
- Do NOT use `threshold: 0.80` — this delays until ~105K, leaving only 23K margin
|
||||
- `max_context_window: 131072` MUST match the model's actual capacity (all GPUs = 128K)
|
||||
- `max_context_window: 131072` MUST stay at the pool floor (NVIDIA hosts 128K; Strix Halo 256K)
|
||||
- See `devops-hermes-compression` skill for full reference
|
||||
|
||||
### Rule 10: Default Model Must Be `syslog-auto` (All Agents)
|
||||
- **Koby Exception**: Per captain ruling 2026-08-11, Koby is a DeepSeek-primary external agent; its primary model remains `deepseek-v4-flash` (via api.deepseek.com to preserve DeepSeek-specific reasoning, while other sections follow Rule 10.
|
||||
- **Hermes agents**: `model.default: syslog-auto`, `custom_providers[0].model: syslog-auto`
|
||||
- **pi agents**: `defaultModel: syslog-auto` in `settings.json`, first model in `models.json`
|
||||
- `syslog-auto` is the LiteLLM routing model — it load-balances between strix-moe
|
||||
and qwen3.6-27B-code, with gemma-4-12b as fallback. Using it protects against:
|
||||
- `syslog-auto` is the LiteLLM routing model — it load-balances across the live pool
|
||||
(see CT 116 `/opt/inference-harness/litellm_config.yaml` for the current members and weights). Using it protects against:
|
||||
- Model name typos that cause 403 errors and silent worker failures
|
||||
- Single GPU downtime (routing falls back automatically)
|
||||
- Key/model authorization mismatches
|
||||
@@ -321,7 +322,8 @@ When an agent shows "context issues" (premature compression, 401s, 504s, DeepSee
|
||||
verify ALL FOUR of these against the live config. They are the only root causes found in production:
|
||||
|
||||
1. **max_context_window correct?** — BOTH `compression.max_context_window` AND `context.max_context_window`
|
||||
MUST be `131072` (all GPUs are 128K). A value of `262144` causes instability near 100K and must NOT be used.
|
||||
MUST be `131072` (the syslog-auto pool floor: NVIDIA hosts are 128K; Strix Halo is 256K).
|
||||
A `262144` client window can route to a 128K NVIDIA host and fail, so it must NOT be used.
|
||||
~83K instead of ~170K. Check: `grep -n max_context_window ~/.hermes/config.yaml`
|
||||
2. **base_url uses authenticated path?** — `custom_providers[0].base_url`, `delegation.base_url`,
|
||||
and ALL `auxiliary.*.base_url` MUST be `http://192.168.68.116/litellm/v1` (Rule 5, canonical)
|
||||
|
||||
@@ -169,16 +169,20 @@ The agent picks up the new key via `infisical run --` at gateway startup.
|
||||
|
||||
**Keys are permanent and use bare agent name aliases.**
|
||||
|
||||
- **Duration**: `null` — keys never expire. This is enforced by `default_key_generate_params` in `litellm_config.yaml`.
|
||||
- **Duration**: `null` — keys never expire. NOT enforced today: CT 116 `litellm_config.yaml` has no `default_key_generate_params` block, and a key generated with no explicit models comes back with an empty models list. OPEN policy question: should agent keys expire by default? (captain security-policy decision, raised separately.)
|
||||
- **Alias convention**: bare agent name only (e.g., `tanko`, `mumuni`, `koby`, `koonimo`). No dates, no versions. The alias IS the identity.
|
||||
- **Rotation triggers**: compromise, personnel departure, or quarterly security hygiene. NOT calendar-driven.
|
||||
- **Max budget**: $100 per key (config default).
|
||||
|
||||
```yaml
|
||||
# In litellm_config.yaml — ensures all future keys inherit these defaults:
|
||||
# NOT currently set in the authority; recommended value. CT 116 litellm_config.yaml has no
|
||||
# default_key_generate_params block today, and a key generated with no explicit models comes back
|
||||
# with an EMPTY models list. `models` is a literal key-generation parameter, so this is a value to
|
||||
# ADD — re-read the live registry at CT 116 /opt/inference-harness/litellm_config.yaml and
|
||||
# re-verify before applying.
|
||||
litellm_settings:
|
||||
default_key_generate_params:
|
||||
models: ["syslog-auto", "qwen3.6-27B-code", "gemma-4-12b"]
|
||||
models: ["syslog-auto", "gpu-dense", "gpu-vision", "strix-moe"]
|
||||
duration: null # ← permanent
|
||||
max_budget: 100
|
||||
metadata:
|
||||
@@ -286,13 +290,13 @@ auxiliary:
|
||||
api_key: sk-<agent-key-from-vault> # ← workaround (get via: infisical secrets get LITELLM_API_KEY --project=agents --env=production --plain)
|
||||
api_key_env: LITELLM_API_KEY
|
||||
base_url: http://192.168.68.116/litellm/v1
|
||||
model: gemma-4-12b
|
||||
model: gpu-vision
|
||||
provider: harness
|
||||
compression:
|
||||
api_key: sk-<agent-key-from-vault> # ← workaround (same as above)
|
||||
api_key_env: LITELLM_API_KEY
|
||||
base_url: http://192.168.68.116/litellm/v1
|
||||
model: gemma-4-12b
|
||||
model: syslog-auto
|
||||
provider: harness
|
||||
```
|
||||
|
||||
|
||||
@@ -47,7 +47,7 @@ connectivity recovery including end-to-end DM validation.
|
||||
|
||||
## Requires
|
||||
|
||||
- SSH access to target host (direct or via amdpve for CTs)
|
||||
- SSH access to target host (direct, or via the guest's Proxmox node for CTs)
|
||||
- Git repo at `https://git.sysloggh.net/SyslogSolution/zulip-platform-plugins.git`
|
||||
- Python 3 with `httpx` installed on target
|
||||
|
||||
@@ -56,7 +56,7 @@ connectivity recovery including end-to-end DM validation.
|
||||
| Host | CT | Proxmox | IP (direct) | Hermes Home | User |
|
||||
|------|-----|---------|-------------|-------------|------|
|
||||
| Tanko | CT112 | amdpve | 192.168.68.122 | /home/jerome/.hermes | jerome | *(DSH since 2026-08-27 — historical, plugin retired on this host)* |
|
||||
| Koby | CT111 | amdpve | 192.168.68.129 | /root/.hermes | root |
|
||||
| Koby | CT111 | storepve | 192.168.68.129 | /root/.hermes | root |
|
||||
| Shumba | — | — | 192.168.68.119 | /home/lucky/.hermes | lucky |
|
||||
|
||||
| Field | Value | Trust |
|
||||
@@ -72,7 +72,7 @@ connectivity recovery including end-to-end DM validation.
|
||||
### Step 1: Resolve Target
|
||||
|
||||
Map `target` to host, CT ID, hermes_home, and user from the live-state table.
|
||||
For CT112 and CT111, route through `ssh root@amdpve` then `pct exec <id>`.
|
||||
For CT112 route through `ssh root@amdpve`; for CT111 route through `ssh root@storepve` — then `pct exec <id>`.
|
||||
|
||||
### Step 2: Pull Latest Plugin Source
|
||||
|
||||
|
||||
@@ -43,7 +43,7 @@ gateway restart, and connection validation.
|
||||
|
||||
## Requires
|
||||
|
||||
- SSH access to target host (direct or via amdpve for CTs)
|
||||
- SSH access to target host (direct, or via the guest's Proxmox node for CTs)
|
||||
- Git repo at `https://git.sysloggh.net/SyslogSolution/zulip-platform-plugins.git`
|
||||
- Python 3 with `httpx` installed on target
|
||||
- Zulip server accessible at `https://chat.sysloggh.net`
|
||||
@@ -52,7 +52,7 @@ gateway restart, and connection validation.
|
||||
|
||||
| Host | CT | Proxmox | IP (direct) | Hermes Home | User |
|
||||
|------|-----|---------|-------------|-------------|------|
|
||||
| Koby | CT111 | amdpve | 192.168.68.129 | /root/.hermes | root |
|
||||
| Koby | CT111 | storepve | 192.168.68.129 | /root/.hermes | root |
|
||||
| Shumba | — | — | 192.168.68.119 | /home/lucky/.hermes | lucky |
|
||||
|
||||
| Field | Value | Trust |
|
||||
@@ -67,7 +67,7 @@ gateway restart, and connection validation.
|
||||
### Step 1: Locate Target
|
||||
|
||||
Map `target` to connectivity parameters from the live-state table above.
|
||||
For CT112 and CT111, route through `ssh root@amdpve` then `pct exec <id>`.
|
||||
For CT112 route through `ssh root@amdpve`; for CT111 route through `ssh root@storepve` — then `pct exec <id>`.
|
||||
|
||||
### Step 2: Deploy Zulip Adapter
|
||||
|
||||
|
||||
@@ -4,7 +4,7 @@ kind: responsibility
|
||||
description: >
|
||||
Optimizes the full Syslog inference stack — LiteLLM routing weights, GPU model
|
||||
assignments, agent context management, and prompt caching — to reduce response
|
||||
times to sub-15s average. All GPUs now at 128K context (stable ceiling).
|
||||
times to sub-15s average. NVIDIA GPUs at 128K context; Strix Halo at 256K (2026-09-12).
|
||||
id: 067NC6KP02RG60S50M40E30928
|
||||
---
|
||||
|
||||
@@ -23,7 +23,7 @@ management, and prompt caching — without sacrificing agent capability.
|
||||
.123, any others on .129/.122) including compression, model, context_window,
|
||||
prompt_caching, memory settings
|
||||
- `gpu-health`: health check response from all 3 GPU backends (strix-moe .15:8080,
|
||||
qwen .8:8080, gemma .110:8080)
|
||||
gpu-dense .8:8080, gpu-vision .110:8080)
|
||||
|
||||
### Maintains
|
||||
|
||||
@@ -56,14 +56,14 @@ duration.
|
||||
**Context is the root cause.** Every ~46K prompt token costs ~87s of
|
||||
prefill time at 532 tok/s. Fix context first, routing second.
|
||||
|
||||
- **Route by task**: qwen for code/standard queries; gemma for
|
||||
compression/auxiliary; strix-moe for compression tasks.
|
||||
- **Route by task**: gpu-dense for code/standard queries; gpu-vision for
|
||||
vision/web-auxiliary; syslog-auto for compression.
|
||||
- **Compress aggressively**: threshold at 40% (not 65%) — a 128K window should
|
||||
compact at 51K, not 85K. Target 15% tail (not 30%).
|
||||
- **Cache everything repeated**: system prompts, skill docs, AGENTS.md — these
|
||||
never change between turns. Single-digit cache hit rate is unacceptable.
|
||||
- **Lower context ceiling**: 128K window is the stable ceiling for agent conversations.
|
||||
GPUs reduced from 256K to 128K (2026-07-17). 128K window should compact at 85K (0.65 threshold). For larger contexts, route to external providers.
|
||||
GPUs reduced from 256K to 128K (2026-07-17) for the NVIDIA hosts; Strix Halo runs 256K (2026-09-12). 128K window should compact at 85K (0.65 threshold). For larger contexts, route to external providers.
|
||||
|
||||
### Shape
|
||||
|
||||
@@ -96,8 +96,10 @@ call enable-prompt-caching
|
||||
hosts: [192.168.68.15, 192.168.68.8, 192.168.68.110]
|
||||
|
||||
-- Phase 5: Verify end-to-end latency
|
||||
-- `models` is a literal verification parameter (a snapshot only): the authoritative registry is
|
||||
-- CT 116 /opt/inference-harness/litellm_config.yaml; re-read it before use.
|
||||
|
||||
call verify-latency
|
||||
host: 192.168.68.116
|
||||
models: [syslog-auto, qwen3.6-27B-code, gemma-4-12b]
|
||||
models: [syslog-auto, gpu-dense, gpu-vision, strix-moe]
|
||||
```
|
||||
|
||||
@@ -16,8 +16,8 @@ description: >
|
||||
**Last verified:** 2026-08-15 — hwepve removed from Tabiri cluster
|
||||
(now 5 nodes: minipve, amdpve, storepve, acerpve, ocupve). hwepve
|
||||
(192.168.68.4) is a standalone PVE node + NetBird routing peer;
|
||||
London relocation pending. CTs 100 (abiba) and 105 (kagentz) moved
|
||||
to minipve.
|
||||
London relocation pending. CT 100 (abiba) is on minipve; CT 105
|
||||
(kagentz) is on amdpve.
|
||||
---
|
||||
|
||||
# Infrastructure Control Pattern
|
||||
@@ -105,16 +105,16 @@ description: >
|
||||
|
||||
| Node | IP | CPU | RAM | VMs/CTs | Role |
|
||||
|------|----|-----|-----|---------|------|
|
||||
| minipve | .12 | 16C | 30GB | abiba, kagentz, authentik, gitea, syslog-api, infisical-vault, jitsi | Auth, git, messaging |
|
||||
| amdpve | .15 | 32C | 62GB | tanko, tdunna, baggy, scottdenya | Agents, compute |
|
||||
| storepve | .6 | 28C | 31GB | docker-vm, ra-h-os, PBS, media, jdownloader, zulip | Docker, storage, chat |
|
||||
| minipve | .12 | 16C | 30GB | abiba, authentik, gitea, syslog-api, infisical-vault, jitsi | Auth, git, messaging |
|
||||
| amdpve | .15 | 32C | 62GB | kagentz, tanko, baggy, scottdenya, adguard2 | Agents, compute |
|
||||
| storepve | .6 | 28C | 31GB | docker-vm, ra-h-os, PBS, media, jdownloader, zulip, tdunna | Docker, storage, chat |
|
||||
| acerpve | .9 | 28C | 31GB | llm-gpu | GPU VMs |
|
||||
| ocupve | .5 | 12C | 14GB | ocu-llm | GPU VMs |
|
||||
|
||||
> **Note:** CTs on storepve include jdownloader (CT 118). AdGuard (CT 102) is on
|
||||
> minipve at .10, not acerpve. Abiba (CT 100) and kagentz (CT 105) are on
|
||||
> minipve (moved from hwepve 2026-08-15). Mumuni runs inside Abiba CT100
|
||||
> (.24); CT 114 (mumuni) no longer exists in the cluster.
|
||||
> **Note:** CTs on storepve include jdownloader (CT 118) and tdunna (CT 111).
|
||||
> AdGuard (CT 102) is on minipve at .10, not acerpve. Abiba (CT 100) is on
|
||||
> minipve (moved from hwepve 2026-08-15); kagentz (CT 105) is on amdpve.
|
||||
> Mumuni runs inside Abiba CT100 (.24); CT 114 (mumuni) no longer exists in the cluster.
|
||||
>
|
||||
> **hwepve (192.168.68.4) — STANDALONE (removed from Tabiri 2026-08-15):**
|
||||
> Huawei MateBook 16 (KLVL-WXX9), pve-manager/9.2.10, kernel 7.0.14-8-pve.
|
||||
@@ -199,7 +199,7 @@ description: >
|
||||
|
||||
### Ecosystem B: CT 116 syslog-api (192.168.68.116)
|
||||
|
||||
11 containers in inference-harness stack (verified live 2026-09-11; LiteLLM upgraded 1.90.0-rc.1 -> 1.99.1):
|
||||
12 containers on CT 116 — 11 in the inference-harness stack + trove-agent-docker (verified live 2026-09-11; LiteLLM upgraded 1.90.0-rc.1 -> 1.99.1; trove-agent-docker added 2026-09-11):
|
||||
|
||||
| Container | Image | Port | Role |
|
||||
|-----------|-------|------|------|
|
||||
@@ -210,6 +210,7 @@ description: >
|
||||
| harness-dashboard | inference-harness-dashboard | :3000 | SyslogAI Harness UI |
|
||||
| harness-grafana | grafana/grafana | :3000→:3001 (direct LAN, not behind nginx) | GPU + Proxmox + Docker dashboards |
|
||||
| harness-prometheus | prom/prometheus | :9090 | Metrics scraper, 6 jobs |
|
||||
| trove-agent-docker | ghcr.io/techdox/trove-agent-docker:latest | outbound agent (no port) | Trove host agent — service inventory + metrics, added 2026-09-11 |
|
||||
|
||||
**Nginx routing**:
|
||||
- `/v1/*` → harness-litellm:4000 (API)
|
||||
@@ -221,9 +222,8 @@ description: >
|
||||
|
||||
**Prometheus targets**:
|
||||
- 192.168.68.8:9400 (RTX 3090 — qwen)
|
||||
- 192.168.68.110:9400 (RTX 5070 — gemma)
|
||||
- 192.168.68.15:9400 (Strix Halo — qwen3.6-35B-udq4)
|
||||
- 192.168.68.24:9401 (Router metrics exporter)
|
||||
- 192.168.68.110:9400 (RTX 5070 — gpu-vision)
|
||||
- 192.168.68.15:9400 (Strix Halo — strix-moe)
|
||||
- harness-litellm:4000 (LiteLLM health)
|
||||
|
||||
### Ecosystem C: Netbird (72.61.0.17 — Hostinger srv1079750.hstgr.cloud)
|
||||
@@ -602,13 +602,13 @@ ssh root@192.168.68.110 "systemctl restart llama-server"
|
||||
| 102 | adguard | **minipve** | **.10** | DNS | ❌ |
|
||||
| 103 | ocu-llm | ocupve | .110 | GPU RTX 5070 | ❌ |
|
||||
| 104 | authentik | minipve | .11 | OIDC | ❌ |
|
||||
| 105 | kagentz | minipve | — | Agent Zero | ✅ |
|
||||
| 105 | kagentz | amdpve | — | Agent Zero | ✅ |
|
||||
| 106 | ra-h-os | storepve | .65 | KG bridge | ✅ MCP |
|
||||
| 107 | pbs | storepve | — | Backups | ❌ |
|
||||
| 108 | media | storepve | — | Media | ❌ |
|
||||
| 109 | docker-vm | storepve | .7 | Docker host | ❌ |
|
||||
| 110 | gitea | minipve | **.17** | Git | ❌ |
|
||||
| 111 | tdunna | amdpve | .129 | Hermes agent | ✅ |
|
||||
| 111 | tdunna | storepve | .129 | Hermes agent — ⛔ REPORT-ONLY (Theo's box, no GC) | ✅ |
|
||||
| 112 | tanko | amdpve | .122 | DSH (DeepSeek Harness) agent | ✅ |
|
||||
| 113 | baggy | amdpve | .114 | Hermes agent | ✅ |
|
||||
| 115 | scottdenya | amdpve | .75 | Denya OneCare | ❌ |
|
||||
@@ -616,6 +616,7 @@ ssh root@192.168.68.110 "systemctl restart llama-server"
|
||||
| 117 | zulip | storepve | .19 | Chat | ❌ |
|
||||
| 118 | jdownloader | storepve | .20 | JDownloader LXC (dedicated, migrated from docker-vm 2026-08-01) | ✅ |
|
||||
| 119 | infisical-vault | minipve | — | Vault | ❌ |
|
||||
| 120 | adguard2 | amdpve | — | DNS (secondary AdGuard) | ❌ |
|
||||
|
||||
## Appendix C: Docker Compose Files Location
|
||||
|
||||
@@ -636,8 +637,8 @@ Source of truth: `/root/scripts/pct-run.sh` or `prose-contracts/scripts/pct-run.
|
||||
| CT | Name | Node | pct-run |
|
||||
|-----|------|------|---------|
|
||||
| 100 | abiba | minipve | `pct-run 100` |
|
||||
| 105 | kagentz | minipve | `pct-run 105` |
|
||||
| 111 | tdunna | amdpve | `pct-run 111` |
|
||||
| 105 | kagentz | amdpve | `pct-run 105` |
|
||||
| 111 | tdunna | storepve | `pct-run 111` (⛔ report-only — no GC) |
|
||||
| 112 | tanko | amdpve | `pct-run 112` |
|
||||
| 113 | baggy | amdpve | `pct-run 113` |
|
||||
| 115 | scottdenya | amdpve | `pct-run 115` |
|
||||
@@ -648,6 +649,9 @@ Source of truth: `/root/scripts/pct-run.sh` or `prose-contracts/scripts/pct-run.
|
||||
| 107 | proxmox-backup | storepve | `pct-run 107` |
|
||||
| 108 | media | storepve | `pct-run 108` |
|
||||
| 117 | zulip | storepve | `pct-run 117` |
|
||||
| 118 | jdownloader | storepve | `pct-run 118` |
|
||||
| 119 | infisical-vault | minipve | `pct-run 119` |
|
||||
| 120 | adguard2 | amdpve | `pct-run 120` |
|
||||
| 102 | adguard | **minipve** | `pct-run 102` |
|
||||
|
||||
GPU bare-metal hosts (.8 acerpve, .110 ocupve, .15 amdpve) are NOT CTs — use SSH directly:
|
||||
|
||||
@@ -129,10 +129,15 @@ the any-HTTP rule. On those — the authenticated Zulip POST and the router
|
||||
pwd -P
|
||||
|
||||
# Zulip API health (POST ping)
|
||||
source /etc/litellm-monitor.env
|
||||
# NOTE: /etc/litellm-monitor.env exists only on CT 116, retrieve keys from CT 116 via:
|
||||
zulip_key=$(ssh root@192.168.68.116 "grep ZULIP_BOT_KEY /etc/litellm-monitor.env | cut -d= -f2")
|
||||
if [ -z "$zulip_key" ]; then
|
||||
echo "credential-missing: ZULIP_BOT_KEY not found in /etc/litellm-monitor.env"
|
||||
else
|
||||
ZULIP_USER="abiba-bot@chat.sysloggh.net"
|
||||
curl -s -o /dev/null -w '%{http_code}' -X POST https://chat.sysloggh.net/api/v1/messages -u "${ZULIP_USER}:${ZULIP_BOT_KEY}"
|
||||
curl -s -o /dev/null -w '%{http_code}' -X POST https://chat.sysloggh.net/api/v1/messages -u "${ZULIP_USER}:${zulip_key}"
|
||||
# Expected: 200 (HTTP 000 = unreachable/cache)
|
||||
fi
|
||||
|
||||
# PM2 process health
|
||||
pm2 jlist
|
||||
|
||||
@@ -12,7 +12,7 @@ triggers:
|
||||
- on "infra update" command
|
||||
- weekly (Sunday 03:00 America/New_York) via Agent Zero scheduler task "weekly-fleet-docker-update" (qSOOVzsU) — implemented 2026-09-08
|
||||
- on security advisory relay from Mumuni
|
||||
version: 1.3.0
|
||||
version: 1.4.0
|
||||
---
|
||||
|
||||
## Maintains
|
||||
@@ -81,6 +81,7 @@ Before ANY update wave:
|
||||
| VM 109 (.7) | Home stack (Pulse, Stirling PDF) — JDownloader moved to CT 118 LXC 2026-08-01 | `cd /opt/home_stack && docker compose pull && docker compose up -d` |
|
||||
| VM 109 (.7) | Audiobookshelf | `cd /opt/audiobookshelf && docker compose pull && docker compose up -d` |
|
||||
| CT 116 (.116) | Inference Harness (LiteLLM, Prometheus, Grafana) | `cd /opt/inference-harness && docker compose pull && docker compose up -d` |
|
||||
| CT 116 (.116, via minipve) | Trove docker agent (trove-agent-docker) | `pct exec 116 -- bash -c 'cd /opt/trove-agent && docker compose pull && docker compose up -d'` |
|
||||
| CT 117 (storepve) | Zulip | `pct exec 117 -- bash -c 'cd /opt/zulip && docker compose pull && docker compose up -d'` (from storepve; compose recreates on zulip_default network) |
|
||||
| CT 117 (storepve) | Jitsi | `pct exec 117 -- bash -c 'cd /opt/jitsi && docker compose pull && docker compose up -d'` (from storepve) |
|
||||
| hwpve (.11) | Authentik (server, worker, postgres) | `ssh root@192.168.68.11 'cd /root && docker compose pull && docker compose up -d'` |
|
||||
@@ -101,6 +102,7 @@ Before ANY update wave:
|
||||
- harness-litellm cold start: allow 3-5 min after recreate — reports unhealthy and :4000 refuses connections while loading config/DB, then recovers to 200 on its own (verified 2026-09-08)
|
||||
- SearXNG test: `curl :8888`
|
||||
- Digest-pin sweep: `grep -rn '@sha256:' /opt/*/docker-compose.y*` on every host — digest-pinned images are INVISIBLE to `docker compose pull` (the pin re-pulls the same digest forever, so new releases never appear). Flag every pin in the run report and propose un-pinning to a floating tag with user approval before editing. Found 2026-09-10: audiobookshelf was digest-pinned at 2.34.0 (container created 2026-07-18) and silently missed by every sweep; dockhand stack was also pinned (stack removed 2026-09-10, unused). After un-pinning audiobookshelf to :latest it updated to 2.36.0 and verified HTTP 200.
|
||||
- Version-pin awareness: a fixed version tag (e.g. `image: ...litellm:1.99.1`) is a no-op for `docker compose pull` just like a digest pin, so the stack silently stops advancing. CT 116 `harness-litellm` is INTENTIONALLY pinned to `1.99.1` (registry `main-stable`/`latest` currently resolve to `1.100.1`, sha256:a3715fa7 — a bleeding-edge jump explicitly declined 2026-09-11). Every run must look up the newest STABLE release tag for any version-pinned image, bump the pin deliberately with user approval, recreate, and re-verify. Never silently revert a pin to a floating tag.
|
||||
|
||||
## Wave 4: Proxmox Kernel Reboot
|
||||
|
||||
|
||||
@@ -67,8 +67,10 @@ description: >
|
||||
4. **If action == "create"**:
|
||||
- Generate new key with key_alias: "{agent_name}" (e.g., "tanko" — bare name, no date)
|
||||
- Set metadata: { "agent": "{agent_name}", "purpose": "agent-inference" }
|
||||
- Duration is null (permanent) — inherited from litellm default_key_generate_params
|
||||
- Set models: ["syslog-auto", "qwen3.6-27B-code", "gemma-4-12b", "strix-moe", "gpu-dense", "gpu-light", "qwen3.6-35B-udq4"]
|
||||
- Duration is whatever the caller passes; NO default enforcement exists today (CT 116 `litellm_config.yaml` has no `default_key_generate_params` block, and a key with no explicit models returns an empty models list). Agent keys are permanent by policy, not by that block. OPEN policy question: should agent keys expire by default? (captain security-policy decision, raised separately.)
|
||||
- Set models: read the live key-scoped set rather than hardcoding one — `/v1/models` is key-scoped,
|
||||
and the authoritative registry is CT 116 `/opt/inference-harness/litellm_config.yaml`. Do not add
|
||||
retired names (`gemma-4-12b`, `gpu-light`, `crew-auto` — all retired 2026-09-12).
|
||||
- Note: `ornith-1.0-35b` is NOT a valid LiteLLM model name (use `strix-moe`, the stable alias). qwen3.6-35B-A3B removed from fleet (was never deployed).
|
||||
- Return the new key
|
||||
5. **If action == "rotate"**:
|
||||
|
||||
@@ -26,9 +26,8 @@ description: >
|
||||
| Model | avg latency | avg TTFT | p-profile (24h) |
|
||||
|---|---|---|---|
|
||||
| syslog-auto | 28.8s | 25.4s | 68 calls took 30-120s; tail to ~300s under load |
|
||||
| qwen3.6-27B-code | 23.0s | — | same backend class as syslog-auto |
|
||||
| strix-moe | 7.5s | — | Strix Halo, healthy |
|
||||
| gemma-4-12b | 2.6s | — | RTX 5070, healthy |
|
||||
| gpu-vision (retired gemma-4-12b, RTX 5070) | 2.6s | — | RTX 5070, healthy |
|
||||
|
||||
Sep 6 incident timeline: failures 04:00-07:00 EDT (0% GPU util = wedged
|
||||
backend), full recovery 07:00-08:00 with ZERO client failures once requests
|
||||
@@ -53,13 +52,11 @@ proxy queuing.
|
||||
|
||||
### 2. Auxiliary tasks — keep template timeouts, one correction
|
||||
|
||||
- vision: 60s (keep), web_extract: 30s (keep) — gemma-4-12b averages 2.6s;
|
||||
these are fine.
|
||||
- vision: 60s (keep), web_extract: 30s (keep) — the 2.6s average was measured on `gemma-4-12b` (retired 2026-09-12); the live RTX 5070 alias is `gpu-vision`.
|
||||
- compression: 300s (keep — this was already raised from 60 per gpu-fleet).
|
||||
- **gpu-dense delegation/x_search: set timeout >= 120s.** The RTX 3090
|
||||
(qwen3.6-27B-code backend, 23.0s avg) is the same speed class as
|
||||
syslog-auto; delegation defaults that assume fast responses will 408 the
|
||||
same way.
|
||||
(gpu-dense backend) is the same speed class as syslog-auto; delegation
|
||||
defaults that assume fast responses will 408 the same way.
|
||||
|
||||
### 3. Retry policy — backoff, not repetition
|
||||
|
||||
@@ -79,8 +76,9 @@ proxy queuing.
|
||||
Send a real model name and use a real (probe-designated) key.
|
||||
- Probe timeout: 30s. A probe that takes longer than 30s IS the alert —
|
||||
report "backend slow (>30s)" rather than hanging.
|
||||
- Probe cadence: at most hourly. The 6h litellm-health cron cadence is the
|
||||
standard; sub-hourly synthetic traffic distorts latency baselines.
|
||||
- Probe cadence: at most hourly. The litellm-health cron cadence (authoritative
|
||||
trigger in contract-registry.yaml) is the standard; sub-hourly synthetic
|
||||
traffic distorts latency baselines.
|
||||
|
||||
### 5. Batch/benchmark jobs — schedule away from 04:00-07:00 EDT and chunk
|
||||
|
||||
|
||||
+89
-35
@@ -1,14 +1,7 @@
|
||||
---
|
||||
kind: function
|
||||
name: litellm-health
|
||||
status: deprecated
|
||||
deprecated_on: 2026-07-09
|
||||
replaced_by: litellm-self-heal.prose.md
|
||||
note: >
|
||||
Consolidated into litellm-self-heal.prose.md to eliminate duplication
|
||||
of architecture diagrams, GPU topology, timeout tables, and container
|
||||
lists. Health check is now § Health Check within litellm-self-heal.
|
||||
This file is retained for reference only — use litellm-self-heal instead.
|
||||
status: active
|
||||
description: >
|
||||
Verifies the LiteLLM inference stack health. Current architecture (2026-07-09):
|
||||
nginx:80 → LiteLLM:4000 → GPU(llama-server) via direct proxy.
|
||||
@@ -51,9 +44,8 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2)
|
||||
**What changed (v3.2.0 → v4.0.0 — 2026-07-08)**:
|
||||
- Router REMOVED from request path — LiteLLM proxies directly to GPU
|
||||
- All GPUs at parallel 2 (was parallel 1)
|
||||
- NVIDIA context reduced 256K→128K to free VRAM — now the stable ceiling across all GPUs (2026-07-17)
|
||||
- LiteLLM timeouts tuned: gemma 25→120s, qwen 40→90s (SUPERSEDED 2026-07-16: qwen 300s, gemma 120s, strix 300s — see litellm-self-heal)
|
||||
- nginx proxy_read_timeout: 600s, LiteLLM request_timeout: 300s
|
||||
- NVIDIA context reduced 256K→128K to free VRAM — the stable NVIDIA ceiling (2026-07-17); Strix Halo runs 256K (2026-09-12)
|
||||
- Timeouts and fallback chains are config state — read them from CT 116 `/opt/inference-harness/litellm_config.yaml`; they are not duplicated here.
|
||||
|
||||
## Parameters
|
||||
|
||||
@@ -75,26 +67,36 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2)
|
||||
|
||||
- SSH key access to backend_host for container checks
|
||||
- Network access to public_url, auth_host, and gpu_dashboard_url
|
||||
- LiteLLM master key for key management endpoints
|
||||
- LiteLLM master key for admin endpoints only (`/key/list`, `/key/generate`, `/key/info`)
|
||||
- The dedicated `monitor` agent key on CT 116 at `/etc/litellm-monitor.env` (root-only 0600) for
|
||||
model inference checks, scoped for every alias step 7 probes (`gpu-dense`, `gpu-vision`,
|
||||
`strix-moe`, `syslog-auto`). Retrieve from the executor's host via:
|
||||
|
||||
```
|
||||
monitor key: ssh root@192.168.68.116 "grep LITELLM_MONITOR_KEY /etc/litellm-monitor.env | cut -d= -f2"
|
||||
master key: ssh root@192.168.68.116 "docker exec harness-litellm printenv LITELLM_MASTER_KEY"
|
||||
```
|
||||
|
||||
If credentials are missing or unreadable, the probe must report `credential-missing` (not bare 401 or "0 keys").
|
||||
The master key must never be used for inference.
|
||||
|
||||
## GPU Fleet Topology
|
||||
|
||||
| Host | IP | Hardware | Models Served | Engine | Context | Parallel |
|
||||
|------|-----|----------|---------------|--------|---------|----------|
|
||||
| llm-gpu | 192.168.68.8 | NVIDIA RTX 3090 (24 GB) | qwen3.6-27B-code | llama-server systemd | **128K** | 2 |
|
||||
| ocu-llm | 192.168.68.110 | NVIDIA RTX 5070 (12 GB) | gemma-4-12b | llama-server systemd | **128K** | 2 |
|
||||
| amdpve | 192.168.68.15 | AMD Strix Halo 64GB UMA | qwen3.6-35B-udq4 (LiteLLM alias: strix-moe) | llama-server systemd (Vulkan) | 128K | 2 |
|
||||
| Host | IP | Hardware | Role |
|
||||
|------|-----|----------|------|
|
||||
| llm-gpu | 192.168.68.8 | NVIDIA RTX 3090 (24 GB) | heavy reasoning (`gpu-dense`) |
|
||||
| ocu-llm | 192.168.68.110 | NVIDIA RTX 5070 (12 GB) | vision / web extract / light tasks (`gpu-vision`) |
|
||||
| amdpve | 192.168.68.15 | AMD Strix Halo 64GB UMA | compression (`strix-moe`) |
|
||||
|
||||
## Model Fallback Chains (LiteLLM)
|
||||
Single source of truth for models, aliases, rpm caps, weights and fallback chains:
|
||||
CT 116 `/opt/inference-harness/litellm_config.yaml`. Do not duplicate those values in
|
||||
contracts — read them there. Do not re-add retired names (`gemma-4-12b`, `gpu-light`,
|
||||
`crew-auto`).
|
||||
|
||||
| Primary | Timeout | Fallback | Timeout |
|
||||
|---------|---------|----------|---------|
|
||||
| qwen3.6-27B-code | 300s | gemma-4-12b | 120s |
|
||||
| gemma-4-12b | 120s | qwen3.6-27B-code | 300s |
|
||||
| qwen3.6-35B-udq4 / strix-moe | 300s | qwen → gemma | — |
|
||||
| syslog-auto (balanced) | 300s | qwen → gemma | — |
|
||||
|
||||
> Global: request_timeout=300s, nginx proxy_read_timeout=600s
|
||||
> Re-scope note (2026-09-12): the earlier plan to restate the live fallback chains and
|
||||
> per-model timeouts in this contract is intentionally superseded — that state is config,
|
||||
> and this contract points at the CT 116 config instead. Step 7 likewise authenticates with
|
||||
> the dedicated `monitor` key, not the master key, which is admin-only.
|
||||
|
||||
## Containers on CT 116
|
||||
|
||||
@@ -112,18 +114,32 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2)
|
||||
|
||||
1. **Read parameters** — Use provided values or defaults
|
||||
|
||||
2. **Check public endpoints**:
|
||||
2. **Check the end-user surfaces** — the public edge and the backend edge serve the SAME
|
||||
app under DIFFERENT paths. They are not interchangeable, so every probe below must name
|
||||
the surface it targets. Never point a check at a path that only resolves on the other
|
||||
surface.
|
||||
|
||||
**Public edge** — `{{public_url}}` (https://litellm.sysloggh.net) serves the app at the
|
||||
ROOT; the `/litellm/` prefix does not exist there and 404s:
|
||||
- GET {{public_url}}/ui/ → expect 200 ("LiteLLM Dashboard")
|
||||
- GET {{public_url}}/docs → expect 200 ("LiteLLM API - Swagger UI")
|
||||
- GET {{public_url}}/litellm/ui/ and {{public_url}}/litellm/docs → expect 404 (not served on this edge)
|
||||
|
||||
**Backend edge** — `http://{{backend_host}}` (port 80) serves the app UNDER `/litellm/`:
|
||||
- GET http://{{backend_host}}/litellm/ui/ → expect 200 ("LiteLLM Dashboard")
|
||||
- GET http://{{backend_host}}/litellm/docs → expect 200 ("LiteLLM API - Swagger UI")
|
||||
- GET http://{{backend_host}}/ui/ and /docs → expect 301 → /litellm/... → 200 (one-hop redirect helpers,
|
||||
added 2026-09-11)
|
||||
|
||||
3. **Check LiteLLM health (no-auth)**:
|
||||
- GET http://{{backend_host}}/litellm/health/liveliness → expect 200
|
||||
|
||||
4. **Check backend container health**:
|
||||
- SSH to {{backend_host}} → `docker ps` → verify 11 containers healthy
|
||||
- SSH to {{backend_host}} → `docker ps` → verify 12 containers healthy (11 harness + trove-agent-docker, added 2026-09-11)
|
||||
- Critical: harness-litellm, harness-nginx, harness-postgres
|
||||
- Monitoring: harness-redis, harness-dashboard, harness-grafana, harness-prometheus,
|
||||
harness-alertmanager, harness-zulip-bridge, harness-docker-stats, harness-pve-exporter
|
||||
harness-alertmanager, harness-zulip-bridge, harness-docker-stats, harness-pve-exporter,
|
||||
trove-agent-docker (ghcr.io/techdox/trove-agent-docker, added 2026-09-11)
|
||||
|
||||
5. **Check GPU fleet health via gpu-monitor** (router decommissioned 2026-09-11):
|
||||
- GET {{gpu_dashboard_url}}/gpu-data → expect 200 with GPU metrics JSON
|
||||
@@ -134,16 +150,54 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2)
|
||||
- Verify GPUs reporting status "healthy"
|
||||
- Check alerts array for active warnings/critical
|
||||
|
||||
7. **Check model inference via LiteLLM** — Test each model:
|
||||
- POST /v1/chat/completions model=gemma-4-12b → expect 200
|
||||
- POST /v1/chat/completions model=qwen3.6-27B-code → expect 200
|
||||
- POST /v1/chat/completions model=strix-moe → expect 200
|
||||
- Use master key for auth
|
||||
7. **Check model inference via LiteLLM** — Test one model on each GPU host. The health
|
||||
check runs on the **backend edge**, not the public edge, so these paths carry the
|
||||
`/litellm/` prefix:
|
||||
- POST http://{{backend_host}}/litellm/v1/chat/completions model=gpu-dense → expect 200 (RTX 3090, .8)
|
||||
- POST http://{{backend_host}}/litellm/v1/chat/completions model=gpu-vision → expect 200 (RTX 5070, .110)
|
||||
- POST http://{{backend_host}}/litellm/v1/chat/completions model=strix-moe → expect 200 (Strix Halo, .15)
|
||||
- Auth uses the dedicated `monitor` agent key. Retrieve via:
|
||||
`ssh root@192.168.68.116 "grep LITELLM_MONITOR_KEY /etc/litellm-monitor.env | cut -d= -f2"`
|
||||
Do NOT use the master key for inference — the master key is for admin endpoints only
|
||||
(`/key/list`, `/key/generate`, `/key/info`). Retrieve master key via:
|
||||
`ssh root@192.168.68.116 "docker exec harness-litellm printenv LITELLM_MASTER_KEY"`
|
||||
- KEY SCOPE: the `monitor` key MUST be scoped for the three probed aliases (`gpu-dense`,
|
||||
`gpu-vision`, `strix-moe`) plus the `syslog-auto` fallback pool, otherwise the probe
|
||||
returns 403 and the host is not covered.
|
||||
If a probe returns 403, widen the monitor key's model list on CT 116 (add the missing
|
||||
alias) and re-run — never drop the host from the probe to make the check pass.
|
||||
- `gemma-4-12b` was retired and returns 400 `Invalid model name` — do not re-add it to
|
||||
this list. The RTX 5070 host now serves `gpu-vision`.
|
||||
- `/v1/models` is **key-scoped**: a model is only visible to keys allowed to use it, so
|
||||
the set depends on the key. Always state which key a model list was read with — a
|
||||
snapshot without its key is not evidence. This probe uses the `monitor` key on the
|
||||
backend surface (`http://{{backend_host}}/litellm/v1/models`). The authoritative model
|
||||
registry is CT 116 `/opt/inference-harness/litellm_config.yaml`; read it there rather
|
||||
than freezing a list here.
|
||||
|
||||
8. **Check agent keys**:
|
||||
- GET /key/list with master key → verify all 6 agents have keys
|
||||
- GET http://{{backend_host}}/litellm/key/list with master key (admin endpoint) → verify all 6 agents have keys
|
||||
- **IMPORTANT**: Run the curl on the CT 116 HOST, not inside the container. The `harness-litellm` container has no curl/wget. Use:
|
||||
`ssh root@192.168.68.116 "curl -s -H 'Authorization: Bearer $MASTER_KEY' http://127.0.0.1:4000/key/list"`
|
||||
- If the response is empty or unparseable, report `admin-call-failed` (not "0 agent keys")
|
||||
|
||||
9. **Check Grafana**:
|
||||
- GET {{grafana_url}}/api/health → expect 200
|
||||
|
||||
10. **Compile and report** — Determine overall_status from individual check results
|
||||
|
||||
## Executor Script (2026-09-13)
|
||||
|
||||
**Run `scripts/litellm-health-check.py` from the clone.** This script implements all 11
|
||||
checks defined above and reports results in a standardized format. Paste its output in
|
||||
the status line.
|
||||
|
||||
- Hand-rolled probes are **not** an acceptable substitute for the script.
|
||||
- Backend-edge checks (steps 2–8) must use `http://192.168.68.116` (internal IP),
|
||||
**not** the public URL `https://litellm.sysloggh.net` (which returns 401 for those paths).
|
||||
- Docker Stats (step 10) must be fetched from the CT 116 host itself (`127.0.0.1:9324/metrics`)
|
||||
because the `harness-docker-stats` container binds to localhost on CT 116.
|
||||
- Admin Key List (step 8) requires the master key expanded locally before SSH, then embedded
|
||||
in the remote curl command with proper quoting.
|
||||
|
||||
Expected output on a healthy fleet: 11/11 passing checks.
|
||||
|
||||
+35
-69
@@ -12,21 +12,21 @@ note: >
|
||||
Reports to /var/log/litellm/health-*.json and Gitea (SyslogSolution/health-logs).
|
||||
GPU monitoring integrated from gpu-monitor on .24:9100.
|
||||
|
||||
Consolidated from litellm-health + litellm-self-heal on 2026-07-09 to eliminate
|
||||
duplication of architecture diagrams, GPU topology, timeout tables, and container
|
||||
lists. Health check is now § Health Check within this contract.
|
||||
Health probes are owned by litellm-health.prose.md (dispatched as
|
||||
`run contract: litellm-health`). This contract owns remediation only — it does not
|
||||
re-specify the probes.
|
||||
|
||||
Source of truth for GPU topology and keys: gpu-fleet.prose.md
|
||||
Last verified: 2026-07-12
|
||||
description: >
|
||||
LiteLLM inference stack health monitoring + self-healing. Verifies the full
|
||||
nginx → LiteLLM → GPU chain, 11 containers on CT 116, 3 GPU hosts, model
|
||||
inference, and agent keys. Applies remediation rules for common failures.
|
||||
LiteLLM inference stack remediation. Applies remediation rules for failures detected
|
||||
by litellm-health.prose.md (nginx → LiteLLM → GPU chain, CT 116 containers, GPU hosts,
|
||||
model inference, and agent keys).
|
||||
Reports every action via Zulip DM and Gitea (SyslogSolution/health-logs).
|
||||
---
|
||||
---
|
||||
|
||||
# LiteLLM Operations — Health Check + Self-Heal
|
||||
# LiteLLM Operations — Self-Heal (Remediation)
|
||||
|
||||
## Architecture (v4.0.0 — Direct: nginx → LiteLLM → GPU)
|
||||
|
||||
@@ -58,41 +58,44 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2)
|
||||
|
||||
## GPU Fleet Topology
|
||||
|
||||
| Host | IP | Hardware | Models Served | Engine | Context | Parallel |
|
||||
|------|-----|----------|---------------|--------|---------|----------|
|
||||
| llm-gpu | 192.168.68.8 | NVIDIA RTX 3090 (24 GB) | qwen3.6-27B-code | llama-server systemd (`/home/llmuser/llama-wrapper.sh`, `-c 131072 --parallel 2 --ngl 99`) | **128K** | 2 |
|
||||
| ocu-llm | 192.168.68.110 | NVIDIA RTX 5070 (12 GB) | gemma-4-12b | llama-server systemd (`/home/llmuser/llama-wrapper.sh`, `--ctx-size 131072 --parallel 2`, IQ4_NL + MTP draft) | **128K** | 2 |
|
||||
| amdpve | 192.168.68.15 | AMD Strix Halo 64GB UMA | qwen3.6-35B-udq4 (LiteLLM alias: `strix-moe`) | llama-server systemd (Vulkan) | 128K | 2 |
|
||||
| Host | IP | Hardware | Role |
|
||||
|------|-----|----------|------|
|
||||
| llm-gpu | 192.168.68.8 | NVIDIA RTX 3090 (24 GB) | heavy reasoning (`gpu-dense`) |
|
||||
| ocu-llm | 192.168.68.110 | NVIDIA RTX 5070 (12 GB) | vision / web extract / light tasks (`gpu-vision`) |
|
||||
| amdpve | 192.168.68.15 | AMD Strix Halo 64GB UMA | compression (`strix-moe`) |
|
||||
|
||||
> Verified on ground 2026-07-16 via `curl /v1/models` on each host + `llama-wrapper.sh`. The AMD host's underlying model is `qwen3.6-35B-udq4`; LiteLLM exposes it under two `model_name`s: `qwen3.6-35B-udq4` and `strix-moe` (rpm 40). The legacy name `ornith-1.0-35b` does NOT exist in LiteLLM and must not be referenced.
|
||||
> Verified on the ground 2026-07-16. The legacy name `ornith-1.0-35b` does NOT exist in LiteLLM and must not be referenced.
|
||||
|
||||
## LiteLLM Model Surface (ground truth — `/opt/inference-harness/litellm_config.yaml` on CT 116)
|
||||
## LiteLLM Model Surface
|
||||
|
||||
`model_name`s served: `qwen3.6-27B-code`, `gemma-4-12b`, `qwen3.6-35B-udq4`, `strix-moe`, `gpu-dense`, `gpu-light`, `syslog-auto`, `crew-auto` (new 2026-08-20).
|
||||
Single source of truth for models, aliases, rpm caps, weights and fallback chains:
|
||||
CT 116 `/opt/inference-harness/litellm_config.yaml`. Do not duplicate those values in
|
||||
contracts — read them there. Do not re-add retired names (`gemma-4-12b`, `gpu-light`,
|
||||
`crew-auto`).
|
||||
|
||||
### Context Cap Split (2026-08-20)
|
||||
### Alias Surface
|
||||
|
||||
| Alias | Serves | Where | Kind |
|
||||
|-------|--------|-------|------|
|
||||
| `gpu-dense` | heavy reasoning | RTX 3090 (192.168.68.8) | direct alias |
|
||||
| `gpu-vision` | vision / web extract / light tasks | RTX 5070 (192.168.68.110) | direct alias AND `syslog-auto` pool member |
|
||||
| `strix-moe` | compression (MoE) | Strix Halo (192.168.68.15) | direct alias |
|
||||
| `syslog-auto` | balanced default | weighted pool across the three GPU hosts | pool router |
|
||||
|
||||
### Context Cap Split (2026-08-20; crew cap RETIRED)
|
||||
|
||||
- **Abiba (firstmate)**: 128K uncapped — unlimited context for primary workloads
|
||||
- **Hermes agents** (mumuni, tanko, koby, koonimo): 128K uncapped
|
||||
- **Crewmates** (ops, tune, verify, auth-keys, build): 64K capped — alias `crew-auto` enforces 64K limit
|
||||
- **Crewmates** (ops, tune, verify, auth-keys, build): the 64K cap was retired together with the `crew-auto` alias; NO context cap is currently in force.
|
||||
|
||||
Preferred implementation: uncap shared pool, add capped alias for crew-only.
|
||||
> The 64K crew cap was retired with `crew-auto` (2026-09-12). No limit is currently in force; reinstating one would need per-key model limits as a separate, deliberately-scoped change.
|
||||
|
||||
- `syslog-auto` is a weighted router model: qwen3.6-27B-code (0.55, rpm 500) + qwen3.6-35B-udq4 (0.30, rpm 60) + gemma-4-12b (0.15, rpm 200).
|
||||
- `gpu-dense` / `gpu-light` are high-rpm aliases (rpm 500) onto qwen3.6-27B-code / gemma-4-12b respectively.
|
||||
- Key scoping: agent keys are restricted to `['syslog-auto','qwen3.6-27B-code','gemma-4-12b','strix-moe','gpu-dense','gpu-light']`. As of 2026-07-16 the `baggy`/`koby`/`mumuni`/`abiba-pi` keys ALSO include `qwen3.6-35B-udq4`; `abiba-pi` additionally includes `deepseek-v4-pro` (cloud fallback). `kagenz0`/`koonimo`/`pi-agents-unified` have the standard 6 only. Agents should still use the stable alias `strix-moe` (not the raw `qwen3.6-35B-udq4`) so model swaps don't break them.
|
||||
- Key scoping: `/v1/models` is key-scoped, so the set a caller sees must be read with a named key rather than assumed — a monitor key, an agent key, and the master key can each return a different set. Agents should use the stable aliases (`strix-moe`, `gpu-vision`, `gpu-dense`) rather than raw model names, so model swaps don't break them.
|
||||
- **NEVER use `litellm_proxy_master_key` (the `sk-litellm-...` master key) for inference.** It is for admin endpoints only (`/key/list`, `/key/generate`, `/key/info`). All inference — agent traffic, health-check model tests, monitor scripts — uses agent-specific keys. The health-check script's model tests use a dedicated `monitor` agent key stored at `/etc/litellm-monitor.env` on CT 116 (root-only, `chmod 600`); `/key/list` is the only call that legitimately uses `$MASTER_KEY`.
|
||||
|
||||
## Model Fallback Chains (LiteLLM)
|
||||
|
||||
| Primary | Timeout | Fallback | Timeout |
|
||||
|---------|---------|----------|---------|
|
||||
| qwen3.6-27B-code | 300s | gemma-4-12b | 120s |
|
||||
| gemma-4-12b | 120s | qwen3.6-27B-code | 300s |
|
||||
| qwen3.6-35B-udq4 / strix-moe | 300s | qwen → gemma | — |
|
||||
| syslog-auto (balanced) | 300s | qwen → gemma | — |
|
||||
|
||||
> Global: request_timeout=300s, nginx proxy_read_timeout=600s
|
||||
> Re-scope note (2026-09-12): the fallback-chain and per-model timeout tables were removed
|
||||
> by the single-source-of-truth re-scope; read those values from the CT 116 config named
|
||||
> above rather than from this contract.
|
||||
|
||||
## Containers on CT 116
|
||||
|
||||
@@ -135,44 +138,7 @@ Preferred implementation: uncap shared pool, add capped alias for crew-only.
|
||||
|
||||
## Health Check
|
||||
|
||||
Run this first on every cycle. Results feed into remediation rules below.
|
||||
|
||||
### 1. Check public endpoints
|
||||
- GET {{public_url}}/ui/ → expect 200 ("LiteLLM Dashboard")
|
||||
- GET {{public_url}}/docs → expect 200 ("LiteLLM API - Swagger UI")
|
||||
|
||||
### 2. Check LiteLLM health (no-auth)
|
||||
- GET http://{{backend_host}}/litellm/health/liveliness → expect 200
|
||||
|
||||
### 3. Check backend container health
|
||||
- SSH to {{backend_host}} → `docker ps` → verify 11 containers healthy
|
||||
- Critical: harness-litellm, harness-nginx, harness-postgres
|
||||
- Monitoring: harness-redis, harness-dashboard, harness-grafana, harness-prometheus,
|
||||
harness-alertmanager, harness-zulip-bridge, harness-docker-stats, harness-pve-exporter
|
||||
- Decommissioned 2026-09-11: harness-router (container, image and config removed)
|
||||
|
||||
### 4. Check GPU fleet health (via fleet dashboard)
|
||||
- GET {{gpu_dashboard_url}}/gpu-data → expect 200 with GPU metrics JSON
|
||||
- Verify GPUs reporting status "healthy"
|
||||
- Check alerts array for active warnings/critical
|
||||
|
||||
### 5. Check model inference via LiteLLM — test each model
|
||||
- POST /v1/chat/completions model=gemma-4-12b → expect 200
|
||||
- POST /v1/chat/completions model=qwen3.6-27B-code → expect 200
|
||||
- POST /v1/chat/completions model=strix-moe → expect 200
|
||||
- Use master key for auth
|
||||
|
||||
### 6. Check agent keys
|
||||
- GET /key/list with master key → verify all 6 agents have keys
|
||||
|
||||
### 7. Check Grafana
|
||||
- GET {{grafana_url}}/api/health → expect 200
|
||||
|
||||
### 8. Compile overall status
|
||||
Determine overall_status from individual check results:
|
||||
- "healthy" — all checks pass
|
||||
- "degraded" — 1-2 non-critical checks fail
|
||||
- "down" — critical checks fail
|
||||
Health probes are owned by `litellm-health.prose.md` (dispatched as `run contract: litellm-health`). This contract owns remediation only — it does not re-specify the probes.
|
||||
|
||||
---
|
||||
---
|
||||
|
||||
@@ -90,8 +90,8 @@ raw data never provided.
|
||||
|
||||
| Worker | Model | Toolsets | Role | Use When |
|
||||
|--------|-------|----------|------|----------|
|
||||
| `syslog-code` | qwen3.6-27B-code | terminal, file, web, memory, skills | Code patches, automation, scripts | Writing/modifying code, creating scripts, debugging, reading/writing files |
|
||||
| `syslog-devops` | qwen3.6-27B-code | terminal, file, web, memory, skills | Infrastructure, DB, bridge, Proxmox | Server ops, SSH, Docker, Proxmox, DB queries, hardware checks |
|
||||
| `syslog-code` | gpu-dense | terminal, file, web, memory, skills | Code patches, automation, scripts | Writing/modifying code, creating scripts, debugging, reading/writing files |
|
||||
| `syslog-devops` | gpu-dense | terminal, file, web, memory, skills | Infrastructure, DB, bridge, Proxmox | Server ops, SSH, Docker, Proxmox, DB queries, hardware checks |
|
||||
| `syslog-email` | strix-moe | terminal, file, web, memory, skills | Email automation, mail operations | Sending/receiving email, inbox management, SMTP operations |
|
||||
| `syslog-research` | strix-moe | terminal, file, web, memory, skills, **browser** | Analysis, classification, data processing | Web research, browser tasks, data analysis, classification, reading docs |
|
||||
| `syslog-review` | strix-moe | terminal, file, web, memory, skills | Verification, QA, audit validation | **ALWAYS** verify worker output before delivery — especially for infra changes, code builds, and research findings |
|
||||
|
||||
@@ -9,7 +9,6 @@ description: >
|
||||
- abiba-telegram: { status: "online", uptime: string, restarts: number }
|
||||
- abiba-zulip: { status: "online", uptime: string, restarts: number }
|
||||
- gitea-runner: { status: "online", uptime: string, restarts: number }
|
||||
- spoton-service: { status: "online", uptime: string, restarts: number }
|
||||
- zulip-watchdog: { status: "online", uptime: string, restarts: number }
|
||||
- last_check: timestamp
|
||||
|
||||
|
||||
@@ -87,7 +87,7 @@ agent: abiba
|
||||
| storepve | 192.168.68.6 | PVE |
|
||||
| acerpve | 192.168.68.9 | PVE (hosts llm-gpu qemu/101) |
|
||||
| minipve | 192.168.68.12 | PVE |
|
||||
| amdpve | 192.168.68.15 | PVE + Strix Halo LLM (qwen3.6-35B-udq4, strix-moe) |
|
||||
| amdpve | 192.168.68.15 | PVE + Strix Halo LLM (strix-moe) |
|
||||
|
||||
## Operations
|
||||
|
||||
|
||||
@@ -0,0 +1,44 @@
|
||||
disk-gc report-only verification — CT 111 / tdunna / 192.168.68.129
|
||||
date: 2026-09-12T19:11:36Z
|
||||
host: abiba (this scanner runs INSIDE CT 100 / abiba)
|
||||
branch head: 6e612ce37b1b9a9688b04e7f816848e7a4185fca
|
||||
command: scripts/disk-gc-plan.py --scan <live fleet scan>
|
||||
|
||||
PURPOSE: prove that on a REAL fleet scan, CT 111 is alerted and NO gc-executor action
|
||||
is emitted for it at any level. No GC command was executed against .129.
|
||||
|
||||
--- live fleet scan (df -P / via scripts/pct-run.sh for LXC, direct SSH for hosts) ---
|
||||
tdunna 84%
|
||||
acerpve 192.168.68.9 77%
|
||||
amdpve 192.168.68.15 76%
|
||||
ocu-llm 192.168.68.110 69%
|
||||
storepve 192.168.68.6 65%
|
||||
kagentz 61%
|
||||
tanko 56%
|
||||
minipve 192.168.68.12 49%
|
||||
authentik 45%
|
||||
infisical-vault 39%
|
||||
adguard 38%
|
||||
ocupve 192.168.68.5 38%
|
||||
scottdenya 35%
|
||||
syslog-api 34%
|
||||
llm-gpu 192.168.68.8 25%
|
||||
abiba 23%
|
||||
baggy 22%
|
||||
jdownloader 21%
|
||||
gitea 16%
|
||||
ra-h-os 13%
|
||||
zulip 12%
|
||||
docker-vm 192.168.68.7 11%
|
||||
adguard2 10%
|
||||
media 9%
|
||||
proxmox-backup-server 4%
|
||||
|
||||
--- planner output (action plan) ---
|
||||
111 AMBER 84.0% -> REPORT-ONLY (no GC) — Theo's box — captain ruling 2026-08-17, re-confirmed 2026-09-10
|
||||
acerpve AMBER 77.0% -> gc-executor
|
||||
amdpve AMBER 76.0% -> gc-executor
|
||||
|
||||
--- verdict ---
|
||||
CT 111 (tdunna) 84% AMBER -> report-only; no gc-executor row emitted; no GC run on .129.
|
||||
Owned hosts acerpve .9 (77%) and amdpve .15 (76%) -> gc-executor (ours).
|
||||
Executable
+227
@@ -0,0 +1,227 @@
|
||||
#!/usr/bin/env bash
|
||||
# capture-dsh-token.sh — refresh the dsh-web login token WITHOUT restarting dsh-web.
|
||||
#
|
||||
# Context (CT 112 / tankodhs.sysloggh.net)
|
||||
# ----------------------------------------
|
||||
# The dsh-web UI (systemd unit `dsh-web.service`, 127.0.0.1:3080) prints a random
|
||||
# launch token to the journal on every start:
|
||||
#
|
||||
# dsh web: http://127.0.0.1:3080/?token=<TOKEN>
|
||||
#
|
||||
# That token is the only way to bootstrap the authority-bound 30-day browser
|
||||
# cookie. It rotates on every dsh-web start, so the Authentik-gated
|
||||
# `location = /dsh-web-login` in /etc/nginx/sites-available/dsh must always
|
||||
# reference the token of the RUNNING process.
|
||||
#
|
||||
# This script:
|
||||
# 1. selects the launch token the RUNNING service actually accepts from the
|
||||
# current systemd invocation — it NEVER stops or starts dsh-web,
|
||||
# 2. records it in /etc/dsh-web/launch-token,
|
||||
# 3. regenerates the nginx include /etc/dsh-web/nginx-login.conf (the
|
||||
# `proxy_pass ...?token=` line consumed by /dsh-web-login),
|
||||
# 4. reloads nginx ONLY when the on-disk include differs from the generated
|
||||
# one or the applied-state stamp does not match the token (the stamp is
|
||||
# written only after a successful reload), rolling the include back on
|
||||
# failure so the next run retries,
|
||||
# 5. removes the legacy unauthenticated :8081 endpoint if it ever reappears.
|
||||
#
|
||||
# Idempotent and safe to run at any time (systemd ExecStartPost or timer).
|
||||
set -euo pipefail
|
||||
umask 077
|
||||
PATH="/usr/local/sbin:/usr/local/bin:/usr/sbin:/usr/bin:/sbin:/bin"
|
||||
|
||||
JOURNAL_UNIT="dsh-web.service"
|
||||
TOKEN_FILE="/etc/dsh-web/launch-token"
|
||||
INCLUDE_FILE="/etc/dsh-web/nginx-login.conf"
|
||||
STAMP_FILE="/etc/dsh-web/nginx-login.conf.applied"
|
||||
PENDING_FILE="/etc/dsh-web/nginx-reload.pending"
|
||||
SITE_ENABLED="/etc/nginx/sites-enabled/dsh"
|
||||
LEGACY_8081="/etc/nginx/sites-enabled/dsh.token"
|
||||
STASH_DIR="/etc/nginx/sites-available"
|
||||
LOCK_FILE="/run/capture-dsh-token.lock"
|
||||
LOGIN_HOST="tankodhs.sysloggh.net"
|
||||
LOGIN_UPSTREAM="http://127.0.0.1:3080"
|
||||
TOKEN_WAIT=120
|
||||
|
||||
log() { printf 'capture-dsh-token: %s\n' "$*" >&2; }
|
||||
die() { printf 'capture-dsh-token: ERROR: %s\n' "$*" >&2; exit 1; }
|
||||
|
||||
[ "$(id -u)" -eq 0 ] || die "must run as root"
|
||||
|
||||
# ── 0. Serialize runs so timer/ExecStartPost/manual runs cannot interleave ──
|
||||
exec 9>"$LOCK_FILE"
|
||||
flock -n 9 || { log "another capture-dsh-token run holds $LOCK_FILE; exiting"; exit 0; }
|
||||
mkdir -p "$(dirname "$PENDING_FILE")"
|
||||
|
||||
# ── 0b. Guarantee the generated include exists before any `nginx -t` ──────
|
||||
# The :80 site includes /etc/dsh-web/nginx-login.conf by literal path, so a
|
||||
# missing include makes every `nginx -t` fail and can wedge recovery. Seed it
|
||||
# from the last known token (or a placeholder); step 4 replaces it.
|
||||
if [ ! -f "$INCLUDE_FILE" ]; then
|
||||
SEED="placeholder"
|
||||
if [ -f "$TOKEN_FILE" ]; then
|
||||
SEED="$(cat "$TOKEN_FILE" 2>/dev/null || true)"
|
||||
[ -n "$SEED" ] || SEED="placeholder"
|
||||
fi
|
||||
printf '%s' "$SEED" | grep -qE '^[A-Za-z0-9._~+/=:@-]+$' || SEED="placeholder"
|
||||
printf 'proxy_pass %s/?token=%s;\n' "$LOGIN_UPSTREAM" "$SEED" > "$INCLUDE_FILE"
|
||||
chmod 600 "$INCLUDE_FILE"
|
||||
log "created missing $INCLUDE_FILE"
|
||||
fi
|
||||
|
||||
# ── 1. Remove the legacy unauthenticated :8081 endpoint, if present ─────────
|
||||
# It bypassed Authentik entirely (listened on 0.0.0.0:8081 with no auth_request)
|
||||
# and must never come back. Stash it rather than delete so it is auditable.
|
||||
if [ -e "$LEGACY_8081" ] || [ -L "$LEGACY_8081" ]; then
|
||||
TS="$(date -u +%Y%m%dT%H%M%SZ)"
|
||||
STASHED="$STASH_DIR/dsh.token.disabled-$TS"
|
||||
mv "$LEGACY_8081" "$STASHED"
|
||||
chmod 600 "$STASHED" 2>/dev/null || true
|
||||
touch "$PENDING_FILE"
|
||||
if ! NGINX_TEST_OUT="$(nginx -t 2>&1)"; then
|
||||
die "nginx config test failed after disabling $LEGACY_8081 (kept disabled at $STASHED): $NGINX_TEST_OUT; a pending reload is recorded so running nginx is reloaded once the config is fixed. The legacy :8081 endpoint will NOT be restored."
|
||||
fi
|
||||
if ! nginx -s reload; then
|
||||
die "nginx reload failed after disabling $LEGACY_8081 (kept disabled at $STASHED); a pending reload is recorded so running nginx is reloaded on the next run. The legacy :8081 endpoint will NOT be restored."
|
||||
fi
|
||||
rm -f "$PENDING_FILE"
|
||||
log "removed legacy :8081 endpoint -> $STASHED"
|
||||
fi
|
||||
|
||||
# ── 1b. Honor a recorded pending reload regardless of token selection ───────
|
||||
# A failed reload leaves PENDING_FILE set so a stashed legacy :8081 file can
|
||||
# never remain loaded in the running nginx while dsh-web is down or not yet
|
||||
# answering. Reconcile it before the token wait.
|
||||
if [ -e "$PENDING_FILE" ]; then
|
||||
if ! NGINX_TEST_OUT="$(nginx -t 2>&1)"; then
|
||||
log "WARNING: pending nginx reload recorded but 'nginx -t' fails: $NGINX_TEST_OUT; continuing so the include can be regenerated; will retry next run"
|
||||
elif ! nginx -s reload; then
|
||||
log "WARNING: pending nginx reload recorded but 'nginx -s reload' failed; will retry next run"
|
||||
else
|
||||
rm -f "$PENDING_FILE"
|
||||
log "completed pending nginx reload"
|
||||
fi
|
||||
fi
|
||||
|
||||
# ── 2. Select the token the RUNNING service actually accepts ────────────────
|
||||
# Re-sample the service's CURRENT systemd invocation on every pass and read
|
||||
# candidates only from it, so a restart that lands during the wait immediately
|
||||
# switches to the new invocation; there is no whole-journal or cross-invocation
|
||||
# fallback, and an empty/unknown invocation just waits. Each candidate is then
|
||||
# functionally verified against the local dsh-web using the public authority,
|
||||
# exactly as the /dsh-web-login proxy does, and the first that answers 303 is
|
||||
# the live token. Candidates are re-probed newest-first on each pass (connection
|
||||
# failures stay eligible) until one is accepted or the wait elapses.
|
||||
journal_tokens() {
|
||||
journalctl -u "$JOURNAL_UNIT" "_SYSTEMD_INVOCATION_ID=$1" --no-pager -o cat 2>/dev/null \
|
||||
| grep -oE 'dsh web: https?://[^[:space:]]+[?&]token=[^[:space:]]+' \
|
||||
| sed -E 's/.*[?&]token=//' \
|
||||
| grep -E '^[A-Za-z0-9._~+/=:@-]+$' \
|
||||
| tac | awk '!seen[$0]++' || true
|
||||
}
|
||||
|
||||
TOKEN=""
|
||||
DEADLINE=$((SECONDS + TOKEN_WAIT))
|
||||
NO_INVOCATION_WARNED=0
|
||||
while [ -z "$TOKEN" ] && [ "$SECONDS" -lt "$DEADLINE" ]; do
|
||||
INVOCATION="$(systemctl show -p InvocationID --value "$JOURNAL_UNIT" 2>/dev/null || true)"
|
||||
if [ -z "$INVOCATION" ] || [ "$INVOCATION" = "n/a" ]; then
|
||||
if [ "$NO_INVOCATION_WARNED" -eq 0 ]; then
|
||||
log "WARNING: no invocation id for $JOURNAL_UNIT; waiting for a live invocation"
|
||||
NO_INVOCATION_WARNED=1
|
||||
fi
|
||||
sleep 2
|
||||
continue
|
||||
fi
|
||||
for cand in $(journal_tokens "$INVOCATION"); do
|
||||
code="$(curl -s -o /dev/null --max-time 5 -w '%{http_code}' \
|
||||
-H "Host: $LOGIN_HOST" "$LOGIN_UPSTREAM/?token=$cand" || true)"
|
||||
if [ "$code" = "303" ]; then
|
||||
TOKEN="$cand"
|
||||
break
|
||||
fi
|
||||
done
|
||||
[ -n "$TOKEN" ] && break
|
||||
sleep 2
|
||||
done
|
||||
|
||||
if [ -z "$TOKEN" ]; then
|
||||
log "no accepted launch token in the current invocation within ${TOKEN_WAIT}s; leaving the include untouched for the next run"
|
||||
[ -e "$PENDING_FILE" ] && die "pending nginx reload could not be completed; will retry next run"
|
||||
exit 0
|
||||
fi
|
||||
|
||||
# ── 3. Record the token (atomic, private) ──────────────────────────────────
|
||||
mkdir -p "$(dirname "$TOKEN_FILE")"
|
||||
if ! printf '%s\n' "$TOKEN" | cmp -s - "$TOKEN_FILE" 2>/dev/null; then
|
||||
printf '%s\n' "$TOKEN" > "$TOKEN_FILE.tmp"
|
||||
chmod 600 "$TOKEN_FILE.tmp"
|
||||
mv "$TOKEN_FILE.tmp" "$TOKEN_FILE"
|
||||
log "recorded live launch token in $TOKEN_FILE"
|
||||
fi
|
||||
chmod 600 "$TOKEN_FILE"
|
||||
|
||||
# ── 4. Regenerate the nginx login include (reload only when it changes) ────
|
||||
NEW_INCLUDE="$(mktemp "$INCLUDE_FILE.XXXXXX")"
|
||||
printf 'proxy_pass %s/?token=%s;\n' "$LOGIN_UPSTREAM" "$TOKEN" > "$NEW_INCLUDE"
|
||||
chmod 600 "$NEW_INCLUDE"
|
||||
|
||||
# The stamp records the token nginx actually loaded. It is written only after a
|
||||
# successful reload, so the early exit is safe only when both the stamp and the
|
||||
# on-disk include agree with the live token; anything else falls through to the
|
||||
# reload path so the include can never silently diverge from what nginx serves.
|
||||
APPLIED=""
|
||||
[ -f "$STAMP_FILE" ] && APPLIED="$(cat "$STAMP_FILE" 2>/dev/null || true)"
|
||||
[ -f "$INCLUDE_FILE" ] && chmod 600 "$INCLUDE_FILE"
|
||||
[ -f "$STAMP_FILE" ] && chmod 600 "$STAMP_FILE"
|
||||
|
||||
if [ "$APPLIED" = "$TOKEN" ] && [ -f "$INCLUDE_FILE" ] && cmp -s "$NEW_INCLUDE" "$INCLUDE_FILE" \
|
||||
&& [ ! -e "$PENDING_FILE" ]; then
|
||||
rm -f "$NEW_INCLUDE"
|
||||
log "token unchanged; nginx not reloaded"
|
||||
exit 0
|
||||
fi
|
||||
|
||||
[ -e "$SITE_ENABLED" ] || { rm -f "$NEW_INCLUDE"; die "$SITE_ENABLED missing; refusing to reload"; }
|
||||
|
||||
RESTORE=""
|
||||
if [ -f "$INCLUDE_FILE" ]; then
|
||||
RESTORE="$(mktemp "$INCLUDE_FILE.bak.XXXXXX")"
|
||||
cp -p "$INCLUDE_FILE" "$RESTORE"
|
||||
chmod 600 "$RESTORE"
|
||||
fi
|
||||
|
||||
mv "$NEW_INCLUDE" "$INCLUDE_FILE"
|
||||
chmod 600 "$INCLUDE_FILE"
|
||||
|
||||
if ! NGINX_TEST_OUT="$(nginx -t 2>&1)"; then
|
||||
if [ -n "$RESTORE" ]; then
|
||||
mv "$RESTORE" "$INCLUDE_FILE"
|
||||
else
|
||||
rm -f "$INCLUDE_FILE"
|
||||
fi
|
||||
die "nginx config test failed: $NGINX_TEST_OUT; previous include restored"
|
||||
fi
|
||||
|
||||
if ! nginx -s reload; then
|
||||
if [ -n "$RESTORE" ]; then
|
||||
mv "$RESTORE" "$INCLUDE_FILE"
|
||||
else
|
||||
rm -f "$INCLUDE_FILE"
|
||||
fi
|
||||
touch "$PENDING_FILE"
|
||||
die "nginx reload failed; previous include restored; a pending reload is recorded so the next run retries"
|
||||
fi
|
||||
|
||||
if [ -n "$RESTORE" ]; then
|
||||
rm -f "$RESTORE"
|
||||
fi
|
||||
|
||||
rm -f "$PENDING_FILE"
|
||||
|
||||
printf '%s\n' "$TOKEN" > "$STAMP_FILE.tmp"
|
||||
chmod 600 "$STAMP_FILE.tmp"
|
||||
mv "$STAMP_FILE.tmp" "$STAMP_FILE"
|
||||
|
||||
log "token changed; nginx reloaded"
|
||||
log "login endpoint: https://$LOGIN_HOST/dsh-web-login (Authentik-gated)"
|
||||
@@ -29,7 +29,13 @@ LITELLM_PUBLIC = "https://litellm.sysloggh.net"
|
||||
LITELLM_BACKEND = "192.168.68.116"
|
||||
AUTH_HOST = "192.168.68.11"
|
||||
|
||||
SYNTHETIC_API_KEY = "sk-U_ydi3B-wfGU-_xESkoU1Q"
|
||||
# Load LiteLLM API key from file (durable, works in cron)
|
||||
LITELLM_KEY_FILE = "/root/.abiba-workspace/secrets/litellm-key.txt"
|
||||
try:
|
||||
with open(LITELLM_KEY_FILE) as f:
|
||||
SYNTHETIC_API_KEY = f.read().strip()
|
||||
except:
|
||||
SYNTHETIC_API_KEY = None # Fail loudly: report "no-key-file" in check
|
||||
|
||||
NOW = datetime.datetime.now()
|
||||
DATE_STR = NOW.strftime("%Y-%m-%d")
|
||||
@@ -68,7 +74,7 @@ def http_get(url, auth=None, timeout=10):
|
||||
cmd = f'curl -sfk --connect-timeout {timeout} -o /dev/null -w "%{{http_code}}" "{url}"'
|
||||
if auth:
|
||||
cmd = cmd.replace('"', '\\"')
|
||||
cmd = f'curl -sfk --connect-timeout {timeout} -u "{auth}" -o /dev/null -w "%{{http_code}}" "{url}"'
|
||||
cmd = f'curl -sfk --connect-timeout {timeout} -H "Authorization: Bearer {auth}" -o /dev/null -w "%{{http_code}}" "{url}"'
|
||||
r = subprocess.run(cmd, shell=True, capture_output=True, text=True, timeout=timeout+2)
|
||||
return r.stdout.strip() or "000"
|
||||
except:
|
||||
@@ -247,6 +253,10 @@ def collect():
|
||||
report["litellm"]["checks"].append({"name": "oidc-auth", "status": "pass" if auth_code in ("200","302") else "fail", "code": auth_code})
|
||||
|
||||
# Check 5: Synthetic API call through LiteLLM
|
||||
if SYNTHETIC_API_KEY is None:
|
||||
report["litellm"]["api_models"] = None
|
||||
report["litellm"]["checks"].append({"name": "api-endpoint", "status": "fail", "code": "no-key-file"})
|
||||
else:
|
||||
api_check = http_get(f"{LITELLM_PUBLIC}/v1/models", auth=SYNTHETIC_API_KEY)
|
||||
report["litellm"]["api_models"] = api_check
|
||||
report["litellm"]["checks"].append({"name": "api-endpoint", "status": "pass" if api_check == "200" else "fail", "code": api_check})
|
||||
|
||||
Executable
+206
@@ -0,0 +1,206 @@
|
||||
#!/usr/bin/env python3
|
||||
"""disk-gc-plan — turn a fleet disk scan into the GC action plan.
|
||||
|
||||
This is the executable side of `disk-gc-threat-response.prose.md`. It exists so the
|
||||
report-only gate is enforced by code that can be tested, rather than by prose the
|
||||
executor might misread.
|
||||
|
||||
THE HARD GATE: guests listed in the contract's `report_only_guests` block are
|
||||
DETECT-AND-REPORT-ONLY at EVERY level (AMBER, RED, CRITICAL). This tool will never
|
||||
emit a `gc-executor` action for one, so no GC command can be constructed for it.
|
||||
|
||||
The gate is keyed on GUEST identity — guest id, hostname, or IP — never on an agent
|
||||
name. An agent-name marker can silently miss the guest it lives on; a guest marker
|
||||
cannot.
|
||||
|
||||
The authoritative exclusion list lives in the contract itself (the fenced ```yaml
|
||||
block containing `report_only_guests:`). This tool reads it from there so there is
|
||||
only ever one copy.
|
||||
|
||||
Usage:
|
||||
disk-gc-plan.py --scan scan.json # [{"id":111,"usage_pct":84}, ...]
|
||||
cat scan.json | disk-gc-plan.py # same, via stdin
|
||||
disk-gc-plan.py --scan scan.json --json # machine-readable plan
|
||||
|
||||
Scan entries may carry any of: id / guest / vmid / ct / ctid, hostname / name, ip.
|
||||
A threshold-crossing entry with no recognizable identity is reported, never GC'd.
|
||||
Exit codes: 0 ok, 1 usage/parse error.
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import json
|
||||
import pathlib
|
||||
import re
|
||||
import sys
|
||||
|
||||
import yaml
|
||||
|
||||
REPO = pathlib.Path(__file__).resolve().parent.parent
|
||||
DEFAULT_CONTRACT = REPO / "disk-gc-threat-response.prose.md"
|
||||
|
||||
AMBER, RED, CRITICAL = 75, 85, 95
|
||||
|
||||
IDENTITY_FIELDS = ("id", "guest", "vmid", "ct", "ctid", "hostname", "name", "ip")
|
||||
IDENTITY_TYPE_PREFIX = re.compile(r"^(?:lxc|qemu)/")
|
||||
UNIDENTIFIED_REASON = "unidentified target - refusing to schedule GC"
|
||||
|
||||
|
||||
def load_report_only_guests(contract_path: pathlib.Path) -> list[dict]:
|
||||
"""Read the authoritative report_only_guests block out of the contract.
|
||||
|
||||
The contract carries it as a fenced ```yaml block. Parsing the declared,
|
||||
machine-readable block is the intended interface — the contract owns the list.
|
||||
"""
|
||||
text = contract_path.read_text(encoding="utf-8")
|
||||
for block in re.findall(r"```yaml\n(.*?)```", text, re.S):
|
||||
if "report_only_guests:" in block:
|
||||
data = yaml.safe_load(block)
|
||||
guests = data.get("report_only_guests") or []
|
||||
if not isinstance(guests, list):
|
||||
raise SystemExit("report_only_guests must be a list")
|
||||
if not guests:
|
||||
raise SystemExit(
|
||||
"report_only_guests is empty or missing - refusing to plan GC "
|
||||
"without the report-only gate"
|
||||
)
|
||||
for guest in guests:
|
||||
if not isinstance(guest, dict) or not _keys(guest):
|
||||
raise SystemExit(
|
||||
"report_only_guests entry has no recognizable identity key "
|
||||
f"(expected one of: {', '.join(IDENTITY_FIELDS)}): {guest!r}"
|
||||
)
|
||||
return guests
|
||||
raise SystemExit(
|
||||
f"no authoritative report_only_guests block found in {contract_path}"
|
||||
)
|
||||
|
||||
|
||||
def _canonical_number(number: float) -> str:
|
||||
if float(number).is_integer():
|
||||
return str(int(number))
|
||||
return str(number).strip().lower()
|
||||
|
||||
|
||||
def _normalize_identity(value: object) -> str:
|
||||
"""Canonicalise a guest identity so differently-encoded ids compare equal:
|
||||
numeric and numeric-string ids collapse to an integer string, Proxmox
|
||||
type prefixes and leading zeros are stripped, and hostnames/IPs are only
|
||||
trimmed and lowercased."""
|
||||
if isinstance(value, bool):
|
||||
return str(value).strip().lower()
|
||||
if isinstance(value, (int, float)):
|
||||
return _canonical_number(float(value))
|
||||
text = str(value).strip().lower()
|
||||
text = IDENTITY_TYPE_PREFIX.sub("", text)
|
||||
try:
|
||||
return _canonical_number(float(text))
|
||||
except ValueError:
|
||||
return text
|
||||
|
||||
|
||||
def _keys(entry: dict) -> set[str]:
|
||||
"""Guest/host identity keys, shared by exclusions and scan entries so the two
|
||||
sides of the gate can never key on different fields."""
|
||||
out: set[str] = set()
|
||||
for field in IDENTITY_FIELDS:
|
||||
value = entry.get(field)
|
||||
if value is None:
|
||||
continue
|
||||
key = _normalize_identity(value)
|
||||
if key:
|
||||
out.add(key)
|
||||
return out
|
||||
|
||||
|
||||
def level_for(pct: float) -> str | None:
|
||||
if pct >= CRITICAL:
|
||||
return "CRITICAL"
|
||||
if pct >= RED:
|
||||
return "RED"
|
||||
if pct >= AMBER:
|
||||
return "AMBER"
|
||||
return None
|
||||
|
||||
|
||||
def build_plan(scan: list[dict], report_only: list[dict]) -> list[dict]:
|
||||
excluded = [(e, _keys(e)) for e in report_only]
|
||||
plan: list[dict] = []
|
||||
for entry in scan:
|
||||
pct = entry.get("usage_pct")
|
||||
if pct is None:
|
||||
continue
|
||||
level = level_for(float(pct))
|
||||
if level is None:
|
||||
continue # GREEN: log only, no action
|
||||
scan_keys = _keys(entry)
|
||||
target = next(
|
||||
(entry.get(k) for k in IDENTITY_FIELDS if entry.get(k) not in (None, "")),
|
||||
"?",
|
||||
)
|
||||
if not scan_keys:
|
||||
plan.append({
|
||||
"target": target,
|
||||
"level": level,
|
||||
"pct": float(pct),
|
||||
"action": "report-only",
|
||||
"reason": UNIDENTIFIED_REASON,
|
||||
})
|
||||
continue
|
||||
match = next((e for e, keys in excluded if keys & scan_keys), None)
|
||||
if match is not None:
|
||||
plan.append({
|
||||
"target": target,
|
||||
"level": level,
|
||||
"pct": float(pct),
|
||||
"action": "report-only",
|
||||
"reason": match.get("reason", "").strip(),
|
||||
})
|
||||
else:
|
||||
plan.append({
|
||||
"target": target,
|
||||
"level": level,
|
||||
"pct": float(pct),
|
||||
"action": "gc-executor",
|
||||
})
|
||||
plan.sort(key=lambda row: row["pct"], reverse=True)
|
||||
return plan
|
||||
|
||||
|
||||
def main() -> int:
|
||||
ap = argparse.ArgumentParser(description="Plan disk GC actions with the report-only gate.")
|
||||
ap.add_argument("--scan", help="JSON file: list of {id|ct|hostname|ip, usage_pct}")
|
||||
ap.add_argument("--contract", default=str(DEFAULT_CONTRACT))
|
||||
ap.add_argument("--json", action="store_true", help="emit the plan as JSON")
|
||||
args = ap.parse_args()
|
||||
|
||||
raw = pathlib.Path(args.scan).read_text() if args.scan else sys.stdin.read()
|
||||
try:
|
||||
scan = json.loads(raw)
|
||||
except json.JSONDecodeError as exc:
|
||||
print(f"invalid scan JSON: {exc}", file=sys.stderr)
|
||||
return 1
|
||||
if not isinstance(scan, list):
|
||||
print("scan must be a JSON list", file=sys.stderr)
|
||||
return 1
|
||||
|
||||
report_only = load_report_only_guests(pathlib.Path(args.contract))
|
||||
plan = build_plan(scan, report_only)
|
||||
|
||||
if args.json:
|
||||
print(json.dumps(plan, indent=2))
|
||||
return 0
|
||||
|
||||
if not plan:
|
||||
print("no threats (nothing at or above 75%)")
|
||||
return 0
|
||||
for row in plan:
|
||||
if row["action"] == "report-only":
|
||||
print(f" {row['target']} {row['level']} {row['pct']}% -> REPORT-ONLY (no GC) — {row['reason']}")
|
||||
else:
|
||||
print(f" {row['target']} {row['level']} {row['pct']}% -> gc-executor")
|
||||
return 0
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
sys.exit(main())
|
||||
Executable
+238
@@ -0,0 +1,238 @@
|
||||
#!/usr/bin/env python3
|
||||
"""
|
||||
LiteLLM Health Check - Contract executor
|
||||
Runs all checks defined in litellm-health.prose.md and reports results.
|
||||
"""
|
||||
|
||||
import subprocess
|
||||
import sys
|
||||
import json
|
||||
import time
|
||||
import random
|
||||
|
||||
# Configuration
|
||||
BACKEND_HOST = "192.168.68.116"
|
||||
GPU_HOSTS = {
|
||||
"gpu-dense": "192.168.68.8",
|
||||
"gpu-vision": "192.168.68.110",
|
||||
"strix-moe": "192.168.68.15"
|
||||
}
|
||||
|
||||
def run_command(cmd, timeout=15):
|
||||
"""Run a command and return (exit_code, stdout, stderr)"""
|
||||
try:
|
||||
result = subprocess.run(
|
||||
cmd,
|
||||
shell=True,
|
||||
capture_output=True,
|
||||
text=True,
|
||||
timeout=timeout
|
||||
)
|
||||
return result.returncode, result.stdout.strip(), result.stderr.strip()
|
||||
except subprocess.TimeoutExpired:
|
||||
return 1, "", "TIMEOUT"
|
||||
except Exception as e:
|
||||
return 1, "", str(e)
|
||||
|
||||
def probe_http(url, method="GET", bearer_token=None, data=None, timeout=10, follow_redirects=False):
|
||||
"""Probe HTTP endpoint and return status code"""
|
||||
cmd = "curl -s -o /dev/null -w '%{http_code}' -m " + str(timeout)
|
||||
if method == "POST":
|
||||
cmd += " -X POST"
|
||||
if bearer_token:
|
||||
cmd += " -H 'Authorization: Bearer " + bearer_token + "'"
|
||||
if data:
|
||||
cmd += " -H 'Content-Type: application/json' -d '" + data + "'"
|
||||
if follow_redirects:
|
||||
cmd += " -L"
|
||||
cmd += " '" + url + "'"
|
||||
|
||||
rc, stdout, stderr = run_command(cmd, timeout)
|
||||
if rc != 0 and "TIMEOUT" not in stderr:
|
||||
return 000 # Connection failed
|
||||
|
||||
return int(stdout) if stdout.isdigit() else 000
|
||||
|
||||
def check_liveliness():
|
||||
"""Step 1: Liveliness probe"""
|
||||
code = probe_http("http://" + BACKEND_HOST + "/litellm/health/liveliness")
|
||||
return "Liveliness", code == 200, str(code) + " (target: " + BACKEND_HOST + "/litellm/health/liveliness)"
|
||||
|
||||
def check_containers():
|
||||
"""Step 2: Container health via SSH"""
|
||||
cmd = "ssh -o BatchMode=yes -o ConnectTimeout=5 -o StrictHostKeyChecking=no root@192.168.68.116 'docker ps --format \"{{.Names}} {{.Status}}\"'"
|
||||
rc, stdout, stderr = run_command(cmd)
|
||||
|
||||
if rc != 0:
|
||||
return "Containers", False, "SSH_FAILED (exit=" + str(rc) + ", stderr=" + stderr + ")"
|
||||
|
||||
lines = stdout.split('\n') if stdout else []
|
||||
container_count = len([l for l in lines if l.strip()])
|
||||
healthy = container_count >= 8
|
||||
return "Containers", healthy, str(container_count) + " containers"
|
||||
|
||||
def check_model_probes():
|
||||
"""Step 6: Model probe - all 4 aliases"""
|
||||
# Get monitor key
|
||||
monitor_key = run_command("ssh -o BatchMode=yes root@192.168.68.116 \"grep LITELLM_MONITOR_KEY /etc/litellm-monitor.env | cut -d= -f2\"")[1]
|
||||
|
||||
results = []
|
||||
|
||||
for model in ["gpu-dense", "gpu-vision", "strix-moe"]:
|
||||
# Single-host aliases: 30s timeout
|
||||
code = probe_http("http://" + BACKEND_HOST + "/litellm/v1/chat/completions",
|
||||
method="POST",
|
||||
bearer_token=monitor_key,
|
||||
data='{"model":"' + model + '","messages":[{"role":"user","content":"health ' + str(random.randint(1000, 9999)) + '"}],"max_tokens":4}',
|
||||
timeout=30)
|
||||
|
||||
results.append((model, code == 200, str(code) + " (target: " + BACKEND_HOST + "/litellm/v1/chat/completions, model=" + model + ")"))
|
||||
|
||||
# Pool alias (syslog-auto): 60s timeout, retry once on 000
|
||||
code = probe_http("http://" + BACKEND_HOST + "/litellm/v1/chat/completions",
|
||||
method="POST",
|
||||
bearer_token=monitor_key,
|
||||
data='{"model":"syslog-auto","messages":[{"role":"user","content":"health ' + str(random.randint(1000, 9999)) + '"}],"max_tokens":4}',
|
||||
timeout=60)
|
||||
|
||||
if code == 000:
|
||||
# Retry once with same timeout
|
||||
time.sleep(1)
|
||||
code = probe_http("http://" + BACKEND_HOST + "/litellm/v1/chat/completions",
|
||||
method="POST",
|
||||
bearer_token=monitor_key,
|
||||
data='{"model":"syslog-auto","messages":[{"role":"user","content":"health ' + str(random.randint(1000, 9999)) + '"}],"max_tokens":4}',
|
||||
timeout=60)
|
||||
|
||||
results.append(("syslog-auto", code == 200, str(code) + " (target: " + BACKEND_HOST + "/litellm/v1/chat/completions, model=syslog-auto)"))
|
||||
|
||||
return results
|
||||
|
||||
def check_admin_key_list():
|
||||
"""Step 8: Admin API key list - use two-step approach"""
|
||||
# Step 1: Get master key
|
||||
mk_cmd = "ssh -o BatchMode=yes root@192.168.68.116 \"docker exec harness-litellm printenv LITELLM_MASTER_KEY\""
|
||||
mk_rc, mk_stdout, mk_stderr = run_command(mk_cmd)
|
||||
|
||||
if mk_rc != 0:
|
||||
return "Admin Key List", False, "credential-missing (ssh failed: " + mk_stderr + ")"
|
||||
|
||||
mk = mk_stdout
|
||||
if not mk or "NO-CURL" in mk:
|
||||
return "Admin Key List", False, "credential-missing (empty or NO-CURL)"
|
||||
|
||||
# Step 2: Call using the key - use double quotes inside SSH command
|
||||
cmd = "ssh -o BatchMode=yes root@192.168.68.116 \"curl -s -H \\\"Authorization: Bearer " + mk + "\\\" http://127.0.0.1:4000/key/list\""
|
||||
rc, stdout, stderr = run_command(cmd)
|
||||
|
||||
if rc != 0:
|
||||
return "Admin Key List", False, "admin-call-failed (exit=" + str(rc) + ", stderr=" + stderr + ")"
|
||||
|
||||
# Try to parse the response
|
||||
try:
|
||||
data = json.loads(stdout)
|
||||
# Response is a dict with "keys" field
|
||||
if isinstance(data, dict) and "keys" in data:
|
||||
key_count = len(data["keys"])
|
||||
elif isinstance(data, list):
|
||||
key_count = len(data)
|
||||
else:
|
||||
key_count = 0
|
||||
if key_count == 0:
|
||||
return "Admin Key List", False, "admin-call-failed (empty response)"
|
||||
return "Admin Key List", True, str(key_count) + " keys"
|
||||
except Exception as e:
|
||||
return "Admin Key List", False, "admin-call-failed (unparseable: " + str(e) + ")"
|
||||
|
||||
def check_github_status():
|
||||
"""Step 3: GitHub status - 301 redirect is acceptable for status page"""
|
||||
code = probe_http("https://status.github.com/api/status.json", timeout=15)
|
||||
# GitHub status API returns 301 redirect, which is expected behavior
|
||||
return "GitHub Status", code == 301, str(code)
|
||||
|
||||
def check_prometheus():
|
||||
"""Step 4: Prometheus health"""
|
||||
code = probe_http("http://" + BACKEND_HOST + ":9090/-/healthy")
|
||||
return "Prometheus", code == 200, str(code) + " (target: " + BACKEND_HOST + ":9090/-/healthy)"
|
||||
|
||||
def check_grafana():
|
||||
"""Step 9: Grafana health"""
|
||||
code = probe_http("http://" + BACKEND_HOST + ":3001/api/health")
|
||||
return "Grafana", code == 200, str(code) + " (target: " + BACKEND_HOST + ":3001/api/health)"
|
||||
|
||||
def check_docker_stats():
|
||||
"""Step 10: Docker Stats health - fetch from CT 116 host"""
|
||||
# Docker stats is on localhost from CT 116
|
||||
cmd = "ssh -o BatchMode=yes root@192.168.68.116 'curl -s http://127.0.0.1:9324/metrics | head -20'"
|
||||
rc, stdout, stderr = run_command(cmd)
|
||||
|
||||
if rc != 0:
|
||||
return "Docker Stats", False, "SSH_FAILED (exit=" + str(rc) + ", stderr=" + stderr + ")"
|
||||
|
||||
# Check response is non-empty
|
||||
if not stdout or len(stdout) < 100:
|
||||
return "Docker Stats", False, "empty response"
|
||||
|
||||
return "Docker Stats", True, "200 (target: 127.0.0.1:9324/metrics from CT 116)"
|
||||
|
||||
def main():
|
||||
print("🏥 LiteLLM Health Check v1.0.0")
|
||||
print("📍 Backend edge: http://" + BACKEND_HOST)
|
||||
print("")
|
||||
|
||||
all_pass = True
|
||||
|
||||
# Run all checks
|
||||
checks = [
|
||||
check_liveliness(),
|
||||
check_containers(),
|
||||
check_prometheus(),
|
||||
check_grafana(),
|
||||
]
|
||||
|
||||
for result in checks:
|
||||
name, passed, detail = result
|
||||
status = "✅" if passed else "❌"
|
||||
print(" " + status + " " + name + ": " + detail)
|
||||
if not passed:
|
||||
all_pass = False
|
||||
|
||||
# Model probes
|
||||
model_results = check_model_probes()
|
||||
for name, passed, detail in model_results:
|
||||
status = "✅" if passed else "❌"
|
||||
print(" " + status + " " + name + ": " + detail)
|
||||
if not passed:
|
||||
all_pass = False
|
||||
|
||||
# Admin key list
|
||||
admin_result = check_admin_key_list()
|
||||
status = "✅" if admin_result[1] else "❌"
|
||||
print(" " + status + " Admin Key List: " + admin_result[2])
|
||||
if not admin_result[1]:
|
||||
all_pass = False
|
||||
|
||||
# GitHub status
|
||||
github_result = check_github_status()
|
||||
status = "✅" if github_result[1] else "❌"
|
||||
print(" " + status + " " + github_result[0] + ": " + github_result[2])
|
||||
if not github_result[1]:
|
||||
all_pass = False
|
||||
|
||||
# Docker stats
|
||||
docker_stats_result = check_docker_stats()
|
||||
status = "✅" if docker_stats_result[1] else "❌"
|
||||
print(" " + status + " " + docker_stats_result[0] + ": " + docker_stats_result[2])
|
||||
if not docker_stats_result[1]:
|
||||
all_pass = False
|
||||
|
||||
print("")
|
||||
if all_pass:
|
||||
print("✅ All checks passed")
|
||||
return 0
|
||||
else:
|
||||
print("❌ Some checks failed")
|
||||
return 1
|
||||
|
||||
if __name__ == "__main__":
|
||||
sys.exit(main())
|
||||
+7
-6
@@ -11,11 +11,13 @@ set -euo pipefail
|
||||
# ── CT ID → PVE Node mapping (maintained HERE, not in prose contracts) ──
|
||||
declare -A CT_NODES=(
|
||||
# amdpve (192.168.68.15)
|
||||
[111]=amdpve # tdunna
|
||||
[105]=amdpve # kagentz (was hwepve — corrected 2026-09-12; live per pvesh)
|
||||
[112]=amdpve # tanko
|
||||
[113]=amdpve # baggy
|
||||
[115]=amdpve # scottdenya
|
||||
[120]=amdpve # adguard2 (added 2026-09-12)
|
||||
# minipve (192.168.68.12)
|
||||
[100]=minipve # abiba (was hwepve)
|
||||
[102]=minipve # adguard (was acerpve)
|
||||
[104]=minipve # authentik
|
||||
[110]=minipve # gitea
|
||||
@@ -25,17 +27,16 @@ declare -A CT_NODES=(
|
||||
[106]=storepve # ra-h-os
|
||||
[107]=storepve # proxmox-backup
|
||||
[108]=storepve # media
|
||||
[111]=storepve # tdunna (was amdpve — corrected 2026-09-12; live per pvesh)
|
||||
[117]=storepve # zulip
|
||||
[118]=storepve # jdownloader
|
||||
# acerpve (192.168.68.9) — no CTs (bare metal GPU .8)
|
||||
[100]=minipve # abiba (was hwepve)
|
||||
[105]=minipve # kagentz (was hwepve)
|
||||
# ocupve (192.168.68.5) — no CTs (bare metal GPU .110)
|
||||
#
|
||||
# REMOVED CTs (migrated to bare metal, decommissioned, or VMs):
|
||||
# 101 llm-gpu → bare metal 192.168.68.8 (RTX 3090)
|
||||
# 103 ocu-llm → bare metal 192.168.68.110 (RTX 5070)
|
||||
# 109 docker-vm → KVM VM 192.168.68.7 (use direct SSH)
|
||||
# 101 llm-gpu → bare metal 192.168.68.8 (RTX 3090) [QEMU VM on acerpve]
|
||||
# 103 ocu-llm → bare metal 192.168.68.110 (RTX 5070) [QEMU VM on ocupve]
|
||||
# 109 docker-vm → KVM VM 192.168.68.7 (use direct SSH) [QEMU VM on storepve]
|
||||
)
|
||||
|
||||
# Each node must be root-accessible via SSH hostname
|
||||
|
||||
@@ -47,17 +47,19 @@ You are a code reviewer for OpenProse infrastructure contracts in the Syslog Sol
|
||||
The infrastructure-control.prose.md contract is the canonical reference for the cluster topology:
|
||||
|
||||
**Proxmox Cluster "Tabiri" (5 nodes):**
|
||||
- amdpve (192.168.68.15): tanko, tdunna, baggy, scottdenya
|
||||
- minipve (192.168.68.12): abiba, kagentz, adguard, authentik, gitea, syslog-api, infisical-vault
|
||||
- storepve (192.168.68.6): docker-vm, ra-h-os, PBS, media, jdownloader, zulip
|
||||
- amdpve (192.168.68.15): kagentz, tanko, baggy, scottdenya, adguard2
|
||||
- minipve (192.168.68.12): abiba, adguard, authentik, gitea, syslog-api, infisical-vault
|
||||
- storepve (192.168.68.6): docker-vm, ra-h-os, PBS, media, jdownloader, zulip, tdunna
|
||||
- acerpve (192.168.68.9): llm-gpu
|
||||
- ocupve (192.168.68.5): ocu-llm
|
||||
|
||||
**CT IDs (verified 2026-07-24 against PVE API):**
|
||||
**CT IDs (verified 2026-09-12 against PVE API):**
|
||||
100:abiba 102:adguard 104:authentik 105:kagentz 106:ra-h-os
|
||||
107:pbs 108:media 110:gitea 111:tdunna 112:tanko
|
||||
113:baggy 115:scottdenya 116:syslog-api 117:zulip
|
||||
118:jdownloader 119:infisical-vault
|
||||
118:jdownloader 119:infisical-vault 120:adguard2
|
||||
|
||||
**CT 111 (tdunna, 192.168.68.129) is REPORT-ONLY — Theo's box; alert only, never garbage-collect.**
|
||||
|
||||
**NO CT 122, CT 123, or .19 exist in the cluster.**
|
||||
|
||||
|
||||
+20
-23
@@ -41,7 +41,9 @@ notify() {
|
||||
# ── Global: Zulip Server ──
|
||||
SERVER_CODE=$(curl -s -o /dev/null -w "%{http_code}" --connect-timeout 10 \
|
||||
https://chat.sysloggh.net/api/v1/server_settings \
|
||||
-u 'abiba-bot@chat.sysloggh.net:cKTDMZAPW08dk3zl05sStzO7HRztzyn8' 2>/dev/null || echo "000")
|
||||
-u 'abiba-bot@chat.sysloggh.net:cKTDMZAPW08dk3zl05sStzO7HRztzyn8' 2>/dev/null) || SERVER_CODE="000"
|
||||
SERVER_CODE=$(printf '%s' "$SERVER_CODE" | tr -d '[:space:]')
|
||||
[ -n "$SERVER_CODE" ] || SERVER_CODE="000"
|
||||
if [ "$SERVER_CODE" != "200" ]; then
|
||||
notify "🔴" "Zulip server returned HTTP $SERVER_CODE"
|
||||
ISSUES=$((ISSUES + 1))
|
||||
@@ -59,7 +61,9 @@ fi
|
||||
# zulip.connected is a PROBE FAILURE: it alerts and NEVER calls pm2 restart.
|
||||
# pm2 restart runs ONLY on affirmative zulip.connected=false.
|
||||
# -- abiba-leg-start (verbatim-extracted by tests/zulip-monitor-abiba.sh)
|
||||
PI_HTTP=$(curl -s -o /dev/null --connect-timeout 5 --max-time 10 -w '%{http_code}' http://localhost:9200/health 2>/dev/null || echo "000")
|
||||
PI_HTTP=$(curl -s -o /dev/null --connect-timeout 5 --max-time 10 -w '%{http_code}' http://localhost:9200/health 2>/dev/null) || PI_HTTP="000"
|
||||
PI_HTTP=$(printf '%s' "$PI_HTTP" | tr -d '[:space:]')
|
||||
[ -n "$PI_HTTP" ] || PI_HTTP="000"
|
||||
PI_BODY=$(curl -s --connect-timeout 5 --max-time 10 http://localhost:9200/health 2>/dev/null || true)
|
||||
PI_STATE=$(printf '%s' "$PI_BODY" | python3 -c '
|
||||
import sys, json
|
||||
@@ -155,31 +159,24 @@ fi
|
||||
# never contact her former host.
|
||||
|
||||
# ── Platform C: Agent Zero (kagentz) ──
|
||||
AZ_A2A=$(ssh -o StrictHostKeyChecking=no -o ConnectTimeout=5 root@192.168.68.14 \
|
||||
"docker exec agent-zero curl -s --connect-timeout 5 http://127.0.0.1:8001/.well-known/agent.json 2>/dev/null" 2>/dev/null || echo "")
|
||||
AZ_ALIVE=$(echo "$AZ_A2A" | python3 -c "import sys,json; d=json.load(sys.stdin); print(d.get('name',''))" 2>/dev/null)
|
||||
AZ_A2A_CODE=$(ssh -o StrictHostKeyChecking=no -o ConnectTimeout=5 root@192.168.68.14 \
|
||||
"docker exec agent-zero curl -s --connect-timeout 5 -o /dev/null -w '%{http_code}' http://127.0.0.1:80/a2a/ 2>/dev/null" 2>/dev/null) || AZ_A2A_CODE="000"
|
||||
AZ_A2A_CODE=$(printf '%s' "$AZ_A2A_CODE" | tr -d '[:space:]')
|
||||
[ -n "$AZ_A2A_CODE" ] || AZ_A2A_CODE="000"
|
||||
|
||||
if [ "$AZ_ALIVE" != "kagentz" ]; then
|
||||
notify "🔴" "kagentz A2A server DOWN — restarting"
|
||||
ssh -o StrictHostKeyChecking=no -o ConnectTimeout=5 root@192.168.68.14 \
|
||||
"docker exec agent-zero bash -c 'pkill -9 -f a2a_agent; sleep 1; cd /a0 && /opt/venv-a0/bin/python3 -u /a0/usr/a2a_agent.py > /tmp/a2a.log 2>&1 &'" 2>/dev/null || true
|
||||
if [ "$AZ_A2A_CODE" = "000" ]; then
|
||||
notify "🔴" "kagentz A2A server DOWN (connection failed)"
|
||||
ISSUES=$((ISSUES + 1))
|
||||
echo " kagentz: ❌ A2A down — restarted" >> "$LOG"
|
||||
echo " kagentz: ❌ A2A down (HTTP 000)" >> "$LOG"
|
||||
else
|
||||
echo " kagentz: ✅ A2A alive" >> "$LOG"
|
||||
|
||||
# Check adapter process
|
||||
AZ_ADAPTER=$(ssh -o StrictHostKeyChecking=no -o ConnectTimeout=5 root@192.168.68.14 \
|
||||
"docker exec agent-zero ps aux 2>/dev/null | grep adapter | grep -v grep | wc -l" 2>/dev/null || echo "0")
|
||||
if [ "$AZ_ADAPTER" -lt 1 ]; then
|
||||
notify "🔴" "kagentz Zulip adapter DOWN — restarting"
|
||||
ssh -o StrictHostKeyChecking=no -o ConnectTimeout=5 root@192.168.68.14 \
|
||||
"docker exec agent-zero bash -c 'cd /a0/usr/kagentz-zulip && ZULIP_SITE=https://chat.sysloggh.net ZULIP_EMAIL=kagentz-bot@chat.sysloggh.net ZULIP_API_KEY=E9q9PXJTxftPYBkb5pBDWupDO7KK21ty ZULIP_AGENT_NAME=kagentz A2A_URL=http://localhost:8001/a2a A2A_TOKEN=8zNgdOEXzYxjQvTl /opt/venv-a0/bin/python3 -u adapter.py > /tmp/zulip-adapter.log 2>&1 &'" 2>/dev/null || true
|
||||
case "$AZ_A2A_CODE" in
|
||||
200|401)
|
||||
echo " kagentz: ✅ A2A alive (HTTP $AZ_A2A_CODE)" >> "$LOG" ;;
|
||||
*)
|
||||
notify "🟡" "kagentz A2A server answered HTTP $AZ_A2A_CODE — running, unexpected status"
|
||||
ISSUES=$((ISSUES + 1))
|
||||
echo " kagentz: ❌ Adapter down — restarted" >> "$LOG"
|
||||
else
|
||||
echo " kagentz: ✅ Adapter running" >> "$LOG"
|
||||
fi
|
||||
echo " kagentz: 🟡 A2A unexpected http=$AZ_A2A_CODE (running, warning)" >> "$LOG" ;;
|
||||
esac
|
||||
fi
|
||||
|
||||
# ── Summary ──
|
||||
|
||||
@@ -0,0 +1,191 @@
|
||||
"""Regression tests for the 2026-09-12 retired-alias sweep in audit-hermes-config.py.
|
||||
|
||||
WHY THIS FILE EXISTS: the executable audit pushed agent configs toward a DEAD alias.
|
||||
Rule 8 required `auxiliary.vision.model == "gpu-light"` and
|
||||
`auxiliary.web_extract.model == "gpu-light"`, but `gpu-light` (and its raw predecessor
|
||||
`gemma-4-12b`) were retired on 2026-09-12 and now return 400 `Invalid model name`; the
|
||||
live RTX 5070 alias is `gpu-vision`. A config that adopted the correct canonical alias
|
||||
therefore FAILED our own audit, so the audit was actively enforcing a broken config.
|
||||
|
||||
These tests execute the real CLI (`python3 audit-hermes-config.py <config>`) and assert
|
||||
observable behaviour — exit code and the emitted rule message — for the live alias and
|
||||
for both retired names. No network, vault, or SSH access is required.
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
import pathlib
|
||||
import subprocess
|
||||
import sys
|
||||
|
||||
ROOT = pathlib.Path(__file__).resolve().parent.parent
|
||||
AUDIT = ROOT / "audit-hermes-config.py"
|
||||
|
||||
BASE = """
|
||||
model:
|
||||
api_key: ""
|
||||
api_key_env: LITELLM_API_KEY
|
||||
base_url: http://192.168.68.116/v1
|
||||
max_tokens: 4096
|
||||
default: syslog-auto
|
||||
provider: harness
|
||||
fallback_providers:
|
||||
provider: deepseek
|
||||
model: deepseek-v4-flash
|
||||
api_key_env: DEEPSEEK_API_KEY
|
||||
compression:
|
||||
model: syslog-auto
|
||||
provider: harness
|
||||
threshold: 0.65
|
||||
max_context_window: 131072
|
||||
auxiliary:
|
||||
vision:
|
||||
model: {alias}
|
||||
provider: harness
|
||||
web_extract:
|
||||
model: {alias}
|
||||
provider: harness
|
||||
compression:
|
||||
model: syslog-auto
|
||||
provider: harness
|
||||
delegation:
|
||||
provider: harness
|
||||
custom_providers:
|
||||
- name: harness
|
||||
key_env: LITELLM_API_KEY
|
||||
base_url: http://192.168.68.116/v1
|
||||
"""
|
||||
|
||||
|
||||
def _run_config(tmp_path, name, text):
|
||||
cfg = tmp_path / name
|
||||
cfg.write_text(text)
|
||||
proc = subprocess.run(
|
||||
[sys.executable, str(AUDIT), str(cfg)],
|
||||
capture_output=True, text=True,
|
||||
)
|
||||
return proc.returncode, proc.stdout
|
||||
|
||||
|
||||
def _run(tmp_path, alias):
|
||||
return _run_config(tmp_path, f"{alias}.yaml", BASE.format(alias=alias))
|
||||
|
||||
|
||||
def test_live_canonical_alias_passes(tmp_path):
|
||||
"""The RTX 5070 alias that actually resolves must satisfy Rule 8."""
|
||||
code, out = _run(tmp_path, "gpu-vision")
|
||||
assert code == 0, out
|
||||
assert "RESULT: PASS" in out
|
||||
|
||||
|
||||
def test_retired_gpu_light_is_rejected(tmp_path):
|
||||
"""A config pinned to the retired alias must fail, not pass."""
|
||||
code, out = _run(tmp_path, "gpu-light")
|
||||
assert code == 1, out
|
||||
assert "auxiliary.vision.model must be gpu-vision" in out
|
||||
assert "RESULT: FAIL" in out
|
||||
|
||||
|
||||
def test_retired_gemma_is_rejected(tmp_path):
|
||||
"""The retired raw model name must fail Rule 8 as well."""
|
||||
code, out = _run(tmp_path, "gemma-4-12b")
|
||||
assert code == 1, out
|
||||
assert "auxiliary.vision.model must be gpu-vision" in out
|
||||
assert "RESULT: FAIL" in out
|
||||
|
||||
|
||||
def test_corrected_compression_example_passes(tmp_path):
|
||||
"""The corrected workaround (vision=gpu-vision, compression=syslog-auto) must PASS."""
|
||||
code, out = _run(tmp_path, "gpu-vision")
|
||||
assert code == 0, out
|
||||
assert "[Rule 7] compression.model must be syslog-auto (got 'syslog-auto')" in out
|
||||
assert "[Rule 7] auxiliary.compression.model must be syslog-auto (got 'syslog-auto')" in out
|
||||
assert "RESULT: PASS" in out
|
||||
|
||||
|
||||
def test_retired_alias_in_delegation_is_rejected(tmp_path):
|
||||
"""delegation.model has no dedicated value rule, so a retired name there used to PASS."""
|
||||
code, out = _run_config(
|
||||
tmp_path,
|
||||
"delegation-gpu-light.yaml",
|
||||
BASE.format(alias="gpu-vision").replace(
|
||||
"delegation:\n provider: harness",
|
||||
"delegation:\n provider: harness\n model: gpu-light",
|
||||
),
|
||||
)
|
||||
assert code == 1, out
|
||||
assert "delegation.model = 'gpu-light' is retired" in out
|
||||
assert "RESULT: FAIL" in out
|
||||
|
||||
|
||||
def test_retired_alias_in_custom_providers_is_rejected(tmp_path):
|
||||
"""custom_providers[*].model is model-bearing; a retired name there must fail."""
|
||||
code, out = _run_config(
|
||||
tmp_path,
|
||||
"custom-provider-gpu-light.yaml",
|
||||
BASE.format(alias="gpu-vision").replace(
|
||||
" - name: harness\n key_env: LITELLM_API_KEY",
|
||||
" - name: harness\n model: gpu-light\n key_env: LITELLM_API_KEY",
|
||||
),
|
||||
)
|
||||
assert code == 1, out
|
||||
assert "custom_providers[0].model = 'gpu-light'" in out
|
||||
assert "RESULT: FAIL" in out
|
||||
|
||||
|
||||
def test_retired_raw_name_fails(tmp_path):
|
||||
"""Retired raw names no longer resolve (400), so they fail; the 2026-09-12 registry change moved qwen3.6-27B-code from raw-but-live to non-resolving."""
|
||||
code, out = _run_config(
|
||||
tmp_path,
|
||||
"retired-qwen.yaml",
|
||||
BASE.format(alias="gpu-vision").replace(
|
||||
"delegation:\n provider: harness",
|
||||
"delegation:\n provider: harness\n model: qwen3.6-27B-code",
|
||||
),
|
||||
)
|
||||
assert code != 0, out
|
||||
assert "delegation.model = 'qwen3.6-27B-code' is retired and no longer resolves" in out
|
||||
assert "use gpu-dense" in out
|
||||
assert "RESULT: FAIL" in out
|
||||
|
||||
|
||||
def test_retired_alias_in_fallback_providers_is_rejected(tmp_path):
|
||||
"""fallback_providers.model is model-bearing; a retired name there must fail."""
|
||||
code, out = _run_config(
|
||||
tmp_path,
|
||||
"fallback-gpu-light.yaml",
|
||||
BASE.format(alias="gpu-vision").replace(" model: deepseek-v4-flash", " model: gpu-light"),
|
||||
)
|
||||
assert code == 1, out
|
||||
assert "fallback_providers.model = 'gpu-light'" in out
|
||||
assert "RESULT: FAIL" in out
|
||||
|
||||
|
||||
def test_retired_alias_in_x_search_is_rejected(tmp_path):
|
||||
"""x_search.model was previously not enumerated; the derivation must catch it."""
|
||||
code, out = _run_config(
|
||||
tmp_path,
|
||||
"x-search-gpu-light.yaml",
|
||||
BASE.format(alias="gpu-vision").replace(
|
||||
"delegation:\n provider: harness",
|
||||
"delegation:\n provider: harness\nx_search:\n model: gpu-light",
|
||||
),
|
||||
)
|
||||
assert code == 1, out
|
||||
assert "x_search.model = 'gpu-light'" in out
|
||||
assert "RESULT: FAIL" in out
|
||||
|
||||
|
||||
def test_retired_alias_in_nested_auxiliary_block_is_rejected(tmp_path):
|
||||
"""A nested auxiliary sub-block outside the named three must still be derived."""
|
||||
code, out = _run_config(
|
||||
tmp_path,
|
||||
"nested-aux-gpu-light.yaml",
|
||||
BASE.format(alias="gpu-vision").replace(
|
||||
" compression:\n model: syslog-auto\n provider: harness\ndelegation:",
|
||||
" compression:\n model: syslog-auto\n provider: harness\n"
|
||||
" tasks:\n summarize:\n model: gpu-light\ndelegation:",
|
||||
),
|
||||
)
|
||||
assert code == 1, out
|
||||
assert "auxiliary.tasks.summarize.model = 'gpu-light'" in out
|
||||
assert "RESULT: FAIL" in out
|
||||
@@ -0,0 +1,185 @@
|
||||
"""Regression tests for the disk-gc report-only gate (CT 111 / tdunna / .129).
|
||||
|
||||
WHY THIS FILE EXISTS: `disk-gc-threat-response.prose.md` defined AMBER as "GC scheduled
|
||||
for next run" and its Execution loop called `gc-executor` for EVERY threat, with no
|
||||
guest-level exclusion. CT 111 (tdunna, 192.168.68.129) belongs to Theo and is
|
||||
report-only per the captain (2026-08-17, re-confirmed 2026-09-10) — so a single AMBER
|
||||
reading on that guest would have scheduled GC commands (apt clean, journal vacuum,
|
||||
log/tmp deletion, snap removal) against someone else's box. The only marker was
|
||||
frontmatter `report_only_agents`, which names an AGENT while the scan unit is a GUEST.
|
||||
|
||||
These tests execute the real planner (`scripts/disk-gc-plan.py`) and assert observable
|
||||
behaviour: an excluded guest never produces a `gc-executor` action at any level, while
|
||||
our own guests still do.
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
import json
|
||||
import pathlib
|
||||
import subprocess
|
||||
import sys
|
||||
|
||||
ROOT = pathlib.Path(__file__).resolve().parent.parent
|
||||
PLAN = ROOT / "scripts" / "disk-gc-plan.py"
|
||||
|
||||
|
||||
def _plan(scan, tmp_path):
|
||||
scan_file = tmp_path / "scan.json"
|
||||
scan_file.write_text(json.dumps(scan))
|
||||
proc = subprocess.run(
|
||||
[sys.executable, str(PLAN), "--scan", str(scan_file), "--json"],
|
||||
capture_output=True, text=True,
|
||||
)
|
||||
assert proc.returncode == 0, proc.stderr
|
||||
return json.loads(proc.stdout)
|
||||
|
||||
|
||||
def _actions_for(plan, target):
|
||||
return [row for row in plan if str(row["target"]) == str(target)]
|
||||
|
||||
|
||||
def test_excluded_guest_never_gets_gc_at_any_level(tmp_path):
|
||||
"""CT 111 at AMBER, RED and CRITICAL — always report-only, never gc-executor."""
|
||||
for pct, level in ((84, "AMBER"), (90, "RED"), (97, "CRITICAL")):
|
||||
plan = _plan([{"id": 111, "hostname": "tdunna", "ip": "192.168.68.129",
|
||||
"usage_pct": pct}], tmp_path)
|
||||
rows = _actions_for(plan, 111)
|
||||
assert rows, f"CT 111 must still be reported at {level}"
|
||||
assert rows[0]["level"] == level
|
||||
assert rows[0]["action"] == "report-only", rows
|
||||
assert not any(r["action"] == "gc-executor" for r in rows)
|
||||
|
||||
|
||||
def test_exclusion_matches_on_any_identity_key(tmp_path):
|
||||
"""The gate is keyed on guest/host, so id, ct, ctid, hostname or IP all match."""
|
||||
for entry in ({"id": 111, "usage_pct": 95},
|
||||
{"ct": 111, "usage_pct": 95},
|
||||
{"ctid": 111, "usage_pct": 95},
|
||||
{"hostname": "tdunna", "usage_pct": 95},
|
||||
{"ip": "192.168.68.129", "usage_pct": 95}):
|
||||
plan = _plan([entry], tmp_path)
|
||||
assert all(r["action"] == "report-only" for r in plan), (entry, plan)
|
||||
|
||||
|
||||
def test_exclusion_matches_encoded_identities(tmp_path):
|
||||
"""A differently-encoded CT 111 id must not slip past the gate to gc-executor."""
|
||||
encodings = ({"id": 111.0},
|
||||
{"id": "111.0"},
|
||||
{"id": "lxc/111"},
|
||||
{"id": "qemu/111"},
|
||||
{"id": "0111"},
|
||||
{"id": " 111 "})
|
||||
for alias in encodings:
|
||||
plan = _plan([{**alias, "usage_pct": 97}], tmp_path)
|
||||
assert plan, alias
|
||||
assert plan[0]["action"] == "report-only", (alias, plan)
|
||||
|
||||
|
||||
def test_our_own_guests_still_get_gc(tmp_path):
|
||||
"""acerpve .9 and amdpve .15 are ours — they must still be acted on."""
|
||||
plan = _plan([{"hostname": "acerpve", "ip": "192.168.68.9", "usage_pct": 77},
|
||||
{"hostname": "amdpve", "ip": "192.168.68.15", "usage_pct": 76}], tmp_path)
|
||||
assert len(plan) == 2
|
||||
assert all(r["action"] == "gc-executor" for r in plan), plan
|
||||
|
||||
|
||||
def test_below_threshold_emits_nothing(tmp_path):
|
||||
"""GREEN guests produce no action at all."""
|
||||
assert _plan([{"id": 111, "usage_pct": 40}], tmp_path) == []
|
||||
|
||||
|
||||
def test_agent_name_alone_does_not_gate_a_guest(tmp_path):
|
||||
"""An agent-name marker must not be the gate: an unrelated guest still gets GC."""
|
||||
plan = _plan([{"id": 999, "hostname": "koby", "usage_pct": 95}], tmp_path)
|
||||
assert plan and plan[0]["action"] == "gc-executor"
|
||||
|
||||
|
||||
def test_unidentified_threat_fails_closed(tmp_path):
|
||||
"""A threshold-crossing entry with no recognized identity must not schedule GC."""
|
||||
plan = _plan([{"usage_pct": 97}], tmp_path)
|
||||
assert plan, "an unidentified threat must still be reported"
|
||||
assert plan[0]["action"] == "report-only", plan
|
||||
assert plan[0]["reason"], plan
|
||||
|
||||
|
||||
def test_exclusion_entry_without_identity_fails_closed(tmp_path):
|
||||
"""A mis-typed exclusion entry must break the run, never silently disable the gate."""
|
||||
contract = tmp_path / "broken.prose.md"
|
||||
contract.write_text(
|
||||
"```yaml\n"
|
||||
"report_only_guests:\n"
|
||||
" - node: storepve\n"
|
||||
" reason: \"typo - no guest identity\"\n"
|
||||
"```\n"
|
||||
)
|
||||
scan_file = tmp_path / "scan.json"
|
||||
scan_file.write_text(json.dumps([{"ct": 111, "usage_pct": 97}]))
|
||||
proc = subprocess.run(
|
||||
[sys.executable, str(PLAN), "--scan", str(scan_file),
|
||||
"--contract", str(contract), "--json"],
|
||||
capture_output=True, text=True,
|
||||
)
|
||||
assert proc.returncode != 0, proc.stdout
|
||||
assert "gc-executor" not in proc.stdout
|
||||
assert "identity" in proc.stderr.lower(), proc.stderr
|
||||
|
||||
|
||||
def test_empty_exclusion_block_fails_closed(tmp_path):
|
||||
"""An emptied report_only_guests list must break the run, not disable the gate."""
|
||||
contract = tmp_path / "empty.prose.md"
|
||||
contract.write_text("```yaml\nreport_only_guests: []\n```\n")
|
||||
scan_file = tmp_path / "scan.json"
|
||||
scan_file.write_text(json.dumps([{"ct": 111, "usage_pct": 97}]))
|
||||
proc = subprocess.run(
|
||||
[sys.executable, str(PLAN), "--scan", str(scan_file),
|
||||
"--contract", str(contract), "--json"],
|
||||
capture_output=True, text=True,
|
||||
)
|
||||
assert proc.returncode != 0, proc.stdout
|
||||
assert "gc-executor" not in proc.stdout
|
||||
assert "report-only gate" in proc.stderr.lower(), proc.stderr
|
||||
|
||||
|
||||
def _plan_with_contract(scan, contract_text, tmp_path, name):
|
||||
contract = tmp_path / name
|
||||
contract.write_text(contract_text)
|
||||
scan_file = tmp_path / f"scan-{name}.json"
|
||||
scan_file.write_text(json.dumps(scan))
|
||||
proc = subprocess.run(
|
||||
[sys.executable, str(PLAN), "--scan", str(scan_file),
|
||||
"--contract", str(contract), "--json"],
|
||||
capture_output=True, text=True,
|
||||
)
|
||||
assert proc.returncode == 0, proc.stderr
|
||||
return json.loads(proc.stdout)
|
||||
|
||||
|
||||
def test_gate_is_read_from_the_contract_block(tmp_path):
|
||||
"""The gate is data-driven by the contract block: the planner excludes the guest
|
||||
when the block names it and acts on it when the block does not. Executes the real
|
||||
planner interface against both fixtures so the behaviour change is observable."""
|
||||
scan = [{"id": 111, "hostname": "tdunna", "ip": "192.168.68.129", "usage_pct": 95}]
|
||||
with_gate = (
|
||||
"```yaml\n"
|
||||
"report_only_guests:\n"
|
||||
" - guest: 111\n"
|
||||
" hostname: tdunna\n"
|
||||
" ip: 192.168.68.129\n"
|
||||
" reason: \"fixture reason\"\n"
|
||||
"```\n"
|
||||
)
|
||||
without_gate = (
|
||||
"```yaml\n"
|
||||
"report_only_guests:\n"
|
||||
" - guest: 999\n"
|
||||
" reason: \"fixture excludes a different guest\"\n"
|
||||
"```\n"
|
||||
)
|
||||
|
||||
gated = _plan_with_contract(scan, with_gate, tmp_path, "gated.prose.md")
|
||||
assert gated[0]["action"] == "report-only", gated
|
||||
assert gated[0]["reason"] == "fixture reason", gated
|
||||
|
||||
ungated = _plan_with_contract(scan, without_gate, tmp_path, "ungated.prose.md")
|
||||
assert ungated[0]["action"] == "gc-executor", ungated
|
||||
assert ungated[0]["action"] != gated[0]["action"]
|
||||
@@ -89,8 +89,7 @@ case "$host" in
|
||||
esac ;;
|
||||
192.168.68.14)
|
||||
case "$cmd" in
|
||||
*agent.json*) printf '%s' "$AZ_A2A" ;;
|
||||
*"ps aux"*) printf '%s\n' "$AZ_PS" ;;
|
||||
*"/a2a/"*) printf '%s' "$AZ_A2A_CODE"; exit "$AZ_A2A_EXIT" ;;
|
||||
esac ;;
|
||||
*)
|
||||
printf 'UNEXPECTED-SSH-HOST %s\n' "$host" >> "$RECORD_DIR/unexpected-ssh" ;;
|
||||
@@ -122,8 +121,7 @@ def _write_exec(path: pathlib.Path, body: str) -> None:
|
||||
|
||||
|
||||
def _run_monitor(tmp_path, *, tanko_svc="active", tanko_http="200",
|
||||
az_a2a='{"name":"kagentz"}',
|
||||
az_ps="root 111 0.1 0.2 /opt/venv-a0/bin/python3 -u adapter.py"):
|
||||
az_a2a_code="401", az_a2a_exit=0):
|
||||
"""Run the shipped monitor in a sandbox; return (proc, record_dir, log_path).
|
||||
|
||||
Only the LOG constant is rewritten (to keep the run inside the worktree).
|
||||
@@ -151,8 +149,8 @@ def _run_monitor(tmp_path, *, tanko_svc="active", tanko_http="200",
|
||||
"RECORD_DIR": str(record),
|
||||
"TANKO_SVC": tanko_svc,
|
||||
"TANKO_HTTP": tanko_http,
|
||||
"AZ_A2A": az_a2a,
|
||||
"AZ_PS": az_ps,
|
||||
"AZ_A2A_CODE": az_a2a_code,
|
||||
"AZ_A2A_EXIT": str(az_a2a_exit),
|
||||
"PI_HTTP": "200",
|
||||
"PI_BODY": CONNECTED_FIXTURE.read_text(),
|
||||
"SERVER_HTTP": "200",
|
||||
@@ -171,8 +169,7 @@ def test_healthy_run_is_quiet_and_never_reaches_mumuni(tmp_path):
|
||||
assert "Server: ✅ HTTP 200" in log
|
||||
assert "Abiba: ✅ Connected" in log
|
||||
assert "Tanko: ✅ service=active http=200" in log
|
||||
assert "kagentz: ✅ A2A alive" in log
|
||||
assert "kagentz: ✅ Adapter running" in log
|
||||
assert "kagentz: ✅ A2A alive (HTTP 401)" in log
|
||||
assert "Result: ✅ All healthy" in log
|
||||
|
||||
# A healthy run emits no notify at all — and certainly no Mumuni one.
|
||||
@@ -214,6 +211,32 @@ def test_failing_run_alerts_on_tanko_but_never_on_mumuni(tmp_path):
|
||||
assert "Result: 🔴 1 issue(s) found" in log
|
||||
|
||||
|
||||
def test_unexpected_a2a_status_is_an_issue_not_healthy(tmp_path):
|
||||
proc, record, log_path = _run_monitor(tmp_path, az_a2a_code="500")
|
||||
assert proc.returncode == 0, proc.stderr
|
||||
log = log_path.read_text()
|
||||
|
||||
assert "kagentz: 🟡 A2A unexpected http=500 (running, warning)" in log
|
||||
assert "kagentz: ✅ A2A alive" not in log
|
||||
assert "Result: 🔴 1 issue(s) found" in log
|
||||
assert "kagentz A2A server answered HTTP 500" in proc.stdout
|
||||
|
||||
|
||||
def test_a2a_connection_failure_is_down_not_unexpected(tmp_path):
|
||||
# curl prints the http_code before failing, so the ssh stub exits non-zero
|
||||
# with "000" on stdout — exercising the real outage path.
|
||||
proc, record, log_path = _run_monitor(tmp_path, az_a2a_code="000",
|
||||
az_a2a_exit=7)
|
||||
assert proc.returncode == 0, proc.stderr
|
||||
log = log_path.read_text()
|
||||
|
||||
assert "kagentz: ❌ A2A down (HTTP 000)" in log
|
||||
assert "kagentz: ✅ A2A alive" not in log
|
||||
assert "unexpected" not in log
|
||||
assert "Result: 🔴 1 issue(s) found" in log
|
||||
assert "kagentz A2A server DOWN (connection failed)" in proc.stdout
|
||||
|
||||
|
||||
# ── scripts/daily-infra-report.py: behavioral digest checks ──────────
|
||||
|
||||
@pytest.fixture(scope="module")
|
||||
@@ -334,7 +357,8 @@ def test_agent_health_roster_has_no_mumuni_entry(ahc):
|
||||
def test_health_contract_retires_mumuni_only_steps():
|
||||
text = HEALTH_CONTRACT.read_text()
|
||||
assert MUMUNI_IP not in text
|
||||
for step in ("**B4:", "**B5:", "**B6:"):
|
||||
for step in ("**B4: Gateway Process**", "**B5: Heartbeat Verification**",
|
||||
"**B6: Response Delivery**"):
|
||||
assert step not in text
|
||||
|
||||
|
||||
|
||||
+191
-28
@@ -1,9 +1,9 @@
|
||||
---
|
||||
kind: responsibility
|
||||
name: zulip-health
|
||||
description: Multi-platform health monitor for the Zulip messaging mesh spanning Platform A (pi/Abiba Zulip bridge), Platform B (Tanko on DSH), and Platform C (Agent Zero Docker). Verifies bot registration, DM delivery, and cross-platform connectivity. Mumuni is no longer monitored from this host — she runs on her own container (kagentz CT 105 on minipve, .14) and is monitored on her side.
|
||||
description: Multi-platform health monitor for the Zulip messaging mesh spanning Platform A (pi/Abiba Zulip bridge), Platform B (Tanko on DSH), and Platform C (Agent Zero Docker). Verifies bot registration, DM delivery, and cross-platform connectivity. The kagentz Zulip adapter leg is retired (its code no longer exists) — Agent Zero is probed for A2A liveness only. Mumuni is no longer monitored from this host — she runs on her own container (kagentz CT 105 on minipve, .14) and is monitored on her side.
|
||||
title: Zulip Mesh Health Monitor — Multi-Platform
|
||||
version: 3.1.0
|
||||
version: 3.3.0
|
||||
runtime_contract: 2
|
||||
agent: abiba
|
||||
report_only_agents:
|
||||
@@ -261,38 +261,203 @@ logged/reported as a warning — reported, never healed on.
|
||||
| HTTP `:3080` connection refused/timeout (`000`) | Same as above |
|
||||
| HTTP status outside the expected set | Log/report as a warning — reported, never healed on |
|
||||
|
||||
**B4: dsh-web Authentication (Tanko — restart-persistent login)**
|
||||
|
||||
The dsh-web UI is token-gated. On every start the process prints a random
|
||||
launch token to the journal:
|
||||
|
||||
```
|
||||
dsh web: http://127.0.0.1:3080/?token=<TOKEN>
|
||||
```
|
||||
|
||||
The token only bootstraps an authority-bound, HMAC-signed browser cookie with a
|
||||
30-day lifetime. The signing secret is durable in
|
||||
`/root/.dsh/.credentials.yaml` (key `client-connection/browser-session`), so a
|
||||
cookie minted once keeps working across `dsh-web` restarts; the launch token
|
||||
itself rotates on every restart.
|
||||
|
||||
**Login endpoint (public, Authentik-gated):**
|
||||
`https://tankodhs.sysloggh.net/dsh-web-login`
|
||||
|
||||
It lives inside the Authentik-gated `:80` server block
|
||||
(`/etc/nginx/sites-available/dsh`, symlinked from
|
||||
`/etc/nginx/sites-enabled/dsh`) as `location = /dsh-web-login`, guarded by
|
||||
`auth_request /outpost.goauthentik.io/auth/nginx`. It proxies to dsh-web with
|
||||
`Host: tankodhs.sysloggh.net`, so the minted cookie is bound to the public
|
||||
authority — never to `127.0.0.1:3080`. The token-dependent line is isolated in
|
||||
the generated include `/etc/dsh-web/nginx-login.conf`:
|
||||
|
||||
```
|
||||
proxy_pass http://127.0.0.1:3080/?token=<TOKEN>;
|
||||
```
|
||||
|
||||
**Token refresh (non-disruptive):**
|
||||
`/opt/deepseek-harness/capture-dsh-token.sh` (source:
|
||||
`scripts/capture-dsh-token.sh`) reads candidate launch tokens from the journal
|
||||
**scoped to the service's current systemd invocation**
|
||||
(`systemctl show -p InvocationID` + `_SYSTEMD_INVOCATION_ID=`), re-sampling the
|
||||
invocation on every pass so a restart that lands during the wait switches to the
|
||||
new invocation; a restarted process's stale token is never considered while its
|
||||
new startup banner is still pending and there is no whole-journal or
|
||||
cross-invocation fallback. Each candidate
|
||||
is then functionally verified against dsh-web with `Host: tankodhs.sysloggh.net`,
|
||||
using the first the running process accepts with `303`. It waits up to 120s for
|
||||
a restarted process to accept a token and re-probes every current-invocation
|
||||
candidate on each pass, so a token that briefly returns `000` while the service
|
||||
is still starting is not disqualified. If none is accepted it leaves the include
|
||||
untouched and exits so the timer retries (exiting non-zero when a pending reload
|
||||
is still outstanding). It writes
|
||||
`/etc/dsh-web/launch-token` and regenerates `/etc/dsh-web/nginx-login.conf`,
|
||||
reloading nginx only when the on-disk include differs from the generated one or
|
||||
the applied-state stamp does not match the token (`nginx -t` guards the reload,
|
||||
and the stamp is written only after a successful `nginx -s reload`, so a failed
|
||||
or interrupted reload is retried on the next run). Any failed reload records a
|
||||
pending-reload marker under `/etc/dsh-web/`; the next run attempts the reload
|
||||
before the token wait, independent of token state, and clears the marker only
|
||||
once the reload succeeds, so a disabled legacy `:8081` file can never leave the
|
||||
running nginx unreloaded. The generated include is recreated before any
|
||||
`nginx -t` if it is missing, so a failed run cannot wedge recovery.
|
||||
Runs are serialized with `flock` on `/run/capture-dsh-token.lock`. It **never
|
||||
stops or starts `dsh-web`**.
|
||||
It is triggered by the `dsh-web.service` drop-in
|
||||
`/etc/systemd/system/dsh-web.service.d/20-token-refresh.conf`
|
||||
(`ExecStartPost=/bin/systemctl --no-block start dsh-web-token.service`) and by
|
||||
`dsh-web-token.timer` every 2 minutes for reconciliation.
|
||||
|
||||
<details><summary>Installed systemd wiring (CT 112)</summary>
|
||||
|
||||
```ini
|
||||
# /etc/systemd/system/dsh-web-token.service
|
||||
[Unit]
|
||||
Description=Refresh the dsh-web launch token for the nginx login endpoint
|
||||
After=dsh-web.service
|
||||
[Service]
|
||||
Type=oneshot
|
||||
TimeoutStartSec=180
|
||||
ExecStart=/opt/deepseek-harness/capture-dsh-token.sh
|
||||
|
||||
# /etc/systemd/system/dsh-web-token.timer
|
||||
[Unit]
|
||||
Description=Periodically refresh the dsh-web login token
|
||||
[Timer]
|
||||
OnBootSec=90s
|
||||
OnUnitActiveSec=120s
|
||||
AccuracySec=10s
|
||||
Persistent=true
|
||||
[Install]
|
||||
WantedBy=timers.target
|
||||
|
||||
# /etc/systemd/system/dsh-web.service.d/20-token-refresh.conf
|
||||
[Service]
|
||||
ExecStartPost=/bin/systemctl --no-block start dsh-web-token.service
|
||||
```
|
||||
|
||||
</details>
|
||||
|
||||
> **Do NOT reintroduce the `:8081` endpoint.** It listened on `0.0.0.0:8081`
|
||||
> with no `auth_request` and was a full Authentik bypass for anyone on the LAN.
|
||||
> The script now removes `/etc/nginx/sites-enabled/dsh.token` automatically if
|
||||
> it ever reappears.
|
||||
|
||||
**Authentication flow:**
|
||||
1. `GET https://tankodhs.sysloggh.net/dsh-web-login`
|
||||
2. Unauthenticated → Authentik sign-in; once authenticated the request reaches
|
||||
dsh-web with `Host: tankodhs.sysloggh.net`.
|
||||
3. dsh-web accepts the launch token on `GET /`, writes the
|
||||
`dsh-auth-<authority-hash>` cookie (30 days, `HttpOnly`, `SameSite=Strict`)
|
||||
and returns `303` to `/`.
|
||||
4. Every later request through `/` presents that cookie; the token is not needed
|
||||
again until the cookie expires or a new browser is used.
|
||||
|
||||
**Verification** (amdpve vantage):
|
||||
```bash
|
||||
# 1. Login endpoint is Authentik-gated: unauthenticated -> 302 (not 200/303).
|
||||
ssh root@192.168.68.15 "pct exec 112 -- curl -s -o /dev/null -w '%{http_code}\n' \
|
||||
-H 'Host: tankodhs.sysloggh.net' http://127.0.0.1/dsh-web-login"
|
||||
# Expected: 302
|
||||
|
||||
# 2. Legacy :8081 endpoint is gone (connection refused -> 000).
|
||||
ssh root@192.168.68.15 "pct exec 112 -- curl -s --max-time 3 -o /dev/null \
|
||||
-w '%{http_code}\n' http://192.168.68.122:8081/"
|
||||
# Expected: 000
|
||||
|
||||
# 3. Backend cookie mint + reuse (exactly what /dsh-web-login proxies to).
|
||||
TOKEN=$(ssh root@192.168.68.15 "pct exec 112 -- cat /etc/dsh-web/launch-token")
|
||||
ssh root@192.168.68.15 "pct exec 112 -- curl -s -c /tmp/dsh.jar -o /dev/null \
|
||||
-H 'Host: tankodhs.sysloggh.net' 'http://127.0.0.1:3080/?token=$TOKEN'"
|
||||
ssh root@192.168.68.15 "pct exec 112 -- curl -s -b /tmp/dsh.jar -o /dev/null \
|
||||
-w '%{http_code}\n' -H 'Host: tankodhs.sysloggh.net' http://127.0.0.1:3080/"
|
||||
# Expected: 200 — the minted dsh-auth-... cookie (authority
|
||||
# tankodhs.sysloggh.net) is replayed on the next request and accepted.
|
||||
|
||||
# 4. Token refresh is non-disruptive and idempotent.
|
||||
ssh root@192.168.68.15 "pct exec 112 -- /opt/deepseek-harness/capture-dsh-token.sh"
|
||||
# Expected: "token unchanged; nginx not reloaded" when nothing changed
|
||||
```
|
||||
|
||||
**Restart durability (acceptance):** after `systemctl restart dsh-web`, (a) the
|
||||
cookie minted before the restart still returns `200` on `/`, and (b) the
|
||||
refreshed `/etc/dsh-web/nginx-login.conf` carries the new token and mints a
|
||||
fresh cookie. Both verified live 2026-09-11.
|
||||
|
||||
```bash
|
||||
# 5. Cookie survives a dsh-web restart, and the new token mints a new cookie.
|
||||
ssh root@192.168.68.15 "pct exec 112 -- systemctl restart dsh-web"
|
||||
# dsh-web is Type=simple: restart returns before :3080 is listening. Bounded-poll
|
||||
# until the socket answers (any status but 000) before asserting the cookie.
|
||||
for i in $(seq 1 60); do
|
||||
UP=$(ssh root@192.168.68.15 "pct exec 112 -- curl -s -o /dev/null -w '%{http_code}' \
|
||||
-H 'Host: tankodhs.sysloggh.net' http://127.0.0.1:3080/")
|
||||
[ "$UP" != "000" ] && break
|
||||
sleep 2
|
||||
done
|
||||
ssh root@192.168.68.15 "pct exec 112 -- curl -s -b /tmp/dsh.jar -o /dev/null \
|
||||
-w '%{http_code}\n' -H 'Host: tankodhs.sysloggh.net' http://127.0.0.1:3080/"
|
||||
# Expected: 200 — the pre-restart cookie is still accepted.
|
||||
# The restart's ExecStartPost (or the 2-minute timer) refreshes the include. A
|
||||
# manual run may no-op on the flock, so poll until the include carries a token
|
||||
# the running process accepts (bounded wait) before the mint+reuse check.
|
||||
for i in $(seq 1 60); do
|
||||
TOKEN=$(ssh root@192.168.68.15 "pct exec 112 -- sed -n 's/.*token=//p' /etc/dsh-web/nginx-login.conf | tr -d ';\n'")
|
||||
CODE=$(ssh root@192.168.68.15 "pct exec 112 -- curl -s -o /dev/null -w '%{http_code}' \
|
||||
-H 'Host: tankodhs.sysloggh.net' 'http://127.0.0.1:3080/?token=$TOKEN'")
|
||||
[ "$CODE" = "303" ] && break
|
||||
sleep 2
|
||||
done
|
||||
# Expected: 303 — the include now holds the token the running process accepts.
|
||||
ssh root@192.168.68.15 "pct exec 112 -- curl -s -c /tmp/dsh-new.jar -o /dev/null \
|
||||
-H 'Host: tankodhs.sysloggh.net' 'http://127.0.0.1:3080/?token=$TOKEN'"
|
||||
ssh root@192.168.68.15 "pct exec 112 -- curl -s -b /tmp/dsh-new.jar -o /dev/null \
|
||||
-w '%{http_code}\n' -H 'Host: tankodhs.sysloggh.net' http://127.0.0.1:3080/"
|
||||
# Expected: 200 — the refreshed token minted a fresh cookie.
|
||||
```
|
||||
|
||||
|
||||
### Step 4: Platform C — Agent Zero (kagentz, CT 105 via Docker host .14)
|
||||
|
||||
> **The kagentz Zulip adapter leg is retired (2026-09-12).** Its code
|
||||
> (`/a0/usr/kagentz-zulip/`) no longer exists in the agent-zero container, so
|
||||
> the former adapter-process and heartbeat/queue checks always failed and the
|
||||
> monitor issued a restart for something that could not start, posting a false
|
||||
> kagentz-adapter-down alert on every run. Do NOT re-add an adapter-process,
|
||||
> heartbeat/queue, or adapter-restart step. Agent Zero is probed for A2A
|
||||
> liveness only, and a probe must never restart a platform.
|
||||
|
||||
**C1: A2A Server Health**
|
||||
|
||||
```bash
|
||||
# A2A server is on :50080 (not :8001) and is auth-gated (401 expected for unauthenticated)
|
||||
ssh root@192.168.68.14 "curl -s --connect-timeout 5 -o /dev/null -w '%{http_code}' http://127.0.0.1:50080/a2a/"
|
||||
# A2A listens on :80 inside the agent-zero container (host-mapped to :50080) and
|
||||
# is auth-gated: an unauthenticated probe gets 401, which means the server is up.
|
||||
ssh root@192.168.68.14 "docker exec agent-zero curl -s --connect-timeout 5 -o /dev/null -w '%{http_code}' http://127.0.0.1:80/a2a/"
|
||||
```
|
||||
|
||||
Expected: `401` (auth-gated, A2A server is up and responding) or `200` (if no auth required). Connection refused (000) → A2A server down.
|
||||
Expected: `401` (auth-gated, A2A server is up and responding) or `200` (if no auth required). Connection refused (`000`) → A2A server down. Any other status → running but unexpected: log/report it, never restart.
|
||||
|
||||
**C2: Adapter Process**
|
||||
**C2: A2A Response Verification**
|
||||
|
||||
```bash
|
||||
ssh root@192.168.68.14 "docker exec agent-zero ps aux | grep adapter | grep -v grep"
|
||||
```
|
||||
|
||||
Adapter should be running. Missing → restart inside container.
|
||||
|
||||
**C3: Heartbeat & Queue**
|
||||
|
||||
```bash
|
||||
ssh root@192.168.68.14 "docker exec agent-zero grep Heartbeat /tmp/zulip-adapter.log | tail -3"
|
||||
```
|
||||
|
||||
Check: `processed=N` incrementing, `silence < 600s`, `reconnects` ≈ 0.
|
||||
|
||||
**C4: A2A Response Verification**
|
||||
|
||||
```bash
|
||||
# A2A server is on :50080 (not :8001) and is auth-gated (401 expected for unauthenticated)
|
||||
ssh root@192.168.68.14 "curl -s -X POST http://127.0.0.1:50080/a2a \
|
||||
# A2A listens on :80 inside the container and is auth-gated (401 expected unauthenticated).
|
||||
ssh root@192.168.68.14 "docker exec agent-zero curl -s -X POST http://127.0.0.1:80/a2a \
|
||||
-H 'Content-Type: application/json' \
|
||||
-H 'Authorization: Bearer $LITELLM_KEY' \
|
||||
-d '{\"jsonrpc\":\"2.0\",\"method\":\"tasks/send\",\"params\":{\"message\":{\"role\":\"user\",\"parts\":[{\"text\":\"ping\"}]}},\"id\":1}'"
|
||||
@@ -304,9 +469,8 @@ Expected: task ID with "working" status. Poll for completion with `tasks/get`. I
|
||||
|
||||
| Condition | Action |
|
||||
|-----------|--------|
|
||||
| A2A `.well-known/agent.json` fails | `docker exec agent-zero bash -c "pkill -9 -f a2a_agent; cd /a0 && /opt/venv-a0/bin/python3 -u /a0/usr/a2a_agent.py > /tmp/a2a.log 2>&1 &"` |
|
||||
| Adapter process missing | Restart adapter inside container with env vars |
|
||||
| Silence > 600s | Restart adapter (auto-reconnect handles BAD_EVENT_QUEUE_ID) |
|
||||
| A2A returns `000` (connection refused/timeout) | Alert only — never restart the platform; investigate the agent-zero container |
|
||||
| A2A returns a status other than `200`/`401` | Log/report as a warning — reported, never healed on |
|
||||
| LiteLLM 401 | Check API key in a2a_agent.py `LITELLM_KEY` |
|
||||
|
||||
### Step 5: Global Checks
|
||||
@@ -316,7 +480,6 @@ Expected: task ID with "working" status. Poll for completion with `tasks/get`. I
|
||||
Check each agent's log for excessive bot-to-bot chatter:
|
||||
- Abiba: `Skipped.*bot msgs` count
|
||||
- Tanko: Repeated DM exchanges between bots
|
||||
- kagentz: Adapter log for bot DMs being processed
|
||||
|
||||
If any bot processes >50 bot-originated messages in 15min → warning.
|
||||
|
||||
|
||||
@@ -10,8 +10,9 @@ description: >
|
||||
|
||||
> **⚠️ RETIRED** — The pi Zulip extension (`~/.pi/agent/extensions/zulip/`) and
|
||||
> PM2 process (`abiba-zulip`) have been decommissioned. All mention/reliability
|
||||
> monitoring now happens through Telegram. Agents (Mumuni on Hermes, Tanko on DSH) and
|
||||
> Agent Zero (kagentz) continue to use Zulip.
|
||||
> monitoring now happens through Telegram. Mumuni (Hermes) and Tanko (DSH)
|
||||
> continue to use Zulip; Agent Zero's Zulip adapter is retired — see
|
||||
> `zulip-health.prose.md` for current Platform C (Agent Zero) state.
|
||||
|
||||
## Maintains
|
||||
|
||||
|
||||
Reference in New Issue
Block a user