Compare commits

..
Author SHA1 Message Date
abiba a2edc2f56f ci(pr-pipeline): fix flaky frontmatter check (grep -q SIGPIPE under pipefail)
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
The validate job runs with bash -e -o pipefail. `echo "$FM" | grep -q '^name:'`
lets grep exit on first match, which can SIGPIPE the echo; pipefail then reports
the pipeline non-zero and the || branch raises a false "Missing name/description".
The flagged file set varied run to run (and included files untouched by the PR)
while a fresh clone of the same commit passes the identical check. Reproduced:
the old form failed 3 of 5 local runs under the same shell flags, the herestring
form passed 5 of 5. Use herestrings so no pipe can be broken.
2026-09-12 17:13:54 +00:00
root 78b501798f no-mistakes(document): Sweep residual gemma labels; align compression rule contradiction
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Failing after 3s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Skipped
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Skipped
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Skipped
2026-09-12 16:54:41 +00:00
root 1f1b47f59d no-mistakes(document): Sweep residual gemma aliases; align compression rule contradiction 2026-09-12 16:51:21 +00:00
root 3d55764799 no-mistakes(review): Split retired-alias audit into fail vs warn; fix key claim 2026-09-12 16:41:35 +00:00
root ce48070f21 no-mistakes(review): Derive model fields recursively; fix historical latency and key claims 2026-09-12 16:33:58 +00:00
root 9e581ab203 no-mistakes(review): Complete retired-alias field coverage; fix misleading example labels 2026-09-12 16:26:14 +00:00
root 9cac3589cf no-mistakes(review): Fix compression example alias; make retired aliases fail audit 2026-09-12 16:19:43 +00:00
abiba 221f9f79f3 fix(audit): reject retired aliases; sweep gpu-light/gemma-4-12b to gpu-vision
audit-hermes-config.py Rule 8 required auxiliary.vision.model and
auxiliary.web_extract.model to equal the retired 'gpu-light', so a config
adopting the live canonical 'gpu-vision' FAILED our own audit - the audit was
enforcing a dead alias (400 Invalid model name). Rule 8 now requires
gpu-vision; retired names gpu-light/crew-auto join the raw-name rejection set;
the guidance message names the live aliases.

Sweep of the remaining references: gpu-self-heal stops canonicalizing
gpu-light; hermes-config-template, hermes-agent-baseline, hermes-key-enforcement,
inference-optimization, litellm-client-timeouts and gpu-fleet now use the live
gpu-vision alias. Where a file restated model/rpm/weight/fallback state it now
points at CT 116 /opt/inference-harness/litellm_config.yaml instead of
duplicating it. koby's .129 config is report-only and recorded, not edited.

Adds tests/test_audit_hermes_config_alias.py: executes the audit CLI and asserts
gpu-vision passes while gpu-light and gemma-4-12b fail.
2026-09-12 16:11:46 +00:00
abiba-bot 1dc040251d Merge pull request 'docs(litellm-health): single source of truth, gpu-vision alias, litellm-health as live owner' (#79) from fix/litellm-health-drift-20260912 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 10s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 0s
2026-09-12 16:05:28 +00:00
root 21f7b6171c no-mistakes(document): Dedupe cadence copies; registry remains authoritative
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 8s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
2026-09-12 16:02:15 +00:00
root bb17c2120f no-mistakes(document): Make litellm-health the live owner; dedupe probes 2026-09-12 16:00:56 +00:00
root dc42ecc235 no-mistakes(document): Align gpu-vision and single-source-of-truth documentation 2026-09-12 15:56:16 +00:00
root baaac9d7c6 no-mistakes(review): delete remaining duplicated timeout and frozen model-list state 2026-09-12 15:43:21 +00:00
root 2f65c38213 no-mistakes(review): delete duplicated config tables, point to CT 116 authority 2026-09-12 15:39:15 +00:00
root d9eb18c024 no-mistakes(review): align sibling contracts on aliases, crew cap, and monitor key 2026-09-12 15:30:14 +00:00
root 3f07b9bbcc no-mistakes(review): fix master-key inference and probe-surface drift in sibling contracts 2026-09-12 15:22:02 +00:00
abiba 1a598d0fcb docs(litellm-health): fix public-vs-backend probe surfaces and stale gemma model list
- Execution step 2 now documents the public edge and the backend edge as two
  distinct surfaces: public serves /ui/ and /docs (404 on the /litellm/ prefix),
  backend http://192.168.68.116 serves /litellm/ui/ and /litellm/docs (with /ui/
  and /docs as 301 helpers). Each probe names its surface.
- GPU topology: ocu-llm RTX 5070 now serves gpu-vision (gemma-4-12b retired).
- Fallback/timeout table rewritten to the live router_settings.fallbacks chains.
- Step 7 model list: gemma-4-12b -> gpu-vision, with a key-scoped /v1/models note
  and the 2026-09-12 master-key registry snapshot.
2026-09-12 15:11:24 +00:00
abiba-bot e3752af162 Merge pull request 'fix(zulip-health): retire the dead kagentz Zulip adapter leg' (#78) from fix/zulip-kagentz-adapter-20260912 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 2s
2026-09-12 14:34:08 +00:00
root 22d2c3acac no-mistakes(document): Reconcile Zulip docs with retired kagentz adapter leg
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 14s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
2026-09-12 14:31:26 +00:00
root 5f582e2c9c no-mistakes(review): fix duplicate-000 probe capture at assignment boundary 2026-09-12 14:23:35 +00:00
root b1b3b4c010 no-mistakes(review): fix kagentz A2A status logging, tests, and contract retirement 2026-09-12 14:19:13 +00:00
root 5d9b9847bc fix: remove kagentz Zulip adapter leg (code no longer exists)
- Remove adapter process check and restart logic
- Keep A2A probe (port 80, HTTP code check)
- The adapter code at /a0/usr/kagentz-zulip/ no longer exists
- Captain's ruling: Zulip communication with agent zero is not priority
2026-09-12 14:09:01 +00:00
root d29da3cc69 Make missing key file fail loudly instead of using dead key fallback
- Drop dead key literal from except block
- Report 'no-key-file' as check status when key file missing/unreadable
2026-09-12 12:49:15 +00:00
root 2ba1016ca8 Fix LiteLLM API key source and auth header
- Load API key from durable file /root/.abiba-workspace/secrets/litellm-key.txt (works in cron)
- Fix http_get to use Bearer token instead of Basic Auth for API endpoint check
- All 6 LiteLLM checks now pass (was 5/6)
2026-09-12 12:45:50 +00:00
abiba-bot 9c8637bcaf Merge pull request 'fix: correct Agent Zero A2A probe port in zulip-monitor.sh' (#76) from fix/a2a-port-20260911 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 9s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 1s
2026-09-11 22:17:37 +00:00
root aec62f7e77 fix: correct Agent Zero A2A probe from stale port 8001 to correct port 80
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 2s
- Changed A2A probe from http://127.0.0.1:8001/.well-known/agent.json to
  http://127.0.0.1:80/a2a/ inside agent-zero container
- Port 8001 does not exist inside container (nothing listens there)
- Port 80 maps to external port 50080; returns 401 (auth-gated, alive by design)
- Updated A2A_URL in adapter.py restart command to use port 80 instead of 8001
- Verified: probe now returns 401 (auth-gated) instead of 000 (connection refused)
- Before: kagentz A2A reported DOWN on every run (false positive due to stale port)
- After: kagentz A2A reports ✅ A2A alive (auth-gated 401 = healthy)
2026-09-11 21:54:13 +00:00
kagentz-bot b326a8944a Merge pull request 'feat(litellm): 1.99.1 update coverage, trove agent, nginx /ui /docs fixes' (#75) from update/litellm-1991-trove-20260911 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 8s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Failing after 5s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Skipped
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Skipped
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Skipped
2026-09-11 19:34:21 +00:00
agent-zero 7730cc7c03 feat(litellm): 1.99.1 update coverage, trove agent, nginx /ui /docs path fixes
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 2s
2026-09-11 15:32:20 -04:00
kagentz-bot af27530edc Merge pull request 'fix(router): decommission legacy GPU router across contracts and fleet scripts' (#74) from fix/decommission-router-20260911 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 3s
2026-09-11 19:19:13 +00:00
abiba-bot 862356bcac Merge pull request 'fix(dsh-web-auth): Authentik-gated :80 login + non-disruptive token capture' (#73) from fm/dsh-web-auth-restart-20260911 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 1s
2026-09-11 18:21:02 +00:00
abiba-bot 767bd22d9c no-mistakes(document): Align zulip-health v3.2.0 registry metadata and B4 formatting
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s
2026-09-11 17:49:25 +00:00
abiba-bot e97145c88f no-mistakes(review): re-sample systemd invocation each pass during token wait 2026-09-11 17:34:53 +00:00
abiba-bot 55f1208eb8 no-mistakes(review): scope dsh token to invocation; log nginx diagnostics 2026-09-11 17:30:33 +00:00
abiba-bot b9adf353ee no-mistakes(review): guard missing include, non-fatal pending reload, chmod stash 2026-09-11 17:24:12 +00:00
abiba-bot aeb66ea22d no-mistakes(review): fix pending-reload path, token modes, restart check 2026-09-11 17:17:26 +00:00
abiba-bot 288f74cf84 no-mistakes(review): reprobe dsh tokens; persist pending nginx reload on failure 2026-09-11 17:12:11 +00:00
abiba-bot c66671dbee no-mistakes(review): simplify dsh token selection and reload state machine 2026-09-11 17:07:00 +00:00
abiba-bot b80d3142aa no-mistakes(review): harden dsh token reload retry, legacy bypass, cookie verification 2026-09-11 16:59:55 +00:00
abiba-bot 266fa1f835 fix(dsh-web-auth): Authentik-gated :80 login + non-disruptive token capture
Correct the dsh-web authentication fix after parent correction:

- Remove the unauthenticated :8081 endpoint (0.0.0.0 bind with no
  auth_request = full Authentik bypass for the LAN). The script now removes
  /etc/nginx/sites-enabled/dsh.token automatically if it reappears.
- Put the login path inside the Authentik-gated :80 server block as
  location = /dsh-web-login; proxy to dsh-web with Host =
  tankodhs.sysloggh.net so the 30-day cookie is bound to the public
  authority, never to 127.0.0.1:3080.
- Isolate the rotating token in a generated include
  /etc/dsh-web/nginx-login.conf; reload nginx only when it changes.
- Replace the disruptive capture (systemctl stop/start dsh-web) with a
  non-disruptive read of the running service's journal, scoped to the
  current systemd invocation so a restarted process's stale token is never
  reused while the new banner is still pending.
- Keep x-dsh-task-board-proxy-token and Host $ak_origin_host intact in '/'.
- Document the corrected design (B4) in zulip-health.prose.md, v3.2.0.

Live-verified 2026-09-11: no auth bypass (302), :8081 refused (000), a
cookie minted before two dsh-web restarts still returns 200, the refreshed
token mints a fresh cookie, and the systemd ExecStartPost/timer refreshes
the token automatically without touching dsh-web.
2026-09-11 16:52:45 +00:00
abiba-bot 85ea1f4f3d no-mistakes(review): harden dsh token capture: loopback bind, atomic nginx config 2026-09-11 15:22:41 +00:00
abiba-bot b2a259fa23 fix: add dsh-web restart-persistent authentication
- Add capture-dsh-token.sh script that captures the dsh-web launch token
- Add login endpoint (/dsh-web-login on :8081) that mints 30-day auth cookie
- Document the authentication flow in zulip-health.prose.md (Platform B4)
- Cookie is authority-bound to 127.0.0.1:3080 with 30-day expiry
- After first login, subsequent requests use the cookie — no token required
2026-09-11 14:56:57 +00:00
26 changed files with 979 additions and 320 deletions
+7 -3
View File
@@ -43,17 +43,21 @@ jobs:
echo "=== Prose Contract Frontmatter Validation ==="
FAILED=0
for f in $(find . -name "*.prose.md" -not -path "./.git/*" -not -path "./runs/*"); do
# NOTE: use herestrings, not `echo "$FM" | grep ...`. Under the runner's
# `-e -o pipefail`, `grep -q` exits on first match and can SIGPIPE the
# producer, making the pipeline report non-zero and raising a false
# "Missing name/description" whose file set varies run to run.
FM=$(sed -n '/^---$/,/^---$/p' "$f" | sed '1d;$d')
[ -z "$FM" ] && { echo " ❌ $f: No YAML frontmatter"; FAILED=$((FAILED+1)); continue; }
KIND=$(echo "$FM" | grep '^kind:' | awk '{print $2}')
KIND=$(grep '^kind:' <<< "$FM" | awk '{print $2}')
case "$KIND" in
function|responsibility|gateway|pattern|test|template|architecture|enforcement) echo " ✅ $f: kind=$KIND" ;;
*) echo " ❌ $f: Invalid kind='$KIND'"; FAILED=$((FAILED+1)) ;;
esac
echo "$FM" | grep -q '^name:' || { echo " ❌ $f: Missing name"; FAILED=$((FAILED+1)); }
echo "$FM" | grep -q '^description:' || { echo " ❌ $f: Missing description"; FAILED=$((FAILED+1)); }
grep -q '^name:' <<< "$FM" || { echo " ❌ $f: Missing name"; FAILED=$((FAILED+1)); }
grep -q '^description:' <<< "$FM" || { echo " ❌ $f: Missing description"; FAILED=$((FAILED+1)); }
done
[ $FAILED -gt 0 ] && { echo "❌ FRONTMATTER FAILED ($FAILED error(s))"; exit 1; }
echo "✅ Frontmatter validation passed"
+3 -3
View File
@@ -88,7 +88,7 @@ prose run memory-audit-maintenance memory_threshold=90 verify_configs=true
prose run hermes-config-template agent_name=syslog-devops default_model=claude-sonnet-4
# Configure an agent with a different auxiliary model
prose run hermes-config-template agent_name=syslog-code default_model=qwen3.6-27B-code auxiliary_model=gemma-4-12b
prose run hermes-config-template agent_name=syslog-code default_model=qwen3.6-27B-code auxiliary_model=gpu-vision
```
### Option B: Manual Execution
@@ -116,7 +116,7 @@ Run on trigger or schedule. Maintain persistent world-model state across runs.
| `zulip-health` | Zulip | Checks Zulip connectivity, message flow, and bot responsiveness. |
| `zulip-mention-reliability` | Zulip | Diagnoses and fixes @mention detection issues in Zulip. |
| `zulip-approval-fix` | Zulip | Fixes broken /approve and /deny slash commands for Hermes agents. |
| `litellm-self-heal` | LiteLLM | Consolidated health check + self-healing for the full nginx → LiteLLM → GPU chain. Verifies 8 containers, 3 GPUs, model inference, and agent keys. Applies 9 remediation rules. (litellm-health merged into this contract 2026-07-09.) |
| `litellm-self-heal` | LiteLLM | Applies remediation rules for LiteLLM stack failures detected by `litellm-health` (full nginx → LiteLLM → GPU chain). 9 remediation rules. |
| `gpu-fleet` | GPU | Manages the GPU inference fleet: model deployment, registration, health checks, LiteLLM sync. |
| `gpu-monitor` | GPU | Comprehensive GPU fleet monitor — polls sidecars, router, LiteLLM every 15s, renders SSE dashboard. |
| `proxmox-monitor` | Infra | Proxmox cluster + Docker monitoring via the existing Grafana/Prometheus stack on CT 116. |
@@ -149,7 +149,7 @@ Called on-demand as single-render tools.
| Contract | Description |
|---|---|
| `litellm-api-keys` | Manages LiteLLM API keys for agent identity. Create, rotate, verify, and list agent keys. References gpu-fleet for current key inventory. |
| `litellm-health` | ⚠️ **DEPRECATED** — consolidated into `litellm-self-heal` (2026-07-09). Retained for reference only. |
| `litellm-health` | LiteLLM health check: public vs backend surfaces, CT 116 containers, GPU fleet, one model per GPU host, and agent keys. Owner of the probes; `litellm-self-heal` owns remediation. |
| `infrastructure-monitoring` | Target-state for Prometheus + GPU exporters + Grafana. Core stack deployed, GPU exporters NOT live. |
| `stirling-pdf-agent-access` | Documents the Stirling-PDF API access pattern for agents — global API key, 12 operations, curl examples. Agents use the `stirling-pdf-api` shared skill for templates. |
| `hello-world` | Minimal test contract — verifies the OpenProse execution pipeline works. |
+85 -17
View File
@@ -36,6 +36,59 @@ def warn(rule, message):
WARNINGS.append(f"[{rule}] {message}")
# Derivation rule: a model name is any scalar under a mapping key named `model` or
# `model_name`, at any depth. The top-level `model:` SECTION is the exception where the model
# name lives under `default`/`model`/`model_name` inside that section, so it is descended
# specially. The only other exception is key `models` (litellm key-generation params carry a
# list of model names). EXTEND THE ALLOWLIST for a new exception; do NOT add another field by
# hand.
MODEL_KEYS = ("model", "model_name")
MODEL_SECTION_KEYS = ("default", "model", "model_name")
MODEL_LIST_KEYS = ("models",)
def _iter_model_values(node, path=""):
"""Yield (path, value) for every model-name-bearing scalar in a config."""
if isinstance(node, dict):
for key, value in node.items():
child = f"{path}.{key}" if path else key
if key in MODEL_KEYS:
if isinstance(value, dict):
for subkey in MODEL_SECTION_KEYS:
subvalue = value.get(subkey)
if isinstance(subvalue, str):
yield (f"{child}.{subkey}", subvalue)
for subkey, subvalue in value.items():
if isinstance(subvalue, (dict, list)):
yield from _iter_model_values(subvalue, f"{child}.{subkey}")
elif isinstance(value, list):
yield from _iter_model_values(value, child)
else:
yield (child, value)
elif key in MODEL_LIST_KEYS:
yield from _iter_model_list(value, child)
elif isinstance(value, (dict, list)):
yield from _iter_model_values(value, child)
elif isinstance(node, list):
for i, item in enumerate(node):
yield from _iter_model_values(item, f"{path}[{i}]")
def _iter_model_list(node, path):
"""Yield scalars under an allowlisted `models` key (list of names or list of dicts)."""
if isinstance(node, list):
for i, item in enumerate(node):
yield from _iter_model_list(item, f"{path}[{i}]")
elif isinstance(node, dict):
for key, value in node.items():
if key in MODEL_KEYS and isinstance(value, str):
yield (f"{path}.{key}", value)
elif isinstance(value, (dict, list)):
yield from _iter_model_list(value, f"{path}.{key}")
else:
yield (path, node)
def audit(path):
with open(path) as f:
cfg = yaml.safe_load(f)
@@ -89,15 +142,16 @@ def audit(path):
)
# --- Rule 8: GPU Workload Distribution ---
# gpu-light (and gemma-4-12b) were retired 2026-09-12; the RTX 5070 stable alias is gpu-vision.
check(
aux.get("vision", {}).get("model") == "gpu-light",
aux.get("vision", {}).get("model") == "gpu-vision",
"Rule 8",
f"auxiliary.vision.model must be gpu-light (got {aux.get('vision', {}).get('model')!r}) — RTX 5070 stable alias",
f"auxiliary.vision.model must be gpu-vision (got {aux.get('vision', {}).get('model')!r}) — RTX 5070 stable alias",
)
check(
aux.get("web_extract", {}).get("model") == "gpu-light",
aux.get("web_extract", {}).get("model") == "gpu-vision",
"Rule 8",
f"auxiliary.web_extract.model must be gpu-light (got {aux.get('web_extract', {}).get('model')!r}) — RTX 5070 stable alias",
f"auxiliary.web_extract.model must be gpu-vision (got {aux.get('web_extract', {}).get('model')!r}) — RTX 5070 stable alias",
)
# --- Rule 9: Compression Threshold ---
@@ -178,21 +232,35 @@ def audit(path):
f"custom_providers[0].base_url must end with /v1 (got {cp.get('base_url')!r})",
)
# --- No raw model names (Rule 7/8 spirit) ---
raw_names = {"gemma-4-12b", "qwen3.6-27B-code", "qwen3.6-35B-udq4", "ornith-1.0-35b"}
for section_path, section_dict in [
("model", model), ("compression", comp),
("auxiliary.vision", aux.get("vision", {})),
("auxiliary.web_extract", aux.get("web_extract", {})),
("auxiliary.compression", aux.get("compression", {})),
("delegation", deleg),
]:
m = section_dict.get("model", "")
if m in raw_names:
# --- Retired/raw model names (Rule 7/8 spirit) ---
# The audit's job is to catch configs that are BROKEN, not to enforce a style preference.
# NON-RESOLVING names (removed 2026-09-12, verified 400/403 via live LiteLLM) must hard-FAIL:
# gpu-light -> gpu-vision ; gemma-4-12b -> gpu-vision
# crew-auto -> syslog-auto (its 64K cap is retired; no cap in force) ; ornith-1.0-35b -> strix-moe
# RESOLVING names (verified 200) are discouraged but working, so they only WARN:
# qwen3.6-27B-code -> gpu-dense ; qwen3.6-35B-udq4 -> strix-moe
# Failing a working alias would reject valid configs - the exact defect this change fixes.
non_resolving = {
"gpu-light": "gpu-vision",
"gemma-4-12b": "gpu-vision",
"crew-auto": "syslog-auto (its 64K cap is retired; no cap in force)",
"ornith-1.0-35b": "strix-moe",
}
raw_but_live = {
"qwen3.6-27B-code": "gpu-dense",
"qwen3.6-35B-udq4": "strix-moe",
}
for field_path, value in _iter_model_values(cfg):
if value in non_resolving:
check(
False,
"Rule 7/8",
f"{field_path} = {value!r} is retired and no longer resolves (2026-09-12) — use {non_resolving[value]}",
)
elif value in raw_but_live:
warn(
"Rule 7/8",
f"{section_path}.model = {m!r} — raw model name, use stable alias instead "
f"(gpu-light, gpu-dense, strix-moe, syslog-auto)",
f"{field_path} = {value!r} is a raw-but-live model name — prefer the stable alias {raw_but_live[value]}",
)
# --- Report ---
+6 -6
View File
@@ -1,5 +1,5 @@
registry_version: 0.1.0
last_updated: '2026-07-13T00:00:00Z'
last_updated: '2026-09-11T00:00:00Z'
updated_by: mumuni
categories:
- compliance
@@ -628,7 +628,7 @@ contracts:
sensitivity: high
status: active
owner: abiba
version: 3.1.0
version: 3.3.0
trigger:
type: scheduled
cadence: '*/15 * * * *'
@@ -693,8 +693,8 @@ contracts:
version: 1.0.0
trigger:
type: scheduled
cadence: '*/10 * * * *'
description: "Every 10 minutes \u2014 LiteLLM proxy health"
cadence: '5 3,7,11,15,19,23 * * *'
description: "4-hourly staggered dispatch via fm-send (run contract litellm-health)"
cron_job_id: null
execution:
agent: abiba
@@ -750,7 +750,7 @@ contracts:
sensitivity: critical
status: active
owner: abiba
version: 1.0.0
version: 1.1.0
trigger:
type: event_driven
description: Triggered by relay message from litellm-health or infrastructure-monitoring
@@ -1365,7 +1365,7 @@ contracts:
sensitivity: high
status: active
owner: ops
version: 1.0.0
version: 1.1.0
trigger:
type: scheduled
cadence: 0 2 * * 0
+6 -2
View File
@@ -2,6 +2,10 @@
Generated: 2026-07-13 20:59:18 ET
> **Point-in-time snapshot.** Schedules and cadences are authoritative in
> `contract-registry.yaml`; any schedule quoted below may be stale. Do not use
> this file as the source of truth for a contract's trigger.
---
## hermes-key-enforcement
@@ -442,7 +446,7 @@ IMPORTANT: If the contract file does not exist in prose-contracts/main, report f
## litellm-health
**Category:** monitoring | **Domain:** litellm | **Owner:** abiba | **Schedule:** */10 * * * *
**Category:** monitoring | **Domain:** litellm | **Owner:** abiba | **Schedule:** see contract-registry.yaml (authoritative)
```
Contract Enforcement: litellm-health
@@ -450,7 +454,7 @@ Contract Enforcement: litellm-health
Category: monitoring
Domain: litellm
Owner: abiba
Schedule: Every 10 minutes — LiteLLM proxy health
Schedule: see contract-registry.yaml (authoritative)
This is a monitoring contract. Execute the monitoring checks defined in the contract. Report any deviations from expected state.
+1 -1
View File
@@ -13,7 +13,7 @@ the description:
1. **What system does this contract touch?** Name the hosts, CTs, containers,
and services explicitly. "The inference fleet" is vague. "GPU .8 (RTX 3090,
qwen), .110 (RTX 5070, gemma), .15 (Strix Halo, strix-moe), and LiteLLM on CT
qwen), .110 (RTX 5070, gpu-vision), .15 (Strix Halo, strix-moe), and LiteLLM on CT
116" is specific.
2. **Who runs this contract, and when?** State the agent, the trigger (cron,
+33 -61
View File
@@ -6,9 +6,10 @@ description: >
registration, health checks, LiteLLM sync, agent key management, GPU
saturation watchdog, Prometheus/Grafana monitoring, and self-healing.
UPDATED 2026-07-15: Stable role-based aliases introduced: strix-moe,
gpu-dense, gpu-light. These never change — only the underlying model does.
gpu-dense, gpu-vision (gpu-light was superseded by gpu-vision on 2026-09-12).
These never change — only the underlying model does.
Strix Halo: strix-moe → unsloth/Qwen3.6-35B-A3B-MTP (UD-Q4_K_M, 22GB).
RTX 5070: gemma-4-12b Q4_K_M → IQ4_NL + MTP draft (122 tok/s, 2x faster).
RTX 5070: gpu-vision — IQ4_NL + MTP draft (~122 tok/s, 2x faster).
UPDATED 2026-07-17: Context reduced fleet-wide from 256K to 128K for stability.
Strix Halo model swapped to qwen3.6-35B-udq4 (22GB, strix-moe alias).
Instability observed near 100K at 256K (now all GPUs at 128K). 128K is the stable ceiling.
@@ -24,7 +25,7 @@ triggers:
## Maintains
- gpu_roster: { models: map, hosts: map } — Single source of truth for all GPU models
- gpu_roster: { models: map, hosts: map } — GPU host/model roster; the authoritative alias/weight/fallback registry is CT 116 `litellm_config.yaml`
- router: { status: "healthy", roster_loaded: bool, models: array }
- litellm: { status: "healthy", keys: array, models: array }
- agent_keys: { agent: api_key } — All agent API keys registered in LiteLLM DB
@@ -79,54 +80,28 @@ triggers:
Agent configs, cron jobs, and workflows MUST use these aliases, never model-specific names.
When a model is swapped on a GPU, ONLY the infrastructure layer changes — agent configs are untouched.
| Alias | GPU | Current Model | Will Route To |
|-------|-----|---------------|---------------|
| `strix-moe` | Strix Halo (.15) | qwen3.6-35B-udq4 | Whatever runs on Strix Halo |
| `gpu-dense` | RTX 3090 (.8) | Qwen3.8-27B-Uncensored-Q4_K_M | Whatever runs on RTX 3090 |
| `gpu-light` | RTX 5070 (.110) | gemma-4-12b | Whatever runs on RTX 5070 |
| Alias | Serves | Where | Kind |
|-------|--------|-------|------|
| `gpu-dense` | heavy reasoning | RTX 3090 (192.168.68.8) | direct alias |
| `gpu-vision` | vision / web extract / light tasks | RTX 5070 (192.168.68.110) | direct alias AND `syslog-auto` pool member |
| `strix-moe` | compression (MoE) | Strix Halo (192.168.68.15) | direct alias |
| `syslog-auto` | balanced default | weighted pool across the three GPU hosts | pool router |
**Backward compatibility**: Old model-specific names (qwen3.6-27B-code, gemma-4-12b, qwen3.5-9b-it) still work
but are deprecated for agent configs. Only the stable aliases survive model swaps.
Single source of truth for models, aliases, rpm caps, weights and fallback chains:
CT 116 `/opt/inference-harness/litellm_config.yaml`. Do not duplicate those values in
contracts — read them there.
## Current Model Assignments (2026-07-15)
| Model | GPU | Host | VRAM | Ctx | KV Cache | Parallel | Batch/Ubatch | Status |
|-------|-----|------|------|-----|----------|----------|-------------|--------|
**Backward compatibility**: Old model-specific names (qwen3.6-27B-code, qwen3.5-9b-it) still work
but are deprecated for agent configs. Only the stable aliases survive model swaps. `gemma-4-12b`
is retired and returns 400 `Invalid model name`.
## Routing Configuration (LiteLLM — July 2026)
### syslog-auto Weighted Pool (Direct GPU — bypasses router)
| Model | GPU | Weight | RPM Cap | Timeout |
|-------|-----|--------|---------|---------|
| Qwen3.8-27B-Uncensored-Q4_K_M | RTX 3090 (.8:8080) | **0.55** | 500 | **300s** |
Model, alias, rpm/weight and fallback values are owned by CT 116
`/opt/inference-harness/litellm_config.yaml` (see § Stable Role-Based Aliases above).
Note: All syslog-auto entries route directly to GPUs with `api_key: not-needed`. The router (port 9000) was decommissioned 2026-09-11 and is NOT in the inference path.
### Direct Model Endpoints
| Model | RPM Cap | Notes |
|-------|---------|-------|
### Stable Aliases (for agent configs — never change)
| Alias | RPM Cap | Routes To | Purpose |
|-------|---------|-----------|---------|
| `strix-moe` | 40 | Strix Halo | Compression tasks (MoE models) |
| `gpu-dense` | 500 | RTX 3090 | Heavy reasoning |
| `gpu-light` | 500 | RTX 5070 | Vision, web extract, light tasks |
### Fallback Chains
- gemma → qwen
- qwen → gemma
- strix-moe → qwen → gemma
- syslog-auto → qwen → gemma → qwen3.6-35B-udq4
### Why Strix Halo RPM Is Capped
- Direct (strix-moe): 40 RPM (tight) — Strix Halo is shared with compression tasks
- Via syslog-auto: 60 RPM (moderate) — prevents flooding when multiple agents use syslog-auto simultaneously
- Combined max: ~100 RPM across both paths — Strix Halo can sustain this at 80°C
## Operations
### add-model
@@ -176,9 +151,7 @@ Show full fleet status: GPUs, models, VRAM, context windows, parallel slots, act
2. Check llama-server processes: `ps aux | grep llama-server` on all 3 hosts
3. Check LiteLLM: `curl http://192.168.68.116/health` (expect "I'm alive!")
4. Check LiteLLM models: `curl -H "Authorization: Bearer $MASTER_KEY" http://192.168.68.116/v1/models`
5. Check LiteLLM timeouts: `grep -n 'timeout:' /opt/inference-harness/litellm_config.yaml`
- Qwen3.5-9B: 120s, qwen3.6-27B-code: 300s, Carnice-Qwen3.6-MoE-35B-A3B/strix-moe: 300s (strix-moe alias retained, legacy name qwen3.6-35B-udq4 deprecated)
- global request_timeout: 300s, nginx proxy_read_timeout: 600s
5. Check LiteLLM timeouts: `grep -n 'timeout:' /opt/inference-harness/litellm_config.yaml` — read the live values from the authority config; do not assert them from this contract.
6. Check AMD metrics: `curl http://192.168.68.15:9400/metrics` (Radeon 8060S, util%, VRAM, temp, power)
7. Check port conflicts: verify only one llama-server on :8080 per host
8. Verify agent keys: 9 keys in LiteLLM DB (`GET /key/list`)
@@ -243,6 +216,7 @@ If no SSH access, send Zulip DM via abiba-bot with vault update instructions.
- **Alert migration**: All alerts now go to `#agent-hub` topics (`alerts-gpu`, `alerts-pm2`, `alerts-infra`) instead of DMs. Cross-agent visibility enabled.
- **tok/s benchmarks**: Measured every 5 min via LiteLLM proxy. Baselines tracked with 30%/50% degradation thresholds.
- **NetBird 502**: Tanko routes through NetBird for litellm.sysloggh.net. Use direct IP if NetBird down.
- **Alias-retirement sweep (2026-09-12)**: `gemma-4-12b`, `gpu-light` and `crew-auto` are retired and replaced by `gpu-vision` / no cap respectively. The agent-facing templates (`hermes-config-template.prose.md`, `hermes-agent-baseline.prose.md`), `litellm-api-keys.prose.md`, `gpu-self-heal.prose.md`, `hermes-key-enforcement.prose.md`, `inference-optimization.prose.md`, `litellm-client-timeouts.prose.md` and the executable `audit-hermes-config.py` were all updated to the live canonical alias in the same change. **koby's config on .129 still names `gpu-light` (and `gemma-4-E4B`); .129 is report-only, so that is recorded for its owner and NOT edited here.**
## GPU Inference Benchmarks (Current)
@@ -265,42 +239,40 @@ History stored at `/root/data/toks-history.json` with 7-day rolling window.
### Stable Aliases — CRITICAL
All agent configs MUST use stable role-based aliases, never model-specific names:
- `compression.model: strix-moe` (NOT `qwen3.6-35B-udq4`)
- `auxiliary.vision.model: gpu-light` (NOT `gemma-4-12b`)
- `delegation.model: gpu-dense` (NOT `qwen3.6-27B-code`)
- `auxiliary.web_extract.model: gpu-light`
- `compression.model: syslog-auto`
- `auxiliary.vision.model: gpu-vision`
- `delegation.model: gpu-dense`
- `auxiliary.web_extract.model: gpu-vision`
When the underlying model is swapped, only the LiteLLM config changes — agent configs are untouched.
### Context Windows
- RTX 3090: **128K** (reduced from 256K 2026-07-17) | RTX 5070: **128K** (reduced from 256K) | Strix Halo: **128K**
- **All agents**: 128K ceiling — stable margin. For >128K workloads, use external providers (deepseek)
- Compression threshold 0.60: fires at ~77K (~51K headroom before 128K ceiling)
- Compression threshold 0.65 (audit Rule 9): fires at ~85K (~43K headroom before 128K ceiling)
- **Pi agents (Abiba)**: `compaction.reserveTokens: 52739` (≈60% of 128K)
- Mumuni compression model alias: `strix-moe` with 300s timeout
- Mumuni compression model alias: `syslog-auto`
### Mumuni Agent Profile
Mumuni (kagentz CT105, 192.168.68.14 — migrated from CT100 2026-08-29) is the primary business assistant. This profile is the reference for all agent configs:
Mumuni (kagentz CT105, 192.168.68.14 — migrated from CT100 2026-08-29) is the primary business assistant. This profile is the reference for all agent configs. The compression values below are the current required values per template Rules 7/9 and `audit-hermes-config.py`; whether Mumuni's LIVE config currently complies is a separate operational question.
| Setting | Value | Notes |
|---------|-------|-------|
| `model.default` | `syslog-auto` | Weighted pool (55% qwen, 30% strix, 15% gemma) |
| `model.default` | `syslog-auto` | Balanced default (pool router) |
| `model.provider` | `custom:litellm` | LiteLLM on CT116 |
| `compression.model` | `strix-moe` | Stable alias — survives model swaps |
| `aux.compression.model` | `strix-moe` | Compression auxiliary model |
| `aux.vision.model` | `gpu-light` | Vision tasks (RTX 5070) |
| `aux.web_extract.model` | `gpu-light` | Web extraction |
| `compression.model` | `syslog-auto` | Rule 7: auto-routing, prevents Strix Halo overload |
| `aux.compression.model` | `syslog-auto` | Must match `compression.model` (Rule 7) |
| `aux.vision.model` | `gpu-vision` | Vision tasks (RTX 5070) |
| `aux.web_extract.model` | `gpu-vision` | Web extraction |
| `delegation.model` | `gpu-dense` | Sub-agent reasoning (RTX 3090) |
| `context.max_context_window` | 131072 (128K) | Reduced from 256K 2026-07-17 — stable 128K ceiling |
| `compression.threshold` | 0.60 | Triggers at ~77K (~60% of 128K) — optimized for 128K context |
| `compression.threshold` | 0.65 | Rule 9: triggers at ~85K for a 128K window |
| `compression.target_ratio` | 0.3 | Compresses to ~38K |
| `compression.protect_last_n` | 40 | Preserves last 40 messages |
| `memory.memory_char_limit` | 800 | Brief memory entries |
| `personalities` | `creative` | Creative assistant personality |
| Platforms | cli, homeassistant, signal, telegram, zulip | All Hermes platforms |
| Main model timeout | 300s | LiteLLM global timeout |
| Compression model timeout | 300s | strix-moe timeout increased from 120s |
### Agent Update Status (2026-07-15)
+2 -2
View File
@@ -27,8 +27,8 @@ agent: abiba
┌──────┐ ┌──────┐ ┌────────┐
│.8:8080│ │.110 │ │.116:80 │
│RTX3090│ │:8080 │ │nginx │
│gemma │ │RTX5070│ │router │
└──────┘ │qwen27B│ │LiteLLM │
│qwen │ │RTX5070│ │router │
└──────┘ │vision │ │LiteLLM │
└──────┘ │dashboard│
└────────┘
```
+10 -9
View File
@@ -12,7 +12,7 @@ description: >
Router (port 9000) DECOMMISSIONED 2026-09-11; references replaced with direct GPU routing.
Benchmark baselines refreshed to live values.
Prometheus exporters removed — not deployed; fall back to direct sidecar probes.
Stable role-based aliases (strix-moe, gpu-dense, gpu-light) from gpu-fleet.
Stable role-based aliases (strix-moe, gpu-dense, gpu-vision) from gpu-fleet.
agent: abiba
depends_on:
- gpu-monitor.prose.md (live data source on .24:9100)
@@ -52,8 +52,8 @@ depends_on:
Key notes:
- All models use direct GPU routing via LiteLLM (`api_key: not-needed`). Router (port 9000) was decommissioned 2026-09-11 and is NOT in the inference path.
- Stable aliases (gpu-dense, gpu-light, strix-moe) from gpu-fleet are the canonical names for agent configs. Model-specific names still work but are deprecated.
- RTX 5070 tok/s is ~145 for Qwen3.5-9B — gpu-light is the fastest endpoint. Route vision/web/light work there first. NOTE: Qwen3.5-9B is multimodal (image+text), NOT text-only like gemma-4-12b was.
- Stable aliases (gpu-dense, gpu-vision, strix-moe) from gpu-fleet are the canonical names for agent configs. Model-specific names still work but are deprecated. The retired names `gpu-light` and `gemma-4-12b` were superseded by `gpu-vision` on 2026-09-12 and no longer resolve (400 `Invalid model name`).
- The RTX 5070 is the fastest endpoint per token — `gpu-vision` is its canonical alias. Route vision/web/light work there first. The RTX 5070 model is multimodal (image+text).
- Strix Halo is 62.9 tok/s (89% of 70.5 baseline) — below optimal but stable. Check for competing workloads.
- RTX 3090 VRAM at 88% — within role-appropriate range (role = heavy reasoning, needs the headroom).
- RTX 5070 VRAM at 83% — role-appropriate for vision/web (smaller batch sizes).
@@ -64,7 +64,7 @@ Key notes:
- **Detect**: Any GPU temp >85°C sustained for 2+ consecutive polls
- **Fix**:
1. Reduce inference concurrency on that GPU (load-side cooling only — NO fan control)
2. Redirect new requests to cooler GPUs via LiteLLM fallback chains (gemma → qwen, qwen → gemma)
2. Redirect new requests to cooler GPUs via LiteLLM fallback chains (gpu-vision → gpu-dense, gpu-dense → gpu-vision)
3. If all GPUs hot, alert about cooling infrastructure
- **Verify**: Temp drops below 80°C within 5 minutes
- **Escalate after**: 3 verification failures → Zulip alert
@@ -157,13 +157,14 @@ Key notes:
### Rule 10: Workload Distribution Optimization (updated 2026-07-18)
- **Detect**: GPU roles misaligned with hardware capabilities
- **Target distribution**:
- RTX 3090 (gpu-dense, 24GB, 74.9 tok/s) → Heavy reasoning, code gen, long conversations (slowest per-token but largest context capacity). Weight: 0.55 (LiteLLM).
- RTX 5070 (gpu-light, 12GB, ~145 tok/s) → Vision (image+text), web search, lightweight tasks. Weight: 0.15 (LiteLLM).
- Strix Halo (strix-moe, 64GB, 62.9 tok/s) → Context compression, summarization, long docs (MoE model). Weight: 0.30 (LiteLLM).
- RTX 3090 (gpu-dense, 24GB, 74.9 tok/s) → Heavy reasoning, code gen, long conversations (slowest per-token but largest context capacity).
- RTX 5070 (gpu-vision, 12GB, ~145 tok/s) → Vision (image+text), web search, lightweight tasks.
- Strix Halo (strix-moe, 64GB, 62.9 tok/s) → Context compression, summarization, long docs (MoE model).
- **Note**: RTX 5070 is the fastest endpoint per token. Route high-volume, low-complexity work there first.
- **Weights are not restated here** — the live `syslog-auto` pool weights and rpm caps live in CT 116 `/opt/inference-harness/litellm_config.yaml`, the single source of truth.
- **Fix**:
- Alert if any GPU is handling workload outside its designated role
- Recommend agent alias updates to match workload to GPU role (use stable aliases: gpu-dense, gpu-light, strix-moe)
- Recommend agent alias updates to match workload to GPU role (use stable aliases: gpu-dense, gpu-vision, strix-moe)
- Track per-GPU request distribution via LiteLLM spend logs
- **Verify**: Each GPU's request pattern matches its designated role within 24h
- **Escalate**: If role mismatch persists >48h → agent alias audit needed
@@ -311,6 +312,6 @@ Pushed to `SyslogSolution/health-logs/gpu/{run_id}.json` — versioned, searchab
If monitor response > 1MB, log a warning and skip the cycle rather than crashing.
### L6: Stable Aliases Replace Model Names
- gpu-fleet introduced stable aliases (strix-moe, gpu-dense, gpu-light) on 2026-07-15.
- gpu-fleet introduced stable aliases (strix-moe, gpu-dense, gpu-light) on 2026-07-15; `gpu-light` was superseded by `gpu-vision` on 2026-09-12.
- Self-heal must use aliases for reporting and alerting, not model-specific names.
- **Rule**: All alert messages and KG nodes use the stable alias as the GPU identifier.
+7 -6
View File
@@ -80,7 +80,7 @@ custom_providers:
auxiliary:
vision:
provider: harness
model: gemma-4-12b # or syslog-auto
model: gpu-vision # RTX 5070 stable alias (Rule 8; do not use syslog-auto for aux)
base_url: http://192.168.68.116/litellm/v1
api_key_env: LITELLM_API_KEY
api_key: <value from: infisical secrets get LITELLM_API_KEY --project=agents --env=production> # ← MANDATORY workaround
@@ -95,7 +95,7 @@ auxiliary:
threshold: 0.65
target_ratio: 0.3
provider: harness
model: syslog-auto # or gemma-4-12b
model: syslog-auto # Rule 7: compression must be syslog-auto
base_url: http://192.168.68.116/litellm/v1
api_key_env: LITELLM_API_KEY
api_key: <value from: infisical secrets get LITELLM_API_KEY --project=agents --env=production> # ← MANDATORY workaround
@@ -185,7 +185,9 @@ Abiba (CT100) runs pi via PM2 with the Zulip extension.
Config files: `~/.pi/agent/models.json`, `~/.pi/agent/settings.json`.
**models.json** — Must only list models authorized for the agent's LiteLLM key.
Key is injected via `infisical run --` wrapper at PM2 startup:
`/v1/models` is key-scoped and the live registry is CT 116
`/opt/inference-harness/litellm_config.yaml`; treat the list below as a snapshot and re-read
the registry before applying. Key is injected via `infisical run --` wrapper at PM2 startup:
```json
{
"providers": {
@@ -197,9 +199,8 @@ Key is injected via `infisical run --` wrapper at PM2 startup:
{ "id": "syslog-auto" },
{ "id": "strix-moe" },
{ "id": "gpu-dense" },
{ "id": "gpu-light" },
{ "id": "qwen3.6-27B-code" },
{ "id": "gemma-4-12b" }
{ "id": "gpu-vision" },
{ "id": "qwen3.6-27B-code" }
]
}
}
+18 -17
View File
@@ -102,7 +102,7 @@ model:
# and falls back to 256K when /v1/models lacks a context
# field (llama-server does). Without this override, agents
# silently run syslog-auto at 256K (verified 2026-08-09).
# Set 65536 if using gemma-4-12b directly (tight VRAM).
# Set 65536 if pinning a single model directly (tight VRAM).
fallback_providers:
provider: deepseek
@@ -143,25 +143,25 @@ compression:
# ─── Auxiliary Tasks (CONSISTENCY RULE) ───
# All auxiliary services MUST use identical model, base_url, and api_key_env:
# model: gpu-light # stable alias (NOT raw "gemma-4-12b")
# model: gpu-vision # stable alias (NOT a raw model name)
# base_url: http://192.168.68.116/litellm/v1 # Rule 5 (2026-08-09): canonical authenticated; /v1 also OK
# api_key_env: LITELLM_API_KEY
# Do NOT use syslog-auto for auxiliary tasks — it routes to the primary GPU.
# gpu-light = RTX 5070 (12B), freeing the Strix Halo for agent reasoning.
# gpu-vision = RTX 5070 (12B), freeing the Strix Halo for agent reasoning.
# Heavy aux (delegation, x_search) use gpu-dense (RTX 3090) instead.
# NEVER use raw model names (gemma-4-12b, qwen3.6-27B-code, qwen3.6-35B-udq4)
# in agent configs — use the stable aliases so model swaps don't break agents.
# NEVER use raw model names (e.g. qwen3.6-27B-code, qwen3.6-35B-udq4; gemma-4-12b is retired
# and no longer resolves) in agent configs — use the stable aliases so model swaps don't break agents.
auxiliary:
vision:
provider: harness
model: gpu-light # stable alias for RTX 5070 (was raw gemma-4-12b)
model: gpu-vision # stable alias for RTX 5070
base_url: http://192.168.68.116/litellm/v1 # Rule 5 (2026-08-09): canonical authenticated; /v1 also OK
api_key_env: LITELLM_API_KEY
timeout: 60
download_timeout: 30
web_extract:
provider: harness
model: gpu-light # stable alias for RTX 5070
model: gpu-vision # stable alias for RTX 5070
base_url: http://192.168.68.116/litellm/v1 # Rule 5 (2026-08-09): canonical authenticated; /v1 also OK
api_key_env: LITELLM_API_KEY
timeout: 30
@@ -252,10 +252,11 @@ The following MUST be identical across ALL profiles:
- For agents needing longer outputs: raise to 8192, but never omit
### Rule 7: Auxiliary Model Consistency (UPDATED 2026-07-16)
- Vision and web_extract use `gemma-4-12b` (RTX 5070 — 12GB, vision-optimized)
- Compression uses `strix-moe` (stable alias for Strix Halo — 64GB, 128K ctx, compression-optimized)
- **`strix-moe` is the only valid compression model name** — LiteLLM does NOT serve `ornith-1.0-35b`
(it serves `strix-moe`, `qwen3.6-35B-udq4`, `gpu-dense`, `gpu-light`, `syslog-auto`, `gemma-4-12b`, `qwen3.6-27B-code`). Old configs with `ornith-1.0-35b` cause 403/model-not-found on compression calls.
- Vision and web_extract use `gpu-vision` (RTX 5070 — 12GB, vision-optimized)
- Compression uses `syslog-auto` — the Strix Halo weighted pool (64GB, 128K ctx, compression-optimized); do NOT pin `compression.model` to `strix-moe` (audit Rule 7 rejects it)
- **`ornith-1.0-35b` is NOT a valid compression model name** — LiteLLM does not serve it
(do not restate the served model list here — CT 116 `/opt/inference-harness/litellm_config.yaml`
is the single source of truth for models, aliases, weights and fallbacks). Old configs with `ornith-1.0-35b` cause 403/model-not-found on compression calls.
- **OPERATIONAL DECISION (2026-07-23): Use `syslog-auto` for compression across all agents.**
The `syslog-auto` alias routes to the Strix Halo, but uses the weighted pool instead of pinning
to `strix-moe` directly. This prevents sustained Strix Halo thermal load because the pool can
@@ -264,7 +265,7 @@ The following MUST be identical across ALL profiles:
- All auxiliary services MUST use identical routing:
- `base_url: http://192.168.68.116/litellm/v1` (Rule 5, 2026-08-09: canonical authenticated; `/v1` also OK)
- `api_key_env: LITELLM_API_KEY`
- **Do NOT use `syslog-auto`** for auxiliary tasks — it routes unpredictably
- **Do NOT use `syslog-auto` for `vision`/`web_extract`** — it routes unpredictably; compression is the deliberate exception (see the OPERATIONAL DECISION above)
- **Compression on Strix Halo**: The strix-moe alias routes to Strix Halo
(64GB UMA, 128K context) — the designated compression GPU. This frees the
RTX 5070 for vision and web search, and the RTX 3090 for heavy reasoning.
@@ -273,11 +274,11 @@ The following MUST be identical across ALL profiles:
### Rule 8: GPU Workload Distribution (UPDATED 2026-07-16)
- **RTX 3090 (24GB, 128K ctx, qwen3.6-27B-code)**: Heavy reasoning, code gen, long conversations
- **RTX 5070 (12GB, 128K ctx, gemma-4-12b)**: Vision, web search, quick tasks, web_extract (IQ4_NL+MTP, ~65% VRAM at 128K)
- **RTX 5070 (12GB, 128K ctx, gpu-vision)**: Vision, web search, quick tasks, web_extract (IQ4_NL+MTP, ~65% VRAM at 128K)
- **Strix Halo (64GB, 128K ctx, syslog-auto)**: Context compression, summarization, long docs
- Agent profiles MUST route auxiliary tasks to the correct GPU:
- `auxiliary.vision.model: gemma-4-12b` (RTX 5070)
- `auxiliary.web_extract.model: gemma-4-12b` (RTX 5070)
- `auxiliary.vision.model: gpu-vision` (RTX 5070)
- `auxiliary.web_extract.model: gpu-vision` (RTX 5070)
- `auxiliary.compression.model: syslog-auto` (Strix Halo)
- Default model (`model.default`) and custom_provider remain `syslog-auto` for auto-routing
- For 128K context window: `threshold: 0.65` (fires at ~85K tokens)
@@ -297,8 +298,8 @@ The following MUST be identical across ALL profiles:
- **Koby Exception**: Per captain ruling 2026-08-11, Koby is a DeepSeek-primary external agent; its primary model remains `deepseek-v4-flash` (via api.deepseek.com to preserve DeepSeek-specific reasoning, while other sections follow Rule 10.
- **Hermes agents**: `model.default: syslog-auto`, `custom_providers[0].model: syslog-auto`
- **pi agents**: `defaultModel: syslog-auto` in `settings.json`, first model in `models.json`
- `syslog-auto` is the LiteLLM routing model — it load-balances between strix-moe
and qwen3.6-27B-code, with gemma-4-12b as fallback. Using it protects against:
- `syslog-auto` is the LiteLLM routing model — it load-balances across the live pool
(see CT 116 `/opt/inference-harness/litellm_config.yaml` for the current members and weights). Using it protects against:
- Model name typos that cause 403 errors and silent worker failures
- Single GPU downtime (routing falls back automatically)
- Key/model authorization mismatches
+9 -5
View File
@@ -169,16 +169,20 @@ The agent picks up the new key via `infisical run --` at gateway startup.
**Keys are permanent and use bare agent name aliases.**
- **Duration**: `null` — keys never expire. This is enforced by `default_key_generate_params` in `litellm_config.yaml`.
- **Duration**: `null` — keys never expire. NOT enforced today: CT 116 `litellm_config.yaml` has no `default_key_generate_params` block, and a key generated with no explicit models comes back with an empty models list. OPEN policy question: should agent keys expire by default? (captain security-policy decision, raised separately.)
- **Alias convention**: bare agent name only (e.g., `tanko`, `mumuni`, `koby`, `koonimo`). No dates, no versions. The alias IS the identity.
- **Rotation triggers**: compromise, personnel departure, or quarterly security hygiene. NOT calendar-driven.
- **Max budget**: $100 per key (config default).
```yaml
# In litellm_config.yaml — ensures all future keys inherit these defaults:
# NOT currently set in the authority; recommended value. CT 116 litellm_config.yaml has no
# default_key_generate_params block today, and a key generated with no explicit models comes back
# with an EMPTY models list. `models` is a literal key-generation parameter, so this is a value to
# ADD — re-read the live registry at CT 116 /opt/inference-harness/litellm_config.yaml and
# re-verify before applying.
litellm_settings:
default_key_generate_params:
models: ["syslog-auto", "qwen3.6-27B-code", "gemma-4-12b"]
models: ["syslog-auto", "qwen3.6-27B-code", "gpu-vision"]
duration: null # ← permanent
max_budget: 100
metadata:
@@ -286,13 +290,13 @@ auxiliary:
api_key: sk-<agent-key-from-vault> # ← workaround (get via: infisical secrets get LITELLM_API_KEY --project=agents --env=production --plain)
api_key_env: LITELLM_API_KEY
base_url: http://192.168.68.116/litellm/v1
model: gemma-4-12b
model: gpu-vision
provider: harness
compression:
api_key: sk-<agent-key-from-vault> # ← workaround (same as above)
api_key_env: LITELLM_API_KEY
base_url: http://192.168.68.116/litellm/v1
model: gemma-4-12b
model: syslog-auto
provider: harness
```
+6 -4
View File
@@ -23,7 +23,7 @@ management, and prompt caching — without sacrificing agent capability.
.123, any others on .129/.122) including compression, model, context_window,
prompt_caching, memory settings
- `gpu-health`: health check response from all 3 GPU backends (strix-moe .15:8080,
qwen .8:8080, gemma .110:8080)
gpu-dense .8:8080, gpu-vision .110:8080)
### Maintains
@@ -56,8 +56,8 @@ duration.
**Context is the root cause.** Every ~46K prompt token costs ~87s of
prefill time at 532 tok/s. Fix context first, routing second.
- **Route by task**: qwen for code/standard queries; gemma for
compression/auxiliary; strix-moe for compression tasks.
- **Route by task**: gpu-dense for code/standard queries; gpu-vision for
vision/web-auxiliary; syslog-auto for compression.
- **Compress aggressively**: threshold at 40% (not 65%) — a 128K window should
compact at 51K, not 85K. Target 15% tail (not 30%).
- **Cache everything repeated**: system prompts, skill docs, AGENTS.md — these
@@ -96,8 +96,10 @@ call enable-prompt-caching
hosts: [192.168.68.15, 192.168.68.8, 192.168.68.110]
-- Phase 5: Verify end-to-end latency
-- `models` is a literal verification parameter (a snapshot only): the authoritative registry is
-- CT 116 /opt/inference-harness/litellm_config.yaml; re-read it before use.
call verify-latency
host: 192.168.68.116
models: [syslog-auto, qwen3.6-27B-code, gemma-4-12b]
models: [syslog-auto, qwen3.6-27B-code, gpu-vision]
```
+3 -3
View File
@@ -199,7 +199,7 @@ description: >
### Ecosystem B: CT 116 syslog-api (192.168.68.116)
11 containers in inference-harness stack (verified live 2026-09-11; LiteLLM upgraded 1.90.0-rc.1 -> 1.99.1):
12 containers on CT 116 — 11 in the inference-harness stack + trove-agent-docker (verified live 2026-09-11; LiteLLM upgraded 1.90.0-rc.1 -> 1.99.1; trove-agent-docker added 2026-09-11):
| Container | Image | Port | Role |
|-----------|-------|------|------|
@@ -210,6 +210,7 @@ description: >
| harness-dashboard | inference-harness-dashboard | :3000 | SyslogAI Harness UI |
| harness-grafana | grafana/grafana | :3000→:3001 (direct LAN, not behind nginx) | GPU + Proxmox + Docker dashboards |
| harness-prometheus | prom/prometheus | :9090 | Metrics scraper, 6 jobs |
| trove-agent-docker | ghcr.io/techdox/trove-agent-docker:latest | outbound agent (no port) | Trove host agent — service inventory + metrics, added 2026-09-11 |
**Nginx routing**:
- `/v1/*` → harness-litellm:4000 (API)
@@ -221,9 +222,8 @@ description: >
**Prometheus targets**:
- 192.168.68.8:9400 (RTX 3090 — qwen)
- 192.168.68.110:9400 (RTX 5070 — gemma)
- 192.168.68.110:9400 (RTX 5070 — gpu-vision)
- 192.168.68.15:9400 (Strix Halo — qwen3.6-35B-udq4)
- 192.168.68.24:9401 (Router metrics exporter)
- harness-litellm:4000 (LiteLLM health)
### Ecosystem C: Netbird (72.61.0.17 — Hostinger srv1079750.hstgr.cloud)
+3 -1
View File
@@ -12,7 +12,7 @@ triggers:
- on "infra update" command
- weekly (Sunday 03:00 America/New_York) via Agent Zero scheduler task "weekly-fleet-docker-update" (qSOOVzsU) — implemented 2026-09-08
- on security advisory relay from Mumuni
version: 1.3.0
version: 1.4.0
---
## Maintains
@@ -81,6 +81,7 @@ Before ANY update wave:
| VM 109 (.7) | Home stack (Pulse, Stirling PDF) — JDownloader moved to CT 118 LXC 2026-08-01 | `cd /opt/home_stack && docker compose pull && docker compose up -d` |
| VM 109 (.7) | Audiobookshelf | `cd /opt/audiobookshelf && docker compose pull && docker compose up -d` |
| CT 116 (.116) | Inference Harness (LiteLLM, Prometheus, Grafana) | `cd /opt/inference-harness && docker compose pull && docker compose up -d` |
| CT 116 (.116, via minipve) | Trove docker agent (trove-agent-docker) | `pct exec 116 -- bash -c 'cd /opt/trove-agent && docker compose pull && docker compose up -d'` |
| CT 117 (storepve) | Zulip | `pct exec 117 -- bash -c 'cd /opt/zulip && docker compose pull && docker compose up -d'` (from storepve; compose recreates on zulip_default network) |
| CT 117 (storepve) | Jitsi | `pct exec 117 -- bash -c 'cd /opt/jitsi && docker compose pull && docker compose up -d'` (from storepve) |
| hwpve (.11) | Authentik (server, worker, postgres) | `ssh root@192.168.68.11 'cd /root && docker compose pull && docker compose up -d'` |
@@ -101,6 +102,7 @@ Before ANY update wave:
- harness-litellm cold start: allow 3-5 min after recreate — reports unhealthy and :4000 refuses connections while loading config/DB, then recovers to 200 on its own (verified 2026-09-08)
- SearXNG test: `curl :8888`
- Digest-pin sweep: `grep -rn '@sha256:' /opt/*/docker-compose.y*` on every host — digest-pinned images are INVISIBLE to `docker compose pull` (the pin re-pulls the same digest forever, so new releases never appear). Flag every pin in the run report and propose un-pinning to a floating tag with user approval before editing. Found 2026-09-10: audiobookshelf was digest-pinned at 2.34.0 (container created 2026-07-18) and silently missed by every sweep; dockhand stack was also pinned (stack removed 2026-09-10, unused). After un-pinning audiobookshelf to :latest it updated to 2.36.0 and verified HTTP 200.
- Version-pin awareness: a fixed version tag (e.g. `image: ...litellm:1.99.1`) is a no-op for `docker compose pull` just like a digest pin, so the stack silently stops advancing. CT 116 `harness-litellm` is INTENTIONALLY pinned to `1.99.1` (registry `main-stable`/`latest` currently resolve to `1.100.1`, sha256:a3715fa7 — a bleeding-edge jump explicitly declined 2026-09-11). Every run must look up the newest STABLE release tag for any version-pinned image, bump the pin deliberately with user approval, recreate, and re-verify. Never silently revert a pin to a floating tag.
## Wave 4: Proxmox Kernel Reboot
+4 -2
View File
@@ -67,8 +67,10 @@ description: >
4. **If action == "create"**:
- Generate new key with key_alias: "{agent_name}" (e.g., "tanko" — bare name, no date)
- Set metadata: { "agent": "{agent_name}", "purpose": "agent-inference" }
- Duration is null (permanent) — inherited from litellm default_key_generate_params
- Set models: ["syslog-auto", "qwen3.6-27B-code", "gemma-4-12b", "strix-moe", "gpu-dense", "gpu-light", "qwen3.6-35B-udq4"]
- Duration is whatever the caller passes; NO default enforcement exists today (CT 116 `litellm_config.yaml` has no `default_key_generate_params` block, and a key with no explicit models returns an empty models list). Agent keys are permanent by policy, not by that block. OPEN policy question: should agent keys expire by default? (captain security-policy decision, raised separately.)
- Set models: read the live key-scoped set rather than hardcoding one — `/v1/models` is key-scoped,
and the authoritative registry is CT 116 `/opt/inference-harness/litellm_config.yaml`. Do not add
retired names (`gemma-4-12b`, `gpu-light`, `crew-auto` — all retired 2026-09-12).
- Note: `ornith-1.0-35b` is NOT a valid LiteLLM model name (use `strix-moe`, the stable alias). qwen3.6-35B-A3B removed from fleet (was never deployed).
- Return the new key
5. **If action == "rotate"**:
+5 -5
View File
@@ -28,7 +28,7 @@ description: >
| syslog-auto | 28.8s | 25.4s | 68 calls took 30-120s; tail to ~300s under load |
| qwen3.6-27B-code | 23.0s | — | same backend class as syslog-auto |
| strix-moe | 7.5s | — | Strix Halo, healthy |
| gemma-4-12b | 2.6s | — | RTX 5070, healthy |
| gemma-4-12b (retired 2026-09-12; RTX 5070 now `gpu-vision`) | 2.6s | — | RTX 5070, healthy |
Sep 6 incident timeline: failures 04:00-07:00 EDT (0% GPU util = wedged
backend), full recovery 07:00-08:00 with ZERO client failures once requests
@@ -53,8 +53,7 @@ proxy queuing.
### 2. Auxiliary tasks — keep template timeouts, one correction
- vision: 60s (keep), web_extract: 30s (keep) — gemma-4-12b averages 2.6s;
these are fine.
- vision: 60s (keep), web_extract: 30s (keep) — the 2.6s average was measured on `gemma-4-12b` (retired 2026-09-12); the live RTX 5070 alias is `gpu-vision`.
- compression: 300s (keep — this was already raised from 60 per gpu-fleet).
- **gpu-dense delegation/x_search: set timeout >= 120s.** The RTX 3090
(qwen3.6-27B-code backend, 23.0s avg) is the same speed class as
@@ -79,8 +78,9 @@ proxy queuing.
Send a real model name and use a real (probe-designated) key.
- Probe timeout: 30s. A probe that takes longer than 30s IS the alert —
report "backend slow (>30s)" rather than hanging.
- Probe cadence: at most hourly. The 6h litellm-health cron cadence is the
standard; sub-hourly synthetic traffic distorts latency baselines.
- Probe cadence: at most hourly. The litellm-health cron cadence (authoritative
trigger in contract-registry.yaml) is the standard; sub-hourly synthetic
traffic distorts latency baselines.
### 5. Batch/benchmark jobs — schedule away from 04:00-07:00 EDT and chunk
+53 -34
View File
@@ -1,14 +1,7 @@
---
kind: function
name: litellm-health
status: deprecated
deprecated_on: 2026-07-09
replaced_by: litellm-self-heal.prose.md
note: >
Consolidated into litellm-self-heal.prose.md to eliminate duplication
of architecture diagrams, GPU topology, timeout tables, and container
lists. Health check is now § Health Check within litellm-self-heal.
This file is retained for reference only — use litellm-self-heal instead.
status: active
description: >
Verifies the LiteLLM inference stack health. Current architecture (2026-07-09):
nginx:80 → LiteLLM:4000 → GPU(llama-server) via direct proxy.
@@ -52,8 +45,7 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2)
- Router REMOVED from request path — LiteLLM proxies directly to GPU
- All GPUs at parallel 2 (was parallel 1)
- NVIDIA context reduced 256K→128K to free VRAM — now the stable ceiling across all GPUs (2026-07-17)
- LiteLLM timeouts tuned: gemma 25→120s, qwen 40→90s (SUPERSEDED 2026-07-16: qwen 300s, gemma 120s, strix 300s — see litellm-self-heal)
- nginx proxy_read_timeout: 600s, LiteLLM request_timeout: 300s
- Timeouts and fallback chains are config state — read them from CT 116 `/opt/inference-harness/litellm_config.yaml`; they are not duplicated here.
## Parameters
@@ -75,26 +67,27 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2)
- SSH key access to backend_host for container checks
- Network access to public_url, auth_host, and gpu_dashboard_url
- LiteLLM master key for key management endpoints
- LiteLLM master key for admin endpoints only (`/key/list`, `/key/generate`, `/key/info`)
- The dedicated `monitor` agent key on CT 116 at `/etc/litellm-monitor.env` (root-only 0600) for
model inference checks — the master key must never be used for inference
## GPU Fleet Topology
| Host | IP | Hardware | Models Served | Engine | Context | Parallel |
|------|-----|----------|---------------|--------|---------|----------|
| llm-gpu | 192.168.68.8 | NVIDIA RTX 3090 (24 GB) | qwen3.6-27B-code | llama-server systemd | **128K** | 2 |
| ocu-llm | 192.168.68.110 | NVIDIA RTX 5070 (12 GB) | gemma-4-12b | llama-server systemd | **128K** | 2 |
| amdpve | 192.168.68.15 | AMD Strix Halo 64GB UMA | qwen3.6-35B-udq4 (LiteLLM alias: strix-moe) | llama-server systemd (Vulkan) | 128K | 2 |
| Host | IP | Hardware | Role |
|------|-----|----------|------|
| llm-gpu | 192.168.68.8 | NVIDIA RTX 3090 (24 GB) | heavy reasoning (`gpu-dense`) |
| ocu-llm | 192.168.68.110 | NVIDIA RTX 5070 (12 GB) | vision / web extract / light tasks (`gpu-vision`) |
| amdpve | 192.168.68.15 | AMD Strix Halo 64GB UMA | compression (`strix-moe`) |
## Model Fallback Chains (LiteLLM)
Single source of truth for models, aliases, rpm caps, weights and fallback chains:
CT 116 `/opt/inference-harness/litellm_config.yaml`. Do not duplicate those values in
contracts — read them there. Do not re-add retired names (`gemma-4-12b`, `gpu-light`,
`crew-auto`).
| Primary | Timeout | Fallback | Timeout |
|---------|---------|----------|---------|
| qwen3.6-27B-code | 300s | gemma-4-12b | 120s |
| gemma-4-12b | 120s | qwen3.6-27B-code | 300s |
| qwen3.6-35B-udq4 / strix-moe | 300s | qwen → gemma | — |
| syslog-auto (balanced) | 300s | qwen → gemma | — |
> Global: request_timeout=300s, nginx proxy_read_timeout=600s
> Re-scope note (2026-09-12): the earlier plan to restate the live fallback chains and
> per-model timeouts in this contract is intentionally superseded — that state is config,
> and this contract points at the CT 116 config instead. Step 7 likewise authenticates with
> the dedicated `monitor` key, not the master key, which is admin-only.
## Containers on CT 116
@@ -112,18 +105,32 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2)
1. **Read parameters** — Use provided values or defaults
2. **Check public endpoints**:
2. **Check the end-user surfaces** — the public edge and the backend edge serve the SAME
app under DIFFERENT paths. They are not interchangeable, so every probe below must name
the surface it targets. Never point a check at a path that only resolves on the other
surface.
**Public edge** — `{{public_url}}` (https://litellm.sysloggh.net) serves the app at the
ROOT; the `/litellm/` prefix does not exist there and 404s:
- GET {{public_url}}/ui/ → expect 200 ("LiteLLM Dashboard")
- GET {{public_url}}/docs → expect 200 ("LiteLLM API - Swagger UI")
- GET {{public_url}}/litellm/ui/ and {{public_url}}/litellm/docs → expect 404 (not served on this edge)
**Backend edge** — `http://{{backend_host}}` (port 80) serves the app UNDER `/litellm/`:
- GET http://{{backend_host}}/litellm/ui/ → expect 200 ("LiteLLM Dashboard")
- GET http://{{backend_host}}/litellm/docs → expect 200 ("LiteLLM API - Swagger UI")
- GET http://{{backend_host}}/ui/ and /docs → expect 301 → /litellm/... → 200 (one-hop redirect helpers,
added 2026-09-11)
3. **Check LiteLLM health (no-auth)**:
- GET http://{{backend_host}}/litellm/health/liveliness → expect 200
4. **Check backend container health**:
- SSH to {{backend_host}} → `docker ps` → verify 11 containers healthy
- SSH to {{backend_host}} → `docker ps` → verify 12 containers healthy (11 harness + trove-agent-docker, added 2026-09-11)
- Critical: harness-litellm, harness-nginx, harness-postgres
- Monitoring: harness-redis, harness-dashboard, harness-grafana, harness-prometheus,
harness-alertmanager, harness-zulip-bridge, harness-docker-stats, harness-pve-exporter
harness-alertmanager, harness-zulip-bridge, harness-docker-stats, harness-pve-exporter,
trove-agent-docker (ghcr.io/techdox/trove-agent-docker, added 2026-09-11)
5. **Check GPU fleet health via gpu-monitor** (router decommissioned 2026-09-11):
- GET {{gpu_dashboard_url}}/gpu-data → expect 200 with GPU metrics JSON
@@ -134,14 +141,26 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2)
- Verify GPUs reporting status "healthy"
- Check alerts array for active warnings/critical
7. **Check model inference via LiteLLM** — Test each model:
- POST /v1/chat/completions model=gemma-4-12b → expect 200
- POST /v1/chat/completions model=qwen3.6-27B-code → expect 200
- POST /v1/chat/completions model=strix-moe → expect 200
- Use master key for auth
7. **Check model inference via LiteLLM** — Test one model on each GPU host. The health
check runs on the **backend edge**, not the public edge, so these paths carry the
`/litellm/` prefix:
- POST http://{{backend_host}}/litellm/v1/chat/completions model=qwen3.6-27B-code → expect 200 (RTX 3090, .8)
- POST http://{{backend_host}}/litellm/v1/chat/completions model=gpu-vision → expect 200 (RTX 5070, .110)
- POST http://{{backend_host}}/litellm/v1/chat/completions model=strix-moe → expect 200 (Strix Halo, .15)
- Auth uses the dedicated `monitor` agent key, read on CT 116 from
`/etc/litellm-monitor.env` (root-only 0600). Do NOT use the master key for inference —
the master key is for admin endpoints only (`/key/list`, `/key/generate`, `/key/info`).
- `gemma-4-12b` was retired and returns 400 `Invalid model name` — do not re-add it to
this list. The RTX 5070 host now serves `gpu-vision`.
- `/v1/models` is **key-scoped**: a model is only visible to keys allowed to use it, so
the set depends on the key. Always state which key a model list was read with — a
snapshot without its key is not evidence. This probe uses the `monitor` key on the
backend surface (`http://{{backend_host}}/litellm/v1/models`). The authoritative model
registry is CT 116 `/opt/inference-harness/litellm_config.yaml`; read it there rather
than freezing a list here.
8. **Check agent keys**:
- GET /key/list with master key → verify all 6 agents have keys
- GET http://{{backend_host}}/litellm/key/list with master key (admin endpoint) → verify all 6 agents have keys
9. **Check Grafana**:
- GET {{grafana_url}}/api/health → expect 200
+37 -71
View File
@@ -11,22 +11,22 @@ note: >
Script: `/opt/inference-harness/scripts/litellm-health-check.sh` on CT 116 (cron `0 */6 * * *`).
Reports to /var/log/litellm/health-*.json and Gitea (SyslogSolution/health-logs).
GPU monitoring integrated from gpu-monitor on .24:9100.
Consolidated from litellm-health + litellm-self-heal on 2026-07-09 to eliminate
duplication of architecture diagrams, GPU topology, timeout tables, and container
lists. Health check is now § Health Check within this contract.
Health probes are owned by litellm-health.prose.md (dispatched as
`run contract: litellm-health`). This contract owns remediation only — it does not
re-specify the probes.
Source of truth for GPU topology and keys: gpu-fleet.prose.md
Last verified: 2026-07-12
description: >
LiteLLM inference stack health monitoring + self-healing. Verifies the full
nginx → LiteLLM → GPU chain, 11 containers on CT 116, 3 GPU hosts, model
inference, and agent keys. Applies remediation rules for common failures.
LiteLLM inference stack remediation. Applies remediation rules for failures detected
by litellm-health.prose.md (nginx → LiteLLM → GPU chain, CT 116 containers, GPU hosts,
model inference, and agent keys).
Reports every action via Zulip DM and Gitea (SyslogSolution/health-logs).
---
---
# LiteLLM Operations — Health Check + Self-Heal
# LiteLLM Operations — Self-Heal (Remediation)
## Architecture (v4.0.0 — Direct: nginx → LiteLLM → GPU)
@@ -58,41 +58,44 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2)
## GPU Fleet Topology
| Host | IP | Hardware | Models Served | Engine | Context | Parallel |
|------|-----|----------|---------------|--------|---------|----------|
| llm-gpu | 192.168.68.8 | NVIDIA RTX 3090 (24 GB) | qwen3.6-27B-code | llama-server systemd (`/home/llmuser/llama-wrapper.sh`, `-c 131072 --parallel 2 --ngl 99`) | **128K** | 2 |
| ocu-llm | 192.168.68.110 | NVIDIA RTX 5070 (12 GB) | gemma-4-12b | llama-server systemd (`/home/llmuser/llama-wrapper.sh`, `--ctx-size 131072 --parallel 2`, IQ4_NL + MTP draft) | **128K** | 2 |
| amdpve | 192.168.68.15 | AMD Strix Halo 64GB UMA | qwen3.6-35B-udq4 (LiteLLM alias: `strix-moe`) | llama-server systemd (Vulkan) | 128K | 2 |
| Host | IP | Hardware | Role |
|------|-----|----------|------|
| llm-gpu | 192.168.68.8 | NVIDIA RTX 3090 (24 GB) | heavy reasoning (`gpu-dense`) |
| ocu-llm | 192.168.68.110 | NVIDIA RTX 5070 (12 GB) | vision / web extract / light tasks (`gpu-vision`) |
| amdpve | 192.168.68.15 | AMD Strix Halo 64GB UMA | compression (`strix-moe`) |
> Verified on ground 2026-07-16 via `curl /v1/models` on each host + `llama-wrapper.sh`. The AMD host's underlying model is `qwen3.6-35B-udq4`; LiteLLM exposes it under two `model_name`s: `qwen3.6-35B-udq4` and `strix-moe` (rpm 40). The legacy name `ornith-1.0-35b` does NOT exist in LiteLLM and must not be referenced.
> Verified on the ground 2026-07-16. The legacy name `ornith-1.0-35b` does NOT exist in LiteLLM and must not be referenced.
## LiteLLM Model Surface (ground truth — `/opt/inference-harness/litellm_config.yaml` on CT 116)
## LiteLLM Model Surface
`model_name`s served: `qwen3.6-27B-code`, `gemma-4-12b`, `qwen3.6-35B-udq4`, `strix-moe`, `gpu-dense`, `gpu-light`, `syslog-auto`, `crew-auto` (new 2026-08-20).
Single source of truth for models, aliases, rpm caps, weights and fallback chains:
CT 116 `/opt/inference-harness/litellm_config.yaml`. Do not duplicate those values in
contracts — read them there. Do not re-add retired names (`gemma-4-12b`, `gpu-light`,
`crew-auto`).
### Context Cap Split (2026-08-20)
### Alias Surface
| Alias | Serves | Where | Kind |
|-------|--------|-------|------|
| `gpu-dense` | heavy reasoning | RTX 3090 (192.168.68.8) | direct alias |
| `gpu-vision` | vision / web extract / light tasks | RTX 5070 (192.168.68.110) | direct alias AND `syslog-auto` pool member |
| `strix-moe` | compression (MoE) | Strix Halo (192.168.68.15) | direct alias |
| `syslog-auto` | balanced default | weighted pool across the three GPU hosts | pool router |
### Context Cap Split (2026-08-20; crew cap RETIRED)
- **Abiba (firstmate)**: 128K uncapped — unlimited context for primary workloads
- **Hermes agents** (mumuni, tanko, koby, koonimo): 128K uncapped
- **Crewmates** (ops, tune, verify, auth-keys, build): 64K capped — alias `crew-auto` enforces 64K limit
- **Crewmates** (ops, tune, verify, auth-keys, build): the 64K cap was retired together with the `crew-auto` alias; NO context cap is currently in force.
Preferred implementation: uncap shared pool, add capped alias for crew-only.
> The 64K crew cap was retired with `crew-auto` (2026-09-12). No limit is currently in force; reinstating one would need per-key model limits as a separate, deliberately-scoped change.
- `syslog-auto` is a weighted router model: qwen3.6-27B-code (0.55, rpm 500) + qwen3.6-35B-udq4 (0.30, rpm 60) + gemma-4-12b (0.15, rpm 200).
- `gpu-dense` / `gpu-light` are high-rpm aliases (rpm 500) onto qwen3.6-27B-code / gemma-4-12b respectively.
- Key scoping: agent keys are restricted to `['syslog-auto','qwen3.6-27B-code','gemma-4-12b','strix-moe','gpu-dense','gpu-light']`. As of 2026-07-16 the `baggy`/`koby`/`mumuni`/`abiba-pi` keys ALSO include `qwen3.6-35B-udq4`; `abiba-pi` additionally includes `deepseek-v4-pro` (cloud fallback). `kagenz0`/`koonimo`/`pi-agents-unified` have the standard 6 only. Agents should still use the stable alias `strix-moe` (not the raw `qwen3.6-35B-udq4`) so model swaps don't break them.
- Key scoping: `/v1/models` is key-scoped, so the set a caller sees must be read with a named key rather than assumed — a monitor key, an agent key, and the master key can each return a different set. Agents should use the stable aliases (`strix-moe`, `gpu-vision`, `gpu-dense`) rather than raw model names, so model swaps don't break them.
- **NEVER use `litellm_proxy_master_key` (the `sk-litellm-...` master key) for inference.** It is for admin endpoints only (`/key/list`, `/key/generate`, `/key/info`). All inference — agent traffic, health-check model tests, monitor scripts — uses agent-specific keys. The health-check script's model tests use a dedicated `monitor` agent key stored at `/etc/litellm-monitor.env` on CT 116 (root-only, `chmod 600`); `/key/list` is the only call that legitimately uses `$MASTER_KEY`.
## Model Fallback Chains (LiteLLM)
| Primary | Timeout | Fallback | Timeout |
|---------|---------|----------|---------|
| qwen3.6-27B-code | 300s | gemma-4-12b | 120s |
| gemma-4-12b | 120s | qwen3.6-27B-code | 300s |
| qwen3.6-35B-udq4 / strix-moe | 300s | qwen → gemma | — |
| syslog-auto (balanced) | 300s | qwen → gemma | — |
> Global: request_timeout=300s, nginx proxy_read_timeout=600s
> Re-scope note (2026-09-12): the fallback-chain and per-model timeout tables were removed
> by the single-source-of-truth re-scope; read those values from the CT 116 config named
> above rather than from this contract.
## Containers on CT 116
@@ -135,44 +138,7 @@ Preferred implementation: uncap shared pool, add capped alias for crew-only.
## Health Check
Run this first on every cycle. Results feed into remediation rules below.
### 1. Check public endpoints
- GET {{public_url}}/ui/ → expect 200 ("LiteLLM Dashboard")
- GET {{public_url}}/docs → expect 200 ("LiteLLM API - Swagger UI")
### 2. Check LiteLLM health (no-auth)
- GET http://{{backend_host}}/litellm/health/liveliness → expect 200
### 3. Check backend container health
- SSH to {{backend_host}} → `docker ps` → verify 11 containers healthy
- Critical: harness-litellm, harness-nginx, harness-postgres
- Monitoring: harness-redis, harness-dashboard, harness-grafana, harness-prometheus,
harness-alertmanager, harness-zulip-bridge, harness-docker-stats, harness-pve-exporter
- Decommissioned 2026-09-11: harness-router (container, image and config removed)
### 4. Check GPU fleet health (via fleet dashboard)
- GET {{gpu_dashboard_url}}/gpu-data → expect 200 with GPU metrics JSON
- Verify GPUs reporting status "healthy"
- Check alerts array for active warnings/critical
### 5. Check model inference via LiteLLM — test each model
- POST /v1/chat/completions model=gemma-4-12b → expect 200
- POST /v1/chat/completions model=qwen3.6-27B-code → expect 200
- POST /v1/chat/completions model=strix-moe → expect 200
- Use master key for auth
### 6. Check agent keys
- GET /key/list with master key → verify all 6 agents have keys
### 7. Check Grafana
- GET {{grafana_url}}/api/health → expect 200
### 8. Compile overall status
Determine overall_status from individual check results:
- "healthy" — all checks pass
- "degraded" — 1-2 non-critical checks fail
- "down" — critical checks fail
Health probes are owned by `litellm-health.prose.md` (dispatched as `run contract: litellm-health`). This contract owns remediation only — it does not re-specify the probes.
---
---
+227
View File
@@ -0,0 +1,227 @@
#!/usr/bin/env bash
# capture-dsh-token.sh — refresh the dsh-web login token WITHOUT restarting dsh-web.
#
# Context (CT 112 / tankodhs.sysloggh.net)
# ----------------------------------------
# The dsh-web UI (systemd unit `dsh-web.service`, 127.0.0.1:3080) prints a random
# launch token to the journal on every start:
#
# dsh web: http://127.0.0.1:3080/?token=<TOKEN>
#
# That token is the only way to bootstrap the authority-bound 30-day browser
# cookie. It rotates on every dsh-web start, so the Authentik-gated
# `location = /dsh-web-login` in /etc/nginx/sites-available/dsh must always
# reference the token of the RUNNING process.
#
# This script:
# 1. selects the launch token the RUNNING service actually accepts from the
# current systemd invocation — it NEVER stops or starts dsh-web,
# 2. records it in /etc/dsh-web/launch-token,
# 3. regenerates the nginx include /etc/dsh-web/nginx-login.conf (the
# `proxy_pass ...?token=` line consumed by /dsh-web-login),
# 4. reloads nginx ONLY when the on-disk include differs from the generated
# one or the applied-state stamp does not match the token (the stamp is
# written only after a successful reload), rolling the include back on
# failure so the next run retries,
# 5. removes the legacy unauthenticated :8081 endpoint if it ever reappears.
#
# Idempotent and safe to run at any time (systemd ExecStartPost or timer).
set -euo pipefail
umask 077
PATH="/usr/local/sbin:/usr/local/bin:/usr/sbin:/usr/bin:/sbin:/bin"
JOURNAL_UNIT="dsh-web.service"
TOKEN_FILE="/etc/dsh-web/launch-token"
INCLUDE_FILE="/etc/dsh-web/nginx-login.conf"
STAMP_FILE="/etc/dsh-web/nginx-login.conf.applied"
PENDING_FILE="/etc/dsh-web/nginx-reload.pending"
SITE_ENABLED="/etc/nginx/sites-enabled/dsh"
LEGACY_8081="/etc/nginx/sites-enabled/dsh.token"
STASH_DIR="/etc/nginx/sites-available"
LOCK_FILE="/run/capture-dsh-token.lock"
LOGIN_HOST="tankodhs.sysloggh.net"
LOGIN_UPSTREAM="http://127.0.0.1:3080"
TOKEN_WAIT=120
log() { printf 'capture-dsh-token: %s\n' "$*" >&2; }
die() { printf 'capture-dsh-token: ERROR: %s\n' "$*" >&2; exit 1; }
[ "$(id -u)" -eq 0 ] || die "must run as root"
# ── 0. Serialize runs so timer/ExecStartPost/manual runs cannot interleave ──
exec 9>"$LOCK_FILE"
flock -n 9 || { log "another capture-dsh-token run holds $LOCK_FILE; exiting"; exit 0; }
mkdir -p "$(dirname "$PENDING_FILE")"
# ── 0b. Guarantee the generated include exists before any `nginx -t` ──────
# The :80 site includes /etc/dsh-web/nginx-login.conf by literal path, so a
# missing include makes every `nginx -t` fail and can wedge recovery. Seed it
# from the last known token (or a placeholder); step 4 replaces it.
if [ ! -f "$INCLUDE_FILE" ]; then
SEED="placeholder"
if [ -f "$TOKEN_FILE" ]; then
SEED="$(cat "$TOKEN_FILE" 2>/dev/null || true)"
[ -n "$SEED" ] || SEED="placeholder"
fi
printf '%s' "$SEED" | grep -qE '^[A-Za-z0-9._~+/=:@-]+$' || SEED="placeholder"
printf 'proxy_pass %s/?token=%s;\n' "$LOGIN_UPSTREAM" "$SEED" > "$INCLUDE_FILE"
chmod 600 "$INCLUDE_FILE"
log "created missing $INCLUDE_FILE"
fi
# ── 1. Remove the legacy unauthenticated :8081 endpoint, if present ─────────
# It bypassed Authentik entirely (listened on 0.0.0.0:8081 with no auth_request)
# and must never come back. Stash it rather than delete so it is auditable.
if [ -e "$LEGACY_8081" ] || [ -L "$LEGACY_8081" ]; then
TS="$(date -u +%Y%m%dT%H%M%SZ)"
STASHED="$STASH_DIR/dsh.token.disabled-$TS"
mv "$LEGACY_8081" "$STASHED"
chmod 600 "$STASHED" 2>/dev/null || true
touch "$PENDING_FILE"
if ! NGINX_TEST_OUT="$(nginx -t 2>&1)"; then
die "nginx config test failed after disabling $LEGACY_8081 (kept disabled at $STASHED): $NGINX_TEST_OUT; a pending reload is recorded so running nginx is reloaded once the config is fixed. The legacy :8081 endpoint will NOT be restored."
fi
if ! nginx -s reload; then
die "nginx reload failed after disabling $LEGACY_8081 (kept disabled at $STASHED); a pending reload is recorded so running nginx is reloaded on the next run. The legacy :8081 endpoint will NOT be restored."
fi
rm -f "$PENDING_FILE"
log "removed legacy :8081 endpoint -> $STASHED"
fi
# ── 1b. Honor a recorded pending reload regardless of token selection ───────
# A failed reload leaves PENDING_FILE set so a stashed legacy :8081 file can
# never remain loaded in the running nginx while dsh-web is down or not yet
# answering. Reconcile it before the token wait.
if [ -e "$PENDING_FILE" ]; then
if ! NGINX_TEST_OUT="$(nginx -t 2>&1)"; then
log "WARNING: pending nginx reload recorded but 'nginx -t' fails: $NGINX_TEST_OUT; continuing so the include can be regenerated; will retry next run"
elif ! nginx -s reload; then
log "WARNING: pending nginx reload recorded but 'nginx -s reload' failed; will retry next run"
else
rm -f "$PENDING_FILE"
log "completed pending nginx reload"
fi
fi
# ── 2. Select the token the RUNNING service actually accepts ────────────────
# Re-sample the service's CURRENT systemd invocation on every pass and read
# candidates only from it, so a restart that lands during the wait immediately
# switches to the new invocation; there is no whole-journal or cross-invocation
# fallback, and an empty/unknown invocation just waits. Each candidate is then
# functionally verified against the local dsh-web using the public authority,
# exactly as the /dsh-web-login proxy does, and the first that answers 303 is
# the live token. Candidates are re-probed newest-first on each pass (connection
# failures stay eligible) until one is accepted or the wait elapses.
journal_tokens() {
journalctl -u "$JOURNAL_UNIT" "_SYSTEMD_INVOCATION_ID=$1" --no-pager -o cat 2>/dev/null \
| grep -oE 'dsh web: https?://[^[:space:]]+[?&]token=[^[:space:]]+' \
| sed -E 's/.*[?&]token=//' \
| grep -E '^[A-Za-z0-9._~+/=:@-]+$' \
| tac | awk '!seen[$0]++' || true
}
TOKEN=""
DEADLINE=$((SECONDS + TOKEN_WAIT))
NO_INVOCATION_WARNED=0
while [ -z "$TOKEN" ] && [ "$SECONDS" -lt "$DEADLINE" ]; do
INVOCATION="$(systemctl show -p InvocationID --value "$JOURNAL_UNIT" 2>/dev/null || true)"
if [ -z "$INVOCATION" ] || [ "$INVOCATION" = "n/a" ]; then
if [ "$NO_INVOCATION_WARNED" -eq 0 ]; then
log "WARNING: no invocation id for $JOURNAL_UNIT; waiting for a live invocation"
NO_INVOCATION_WARNED=1
fi
sleep 2
continue
fi
for cand in $(journal_tokens "$INVOCATION"); do
code="$(curl -s -o /dev/null --max-time 5 -w '%{http_code}' \
-H "Host: $LOGIN_HOST" "$LOGIN_UPSTREAM/?token=$cand" || true)"
if [ "$code" = "303" ]; then
TOKEN="$cand"
break
fi
done
[ -n "$TOKEN" ] && break
sleep 2
done
if [ -z "$TOKEN" ]; then
log "no accepted launch token in the current invocation within ${TOKEN_WAIT}s; leaving the include untouched for the next run"
[ -e "$PENDING_FILE" ] && die "pending nginx reload could not be completed; will retry next run"
exit 0
fi
# ── 3. Record the token (atomic, private) ──────────────────────────────────
mkdir -p "$(dirname "$TOKEN_FILE")"
if ! printf '%s\n' "$TOKEN" | cmp -s - "$TOKEN_FILE" 2>/dev/null; then
printf '%s\n' "$TOKEN" > "$TOKEN_FILE.tmp"
chmod 600 "$TOKEN_FILE.tmp"
mv "$TOKEN_FILE.tmp" "$TOKEN_FILE"
log "recorded live launch token in $TOKEN_FILE"
fi
chmod 600 "$TOKEN_FILE"
# ── 4. Regenerate the nginx login include (reload only when it changes) ────
NEW_INCLUDE="$(mktemp "$INCLUDE_FILE.XXXXXX")"
printf 'proxy_pass %s/?token=%s;\n' "$LOGIN_UPSTREAM" "$TOKEN" > "$NEW_INCLUDE"
chmod 600 "$NEW_INCLUDE"
# The stamp records the token nginx actually loaded. It is written only after a
# successful reload, so the early exit is safe only when both the stamp and the
# on-disk include agree with the live token; anything else falls through to the
# reload path so the include can never silently diverge from what nginx serves.
APPLIED=""
[ -f "$STAMP_FILE" ] && APPLIED="$(cat "$STAMP_FILE" 2>/dev/null || true)"
[ -f "$INCLUDE_FILE" ] && chmod 600 "$INCLUDE_FILE"
[ -f "$STAMP_FILE" ] && chmod 600 "$STAMP_FILE"
if [ "$APPLIED" = "$TOKEN" ] && [ -f "$INCLUDE_FILE" ] && cmp -s "$NEW_INCLUDE" "$INCLUDE_FILE" \
&& [ ! -e "$PENDING_FILE" ]; then
rm -f "$NEW_INCLUDE"
log "token unchanged; nginx not reloaded"
exit 0
fi
[ -e "$SITE_ENABLED" ] || { rm -f "$NEW_INCLUDE"; die "$SITE_ENABLED missing; refusing to reload"; }
RESTORE=""
if [ -f "$INCLUDE_FILE" ]; then
RESTORE="$(mktemp "$INCLUDE_FILE.bak.XXXXXX")"
cp -p "$INCLUDE_FILE" "$RESTORE"
chmod 600 "$RESTORE"
fi
mv "$NEW_INCLUDE" "$INCLUDE_FILE"
chmod 600 "$INCLUDE_FILE"
if ! NGINX_TEST_OUT="$(nginx -t 2>&1)"; then
if [ -n "$RESTORE" ]; then
mv "$RESTORE" "$INCLUDE_FILE"
else
rm -f "$INCLUDE_FILE"
fi
die "nginx config test failed: $NGINX_TEST_OUT; previous include restored"
fi
if ! nginx -s reload; then
if [ -n "$RESTORE" ]; then
mv "$RESTORE" "$INCLUDE_FILE"
else
rm -f "$INCLUDE_FILE"
fi
touch "$PENDING_FILE"
die "nginx reload failed; previous include restored; a pending reload is recorded so the next run retries"
fi
if [ -n "$RESTORE" ]; then
rm -f "$RESTORE"
fi
rm -f "$PENDING_FILE"
printf '%s\n' "$TOKEN" > "$STAMP_FILE.tmp"
chmod 600 "$STAMP_FILE.tmp"
mv "$STAMP_FILE.tmp" "$STAMP_FILE"
log "token changed; nginx reloaded"
log "login endpoint: https://$LOGIN_HOST/dsh-web-login (Authentik-gated)"
+15 -5
View File
@@ -29,7 +29,13 @@ LITELLM_PUBLIC = "https://litellm.sysloggh.net"
LITELLM_BACKEND = "192.168.68.116"
AUTH_HOST = "192.168.68.11"
SYNTHETIC_API_KEY = "sk-U_ydi3B-wfGU-_xESkoU1Q"
# Load LiteLLM API key from file (durable, works in cron)
LITELLM_KEY_FILE = "/root/.abiba-workspace/secrets/litellm-key.txt"
try:
with open(LITELLM_KEY_FILE) as f:
SYNTHETIC_API_KEY = f.read().strip()
except:
SYNTHETIC_API_KEY = None # Fail loudly: report "no-key-file" in check
NOW = datetime.datetime.now()
DATE_STR = NOW.strftime("%Y-%m-%d")
@@ -68,7 +74,7 @@ def http_get(url, auth=None, timeout=10):
cmd = f'curl -sfk --connect-timeout {timeout} -o /dev/null -w "%{{http_code}}" "{url}"'
if auth:
cmd = cmd.replace('"', '\\"')
cmd = f'curl -sfk --connect-timeout {timeout} -u "{auth}" -o /dev/null -w "%{{http_code}}" "{url}"'
cmd = f'curl -sfk --connect-timeout {timeout} -H "Authorization: Bearer {auth}" -o /dev/null -w "%{{http_code}}" "{url}"'
r = subprocess.run(cmd, shell=True, capture_output=True, text=True, timeout=timeout+2)
return r.stdout.strip() or "000"
except:
@@ -247,9 +253,13 @@ def collect():
report["litellm"]["checks"].append({"name": "oidc-auth", "status": "pass" if auth_code in ("200","302") else "fail", "code": auth_code})
# Check 5: Synthetic API call through LiteLLM
api_check = http_get(f"{LITELLM_PUBLIC}/v1/models", auth=SYNTHETIC_API_KEY)
report["litellm"]["api_models"] = api_check
report["litellm"]["checks"].append({"name": "api-endpoint", "status": "pass" if api_check == "200" else "fail", "code": api_check})
if SYNTHETIC_API_KEY is None:
report["litellm"]["api_models"] = None
report["litellm"]["checks"].append({"name": "api-endpoint", "status": "fail", "code": "no-key-file"})
else:
api_check = http_get(f"{LITELLM_PUBLIC}/v1/models", auth=SYNTHETIC_API_KEY)
report["litellm"]["api_models"] = api_check
report["litellm"]["checks"].append({"name": "api-endpoint", "status": "pass" if api_check == "200" else "fail", "code": api_check})
# ── NFS Mounts ──
nfs = ssh("192.168.68.7", "df -h /media/storage /media/mediastore 2>/dev/null | tail -n +2")
+21 -24
View File
@@ -41,7 +41,9 @@ notify() {
# ── Global: Zulip Server ──
SERVER_CODE=$(curl -s -o /dev/null -w "%{http_code}" --connect-timeout 10 \
https://chat.sysloggh.net/api/v1/server_settings \
-u 'abiba-bot@chat.sysloggh.net:cKTDMZAPW08dk3zl05sStzO7HRztzyn8' 2>/dev/null || echo "000")
-u 'abiba-bot@chat.sysloggh.net:cKTDMZAPW08dk3zl05sStzO7HRztzyn8' 2>/dev/null) || SERVER_CODE="000"
SERVER_CODE=$(printf '%s' "$SERVER_CODE" | tr -d '[:space:]')
[ -n "$SERVER_CODE" ] || SERVER_CODE="000"
if [ "$SERVER_CODE" != "200" ]; then
notify "🔴" "Zulip server returned HTTP $SERVER_CODE"
ISSUES=$((ISSUES + 1))
@@ -59,7 +61,9 @@ fi
# zulip.connected is a PROBE FAILURE: it alerts and NEVER calls pm2 restart.
# pm2 restart runs ONLY on affirmative zulip.connected=false.
# -- abiba-leg-start (verbatim-extracted by tests/zulip-monitor-abiba.sh)
PI_HTTP=$(curl -s -o /dev/null --connect-timeout 5 --max-time 10 -w '%{http_code}' http://localhost:9200/health 2>/dev/null || echo "000")
PI_HTTP=$(curl -s -o /dev/null --connect-timeout 5 --max-time 10 -w '%{http_code}' http://localhost:9200/health 2>/dev/null) || PI_HTTP="000"
PI_HTTP=$(printf '%s' "$PI_HTTP" | tr -d '[:space:]')
[ -n "$PI_HTTP" ] || PI_HTTP="000"
PI_BODY=$(curl -s --connect-timeout 5 --max-time 10 http://localhost:9200/health 2>/dev/null || true)
PI_STATE=$(printf '%s' "$PI_BODY" | python3 -c '
import sys, json
@@ -155,31 +159,24 @@ fi
# never contact her former host.
# ── Platform C: Agent Zero (kagentz) ──
AZ_A2A=$(ssh -o StrictHostKeyChecking=no -o ConnectTimeout=5 root@192.168.68.14 \
"docker exec agent-zero curl -s --connect-timeout 5 http://127.0.0.1:8001/.well-known/agent.json 2>/dev/null" 2>/dev/null || echo "")
AZ_ALIVE=$(echo "$AZ_A2A" | python3 -c "import sys,json; d=json.load(sys.stdin); print(d.get('name',''))" 2>/dev/null)
AZ_A2A_CODE=$(ssh -o StrictHostKeyChecking=no -o ConnectTimeout=5 root@192.168.68.14 \
"docker exec agent-zero curl -s --connect-timeout 5 -o /dev/null -w '%{http_code}' http://127.0.0.1:80/a2a/ 2>/dev/null" 2>/dev/null) || AZ_A2A_CODE="000"
AZ_A2A_CODE=$(printf '%s' "$AZ_A2A_CODE" | tr -d '[:space:]')
[ -n "$AZ_A2A_CODE" ] || AZ_A2A_CODE="000"
if [ "$AZ_ALIVE" != "kagentz" ]; then
notify "🔴" "kagentz A2A server DOWN — restarting"
ssh -o StrictHostKeyChecking=no -o ConnectTimeout=5 root@192.168.68.14 \
"docker exec agent-zero bash -c 'pkill -9 -f a2a_agent; sleep 1; cd /a0 && /opt/venv-a0/bin/python3 -u /a0/usr/a2a_agent.py > /tmp/a2a.log 2>&1 &'" 2>/dev/null || true
if [ "$AZ_A2A_CODE" = "000" ]; then
notify "🔴" "kagentz A2A server DOWN (connection failed)"
ISSUES=$((ISSUES + 1))
echo " kagentz: ❌ A2A down — restarted" >> "$LOG"
echo " kagentz: ❌ A2A down (HTTP 000)" >> "$LOG"
else
echo " kagentz: ✅ A2A alive" >> "$LOG"
# Check adapter process
AZ_ADAPTER=$(ssh -o StrictHostKeyChecking=no -o ConnectTimeout=5 root@192.168.68.14 \
"docker exec agent-zero ps aux 2>/dev/null | grep adapter | grep -v grep | wc -l" 2>/dev/null || echo "0")
if [ "$AZ_ADAPTER" -lt 1 ]; then
notify "🔴" "kagentz Zulip adapter DOWN — restarting"
ssh -o StrictHostKeyChecking=no -o ConnectTimeout=5 root@192.168.68.14 \
"docker exec agent-zero bash -c 'cd /a0/usr/kagentz-zulip && ZULIP_SITE=https://chat.sysloggh.net ZULIP_EMAIL=kagentz-bot@chat.sysloggh.net ZULIP_API_KEY=E9q9PXJTxftPYBkb5pBDWupDO7KK21ty ZULIP_AGENT_NAME=kagentz A2A_URL=http://localhost:8001/a2a A2A_TOKEN=8zNgdOEXzYxjQvTl /opt/venv-a0/bin/python3 -u adapter.py > /tmp/zulip-adapter.log 2>&1 &'" 2>/dev/null || true
ISSUES=$((ISSUES + 1))
echo " kagentz: ❌ Adapter down — restarted" >> "$LOG"
else
echo " kagentz: ✅ Adapter running" >> "$LOG"
fi
case "$AZ_A2A_CODE" in
200|401)
echo " kagentz: ✅ A2A alive (HTTP $AZ_A2A_CODE)" >> "$LOG" ;;
*)
notify "🟡" "kagentz A2A server answered HTTP $AZ_A2A_CODE — running, unexpected status"
ISSUES=$((ISSUES + 1))
echo " kagentz: 🟡 A2A unexpected http=$AZ_A2A_CODE (running, warning)" >> "$LOG" ;;
esac
fi
# ── Summary ──
+191
View File
@@ -0,0 +1,191 @@
"""Regression tests for the 2026-09-12 retired-alias sweep in audit-hermes-config.py.
WHY THIS FILE EXISTS: the executable audit pushed agent configs toward a DEAD alias.
Rule 8 required `auxiliary.vision.model == "gpu-light"` and
`auxiliary.web_extract.model == "gpu-light"`, but `gpu-light` (and its raw predecessor
`gemma-4-12b`) were retired on 2026-09-12 and now return 400 `Invalid model name`; the
live RTX 5070 alias is `gpu-vision`. A config that adopted the correct canonical alias
therefore FAILED our own audit, so the audit was actively enforcing a broken config.
These tests execute the real CLI (`python3 audit-hermes-config.py <config>`) and assert
observable behaviour — exit code and the emitted rule message — for the live alias and
for both retired names. No network, vault, or SSH access is required.
"""
from __future__ import annotations
import pathlib
import subprocess
import sys
ROOT = pathlib.Path(__file__).resolve().parent.parent
AUDIT = ROOT / "audit-hermes-config.py"
BASE = """
model:
api_key: ""
api_key_env: LITELLM_API_KEY
base_url: http://192.168.68.116/v1
max_tokens: 4096
default: syslog-auto
provider: harness
fallback_providers:
provider: deepseek
model: deepseek-v4-flash
api_key_env: DEEPSEEK_API_KEY
compression:
model: syslog-auto
provider: harness
threshold: 0.65
max_context_window: 131072
auxiliary:
vision:
model: {alias}
provider: harness
web_extract:
model: {alias}
provider: harness
compression:
model: syslog-auto
provider: harness
delegation:
provider: harness
custom_providers:
- name: harness
key_env: LITELLM_API_KEY
base_url: http://192.168.68.116/v1
"""
def _run_config(tmp_path, name, text):
cfg = tmp_path / name
cfg.write_text(text)
proc = subprocess.run(
[sys.executable, str(AUDIT), str(cfg)],
capture_output=True, text=True,
)
return proc.returncode, proc.stdout
def _run(tmp_path, alias):
return _run_config(tmp_path, f"{alias}.yaml", BASE.format(alias=alias))
def test_live_canonical_alias_passes(tmp_path):
"""The RTX 5070 alias that actually resolves must satisfy Rule 8."""
code, out = _run(tmp_path, "gpu-vision")
assert code == 0, out
assert "RESULT: PASS" in out
def test_retired_gpu_light_is_rejected(tmp_path):
"""A config pinned to the retired alias must fail, not pass."""
code, out = _run(tmp_path, "gpu-light")
assert code == 1, out
assert "auxiliary.vision.model must be gpu-vision" in out
assert "RESULT: FAIL" in out
def test_retired_gemma_is_rejected(tmp_path):
"""The retired raw model name must fail Rule 8 as well."""
code, out = _run(tmp_path, "gemma-4-12b")
assert code == 1, out
assert "auxiliary.vision.model must be gpu-vision" in out
assert "RESULT: FAIL" in out
def test_corrected_compression_example_passes(tmp_path):
"""The corrected workaround (vision=gpu-vision, compression=syslog-auto) must PASS."""
code, out = _run(tmp_path, "gpu-vision")
assert code == 0, out
assert "[Rule 7] compression.model must be syslog-auto (got 'syslog-auto')" in out
assert "[Rule 7] auxiliary.compression.model must be syslog-auto (got 'syslog-auto')" in out
assert "RESULT: PASS" in out
def test_retired_alias_in_delegation_is_rejected(tmp_path):
"""delegation.model has no dedicated value rule, so a retired name there used to PASS."""
code, out = _run_config(
tmp_path,
"delegation-gpu-light.yaml",
BASE.format(alias="gpu-vision").replace(
"delegation:\n provider: harness",
"delegation:\n provider: harness\n model: gpu-light",
),
)
assert code == 1, out
assert "delegation.model = 'gpu-light' is retired" in out
assert "RESULT: FAIL" in out
def test_retired_alias_in_custom_providers_is_rejected(tmp_path):
"""custom_providers[*].model is model-bearing; a retired name there must fail."""
code, out = _run_config(
tmp_path,
"custom-provider-gpu-light.yaml",
BASE.format(alias="gpu-vision").replace(
" - name: harness\n key_env: LITELLM_API_KEY",
" - name: harness\n model: gpu-light\n key_env: LITELLM_API_KEY",
),
)
assert code == 1, out
assert "custom_providers[0].model = 'gpu-light'" in out
assert "RESULT: FAIL" in out
def test_raw_but_live_alias_warns_but_passes(tmp_path):
"""Raw-but-live names resolve (200), so they warn only; failing them rejects valid configs."""
code, out = _run_config(
tmp_path,
"raw-qwen.yaml",
BASE.format(alias="gpu-vision").replace(
"delegation:\n provider: harness",
"delegation:\n provider: harness\n model: qwen3.6-27B-code",
),
)
assert code == 0, out
assert "delegation.model = 'qwen3.6-27B-code' is a raw-but-live model name" in out
assert "prefer the stable alias gpu-dense" in out
assert "RESULT: PASS" in out
def test_retired_alias_in_fallback_providers_is_rejected(tmp_path):
"""fallback_providers.model is model-bearing; a retired name there must fail."""
code, out = _run_config(
tmp_path,
"fallback-gpu-light.yaml",
BASE.format(alias="gpu-vision").replace(" model: deepseek-v4-flash", " model: gpu-light"),
)
assert code == 1, out
assert "fallback_providers.model = 'gpu-light'" in out
assert "RESULT: FAIL" in out
def test_retired_alias_in_x_search_is_rejected(tmp_path):
"""x_search.model was previously not enumerated; the derivation must catch it."""
code, out = _run_config(
tmp_path,
"x-search-gpu-light.yaml",
BASE.format(alias="gpu-vision").replace(
"delegation:\n provider: harness",
"delegation:\n provider: harness\nx_search:\n model: gpu-light",
),
)
assert code == 1, out
assert "x_search.model = 'gpu-light'" in out
assert "RESULT: FAIL" in out
def test_retired_alias_in_nested_auxiliary_block_is_rejected(tmp_path):
"""A nested auxiliary sub-block outside the named three must still be derived."""
code, out = _run_config(
tmp_path,
"nested-aux-gpu-light.yaml",
BASE.format(alias="gpu-vision").replace(
" compression:\n model: syslog-auto\n provider: harness\ndelegation:",
" compression:\n model: syslog-auto\n provider: harness\n"
" tasks:\n summarize:\n model: gpu-light\ndelegation:",
),
)
assert code == 1, out
assert "auxiliary.tasks.summarize.model = 'gpu-light'" in out
assert "RESULT: FAIL" in out
+33 -9
View File
@@ -89,8 +89,7 @@ case "$host" in
esac ;;
192.168.68.14)
case "$cmd" in
*agent.json*) printf '%s' "$AZ_A2A" ;;
*"ps aux"*) printf '%s\n' "$AZ_PS" ;;
*"/a2a/"*) printf '%s' "$AZ_A2A_CODE"; exit "$AZ_A2A_EXIT" ;;
esac ;;
*)
printf 'UNEXPECTED-SSH-HOST %s\n' "$host" >> "$RECORD_DIR/unexpected-ssh" ;;
@@ -122,8 +121,7 @@ def _write_exec(path: pathlib.Path, body: str) -> None:
def _run_monitor(tmp_path, *, tanko_svc="active", tanko_http="200",
az_a2a='{"name":"kagentz"}',
az_ps="root 111 0.1 0.2 /opt/venv-a0/bin/python3 -u adapter.py"):
az_a2a_code="401", az_a2a_exit=0):
"""Run the shipped monitor in a sandbox; return (proc, record_dir, log_path).
Only the LOG constant is rewritten (to keep the run inside the worktree).
@@ -151,8 +149,8 @@ def _run_monitor(tmp_path, *, tanko_svc="active", tanko_http="200",
"RECORD_DIR": str(record),
"TANKO_SVC": tanko_svc,
"TANKO_HTTP": tanko_http,
"AZ_A2A": az_a2a,
"AZ_PS": az_ps,
"AZ_A2A_CODE": az_a2a_code,
"AZ_A2A_EXIT": str(az_a2a_exit),
"PI_HTTP": "200",
"PI_BODY": CONNECTED_FIXTURE.read_text(),
"SERVER_HTTP": "200",
@@ -171,8 +169,7 @@ def test_healthy_run_is_quiet_and_never_reaches_mumuni(tmp_path):
assert "Server: ✅ HTTP 200" in log
assert "Abiba: ✅ Connected" in log
assert "Tanko: ✅ service=active http=200" in log
assert "kagentz: ✅ A2A alive" in log
assert "kagentz: ✅ Adapter running" in log
assert "kagentz: ✅ A2A alive (HTTP 401)" in log
assert "Result: ✅ All healthy" in log
# A healthy run emits no notify at all — and certainly no Mumuni one.
@@ -214,6 +211,32 @@ def test_failing_run_alerts_on_tanko_but_never_on_mumuni(tmp_path):
assert "Result: 🔴 1 issue(s) found" in log
def test_unexpected_a2a_status_is_an_issue_not_healthy(tmp_path):
proc, record, log_path = _run_monitor(tmp_path, az_a2a_code="500")
assert proc.returncode == 0, proc.stderr
log = log_path.read_text()
assert "kagentz: 🟡 A2A unexpected http=500 (running, warning)" in log
assert "kagentz: ✅ A2A alive" not in log
assert "Result: 🔴 1 issue(s) found" in log
assert "kagentz A2A server answered HTTP 500" in proc.stdout
def test_a2a_connection_failure_is_down_not_unexpected(tmp_path):
# curl prints the http_code before failing, so the ssh stub exits non-zero
# with "000" on stdout — exercising the real outage path.
proc, record, log_path = _run_monitor(tmp_path, az_a2a_code="000",
az_a2a_exit=7)
assert proc.returncode == 0, proc.stderr
log = log_path.read_text()
assert "kagentz: ❌ A2A down (HTTP 000)" in log
assert "kagentz: ✅ A2A alive" not in log
assert "unexpected" not in log
assert "Result: 🔴 1 issue(s) found" in log
assert "kagentz A2A server DOWN (connection failed)" in proc.stdout
# ── scripts/daily-infra-report.py: behavioral digest checks ──────────
@pytest.fixture(scope="module")
@@ -334,7 +357,8 @@ def test_agent_health_roster_has_no_mumuni_entry(ahc):
def test_health_contract_retires_mumuni_only_steps():
text = HEALTH_CONTRACT.read_text()
assert MUMUNI_IP not in text
for step in ("**B4:", "**B5:", "**B6:"):
for step in ("**B4: Gateway Process**", "**B5: Heartbeat Verification**",
"**B6: Response Delivery**"):
assert step not in text
+191 -28
View File
@@ -1,9 +1,9 @@
---
kind: responsibility
name: zulip-health
description: Multi-platform health monitor for the Zulip messaging mesh spanning Platform A (pi/Abiba Zulip bridge), Platform B (Tanko on DSH), and Platform C (Agent Zero Docker). Verifies bot registration, DM delivery, and cross-platform connectivity. Mumuni is no longer monitored from this host — she runs on her own container (kagentz CT 105 on minipve, .14) and is monitored on her side.
description: Multi-platform health monitor for the Zulip messaging mesh spanning Platform A (pi/Abiba Zulip bridge), Platform B (Tanko on DSH), and Platform C (Agent Zero Docker). Verifies bot registration, DM delivery, and cross-platform connectivity. The kagentz Zulip adapter leg is retired (its code no longer exists) — Agent Zero is probed for A2A liveness only. Mumuni is no longer monitored from this host — she runs on her own container (kagentz CT 105 on minipve, .14) and is monitored on her side.
title: Zulip Mesh Health Monitor — Multi-Platform
version: 3.1.0
version: 3.3.0
runtime_contract: 2
agent: abiba
report_only_agents:
@@ -261,38 +261,203 @@ logged/reported as a warning — reported, never healed on.
| HTTP `:3080` connection refused/timeout (`000`) | Same as above |
| HTTP status outside the expected set | Log/report as a warning — reported, never healed on |
**B4: dsh-web Authentication (Tanko — restart-persistent login)**
The dsh-web UI is token-gated. On every start the process prints a random
launch token to the journal:
```
dsh web: http://127.0.0.1:3080/?token=<TOKEN>
```
The token only bootstraps an authority-bound, HMAC-signed browser cookie with a
30-day lifetime. The signing secret is durable in
`/root/.dsh/.credentials.yaml` (key `client-connection/browser-session`), so a
cookie minted once keeps working across `dsh-web` restarts; the launch token
itself rotates on every restart.
**Login endpoint (public, Authentik-gated):**
`https://tankodhs.sysloggh.net/dsh-web-login`
It lives inside the Authentik-gated `:80` server block
(`/etc/nginx/sites-available/dsh`, symlinked from
`/etc/nginx/sites-enabled/dsh`) as `location = /dsh-web-login`, guarded by
`auth_request /outpost.goauthentik.io/auth/nginx`. It proxies to dsh-web with
`Host: tankodhs.sysloggh.net`, so the minted cookie is bound to the public
authority — never to `127.0.0.1:3080`. The token-dependent line is isolated in
the generated include `/etc/dsh-web/nginx-login.conf`:
```
proxy_pass http://127.0.0.1:3080/?token=<TOKEN>;
```
**Token refresh (non-disruptive):**
`/opt/deepseek-harness/capture-dsh-token.sh` (source:
`scripts/capture-dsh-token.sh`) reads candidate launch tokens from the journal
**scoped to the service's current systemd invocation**
(`systemctl show -p InvocationID` + `_SYSTEMD_INVOCATION_ID=`), re-sampling the
invocation on every pass so a restart that lands during the wait switches to the
new invocation; a restarted process's stale token is never considered while its
new startup banner is still pending and there is no whole-journal or
cross-invocation fallback. Each candidate
is then functionally verified against dsh-web with `Host: tankodhs.sysloggh.net`,
using the first the running process accepts with `303`. It waits up to 120s for
a restarted process to accept a token and re-probes every current-invocation
candidate on each pass, so a token that briefly returns `000` while the service
is still starting is not disqualified. If none is accepted it leaves the include
untouched and exits so the timer retries (exiting non-zero when a pending reload
is still outstanding). It writes
`/etc/dsh-web/launch-token` and regenerates `/etc/dsh-web/nginx-login.conf`,
reloading nginx only when the on-disk include differs from the generated one or
the applied-state stamp does not match the token (`nginx -t` guards the reload,
and the stamp is written only after a successful `nginx -s reload`, so a failed
or interrupted reload is retried on the next run). Any failed reload records a
pending-reload marker under `/etc/dsh-web/`; the next run attempts the reload
before the token wait, independent of token state, and clears the marker only
once the reload succeeds, so a disabled legacy `:8081` file can never leave the
running nginx unreloaded. The generated include is recreated before any
`nginx -t` if it is missing, so a failed run cannot wedge recovery.
Runs are serialized with `flock` on `/run/capture-dsh-token.lock`. It **never
stops or starts `dsh-web`**.
It is triggered by the `dsh-web.service` drop-in
`/etc/systemd/system/dsh-web.service.d/20-token-refresh.conf`
(`ExecStartPost=/bin/systemctl --no-block start dsh-web-token.service`) and by
`dsh-web-token.timer` every 2 minutes for reconciliation.
<details><summary>Installed systemd wiring (CT 112)</summary>
```ini
# /etc/systemd/system/dsh-web-token.service
[Unit]
Description=Refresh the dsh-web launch token for the nginx login endpoint
After=dsh-web.service
[Service]
Type=oneshot
TimeoutStartSec=180
ExecStart=/opt/deepseek-harness/capture-dsh-token.sh
# /etc/systemd/system/dsh-web-token.timer
[Unit]
Description=Periodically refresh the dsh-web login token
[Timer]
OnBootSec=90s
OnUnitActiveSec=120s
AccuracySec=10s
Persistent=true
[Install]
WantedBy=timers.target
# /etc/systemd/system/dsh-web.service.d/20-token-refresh.conf
[Service]
ExecStartPost=/bin/systemctl --no-block start dsh-web-token.service
```
</details>
> **Do NOT reintroduce the `:8081` endpoint.** It listened on `0.0.0.0:8081`
> with no `auth_request` and was a full Authentik bypass for anyone on the LAN.
> The script now removes `/etc/nginx/sites-enabled/dsh.token` automatically if
> it ever reappears.
**Authentication flow:**
1. `GET https://tankodhs.sysloggh.net/dsh-web-login`
2. Unauthenticated → Authentik sign-in; once authenticated the request reaches
dsh-web with `Host: tankodhs.sysloggh.net`.
3. dsh-web accepts the launch token on `GET /`, writes the
`dsh-auth-<authority-hash>` cookie (30 days, `HttpOnly`, `SameSite=Strict`)
and returns `303` to `/`.
4. Every later request through `/` presents that cookie; the token is not needed
again until the cookie expires or a new browser is used.
**Verification** (amdpve vantage):
```bash
# 1. Login endpoint is Authentik-gated: unauthenticated -> 302 (not 200/303).
ssh root@192.168.68.15 "pct exec 112 -- curl -s -o /dev/null -w '%{http_code}\n' \
-H 'Host: tankodhs.sysloggh.net' http://127.0.0.1/dsh-web-login"
# Expected: 302
# 2. Legacy :8081 endpoint is gone (connection refused -> 000).
ssh root@192.168.68.15 "pct exec 112 -- curl -s --max-time 3 -o /dev/null \
-w '%{http_code}\n' http://192.168.68.122:8081/"
# Expected: 000
# 3. Backend cookie mint + reuse (exactly what /dsh-web-login proxies to).
TOKEN=$(ssh root@192.168.68.15 "pct exec 112 -- cat /etc/dsh-web/launch-token")
ssh root@192.168.68.15 "pct exec 112 -- curl -s -c /tmp/dsh.jar -o /dev/null \
-H 'Host: tankodhs.sysloggh.net' 'http://127.0.0.1:3080/?token=$TOKEN'"
ssh root@192.168.68.15 "pct exec 112 -- curl -s -b /tmp/dsh.jar -o /dev/null \
-w '%{http_code}\n' -H 'Host: tankodhs.sysloggh.net' http://127.0.0.1:3080/"
# Expected: 200 — the minted dsh-auth-... cookie (authority
# tankodhs.sysloggh.net) is replayed on the next request and accepted.
# 4. Token refresh is non-disruptive and idempotent.
ssh root@192.168.68.15 "pct exec 112 -- /opt/deepseek-harness/capture-dsh-token.sh"
# Expected: "token unchanged; nginx not reloaded" when nothing changed
```
**Restart durability (acceptance):** after `systemctl restart dsh-web`, (a) the
cookie minted before the restart still returns `200` on `/`, and (b) the
refreshed `/etc/dsh-web/nginx-login.conf` carries the new token and mints a
fresh cookie. Both verified live 2026-09-11.
```bash
# 5. Cookie survives a dsh-web restart, and the new token mints a new cookie.
ssh root@192.168.68.15 "pct exec 112 -- systemctl restart dsh-web"
# dsh-web is Type=simple: restart returns before :3080 is listening. Bounded-poll
# until the socket answers (any status but 000) before asserting the cookie.
for i in $(seq 1 60); do
UP=$(ssh root@192.168.68.15 "pct exec 112 -- curl -s -o /dev/null -w '%{http_code}' \
-H 'Host: tankodhs.sysloggh.net' http://127.0.0.1:3080/")
[ "$UP" != "000" ] && break
sleep 2
done
ssh root@192.168.68.15 "pct exec 112 -- curl -s -b /tmp/dsh.jar -o /dev/null \
-w '%{http_code}\n' -H 'Host: tankodhs.sysloggh.net' http://127.0.0.1:3080/"
# Expected: 200 — the pre-restart cookie is still accepted.
# The restart's ExecStartPost (or the 2-minute timer) refreshes the include. A
# manual run may no-op on the flock, so poll until the include carries a token
# the running process accepts (bounded wait) before the mint+reuse check.
for i in $(seq 1 60); do
TOKEN=$(ssh root@192.168.68.15 "pct exec 112 -- sed -n 's/.*token=//p' /etc/dsh-web/nginx-login.conf | tr -d ';\n'")
CODE=$(ssh root@192.168.68.15 "pct exec 112 -- curl -s -o /dev/null -w '%{http_code}' \
-H 'Host: tankodhs.sysloggh.net' 'http://127.0.0.1:3080/?token=$TOKEN'")
[ "$CODE" = "303" ] && break
sleep 2
done
# Expected: 303 — the include now holds the token the running process accepts.
ssh root@192.168.68.15 "pct exec 112 -- curl -s -c /tmp/dsh-new.jar -o /dev/null \
-H 'Host: tankodhs.sysloggh.net' 'http://127.0.0.1:3080/?token=$TOKEN'"
ssh root@192.168.68.15 "pct exec 112 -- curl -s -b /tmp/dsh-new.jar -o /dev/null \
-w '%{http_code}\n' -H 'Host: tankodhs.sysloggh.net' http://127.0.0.1:3080/"
# Expected: 200 — the refreshed token minted a fresh cookie.
```
### Step 4: Platform C — Agent Zero (kagentz, CT 105 via Docker host .14)
> **The kagentz Zulip adapter leg is retired (2026-09-12).** Its code
> (`/a0/usr/kagentz-zulip/`) no longer exists in the agent-zero container, so
> the former adapter-process and heartbeat/queue checks always failed and the
> monitor issued a restart for something that could not start, posting a false
> kagentz-adapter-down alert on every run. Do NOT re-add an adapter-process,
> heartbeat/queue, or adapter-restart step. Agent Zero is probed for A2A
> liveness only, and a probe must never restart a platform.
**C1: A2A Server Health**
```bash
# A2A server is on :50080 (not :8001) and is auth-gated (401 expected for unauthenticated)
ssh root@192.168.68.14 "curl -s --connect-timeout 5 -o /dev/null -w '%{http_code}' http://127.0.0.1:50080/a2a/"
# A2A listens on :80 inside the agent-zero container (host-mapped to :50080) and
# is auth-gated: an unauthenticated probe gets 401, which means the server is up.
ssh root@192.168.68.14 "docker exec agent-zero curl -s --connect-timeout 5 -o /dev/null -w '%{http_code}' http://127.0.0.1:80/a2a/"
```
Expected: `401` (auth-gated, A2A server is up and responding) or `200` (if no auth required). Connection refused (000) → A2A server down.
Expected: `401` (auth-gated, A2A server is up and responding) or `200` (if no auth required). Connection refused (`000`) → A2A server down. Any other status → running but unexpected: log/report it, never restart.
**C2: Adapter Process**
**C2: A2A Response Verification**
```bash
ssh root@192.168.68.14 "docker exec agent-zero ps aux | grep adapter | grep -v grep"
```
Adapter should be running. Missing → restart inside container.
**C3: Heartbeat & Queue**
```bash
ssh root@192.168.68.14 "docker exec agent-zero grep Heartbeat /tmp/zulip-adapter.log | tail -3"
```
Check: `processed=N` incrementing, `silence < 600s`, `reconnects` ≈ 0.
**C4: A2A Response Verification**
```bash
# A2A server is on :50080 (not :8001) and is auth-gated (401 expected for unauthenticated)
ssh root@192.168.68.14 "curl -s -X POST http://127.0.0.1:50080/a2a \
# A2A listens on :80 inside the container and is auth-gated (401 expected unauthenticated).
ssh root@192.168.68.14 "docker exec agent-zero curl -s -X POST http://127.0.0.1:80/a2a \
-H 'Content-Type: application/json' \
-H 'Authorization: Bearer $LITELLM_KEY' \
-d '{\"jsonrpc\":\"2.0\",\"method\":\"tasks/send\",\"params\":{\"message\":{\"role\":\"user\",\"parts\":[{\"text\":\"ping\"}]}},\"id\":1}'"
@@ -304,9 +469,8 @@ Expected: task ID with "working" status. Poll for completion with `tasks/get`. I
| Condition | Action |
|-----------|--------|
| A2A `.well-known/agent.json` fails | `docker exec agent-zero bash -c "pkill -9 -f a2a_agent; cd /a0 && /opt/venv-a0/bin/python3 -u /a0/usr/a2a_agent.py > /tmp/a2a.log 2>&1 &"` |
| Adapter process missing | Restart adapter inside container with env vars |
| Silence > 600s | Restart adapter (auto-reconnect handles BAD_EVENT_QUEUE_ID) |
| A2A returns `000` (connection refused/timeout) | Alert only — never restart the platform; investigate the agent-zero container |
| A2A returns a status other than `200`/`401` | Log/report as a warning — reported, never healed on |
| LiteLLM 401 | Check API key in a2a_agent.py `LITELLM_KEY` |
### Step 5: Global Checks
@@ -316,7 +480,6 @@ Expected: task ID with "working" status. Poll for completion with `tasks/get`. I
Check each agent's log for excessive bot-to-bot chatter:
- Abiba: `Skipped.*bot msgs` count
- Tanko: Repeated DM exchanges between bots
- kagentz: Adapter log for bot DMs being processed
If any bot processes >50 bot-originated messages in 15min → warning.
+3 -2
View File
@@ -10,8 +10,9 @@ description: >
> **⚠️ RETIRED** — The pi Zulip extension (`~/.pi/agent/extensions/zulip/`) and
> PM2 process (`abiba-zulip`) have been decommissioned. All mention/reliability
> monitoring now happens through Telegram. Agents (Mumuni on Hermes, Tanko on DSH) and
> Agent Zero (kagentz) continue to use Zulip.
> monitoring now happens through Telegram. Mumuni (Hermes) and Tanko (DSH)
> continue to use Zulip; Agent Zero's Zulip adapter is retired — see
> `zulip-health.prose.md` for current Platform C (Agent Zero) state.
## Maintains