From bd0065bb310d466c19344c47204b2464bd3b7305 Mon Sep 17 00:00:00 2001 From: root Date: Sat, 12 Sep 2026 22:02:58 +0000 Subject: [PATCH] no-mistakes(review): Fix monitor-key scope, Strix context, retired delegation names --- gpu-fleet.prose.md | 16 ++++++++-------- hermes-agent-baseline.prose.md | 2 +- hermes-config-template.prose.md | 17 +++++++++-------- inference-optimization.prose.md | 4 ++-- litellm-health.prose.md | 13 +++++++------ mumuni-delegation-prose-contract.prose.md | 4 ++-- 6 files changed, 29 insertions(+), 27 deletions(-) diff --git a/gpu-fleet.prose.md b/gpu-fleet.prose.md index 83c97c1..852717a 100644 --- a/gpu-fleet.prose.md +++ b/gpu-fleet.prose.md @@ -8,12 +8,12 @@ description: > UPDATED 2026-07-15: Stable role-based aliases introduced: strix-moe, gpu-dense, gpu-vision (gpu-light was superseded by gpu-vision on 2026-09-12). These never change — only the underlying model does. - Strix Halo: strix-moe → unsloth/Qwen3.6-35B-A3B-MTP (UD-Q4_K_M, 22GB). + Strix Halo: strix-moe → Carnice-Qwen3.6-MoE-35B-A3B-Q4_K_M.gguf (22GB, 256K ctx). RTX 5070: gpu-vision — IQ4_NL + MTP draft (~122 tok/s, 2x faster). - UPDATED 2026-07-17: Context reduced fleet-wide from 256K to 128K for stability. + UPDATED 2026-07-17: NVIDIA host context reduced from 256K to 128K for stability. Strix Halo model: Carnice-Qwen3.6-MoE-35B-A3B-Q4_K_M.gguf (alias strix-moe, 256K context). - Instability observed near 100K at 256K (now all GPUs at 128K). 128K is the stable ceiling. - For larger context needs → fall back to external providers (deepseek). + Strix Halo runs 256K (n_ctx 262144, --kv-unified); RTX 3090 and RTX 5070 remain at 128K. + For >128K on NVIDIA hosts → fall back to external providers (deepseek). VRAM headroom improved: RTX 3090 ~70%, RTX 5070 ~65%. agent: abiba triggers: @@ -190,7 +190,7 @@ If no SSH access, send Zulip DM via abiba-bot with vault update instructions. | `/root/scripts/gpu-saturation-watchdog.py` | pi (.24) | Auto-restart stuck llama-server | | `/root/dashboard/gpu-fleet.html` | pi (.24) | Live HTML dashboard | | `/etc/systemd/system/llama-server.service` | .8, .110 | llama-server daemons (Nvidia GPUs) | -| `/etc/systemd/system/strix-server.service` | .15 (amdpve) | llama-server daemon (Vulkan, Strix Halo) running unsloth/Qwen3.6-35B-A3B-MTP-GGUF. Note: `llama-server.service` and `llama-server@.service` are **masked** on .15 to prevent port 8080 collisions. | +| `/etc/systemd/system/strix-server.service` | .15 (amdpve) | llama-server daemon (Vulkan, Strix Halo) running Carnice-Qwen3.6-MoE-35B-A3B-Q4_K_M.gguf (256K ctx). Note: `llama-server.service` and `llama-server@.service` are **masked** on .15 to prevent port 8080 collisions. | ## Prometheus & Grafana @@ -247,8 +247,8 @@ All agent configs MUST use stable role-based aliases, never model-specific names When the underlying model is swapped, only the LiteLLM config changes — agent configs are untouched. ### Context Windows -- RTX 3090: **128K** (reduced from 256K 2026-07-17) | RTX 5070: **128K** (reduced from 256K) | Strix Halo: **128K** -- **All agents**: 128K ceiling — stable margin. For >128K workloads, use external providers (deepseek) +- RTX 3090: **128K** (reduced from 256K 2026-07-17) | RTX 5070: **128K** (reduced from 256K) | Strix Halo: **256K** (2026-09-12) +- **Agents via `syslog-auto`**: 128K ceiling — the pool's safe floor (NVIDIA hosts are 128K). For >128K workloads, use external providers (deepseek) - Compression threshold 0.65 (audit Rule 9): fires at ~85K (~43K headroom before 128K ceiling) - **Pi agents (Abiba)**: `compaction.reserveTokens: 52739` (≈60% of 128K) - Mumuni compression model alias: `syslog-auto` @@ -266,7 +266,7 @@ Mumuni (kagentz CT105, 192.168.68.14 — migrated from CT100 2026-08-29) is the | `aux.vision.model` | `gpu-vision` | Vision tasks (RTX 5070) | | `aux.web_extract.model` | `gpu-vision` | Web extraction | | `delegation.model` | `gpu-dense` | Sub-agent reasoning (RTX 3090) | -| `context.max_context_window` | 131072 (128K) | Reduced from 256K 2026-07-17 — stable 128K ceiling | +| `context.max_context_window` | 131072 (128K) | Conservative `syslog-auto` pool floor (NVIDIA hosts 128K; Strix Halo 256K) | | `compression.threshold` | 0.65 | Rule 9: triggers at ~85K for a 128K window | | `compression.target_ratio` | 0.3 | Compresses to ~38K | | `compression.protect_last_n` | 40 | Preserves last 40 messages | diff --git a/hermes-agent-baseline.prose.md b/hermes-agent-baseline.prose.md index 2e2eb9a..fba3a89 100644 --- a/hermes-agent-baseline.prose.md +++ b/hermes-agent-baseline.prose.md @@ -5,7 +5,7 @@ version: 1.0.0 description: > Canonical known-good baseline for all Syslog Hermes agents. Captures the exact configuration state, keys, workarounds, and audit procedure. When an agent's - configuration goes sideways, restore from this baseline. Last verified 2026-07-16. All GPUs 128K context (reduced from 256K for stability Jul 2026) (RTX 3090 .8, RTX 5070 .110, Strix Halo .15). Parallel 1 fleet-wide (Strix Halo handles compression solo). + configuration goes sideways, restore from this baseline. Last verified 2026-07-16. RTX 3090/5070 at 128K (reduced from 256K for stability Jul 2026); Strix Halo at 256K (2026-09-12) (RTX 3090 .8, RTX 5070 .110, Strix Halo .15). Parallel 1 fleet-wide (Strix Halo handles compression solo). author: Abiba (pi agent) --- diff --git a/hermes-config-template.prose.md b/hermes-config-template.prose.md index 1a98545..3f36d62 100644 --- a/hermes-config-template.prose.md +++ b/hermes-config-template.prose.md @@ -97,7 +97,7 @@ model: base_url: http://192.168.68.116/litellm/v1 # Rule 5 (2026-08-09): canonical authenticated path; /v1 also OK api_key_env: LITELLM_API_KEY # Injected via infisical run -- wrapper max_tokens: 4096 # ⚠️ CRITICAL: Prevents unbounded generation - context_length: 131072 # For syslog-auto (all GPUs at 128K for stability). + context_length: 131072 # Conservative floor for syslog-auto (NVIDIA hosts 128K; Strix Halo 256K). # ⚠️ MANDATORY: Hermes probes unknown models from 256K # and falls back to 256K when /v1/models lacks a context # field (llama-server does). Without this override, agents @@ -133,7 +133,7 @@ compression: enabled: true model: syslog-auto # ⚠️ Must match auxiliary.compression.model. Stable alias (gpu-fleet § Stable Role-Based Aliases). NOT ornith-1.0-35b (LiteLLM does not serve that name). provider: harness - max_context_window: 131072 # MUST match actual GPU capacity. All 3 GPUs are 128K (Jul 17). + max_context_window: 131072 # MUST stay at the syslog-auto pool floor: NVIDIA hosts are 128K, Strix Halo 256K (2026-09-12). threshold: 0.65 # Fires at ~170K for 262K window, ~85K for 128K target_ratio: 0.30 protect_last_n: 40 @@ -253,7 +253,7 @@ The following MUST be identical across ALL profiles: ### Rule 7: Auxiliary Model Consistency (UPDATED 2026-07-16) - Vision and web_extract use `gpu-vision` (RTX 5070 — 12GB, vision-optimized) -- Compression uses `syslog-auto` — the Strix Halo weighted pool (64GB, 128K ctx, compression-optimized); do NOT pin `compression.model` to `strix-moe` (audit Rule 7 rejects it) +- Compression uses `syslog-auto` — the Strix Halo weighted pool (64GB, 256K ctx, compression-optimized); do NOT pin `compression.model` to `strix-moe` (audit Rule 7 rejects it) - **`ornith-1.0-35b` is NOT a valid compression model name** — LiteLLM does not serve it (do not restate the served model list here — CT 116 `/opt/inference-harness/litellm_config.yaml` is the single source of truth for models, aliases, weights and fallbacks). Old configs with `ornith-1.0-35b` cause 403/model-not-found on compression calls. @@ -267,15 +267,15 @@ The following MUST be identical across ALL profiles: - `api_key_env: LITELLM_API_KEY` - **Do NOT use `syslog-auto` for `vision`/`web_extract`** — it routes unpredictably; compression is the deliberate exception (see the OPERATIONAL DECISION above) - **Compression on Strix Halo**: The strix-moe alias routes to Strix Halo - (64GB UMA, 128K context) — the designated compression GPU. This frees the + (64GB UMA, 256K context) — the designated compression GPU. This frees the RTX 5070 for vision and web search, and the RTX 3090 for heavy reasoning. - The `compression:` block's `model` MUST match `auxiliary: compression: model` -- The `compression: max_context_window: 131072` MUST match actual GPU capacity (128K) +- The `compression: max_context_window: 131072` MUST stay at the syslog-auto pool floor (NVIDIA hosts 128K; Strix Halo 256K) ### Rule 8: GPU Workload Distribution (UPDATED 2026-07-16) - **RTX 3090 (24GB, 128K ctx, gpu-dense)**: Heavy reasoning, code gen, long conversations - **RTX 5070 (12GB, 128K ctx, gpu-vision)**: Vision, web search, quick tasks, web_extract (IQ4_NL+MTP, ~65% VRAM at 128K) -- **Strix Halo (64GB, 128K ctx, syslog-auto)**: Context compression, summarization, long docs +- **Strix Halo (64GB, 256K ctx, syslog-auto)**: Context compression, summarization, long docs - Agent profiles MUST route auxiliary tasks to the correct GPU: - `auxiliary.vision.model: gpu-vision` (RTX 5070) - `auxiliary.web_extract.model: gpu-vision` (RTX 5070) @@ -291,7 +291,7 @@ The following MUST be identical across ALL profiles: - For 128K context window: `threshold: 0.65` (fires at ~85K tokens) - Do NOT use `threshold: 0.25` — this fires at 65K, causing premature context loss - Do NOT use `threshold: 0.80` — this delays until ~105K, leaving only 23K margin -- `max_context_window: 131072` MUST match the model's actual capacity (all GPUs = 128K) +- `max_context_window: 131072` MUST stay at the pool floor (NVIDIA hosts 128K; Strix Halo 256K) - See `devops-hermes-compression` skill for full reference ### Rule 10: Default Model Must Be `syslog-auto` (All Agents) @@ -322,7 +322,8 @@ When an agent shows "context issues" (premature compression, 401s, 504s, DeepSee verify ALL FOUR of these against the live config. They are the only root causes found in production: 1. **max_context_window correct?** — BOTH `compression.max_context_window` AND `context.max_context_window` - MUST be `131072` (all GPUs are 128K). A value of `262144` causes instability near 100K and must NOT be used. + MUST be `131072` (the syslog-auto pool floor: NVIDIA hosts are 128K; Strix Halo is 256K). + A `262144` client window can route to a 128K NVIDIA host and fail, so it must NOT be used. ~83K instead of ~170K. Check: `grep -n max_context_window ~/.hermes/config.yaml` 2. **base_url uses authenticated path?** — `custom_providers[0].base_url`, `delegation.base_url`, and ALL `auxiliary.*.base_url` MUST be `http://192.168.68.116/litellm/v1` (Rule 5, canonical) diff --git a/inference-optimization.prose.md b/inference-optimization.prose.md index b0ee937..ce660c5 100644 --- a/inference-optimization.prose.md +++ b/inference-optimization.prose.md @@ -4,7 +4,7 @@ kind: responsibility description: > Optimizes the full Syslog inference stack — LiteLLM routing weights, GPU model assignments, agent context management, and prompt caching — to reduce response - times to sub-15s average. All GPUs now at 128K context (stable ceiling). + times to sub-15s average. NVIDIA GPUs at 128K context; Strix Halo at 256K (2026-09-12). id: 067NC6KP02RG60S50M40E30928 --- @@ -63,7 +63,7 @@ prefill time at 532 tok/s. Fix context first, routing second. - **Cache everything repeated**: system prompts, skill docs, AGENTS.md — these never change between turns. Single-digit cache hit rate is unacceptable. - **Lower context ceiling**: 128K window is the stable ceiling for agent conversations. - GPUs reduced from 256K to 128K (2026-07-17). 128K window should compact at 85K (0.65 threshold). For larger contexts, route to external providers. + GPUs reduced from 256K to 128K (2026-07-17) for the NVIDIA hosts; Strix Halo runs 256K (2026-09-12). 128K window should compact at 85K (0.65 threshold). For larger contexts, route to external providers. ### Shape diff --git a/litellm-health.prose.md b/litellm-health.prose.md index 386e718..5e47053 100644 --- a/litellm-health.prose.md +++ b/litellm-health.prose.md @@ -44,7 +44,7 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2) **What changed (v3.2.0 → v4.0.0 — 2026-07-08)**: - Router REMOVED from request path — LiteLLM proxies directly to GPU - All GPUs at parallel 2 (was parallel 1) -- NVIDIA context reduced 256K→128K to free VRAM — now the stable ceiling across all GPUs (2026-07-17) +- NVIDIA context reduced 256K→128K to free VRAM — the stable NVIDIA ceiling (2026-07-17); Strix Halo runs 256K (2026-09-12) - Timeouts and fallback chains are config state — read them from CT 116 `/opt/inference-harness/litellm_config.yaml`; they are not duplicated here. ## Parameters @@ -69,7 +69,8 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2) - Network access to public_url, auth_host, and gpu_dashboard_url - LiteLLM master key for admin endpoints only (`/key/list`, `/key/generate`, `/key/info`) - The dedicated `monitor` agent key on CT 116 at `/etc/litellm-monitor.env` (root-only 0600) for - model inference checks — the master key must never be used for inference + model inference checks, scoped for every alias step 7 probes (`gpu-dense`, `gpu-vision`, + `strix-moe`) — the master key must never be used for inference ## GPU Fleet Topology @@ -150,10 +151,10 @@ contracts — read them there. Do not re-add retired names (`gemma-4-12b`, `gpu- - Auth uses the dedicated `monitor` agent key, read on CT 116 from `/etc/litellm-monitor.env` (root-only 0600). Do NOT use the master key for inference — the master key is for admin endpoints only (`/key/list`, `/key/generate`, `/key/info`). - - NOTE: the monitor key is scoped for gpu-vision, strix-moe, syslog-auto but NOT gpu-dense. - The gpu-dense probe above requires a key with gpu-dense access; if unavailable, the probe - should be run with a different key or the contract should note the gap rather than silently - dropping host coverage. + - KEY SCOPE: the `monitor` key MUST be scoped for all three probed aliases (`gpu-dense`, + `gpu-vision`, `strix-moe`), otherwise the probe returns 403 and the host is not covered. + If a probe returns 403, widen the monitor key's model list on CT 116 (add the missing + alias) and re-run — never drop the host from the probe to make the check pass. - `gemma-4-12b` was retired and returns 400 `Invalid model name` — do not re-add it to this list. The RTX 5070 host now serves `gpu-vision`. - `/v1/models` is **key-scoped**: a model is only visible to keys allowed to use it, so diff --git a/mumuni-delegation-prose-contract.prose.md b/mumuni-delegation-prose-contract.prose.md index 38e15cb..48c425a 100644 --- a/mumuni-delegation-prose-contract.prose.md +++ b/mumuni-delegation-prose-contract.prose.md @@ -90,8 +90,8 @@ raw data never provided. | Worker | Model | Toolsets | Role | Use When | |--------|-------|----------|------|----------| -| `syslog-code` | qwen3.6-27B-code | terminal, file, web, memory, skills | Code patches, automation, scripts | Writing/modifying code, creating scripts, debugging, reading/writing files | -| `syslog-devops` | qwen3.6-27B-code | terminal, file, web, memory, skills | Infrastructure, DB, bridge, Proxmox | Server ops, SSH, Docker, Proxmox, DB queries, hardware checks | +| `syslog-code` | gpu-dense | terminal, file, web, memory, skills | Code patches, automation, scripts | Writing/modifying code, creating scripts, debugging, reading/writing files | +| `syslog-devops` | gpu-dense | terminal, file, web, memory, skills | Infrastructure, DB, bridge, Proxmox | Server ops, SSH, Docker, Proxmox, DB queries, hardware checks | | `syslog-email` | strix-moe | terminal, file, web, memory, skills | Email automation, mail operations | Sending/receiving email, inbox management, SMTP operations | | `syslog-research` | strix-moe | terminal, file, web, memory, skills, **browser** | Analysis, classification, data processing | Web research, browser tasks, data analysis, classification, reading docs | | `syslog-review` | strix-moe | terminal, file, web, memory, skills | Verification, QA, audit validation | **ALWAYS** verify worker output before delivery — especially for infra changes, code builds, and research findings |