diff --git a/README.md b/README.md index c3860c2..c850817 100644 --- a/README.md +++ b/README.md @@ -88,7 +88,7 @@ prose run memory-audit-maintenance memory_threshold=90 verify_configs=true prose run hermes-config-template agent_name=syslog-devops default_model=claude-sonnet-4 # Configure an agent with a different auxiliary model -prose run hermes-config-template agent_name=syslog-code default_model=qwen3.6-27B-code auxiliary_model=gpu-vision +prose run hermes-config-template agent_name=syslog-code default_model=gpu-dense auxiliary_model=gpu-vision ``` ### Option B: Manual Execution diff --git a/audit-hermes-config.py b/audit-hermes-config.py index 2a69e1b..2c3d109 100644 --- a/audit-hermes-config.py +++ b/audit-hermes-config.py @@ -163,7 +163,7 @@ def audit(path): check( comp.get("max_context_window") == 131072, "Rule 9", - f"compression.max_context_window must be 131072 (got {comp.get('max_context_window')!r}) — matches 128K GPU capacity", + f"compression.max_context_window must be 131072 (got {comp.get('max_context_window')!r}) — syslog-auto pool floor (NVIDIA hosts 128K; Strix Halo 256K)", ) # --- Rule 10: Default Model Must Be syslog-auto --- diff --git a/gpu-fleet.prose.md b/gpu-fleet.prose.md index 490649d..f2f5db7 100644 --- a/gpu-fleet.prose.md +++ b/gpu-fleet.prose.md @@ -8,12 +8,12 @@ description: > UPDATED 2026-07-15: Stable role-based aliases introduced: strix-moe, gpu-dense, gpu-vision (gpu-light was superseded by gpu-vision on 2026-09-12). These never change — only the underlying model does. - Strix Halo: strix-moe → unsloth/Qwen3.6-35B-A3B-MTP (UD-Q4_K_M, 22GB). + Strix Halo: strix-moe → Carnice-Qwen3.6-MoE-35B-A3B-Q4_K_M.gguf (22GB, 256K ctx). RTX 5070: gpu-vision — IQ4_NL + MTP draft (~122 tok/s, 2x faster). - UPDATED 2026-07-17: Context reduced fleet-wide from 256K to 128K for stability. - Strix Halo model swapped to qwen3.6-35B-udq4 (22GB, strix-moe alias). - Instability observed near 100K at 256K (now all GPUs at 128K). 128K is the stable ceiling. - For larger context needs → fall back to external providers (deepseek). + UPDATED 2026-07-17: NVIDIA host context reduced from 256K to 128K for stability. + Strix Halo model: Carnice-Qwen3.6-MoE-35B-A3B-Q4_K_M.gguf (alias strix-moe, 256K context). + Strix Halo runs 256K (n_ctx 262144, --kv-unified); RTX 3090 and RTX 5070 remain at 128K. + For >128K on NVIDIA hosts → fall back to external providers (deepseek). VRAM headroom improved: RTX 3090 ~70%, RTX 5070 ~65%. agent: abiba triggers: @@ -92,7 +92,7 @@ CT 116 `/opt/inference-harness/litellm_config.yaml`. Do not duplicate those valu contracts — read them there. **No backward compatibility**: Old model-specific names (qwen3.6-27B-code, qwen3.5-9b-it) are retired as of 2026-09-12 -but are deprecated for agent configs. Only the stable aliases survive model swaps. `gemma-4-12b` +and no longer resolve; do not use them in agent configs. Only the stable aliases survive model swaps. `gemma-4-12b` is retired and returns 400 `Invalid model name`. ## Routing Configuration (LiteLLM — July 2026) @@ -190,7 +190,7 @@ If no SSH access, send Zulip DM via abiba-bot with vault update instructions. | `/root/scripts/gpu-saturation-watchdog.py` | pi (.24) | Auto-restart stuck llama-server | | `/root/dashboard/gpu-fleet.html` | pi (.24) | Live HTML dashboard | | `/etc/systemd/system/llama-server.service` | .8, .110 | llama-server daemons (Nvidia GPUs) | -| `/etc/systemd/system/strix-server.service` | .15 (amdpve) | llama-server daemon (Vulkan, Strix Halo) running unsloth/Qwen3.6-35B-A3B-MTP-GGUF. Note: `llama-server.service` and `llama-server@.service` are **masked** on .15 to prevent port 8080 collisions. | +| `/etc/systemd/system/strix-server.service` | .15 (amdpve) | llama-server daemon (Vulkan, Strix Halo) running Carnice-Qwen3.6-MoE-35B-A3B-Q4_K_M.gguf (256K ctx). Note: `llama-server.service` and `llama-server@.service` are **masked** on .15 to prevent port 8080 collisions. | ## Prometheus & Grafana @@ -247,8 +247,8 @@ All agent configs MUST use stable role-based aliases, never model-specific names When the underlying model is swapped, only the LiteLLM config changes — agent configs are untouched. ### Context Windows -- RTX 3090: **128K** (reduced from 256K 2026-07-17) | RTX 5070: **128K** (reduced from 256K) | Strix Halo: **128K** -- **All agents**: 128K ceiling — stable margin. For >128K workloads, use external providers (deepseek) +- RTX 3090: **128K** (reduced from 256K 2026-07-17) | RTX 5070: **128K** (reduced from 256K) | Strix Halo: **256K** (2026-09-12) +- **Agents via `syslog-auto`**: 128K ceiling — the pool's safe floor (NVIDIA hosts are 128K). For >128K workloads, use external providers (deepseek) - Compression threshold 0.65 (audit Rule 9): fires at ~85K (~43K headroom before 128K ceiling) - **Pi agents (Abiba)**: `compaction.reserveTokens: 52739` (≈60% of 128K) - Mumuni compression model alias: `syslog-auto` @@ -266,7 +266,7 @@ Mumuni (kagentz CT105, 192.168.68.14 — migrated from CT100 2026-08-29) is the | `aux.vision.model` | `gpu-vision` | Vision tasks (RTX 5070) | | `aux.web_extract.model` | `gpu-vision` | Web extraction | | `delegation.model` | `gpu-dense` | Sub-agent reasoning (RTX 3090) | -| `context.max_context_window` | 131072 (128K) | Reduced from 256K 2026-07-17 — stable 128K ceiling | +| `context.max_context_window` | 131072 (128K) | Conservative `syslog-auto` pool floor (NVIDIA hosts 128K; Strix Halo 256K) | | `compression.threshold` | 0.65 | Rule 9: triggers at ~85K for a 128K window | | `compression.target_ratio` | 0.3 | Compresses to ~38K | | `compression.protect_last_n` | 40 | Preserves last 40 messages | diff --git a/gpu-self-heal.prose.md b/gpu-self-heal.prose.md index 89bbd1b..c7c26ae 100644 --- a/gpu-self-heal.prose.md +++ b/gpu-self-heal.prose.md @@ -52,7 +52,7 @@ depends_on: Key notes: - All models use direct GPU routing via LiteLLM (`api_key: not-needed`). Router (port 9000) was decommissioned 2026-09-11 and is NOT in the inference path. -- Stable aliases (gpu-dense, gpu-vision, strix-moe) from gpu-fleet are the canonical names for agent configs. Model-specific names still work but are deprecated. The retired names `gpu-light` and `gemma-4-12b` were superseded by `gpu-vision` on 2026-09-12 and no longer resolve (400 `Invalid model name`). +- Stable aliases (gpu-dense, gpu-vision, strix-moe) from gpu-fleet are the canonical names for agent configs — use these, not model-specific names. The retired names `gpu-light` and `gemma-4-12b` were superseded by `gpu-vision` on 2026-09-12 and no longer resolve (400 `Invalid model name`). - The RTX 5070 is the fastest endpoint per token — `gpu-vision` is its canonical alias. Route vision/web/light work there first. The RTX 5070 model is multimodal (image+text). - Strix Halo is 62.9 tok/s (89% of 70.5 baseline) — below optimal but stable. Check for competing workloads. - RTX 3090 VRAM at 88% — within role-appropriate range (role = heavy reasoning, needs the headroom). @@ -142,7 +142,7 @@ Key notes: - **Escalate**: If Tier 2 triggers and temp still rising after 5 min → possible hardware failure ### Rule 9: Context Window Optimization -- **Detect**: Benchmark tok/s vs baseline for each GPU at current context (all 128K) +- **Detect**: Benchmark tok/s vs baseline for each GPU at its current context - RTX 3090 (128K ctx, ThinkingCap): baseline 74.8 tok/s — currently at 74.9 (100%) - RTX 5070 (128K ctx, HauhauCS QAT): baseline 165.2 tok/s — currently at 169.6 (103%) - Strix Halo (256K ctx, Carnice-Qwen3.6-MoE-35B-A3B-Q4_K_M.gguf): baseline 70.5 tok/s — currently at 62.9 (89%) diff --git a/hermes-agent-baseline.prose.md b/hermes-agent-baseline.prose.md index 2e2eb9a..fba3a89 100644 --- a/hermes-agent-baseline.prose.md +++ b/hermes-agent-baseline.prose.md @@ -5,7 +5,7 @@ version: 1.0.0 description: > Canonical known-good baseline for all Syslog Hermes agents. Captures the exact configuration state, keys, workarounds, and audit procedure. When an agent's - configuration goes sideways, restore from this baseline. Last verified 2026-07-16. All GPUs 128K context (reduced from 256K for stability Jul 2026) (RTX 3090 .8, RTX 5070 .110, Strix Halo .15). Parallel 1 fleet-wide (Strix Halo handles compression solo). + configuration goes sideways, restore from this baseline. Last verified 2026-07-16. RTX 3090/5070 at 128K (reduced from 256K for stability Jul 2026); Strix Halo at 256K (2026-09-12) (RTX 3090 .8, RTX 5070 .110, Strix Halo .15). Parallel 1 fleet-wide (Strix Halo handles compression solo). author: Abiba (pi agent) --- diff --git a/hermes-config-template.prose.md b/hermes-config-template.prose.md index 1a98545..a6e3d29 100644 --- a/hermes-config-template.prose.md +++ b/hermes-config-template.prose.md @@ -97,7 +97,7 @@ model: base_url: http://192.168.68.116/litellm/v1 # Rule 5 (2026-08-09): canonical authenticated path; /v1 also OK api_key_env: LITELLM_API_KEY # Injected via infisical run -- wrapper max_tokens: 4096 # ⚠️ CRITICAL: Prevents unbounded generation - context_length: 131072 # For syslog-auto (all GPUs at 128K for stability). + context_length: 131072 # Conservative floor for syslog-auto (NVIDIA hosts 128K; Strix Halo 256K). # ⚠️ MANDATORY: Hermes probes unknown models from 256K # and falls back to 256K when /v1/models lacks a context # field (llama-server does). Without this override, agents @@ -133,7 +133,7 @@ compression: enabled: true model: syslog-auto # ⚠️ Must match auxiliary.compression.model. Stable alias (gpu-fleet § Stable Role-Based Aliases). NOT ornith-1.0-35b (LiteLLM does not serve that name). provider: harness - max_context_window: 131072 # MUST match actual GPU capacity. All 3 GPUs are 128K (Jul 17). + max_context_window: 131072 # MUST stay at the syslog-auto pool floor: NVIDIA hosts are 128K, Strix Halo 256K (2026-09-12). threshold: 0.65 # Fires at ~170K for 262K window, ~85K for 128K target_ratio: 0.30 protect_last_n: 40 @@ -253,7 +253,7 @@ The following MUST be identical across ALL profiles: ### Rule 7: Auxiliary Model Consistency (UPDATED 2026-07-16) - Vision and web_extract use `gpu-vision` (RTX 5070 — 12GB, vision-optimized) -- Compression uses `syslog-auto` — the Strix Halo weighted pool (64GB, 128K ctx, compression-optimized); do NOT pin `compression.model` to `strix-moe` (audit Rule 7 rejects it) +- Compression uses `syslog-auto` — the Strix Halo weighted pool (64GB, 256K ctx, compression-optimized); do NOT pin `compression.model` to `strix-moe` (audit Rule 7 rejects it) - **`ornith-1.0-35b` is NOT a valid compression model name** — LiteLLM does not serve it (do not restate the served model list here — CT 116 `/opt/inference-harness/litellm_config.yaml` is the single source of truth for models, aliases, weights and fallbacks). Old configs with `ornith-1.0-35b` cause 403/model-not-found on compression calls. @@ -267,15 +267,15 @@ The following MUST be identical across ALL profiles: - `api_key_env: LITELLM_API_KEY` - **Do NOT use `syslog-auto` for `vision`/`web_extract`** — it routes unpredictably; compression is the deliberate exception (see the OPERATIONAL DECISION above) - **Compression on Strix Halo**: The strix-moe alias routes to Strix Halo - (64GB UMA, 128K context) — the designated compression GPU. This frees the + (64GB UMA, 256K context) — the designated compression GPU. This frees the RTX 5070 for vision and web search, and the RTX 3090 for heavy reasoning. - The `compression:` block's `model` MUST match `auxiliary: compression: model` -- The `compression: max_context_window: 131072` MUST match actual GPU capacity (128K) +- The `compression: max_context_window: 131072` MUST stay at the syslog-auto pool floor (NVIDIA hosts 128K; Strix Halo 256K) ### Rule 8: GPU Workload Distribution (UPDATED 2026-07-16) - **RTX 3090 (24GB, 128K ctx, gpu-dense)**: Heavy reasoning, code gen, long conversations - **RTX 5070 (12GB, 128K ctx, gpu-vision)**: Vision, web search, quick tasks, web_extract (IQ4_NL+MTP, ~65% VRAM at 128K) -- **Strix Halo (64GB, 128K ctx, syslog-auto)**: Context compression, summarization, long docs +- **Strix Halo (64GB, 256K ctx, syslog-auto)**: Context compression, summarization, long docs - Agent profiles MUST route auxiliary tasks to the correct GPU: - `auxiliary.vision.model: gpu-vision` (RTX 5070) - `auxiliary.web_extract.model: gpu-vision` (RTX 5070) @@ -284,14 +284,14 @@ The following MUST be identical across ALL profiles: - For 128K context window: `threshold: 0.65` (fires at ~85K tokens) - Do NOT use `threshold: 0.25` — this fires at 65K, causing premature context loss - Do NOT use `threshold: 0.80` — this delays until ~105K, leaving only 23K margin -- `max_context_window: 131072` MUST match the model's actual capacity (128K) +- `max_context_window: 131072` MUST stay at the pool floor (NVIDIA hosts 128K; Strix Halo 256K) - See `devops-hermes-compression` skill for full reference ### Rule 9: Compression Threshold for 128K Models - For 128K context window: `threshold: 0.65` (fires at ~85K tokens) - Do NOT use `threshold: 0.25` — this fires at 65K, causing premature context loss - Do NOT use `threshold: 0.80` — this delays until ~105K, leaving only 23K margin -- `max_context_window: 131072` MUST match the model's actual capacity (all GPUs = 128K) +- `max_context_window: 131072` MUST stay at the pool floor (NVIDIA hosts 128K; Strix Halo 256K) - See `devops-hermes-compression` skill for full reference ### Rule 10: Default Model Must Be `syslog-auto` (All Agents) @@ -322,7 +322,8 @@ When an agent shows "context issues" (premature compression, 401s, 504s, DeepSee verify ALL FOUR of these against the live config. They are the only root causes found in production: 1. **max_context_window correct?** — BOTH `compression.max_context_window` AND `context.max_context_window` - MUST be `131072` (all GPUs are 128K). A value of `262144` causes instability near 100K and must NOT be used. + MUST be `131072` (the syslog-auto pool floor: NVIDIA hosts are 128K; Strix Halo is 256K). + A `262144` client window can route to a 128K NVIDIA host and fail, so it must NOT be used. ~83K instead of ~170K. Check: `grep -n max_context_window ~/.hermes/config.yaml` 2. **base_url uses authenticated path?** — `custom_providers[0].base_url`, `delegation.base_url`, and ALL `auxiliary.*.base_url` MUST be `http://192.168.68.116/litellm/v1` (Rule 5, canonical) diff --git a/inference-optimization.prose.md b/inference-optimization.prose.md index b0ee937..ce660c5 100644 --- a/inference-optimization.prose.md +++ b/inference-optimization.prose.md @@ -4,7 +4,7 @@ kind: responsibility description: > Optimizes the full Syslog inference stack — LiteLLM routing weights, GPU model assignments, agent context management, and prompt caching — to reduce response - times to sub-15s average. All GPUs now at 128K context (stable ceiling). + times to sub-15s average. NVIDIA GPUs at 128K context; Strix Halo at 256K (2026-09-12). id: 067NC6KP02RG60S50M40E30928 --- @@ -63,7 +63,7 @@ prefill time at 532 tok/s. Fix context first, routing second. - **Cache everything repeated**: system prompts, skill docs, AGENTS.md — these never change between turns. Single-digit cache hit rate is unacceptable. - **Lower context ceiling**: 128K window is the stable ceiling for agent conversations. - GPUs reduced from 256K to 128K (2026-07-17). 128K window should compact at 85K (0.65 threshold). For larger contexts, route to external providers. + GPUs reduced from 256K to 128K (2026-07-17) for the NVIDIA hosts; Strix Halo runs 256K (2026-09-12). 128K window should compact at 85K (0.65 threshold). For larger contexts, route to external providers. ### Shape diff --git a/infrastructure-control.prose.md b/infrastructure-control.prose.md index 5c40667..9734a1c 100644 --- a/infrastructure-control.prose.md +++ b/infrastructure-control.prose.md @@ -223,7 +223,7 @@ description: > **Prometheus targets**: - 192.168.68.8:9400 (RTX 3090 — qwen) - 192.168.68.110:9400 (RTX 5070 — gpu-vision) -- 192.168.68.15:9400 (Strix Halo — qwen3.6-35B-udq4) +- 192.168.68.15:9400 (Strix Halo — strix-moe) - harness-litellm:4000 (LiteLLM health) ### Ecosystem C: Netbird (72.61.0.17 — Hostinger srv1079750.hstgr.cloud) diff --git a/litellm-health.prose.md b/litellm-health.prose.md index 2061843..5e47053 100644 --- a/litellm-health.prose.md +++ b/litellm-health.prose.md @@ -44,7 +44,7 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2) **What changed (v3.2.0 → v4.0.0 — 2026-07-08)**: - Router REMOVED from request path — LiteLLM proxies directly to GPU - All GPUs at parallel 2 (was parallel 1) -- NVIDIA context reduced 256K→128K to free VRAM — now the stable ceiling across all GPUs (2026-07-17) +- NVIDIA context reduced 256K→128K to free VRAM — the stable NVIDIA ceiling (2026-07-17); Strix Halo runs 256K (2026-09-12) - Timeouts and fallback chains are config state — read them from CT 116 `/opt/inference-harness/litellm_config.yaml`; they are not duplicated here. ## Parameters @@ -69,7 +69,8 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2) - Network access to public_url, auth_host, and gpu_dashboard_url - LiteLLM master key for admin endpoints only (`/key/list`, `/key/generate`, `/key/info`) - The dedicated `monitor` agent key on CT 116 at `/etc/litellm-monitor.env` (root-only 0600) for - model inference checks — the master key must never be used for inference + model inference checks, scoped for every alias step 7 probes (`gpu-dense`, `gpu-vision`, + `strix-moe`) — the master key must never be used for inference ## GPU Fleet Topology @@ -144,12 +145,16 @@ contracts — read them there. Do not re-add retired names (`gemma-4-12b`, `gpu- 7. **Check model inference via LiteLLM** — Test one model on each GPU host. The health check runs on the **backend edge**, not the public edge, so these paths carry the `/litellm/` prefix: - - POST http://{{backend_host}}/litellm/v1/chat/completions model=gpu-vision → expect 200 (RTX 5070, .110) + - POST http://{{backend_host}}/litellm/v1/chat/completions model=gpu-dense → expect 200 (RTX 3090, .8) - POST http://{{backend_host}}/litellm/v1/chat/completions model=gpu-vision → expect 200 (RTX 5070, .110) - POST http://{{backend_host}}/litellm/v1/chat/completions model=strix-moe → expect 200 (Strix Halo, .15) - Auth uses the dedicated `monitor` agent key, read on CT 116 from `/etc/litellm-monitor.env` (root-only 0600). Do NOT use the master key for inference — the master key is for admin endpoints only (`/key/list`, `/key/generate`, `/key/info`). + - KEY SCOPE: the `monitor` key MUST be scoped for all three probed aliases (`gpu-dense`, + `gpu-vision`, `strix-moe`), otherwise the probe returns 403 and the host is not covered. + If a probe returns 403, widen the monitor key's model list on CT 116 (add the missing + alias) and re-run — never drop the host from the probe to make the check pass. - `gemma-4-12b` was retired and returns 400 `Invalid model name` — do not re-add it to this list. The RTX 5070 host now serves `gpu-vision`. - `/v1/models` is **key-scoped**: a model is only visible to keys allowed to use it, so diff --git a/mumuni-delegation-prose-contract.prose.md b/mumuni-delegation-prose-contract.prose.md index 38e15cb..48c425a 100644 --- a/mumuni-delegation-prose-contract.prose.md +++ b/mumuni-delegation-prose-contract.prose.md @@ -90,8 +90,8 @@ raw data never provided. | Worker | Model | Toolsets | Role | Use When | |--------|-------|----------|------|----------| -| `syslog-code` | qwen3.6-27B-code | terminal, file, web, memory, skills | Code patches, automation, scripts | Writing/modifying code, creating scripts, debugging, reading/writing files | -| `syslog-devops` | qwen3.6-27B-code | terminal, file, web, memory, skills | Infrastructure, DB, bridge, Proxmox | Server ops, SSH, Docker, Proxmox, DB queries, hardware checks | +| `syslog-code` | gpu-dense | terminal, file, web, memory, skills | Code patches, automation, scripts | Writing/modifying code, creating scripts, debugging, reading/writing files | +| `syslog-devops` | gpu-dense | terminal, file, web, memory, skills | Infrastructure, DB, bridge, Proxmox | Server ops, SSH, Docker, Proxmox, DB queries, hardware checks | | `syslog-email` | strix-moe | terminal, file, web, memory, skills | Email automation, mail operations | Sending/receiving email, inbox management, SMTP operations | | `syslog-research` | strix-moe | terminal, file, web, memory, skills, **browser** | Analysis, classification, data processing | Web research, browser tasks, data analysis, classification, reading docs | | `syslog-review` | strix-moe | terminal, file, web, memory, skills | Verification, QA, audit validation | **ALWAYS** verify worker output before delivery — especially for infra changes, code builds, and research findings | diff --git a/proxmox-monitor.prose.md b/proxmox-monitor.prose.md index 39e81d6..a7904e1 100644 --- a/proxmox-monitor.prose.md +++ b/proxmox-monitor.prose.md @@ -87,7 +87,7 @@ agent: abiba | storepve | 192.168.68.6 | PVE | | acerpve | 192.168.68.9 | PVE (hosts llm-gpu qemu/101) | | minipve | 192.168.68.12 | PVE | -| amdpve | 192.168.68.15 | PVE + Strix Halo LLM (qwen3.6-35B-udq4, strix-moe) | +| amdpve | 192.168.68.15 | PVE + Strix Halo LLM (strix-moe) | ## Operations