diff --git a/gpu-fleet.prose.md b/gpu-fleet.prose.md index 1bef170..3d4d13e 100644 --- a/gpu-fleet.prose.md +++ b/gpu-fleet.prose.md @@ -239,7 +239,7 @@ History stored at `/root/data/toks-history.json` with 7-day rolling window. ### Stable Aliases — CRITICAL All agent configs MUST use stable role-based aliases, never model-specific names: -- `compression.model: strix-moe` +- `compression.model: syslog-auto` - `auxiliary.vision.model: gpu-vision` - `delegation.model: gpu-dense` - `auxiliary.web_extract.model: gpu-vision` @@ -249,25 +249,25 @@ When the underlying model is swapped, only the LiteLLM config changes — agent ### Context Windows - RTX 3090: **128K** (reduced from 256K 2026-07-17) | RTX 5070: **128K** (reduced from 256K) | Strix Halo: **128K** - **All agents**: 128K ceiling — stable margin. For >128K workloads, use external providers (deepseek) -- Compression threshold 0.60: fires at ~77K (~51K headroom before 128K ceiling) +- Compression threshold 0.65 (audit Rule 9): fires at ~85K (~43K headroom before 128K ceiling) - **Pi agents (Abiba)**: `compaction.reserveTokens: 52739` (≈60% of 128K) -- Mumuni compression model alias: `strix-moe` +- Mumuni compression model alias: `syslog-auto` ### Mumuni Agent Profile -Mumuni (kagentz CT105, 192.168.68.14 — migrated from CT100 2026-08-29) is the primary business assistant. This profile is the reference for all agent configs: +Mumuni (kagentz CT105, 192.168.68.14 — migrated from CT100 2026-08-29) is the primary business assistant. This profile is the reference for all agent configs. The compression values below are the current required values per template Rules 7/9 and `audit-hermes-config.py`; whether Mumuni's LIVE config currently complies is a separate operational question. | Setting | Value | Notes | |---------|-------|-------| | `model.default` | `syslog-auto` | Balanced default (pool router) | | `model.provider` | `custom:litellm` | LiteLLM on CT116 | -| `compression.model` | `strix-moe` | Stable alias — survives model swaps | -| `aux.compression.model` | `strix-moe` | Compression auxiliary model | +| `compression.model` | `syslog-auto` | Rule 7: auto-routing, prevents Strix Halo overload | +| `aux.compression.model` | `syslog-auto` | Must match `compression.model` (Rule 7) | | `aux.vision.model` | `gpu-vision` | Vision tasks (RTX 5070) | | `aux.web_extract.model` | `gpu-vision` | Web extraction | | `delegation.model` | `gpu-dense` | Sub-agent reasoning (RTX 3090) | | `context.max_context_window` | 131072 (128K) | Reduced from 256K 2026-07-17 — stable 128K ceiling | -| `compression.threshold` | 0.60 | Triggers at ~77K (~60% of 128K) — optimized for 128K context | +| `compression.threshold` | 0.65 | Rule 9: triggers at ~85K for a 128K window | | `compression.target_ratio` | 0.3 | Compresses to ~38K | | `compression.protect_last_n` | 40 | Preserves last 40 messages | | `memory.memory_char_limit` | 800 | Brief memory entries | diff --git a/gpu-monitor.prose.md b/gpu-monitor.prose.md index a6808e1..72bad0d 100644 --- a/gpu-monitor.prose.md +++ b/gpu-monitor.prose.md @@ -27,8 +27,8 @@ agent: abiba ┌──────┐ ┌──────┐ ┌────────┐ │.8:8080│ │.110 │ │.116:80 │ │RTX3090│ │:8080 │ │nginx │ -│gemma │ │RTX5070│ │router │ -└──────┘ │qwen27B│ │LiteLLM │ +│qwen │ │RTX5070│ │router │ +└──────┘ │vision │ │LiteLLM │ └──────┘ │dashboard│ └────────┘ ``` diff --git a/infrastructure-control.prose.md b/infrastructure-control.prose.md index b9b682b..d7a4a36 100644 --- a/infrastructure-control.prose.md +++ b/infrastructure-control.prose.md @@ -222,7 +222,7 @@ description: > **Prometheus targets**: - 192.168.68.8:9400 (RTX 3090 — qwen) -- 192.168.68.110:9400 (RTX 5070 — gemma) +- 192.168.68.110:9400 (RTX 5070 — gpu-vision) - 192.168.68.15:9400 (Strix Halo — qwen3.6-35B-udq4) - harness-litellm:4000 (LiteLLM health)