no-mistakes(document): Sweep residual gemma labels; align compression rule contradiction
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Failing after 3s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Skipped
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Skipped
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Skipped

This commit is contained in:
root
2026-09-12 16:54:41 +00:00
parent 1f1b47f59d
commit 78b501798f
3 changed files with 10 additions and 10 deletions
+7 -7
View File
@@ -239,7 +239,7 @@ History stored at `/root/data/toks-history.json` with 7-day rolling window.
### Stable Aliases — CRITICAL ### Stable Aliases — CRITICAL
All agent configs MUST use stable role-based aliases, never model-specific names: All agent configs MUST use stable role-based aliases, never model-specific names:
- `compression.model: strix-moe` - `compression.model: syslog-auto`
- `auxiliary.vision.model: gpu-vision` - `auxiliary.vision.model: gpu-vision`
- `delegation.model: gpu-dense` - `delegation.model: gpu-dense`
- `auxiliary.web_extract.model: gpu-vision` - `auxiliary.web_extract.model: gpu-vision`
@@ -249,25 +249,25 @@ When the underlying model is swapped, only the LiteLLM config changes — agent
### Context Windows ### Context Windows
- RTX 3090: **128K** (reduced from 256K 2026-07-17) | RTX 5070: **128K** (reduced from 256K) | Strix Halo: **128K** - RTX 3090: **128K** (reduced from 256K 2026-07-17) | RTX 5070: **128K** (reduced from 256K) | Strix Halo: **128K**
- **All agents**: 128K ceiling — stable margin. For >128K workloads, use external providers (deepseek) - **All agents**: 128K ceiling — stable margin. For >128K workloads, use external providers (deepseek)
- Compression threshold 0.60: fires at ~77K (~51K headroom before 128K ceiling) - Compression threshold 0.65 (audit Rule 9): fires at ~85K (~43K headroom before 128K ceiling)
- **Pi agents (Abiba)**: `compaction.reserveTokens: 52739` (≈60% of 128K) - **Pi agents (Abiba)**: `compaction.reserveTokens: 52739` (≈60% of 128K)
- Mumuni compression model alias: `strix-moe` - Mumuni compression model alias: `syslog-auto`
### Mumuni Agent Profile ### Mumuni Agent Profile
Mumuni (kagentz CT105, 192.168.68.14 — migrated from CT100 2026-08-29) is the primary business assistant. This profile is the reference for all agent configs: Mumuni (kagentz CT105, 192.168.68.14 — migrated from CT100 2026-08-29) is the primary business assistant. This profile is the reference for all agent configs. The compression values below are the current required values per template Rules 7/9 and `audit-hermes-config.py`; whether Mumuni's LIVE config currently complies is a separate operational question.
| Setting | Value | Notes | | Setting | Value | Notes |
|---------|-------|-------| |---------|-------|-------|
| `model.default` | `syslog-auto` | Balanced default (pool router) | | `model.default` | `syslog-auto` | Balanced default (pool router) |
| `model.provider` | `custom:litellm` | LiteLLM on CT116 | | `model.provider` | `custom:litellm` | LiteLLM on CT116 |
| `compression.model` | `strix-moe` | Stable alias — survives model swaps | | `compression.model` | `syslog-auto` | Rule 7: auto-routing, prevents Strix Halo overload |
| `aux.compression.model` | `strix-moe` | Compression auxiliary model | | `aux.compression.model` | `syslog-auto` | Must match `compression.model` (Rule 7) |
| `aux.vision.model` | `gpu-vision` | Vision tasks (RTX 5070) | | `aux.vision.model` | `gpu-vision` | Vision tasks (RTX 5070) |
| `aux.web_extract.model` | `gpu-vision` | Web extraction | | `aux.web_extract.model` | `gpu-vision` | Web extraction |
| `delegation.model` | `gpu-dense` | Sub-agent reasoning (RTX 3090) | | `delegation.model` | `gpu-dense` | Sub-agent reasoning (RTX 3090) |
| `context.max_context_window` | 131072 (128K) | Reduced from 256K 2026-07-17 — stable 128K ceiling | | `context.max_context_window` | 131072 (128K) | Reduced from 256K 2026-07-17 — stable 128K ceiling |
| `compression.threshold` | 0.60 | Triggers at ~77K (~60% of 128K) — optimized for 128K context | | `compression.threshold` | 0.65 | Rule 9: triggers at ~85K for a 128K window |
| `compression.target_ratio` | 0.3 | Compresses to ~38K | | `compression.target_ratio` | 0.3 | Compresses to ~38K |
| `compression.protect_last_n` | 40 | Preserves last 40 messages | | `compression.protect_last_n` | 40 | Preserves last 40 messages |
| `memory.memory_char_limit` | 800 | Brief memory entries | | `memory.memory_char_limit` | 800 | Brief memory entries |
+2 -2
View File
@@ -27,8 +27,8 @@ agent: abiba
┌──────┐ ┌──────┐ ┌────────┐ ┌──────┐ ┌──────┐ ┌────────┐
│.8:8080│ │.110 │ │.116:80 │ │.8:8080│ │.110 │ │.116:80 │
│RTX3090│ │:8080 │ │nginx │ │RTX3090│ │:8080 │ │nginx │
│gemma │ │RTX5070│ │router │ │qwen │ │RTX5070│ │router │
└──────┘ │qwen27B│ │LiteLLM │ └──────┘ │vision │ │LiteLLM │
└──────┘ │dashboard│ └──────┘ │dashboard│
└────────┘ └────────┘
``` ```
+1 -1
View File
@@ -222,7 +222,7 @@ description: >
**Prometheus targets**: **Prometheus targets**:
- 192.168.68.8:9400 (RTX 3090 — qwen) - 192.168.68.8:9400 (RTX 3090 — qwen)
- 192.168.68.110:9400 (RTX 5070 — gemma) - 192.168.68.110:9400 (RTX 5070 — gpu-vision)
- 192.168.68.15:9400 (Strix Halo — qwen3.6-35B-udq4) - 192.168.68.15:9400 (Strix Halo — qwen3.6-35B-udq4)
- harness-litellm:4000 (LiteLLM health) - harness-litellm:4000 (LiteLLM health)