no-mistakes(document): Sweep residual gemma labels; align compression rule contradiction
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Failing after 3s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Skipped
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Skipped
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Skipped
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Failing after 3s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Skipped
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Skipped
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Skipped
This commit is contained in:
+7
-7
@@ -239,7 +239,7 @@ History stored at `/root/data/toks-history.json` with 7-day rolling window.
|
|||||||
### Stable Aliases — CRITICAL
|
### Stable Aliases — CRITICAL
|
||||||
|
|
||||||
All agent configs MUST use stable role-based aliases, never model-specific names:
|
All agent configs MUST use stable role-based aliases, never model-specific names:
|
||||||
- `compression.model: strix-moe`
|
- `compression.model: syslog-auto`
|
||||||
- `auxiliary.vision.model: gpu-vision`
|
- `auxiliary.vision.model: gpu-vision`
|
||||||
- `delegation.model: gpu-dense`
|
- `delegation.model: gpu-dense`
|
||||||
- `auxiliary.web_extract.model: gpu-vision`
|
- `auxiliary.web_extract.model: gpu-vision`
|
||||||
@@ -249,25 +249,25 @@ When the underlying model is swapped, only the LiteLLM config changes — agent
|
|||||||
### Context Windows
|
### Context Windows
|
||||||
- RTX 3090: **128K** (reduced from 256K 2026-07-17) | RTX 5070: **128K** (reduced from 256K) | Strix Halo: **128K**
|
- RTX 3090: **128K** (reduced from 256K 2026-07-17) | RTX 5070: **128K** (reduced from 256K) | Strix Halo: **128K**
|
||||||
- **All agents**: 128K ceiling — stable margin. For >128K workloads, use external providers (deepseek)
|
- **All agents**: 128K ceiling — stable margin. For >128K workloads, use external providers (deepseek)
|
||||||
- Compression threshold 0.60: fires at ~77K (~51K headroom before 128K ceiling)
|
- Compression threshold 0.65 (audit Rule 9): fires at ~85K (~43K headroom before 128K ceiling)
|
||||||
- **Pi agents (Abiba)**: `compaction.reserveTokens: 52739` (≈60% of 128K)
|
- **Pi agents (Abiba)**: `compaction.reserveTokens: 52739` (≈60% of 128K)
|
||||||
- Mumuni compression model alias: `strix-moe`
|
- Mumuni compression model alias: `syslog-auto`
|
||||||
|
|
||||||
### Mumuni Agent Profile
|
### Mumuni Agent Profile
|
||||||
|
|
||||||
Mumuni (kagentz CT105, 192.168.68.14 — migrated from CT100 2026-08-29) is the primary business assistant. This profile is the reference for all agent configs:
|
Mumuni (kagentz CT105, 192.168.68.14 — migrated from CT100 2026-08-29) is the primary business assistant. This profile is the reference for all agent configs. The compression values below are the current required values per template Rules 7/9 and `audit-hermes-config.py`; whether Mumuni's LIVE config currently complies is a separate operational question.
|
||||||
|
|
||||||
| Setting | Value | Notes |
|
| Setting | Value | Notes |
|
||||||
|---------|-------|-------|
|
|---------|-------|-------|
|
||||||
| `model.default` | `syslog-auto` | Balanced default (pool router) |
|
| `model.default` | `syslog-auto` | Balanced default (pool router) |
|
||||||
| `model.provider` | `custom:litellm` | LiteLLM on CT116 |
|
| `model.provider` | `custom:litellm` | LiteLLM on CT116 |
|
||||||
| `compression.model` | `strix-moe` | Stable alias — survives model swaps |
|
| `compression.model` | `syslog-auto` | Rule 7: auto-routing, prevents Strix Halo overload |
|
||||||
| `aux.compression.model` | `strix-moe` | Compression auxiliary model |
|
| `aux.compression.model` | `syslog-auto` | Must match `compression.model` (Rule 7) |
|
||||||
| `aux.vision.model` | `gpu-vision` | Vision tasks (RTX 5070) |
|
| `aux.vision.model` | `gpu-vision` | Vision tasks (RTX 5070) |
|
||||||
| `aux.web_extract.model` | `gpu-vision` | Web extraction |
|
| `aux.web_extract.model` | `gpu-vision` | Web extraction |
|
||||||
| `delegation.model` | `gpu-dense` | Sub-agent reasoning (RTX 3090) |
|
| `delegation.model` | `gpu-dense` | Sub-agent reasoning (RTX 3090) |
|
||||||
| `context.max_context_window` | 131072 (128K) | Reduced from 256K 2026-07-17 — stable 128K ceiling |
|
| `context.max_context_window` | 131072 (128K) | Reduced from 256K 2026-07-17 — stable 128K ceiling |
|
||||||
| `compression.threshold` | 0.60 | Triggers at ~77K (~60% of 128K) — optimized for 128K context |
|
| `compression.threshold` | 0.65 | Rule 9: triggers at ~85K for a 128K window |
|
||||||
| `compression.target_ratio` | 0.3 | Compresses to ~38K |
|
| `compression.target_ratio` | 0.3 | Compresses to ~38K |
|
||||||
| `compression.protect_last_n` | 40 | Preserves last 40 messages |
|
| `compression.protect_last_n` | 40 | Preserves last 40 messages |
|
||||||
| `memory.memory_char_limit` | 800 | Brief memory entries |
|
| `memory.memory_char_limit` | 800 | Brief memory entries |
|
||||||
|
|||||||
@@ -27,8 +27,8 @@ agent: abiba
|
|||||||
┌──────┐ ┌──────┐ ┌────────┐
|
┌──────┐ ┌──────┐ ┌────────┐
|
||||||
│.8:8080│ │.110 │ │.116:80 │
|
│.8:8080│ │.110 │ │.116:80 │
|
||||||
│RTX3090│ │:8080 │ │nginx │
|
│RTX3090│ │:8080 │ │nginx │
|
||||||
│gemma │ │RTX5070│ │router │
|
│qwen │ │RTX5070│ │router │
|
||||||
└──────┘ │qwen27B│ │LiteLLM │
|
└──────┘ │vision │ │LiteLLM │
|
||||||
└──────┘ │dashboard│
|
└──────┘ │dashboard│
|
||||||
└────────┘
|
└────────┘
|
||||||
```
|
```
|
||||||
|
|||||||
@@ -222,7 +222,7 @@ description: >
|
|||||||
|
|
||||||
**Prometheus targets**:
|
**Prometheus targets**:
|
||||||
- 192.168.68.8:9400 (RTX 3090 — qwen)
|
- 192.168.68.8:9400 (RTX 3090 — qwen)
|
||||||
- 192.168.68.110:9400 (RTX 5070 — gemma)
|
- 192.168.68.110:9400 (RTX 5070 — gpu-vision)
|
||||||
- 192.168.68.15:9400 (Strix Halo — qwen3.6-35B-udq4)
|
- 192.168.68.15:9400 (Strix Halo — qwen3.6-35B-udq4)
|
||||||
- harness-litellm:4000 (LiteLLM health)
|
- harness-litellm:4000 (LiteLLM health)
|
||||||
|
|
||||||
|
|||||||
Reference in New Issue
Block a user