Item 6: Auto-compaction at ~60% — update thresholds
- gpu-fleet.prose.md: Update compression.threshold to 0.60 (~77K triggers) - Add pi compaction.reserveTokens: 52739 (≈60% of 128K) - Remove stale 256K hardcode references
This commit is contained in:
+3
-2
@@ -296,7 +296,8 @@ When the underlying model is swapped, only the LiteLLM config changes — agent
|
|||||||
### Context Windows
|
### Context Windows
|
||||||
- RTX 3090: **128K** (reduced from 256K 2026-07-17) | RTX 5070: **128K** (reduced from 256K) | Strix Halo: **128K**
|
- RTX 3090: **128K** (reduced from 256K 2026-07-17) | RTX 5070: **128K** (reduced from 256K) | Strix Halo: **128K**
|
||||||
- **All agents**: 128K ceiling — stable margin. For >128K workloads, use external providers (deepseek)
|
- **All agents**: 128K ceiling — stable margin. For >128K workloads, use external providers (deepseek)
|
||||||
- Compression threshold 0.65: fires at ~85K (~43K headroom before 128K ceiling)
|
- Compression threshold 0.60: fires at ~77K (~51K headroom before 128K ceiling)
|
||||||
|
- **Pi agents (Abiba)**: `compaction.reserveTokens: 52739` (≈60% of 128K)
|
||||||
- Mumuni compression model alias: `strix-moe` with 300s timeout
|
- Mumuni compression model alias: `strix-moe` with 300s timeout
|
||||||
|
|
||||||
### Mumuni Agent Profile
|
### Mumuni Agent Profile
|
||||||
@@ -313,7 +314,7 @@ Mumuni (CT100/abiba, 192.168.68.24) is the primary business assistant. This prof
|
|||||||
| `aux.web_extract.model` | `gpu-light` | Web extraction |
|
| `aux.web_extract.model` | `gpu-light` | Web extraction |
|
||||||
| `delegation.model` | `gpu-dense` | Sub-agent reasoning (RTX 3090) |
|
| `delegation.model` | `gpu-dense` | Sub-agent reasoning (RTX 3090) |
|
||||||
| `context.max_context_window` | 131072 (128K) | Reduced from 256K 2026-07-17 — stable 128K ceiling |
|
| `context.max_context_window` | 131072 (128K) | Reduced from 256K 2026-07-17 — stable 128K ceiling |
|
||||||
| `compression.threshold` | 0.65 | Triggers at ~85K |
|
| `compression.threshold` | 0.60 | Triggers at ~77K (~60% of 128K) — optimized for 128K context |
|
||||||
| `compression.target_ratio` | 0.3 | Compresses to ~38K |
|
| `compression.target_ratio` | 0.3 | Compresses to ~38K |
|
||||||
| `compression.protect_last_n` | 40 | Preserves last 40 messages |
|
| `compression.protect_last_n` | 40 | Preserves last 40 messages |
|
||||||
| `memory.memory_char_limit` | 800 | Brief memory entries |
|
| `memory.memory_char_limit` | 800 | Brief memory entries |
|
||||||
|
|||||||
Reference in New Issue
Block a user