From 66f94d14fd4139365d53f462933f93beec79a420 Mon Sep 17 00:00:00 2001 From: root Date: Thu, 20 Aug 2026 07:54:36 +0000 Subject: [PATCH] =?UTF-8?q?Item=206:=20Auto-compaction=20at=20~60%=20?= =?UTF-8?q?=E2=80=94=20update=20thresholds?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit - gpu-fleet.prose.md: Update compression.threshold to 0.60 (~77K triggers) - Add pi compaction.reserveTokens: 52739 (≈60% of 128K) - Remove stale 256K hardcode references --- gpu-fleet.prose.md | 5 +++-- 1 file changed, 3 insertions(+), 2 deletions(-) diff --git a/gpu-fleet.prose.md b/gpu-fleet.prose.md index 2cf6da8..13d4066 100644 --- a/gpu-fleet.prose.md +++ b/gpu-fleet.prose.md @@ -296,7 +296,8 @@ When the underlying model is swapped, only the LiteLLM config changes — agent ### Context Windows - RTX 3090: **128K** (reduced from 256K 2026-07-17) | RTX 5070: **128K** (reduced from 256K) | Strix Halo: **128K** - **All agents**: 128K ceiling — stable margin. For >128K workloads, use external providers (deepseek) -- Compression threshold 0.65: fires at ~85K (~43K headroom before 128K ceiling) +- Compression threshold 0.60: fires at ~77K (~51K headroom before 128K ceiling) +- **Pi agents (Abiba)**: `compaction.reserveTokens: 52739` (≈60% of 128K) - Mumuni compression model alias: `strix-moe` with 300s timeout ### Mumuni Agent Profile @@ -313,7 +314,7 @@ Mumuni (CT100/abiba, 192.168.68.24) is the primary business assistant. This prof | `aux.web_extract.model` | `gpu-light` | Web extraction | | `delegation.model` | `gpu-dense` | Sub-agent reasoning (RTX 3090) | | `context.max_context_window` | 131072 (128K) | Reduced from 256K 2026-07-17 — stable 128K ceiling | -| `compression.threshold` | 0.65 | Triggers at ~85K | +| `compression.threshold` | 0.60 | Triggers at ~77K (~60% of 128K) — optimized for 128K context | | `compression.target_ratio` | 0.3 | Compresses to ~38K | | `compression.protect_last_n` | 40 | Preserves last 40 messages | | `memory.memory_char_limit` | 800 | Brief memory entries |