From 39e2fa0cfa442bd3674d58136b330050e79b48c5 Mon Sep 17 00:00:00 2001 From: Abiba Date: Thu, 16 Jul 2026 13:57:31 +0000 Subject: [PATCH] =?UTF-8?q?Contracts:=20Mumuni=20context-fix=20(WAL=20#130?= =?UTF-8?q?0)=20=E2=80=94=20strix-moe,=20256K=20all=20GPUs,=20Rule=2012/13?= =?UTF-8?q?,=20machine-identity=20vault=20procedure?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit --- gpu-fleet.prose.md | 6 +-- hermes-config-template.prose.md | 83 +++++++++++++++++++++++++-------- litellm-api-keys.prose.md | 48 ++++++++++++++++++- 3 files changed, 113 insertions(+), 24 deletions(-) diff --git a/gpu-fleet.prose.md b/gpu-fleet.prose.md index 0704ff7..72a6fa6 100644 --- a/gpu-fleet.prose.md +++ b/gpu-fleet.prose.md @@ -67,7 +67,7 @@ triggers: │ CT 8 │ │ CT 110 │ │ CT 15 │ │ pi (.24) │ │ RTX 3090 │ │ RTX 5070 │ │ Strix Halo│ │ GPU Monitor │ │ 24GB │ │ 12GB │ │ 64GB UMA │ │ :9100 │ -│ 256K ctx │ │ 131K ctx │ │ 256K ctx │ │ Watchdog │ +│ 256K ctx │ │ 256K ctx │ │ 256K ctx │ │ Watchdog │ │ qwen3.6 │ │ gemma-4-12b │ │ ornith35B │ │ Prometheus │ │ 27B-code │ │ :8080 │ │ :8080 │ │ exporter │ │ :8080 │ │ :9400 (exp) │ │ :9400(exp)│ │ :9401 │ @@ -94,7 +94,7 @@ but are deprecated for agent configs. Only the stable aliases survive model swap | Model | GPU | Host | VRAM | Ctx | KV Cache | Parallel | Batch/Ubatch | Status | |-------|-----|------|------|-----|----------|----------|-------------|--------| | qwen3.6-27B-code (MTP) | RTX 3090 | .8 (llm-gpu) | 22.2/24.6GB (90%) | **256K** 🚀 | turbo4 | 2 | default | ✅ 63 tok/s | -| gemma-4-12b | RTX 5070 | .110 (ocu-llm) | 10.0/12.2GB (82%) | 131K | q4_0 | 2 | 2048/1024 | ✅ healthy | +| gemma-4-12b | RTX 5070 | .110 (ocu-llm) | 10.0/12.2GB (82%) | 256K | q4_0 | 2 | 2048/1024 | ✅ healthy | | qwen3.6-35B-udq4 | Strix Halo Vulkan | .15 (amdpve) | ~9GB/64GB | 256K | q8_0 | 2 | 2048/512 | ✅ healthy | ## Routing Configuration (LiteLLM — July 2026) @@ -306,7 +306,7 @@ Mumuni (CT114, 192.168.68.123) is the primary business assistant. This profile i | `aux.vision.model` | `gpu-light` | Vision tasks (RTX 5070) | | `aux.web_extract.model` | `gpu-light` | Web extraction | | `delegation.model` | `gpu-dense` | Sub-agent reasoning (RTX 3090) | -| `context.max_context_window` | 131072 (128K) | Kept at 128K to avoid compression timeout | +| `context.max_context_window` | 262144 (256K) | Fixed 2026-07-16 (was 131072 — caused premature compression, WAL #1300) | | `compression.threshold` | 0.65 | Triggers at ~85K | | `compression.target_ratio` | 0.3 | Compresses to ~38K | | `compression.protect_last_n` | 40 | Preserves last 40 messages | diff --git a/hermes-config-template.prose.md b/hermes-config-template.prose.md index 275e846..fda8075 100644 --- a/hermes-config-template.prose.md +++ b/hermes-config-template.prose.md @@ -5,10 +5,12 @@ description: > Standard Hermes configuration template for Syslog Solution LLC agents. Enforces shared infrastructure setup (Firecrawl, SearXNG, local models, RA-H OS MCP) while keeping agent-specific API keys and model choices. - UPDATED 2026-07-12: GPU workload redistributed. Compression → Strix Halo (ornith-1.0-35b). - RTX 3090 context verified at 256K (was incorrectly documented as 128K). - RTX 5070 stays at 131K for vision/web. Infisical .env fallback required (Rule 3). - Parallel counts corrected (RTX 3090=1, RTX 5070=2, Strix=2). + UPDATED 2026-07-16: Compression model is the stable alias `strix-moe` (NOT `ornith-1.0-35b`, + which LiteLLM does not serve). All 3 GPUs verified 256K (RTX 5070 bumped 131K→256K on Jul 15). + Added Rule 12 (Context-Issue Diagnostic) + Rule 13 (.env fallback enforcement) from the + 2026-07-16 Mumuni root-cause investigation (WAL #1300). + UPDATED 2026-07-12: GPU workload redistributed. Compression → Strix Halo. RTX 3090 context + verified at 256K. Infisical .env fallback required (Rule 3/13). --- ## Maintains @@ -31,7 +33,7 @@ Sub-agent profiles inherit auth from the main config — no separate keys needed | Agent | Key Alias | Host | SSH | Sub-Agents | |-------|-----------|------|-----|-----------| | Tanko | `tanko-*` | 192.168.68.122 | jerome@.122 | — | -| Mumuni | `mumuni-jul2026` | 192.168.68.123 | root@.123 | 6 profiles ✱ | +| Mumuni | `mumuni` | 192.168.68.123 | root@.123 | 6 profiles ✱ | | Abiba | `abiba-pi` | 192.168.68.24 | local | — | | Koby | `koby` | CT 111 (tdunna) | Zulip | — | | Koonimo | `koonimo` | CT 113 (baggy) | Zulip | — | @@ -129,10 +131,9 @@ mcp_servers: # ─── Compression ─── compression: enabled: true - model: ornith-1.0-35b # ⚠️ Must match auxiliary.compression.model + model: strix-moe # ⚠️ Must match auxiliary.compression.model. Stable alias (gpu-fleet § Stable Role-Based Aliases). NOT ornith-1.0-35b (LiteLLM does not serve that name). provider: harness - max_context_window: 262144 # For Strix Halo compression (256K ctx). - # Set 131072 if using gemma directly. + max_context_window: 262144 # MUST match actual GPU capacity. All 3 GPUs are 256K (Jul 15). threshold: 0.65 # Fires at ~170K for 262K window, ~85K for 128K target_ratio: 0.30 protect_last_n: 40 @@ -164,10 +165,10 @@ auxiliary: timeout: 30 compression: provider: harness - model: ornith-1.0-35b - base_url: http://192.168.68.116/v1 + model: strix-moe # MUST match compression.model above. Stable alias for Strix Halo. + base_url: http://192.168.68.116/v1 # Rule 5: /v1 NOT /litellm/v1 api_key_env: LITELLM_API_KEY - timeout: 60 + timeout: 300 # gpu-fleet: 300s for large-history summarization (was 60) # ─── Custom Provider ─── custom_providers: @@ -236,23 +237,25 @@ The following MUST be identical across ALL profiles: - Apply to BOTH main config AND all sub-agent profiles - For agents needing longer outputs: raise to 8192, but never omit -### Rule 7: Auxiliary Model Consistency (UPDATED July 2026) +### Rule 7: Auxiliary Model Consistency (UPDATED 2026-07-16) - Vision and web_extract use `gemma-4-12b` (RTX 5070 — 12GB, vision-optimized) -- Compression uses `ornith-1.0-35b` (Strix Halo — 64GB, 256K ctx, compression-optimized) +- Compression uses `strix-moe` (stable alias for Strix Halo — 64GB, 256K ctx, compression-optimized) +- **`strix-moe` is the only valid compression model name** — LiteLLM does NOT serve `ornith-1.0-35b` + (it serves `strix-moe`, `qwen3.6-35B-udq4`, `gpu-dense`, `gpu-light`, `syslog-auto`, `gemma-4-12b`, `qwen3.6-27B-code`). Old configs with `ornith-1.0-35b` cause 403/model-not-found on compression calls. - All auxiliary services MUST use identical routing: - - `base_url: http://192.168.68.116/v1` + - `base_url: http://192.168.68.116/v1` (Rule 5: `/v1`, NOT `/litellm/v1`) - `api_key_env: LITELLM_API_KEY` - **Do NOT use `syslog-auto`** for auxiliary tasks — it routes unpredictably -- **Compression moved to Strix Halo (July 2026)**: The ornith-1.0-35b model on Strix Halo - (64GB UMA, 256K context, 72.4 tok/s) is the designated compression GPU. This frees the +- **Compression on Strix Halo**: The strix-moe alias routes to Strix Halo + (64GB UMA, 256K context) — the designated compression GPU. This frees the RTX 5070 for vision and web search, and the RTX 3090 for heavy reasoning. - The `compression:` block's `model` MUST match `auxiliary: compression: model` - The `compression: max_context_window: 262144` MUST match Strix Halo's actual capacity -### Rule 8: GPU Workload Distribution (July 2026) +### Rule 8: GPU Workload Distribution (UPDATED 2026-07-16) - **RTX 3090 (24GB, 256K ctx, qwen3.6-27B-code)**: Heavy reasoning, code gen, long conversations -- **RTX 5070 (12GB, 131K ctx, gemma-4-12b)**: Vision, web search, quick tasks, web_extract -- **Strix Halo (64GB, 256K ctx, ornith-1.0-35b)**: Context compression, summarization, long docs +- **RTX 5070 (12GB, 256K ctx, gemma-4-12b)**: Vision, web search, quick tasks, web_extract (bumped 131K→256K Jul 15; IQ4_NL+MTP, 88% VRAM) +- **Strix Halo (64GB, 256K ctx, strix-moe)**: Context compression, summarization, long docs - Agent profiles MUST route auxiliary tasks to the correct GPU: - `auxiliary.vision.model: gemma-4-12b` (RTX 5070) - `auxiliary.web_extract.model: gemma-4-12b` (RTX 5070) @@ -293,6 +296,48 @@ The following MUST be identical across ALL profiles: - A non-existent model ID causes 403 errors that silently break the pi RPC worker (no `agent_end` emitted, worker stays "busy", Zulip messages pile up unprocessed) +### Rule 12: Context-Issue Diagnostic Checklist (ADDED 2026-07-16, WAL #1300) +When an agent shows "context issues" (premature compression, 401s, 504s, DeepSeek fallback), +verify ALL FOUR of these against the live config. They are the only root causes found in production: + +1. **max_context_window correct?** — BOTH `compression.max_context_window` AND `context.max_context_window` + MUST be `262144` (all GPUs are 256K). A value of `131072` causes premature compression at + ~83K instead of ~170K. Check: `grep -n max_context_window ~/.hermes/config.yaml` +2. **base_url uses /v1 NOT /litellm/v1?** — `custom_providers[0].base_url`, `delegation.base_url`, + and ALL `auxiliary.*.base_url` MUST be `http://192.168.68.116/v1` (Rule 5). nginx `/litellm/` + has a 60s default timeout → 504 on any inference >60s; `/v1/` has 600s. + Check: `grep -n 'litellm/v1' ~/.hermes/config.yaml` (must return NOTHING) +3. **LITELLM_API_KEY valid?** — The key must be a real LiteLLM key (`sk-` + 64 hex, 67 chars). + Malformed values (e.g. `sk-_SWAl_Vu_…`, 47 chars) return 401 → DeepSeek fallback. + Verify: `curl -s -o /dev/null -w '%{http_code}' -H "Authorization: Bearer $LITELLM_API_KEY" http://192.168.68.116/v1/models` (must be 200) +4. **custom_providers aligned?** — `model: syslog-auto` (Rule 10), `api_mode: chat_completions` + (NOT `responses`). A wrong api_mode causes silent request failures. + +One-line agent health check (run on the agent host): +```bash +PID=$(pgrep -f "python -m hermes_cli.main gateway run" | head -1) +cat /proc/$PID/environ | tr '\0' '\n' | grep ^LITELLM_API_KEY= | sed 's/=.*//' +curl -s -o /dev/null -w 'key_health: %{http_code}\n' -H "Authorization: Bearer $(cat /proc/$PID/environ | tr '\0' '\n' | grep ^LITELLM_API_KEY= | cut -d= -f2)" http://192.168.68.116/v1/models +``` + +### Rule 13: .env Fallback Enforcement (ADDED 2026-07-16, WAL #1300) +- The infisical-gateway.sh wrapper MUST source `~/.hermes/.env` BEFORE exporting `LITELLM_API_KEY`, + so the local .env overrides a stale Infisical vault value: + ```bash + -- bash -c " + . /root/.hermes/.env # [FALLBACK Rule 3/13] local .env overrides stale vault value + export LITELLM_API_KEY=\"\$MUMUNI_LITELLM_API_KEY\" + ... + ``` +- `~/.hermes/.env` MUST contain `MUMUNI_LITELLM_API_KEY=sk-` (or the agent's equivalent). +- This is sanctioned by Rule 3: "Infisical service tokens can expire/404; the .env fallback prevents + agents from running without keys." +- The systemd unit `hermes-gateway.service` ExecStart MUST be `/root/.hermes/infisical-gateway.sh` + (the wrapper), NOT `python -m hermes_cli.main gateway run` directly (direct exec breaks ALL env + injection → "No messaging platforms enabled"). +- To sync the Infisical vault (when write access is available): see `litellm-api-keys.prose.md` + § Machine Identity for Vault Writes. + ## Execution 1. **Check current config** — Read the target agent's config.yaml diff --git a/litellm-api-keys.prose.md b/litellm-api-keys.prose.md index d3fc3e7..4ed734c 100644 --- a/litellm-api-keys.prose.md +++ b/litellm-api-keys.prose.md @@ -51,8 +51,8 @@ description: > - Generate new key with key_alias: "{agent_name}" (e.g., "tanko" — bare name, no date) - Set metadata: { "agent": "{agent_name}", "purpose": "agent-inference" } - Duration is null (permanent) — inherited from litellm default_key_generate_params - - Set models: ["syslog-auto", "qwen3.6-27B-code", "gemma-4-12b", "ornith-1.0-35b"] - - Note: qwen3.6-35B-A3B removed from fleet (was never deployed on any GPU) + - Set models: ["syslog-auto", "qwen3.6-27B-code", "gemma-4-12b", "strix-moe", "gpu-dense", "gpu-light", "qwen3.6-35B-udq4"] + - Note: `ornith-1.0-35b` is NOT a valid LiteLLM model name (use `strix-moe`, the stable alias). qwen3.6-35B-A3B removed from fleet (was never deployed). - Return the new key 5. **If action == "rotate"**: - Generate new key with same alias (LiteLLM replaces the old key) @@ -67,3 +67,47 @@ description: > - Test the key against LiteLLM /v1/models - Confirm key alias matches agent_name in LiteLLM key list - Verify agent gateway uses vault wrapper: `cat /proc//cmdline` shows `infisical run` + +## Machine Identity for Vault Writes (ADDED 2026-07-16, WAL #1300) + +**Problem:** The infisical CLI on agent hosts is logged in as a user session (jerome@sysloggh.com). +In CLI v0.38.0, `infisical secrets set` / `infisical export` fail with "project id missing" / "workspace +key 404" — a known bug where user-session auth works for `run` but NOT for `secrets set`. The apt +repo only ships 0.38.0, so `apt upgrade` does not help. + +**Proper fix — Machine Identity (Infisical automation best practice):** +Create a machine identity with READ+WRITE scope on the `agents` project (project_id= +`322fceab-39da-4854-a55a-568e76c0f13f`, env `prod`). Store client_id + client_secret securely. +Then vault writes work from any host: +```bash +# Get a machine-identity access token +TOKEN=$(curl -fsSL -X POST https://vault.sysloggh.net/api/v1/auth/universal-auth/login \ + -H 'Content-Type: application/json' \ + -d '{"clientId":"","clientSecret":""}' | jq -r .accessToken) +# Write a secret via REST API v3 + curl -fsSL -X PATCH https://vault.sysloggh.net/api/v3/secrets/MUMUNI_LITELLM_API_KEY \ + -H "Authorization: Bearer $TOKEN" -H 'Content-Type: application/json' \ + -d '{"environment":"prod","secretValue":"sk-","workspaceId":"","type":"shared"}' +# OR via CLI: infisical secrets set --token=$TOKEN --projectId=322fceab... --env=prod ... +``` +Creation requires the Infisical web UI (https://vault.sysloggh.net) under Project Settings → +Machine Identities, or an admin API call. **TODO: create `abiba-automation` machine identity +and store its credentials in the vault itself (or a root-only file).** + +**Interim (working now):** the `.env` fallback (hermes-config-template Rule 3/13). The +infisical-gateway.sh wrapper sources `~/.hermes/.env`, so its `MUMUNI_LITELLM_API_KEY` overrides +the stale vault value. This is fully functional and contract-sanctioned — the vault sync above is +only for consistency so .env and vault never drift. + +## Key Rotation Log + +| Date | Agent | Action | Notes | +|------|-------|--------|-------| +| 2026-07-16 | mumuni | rotate | Old key malformed (sk-_SWAl_Vu_, 47 chars, not LiteLLM format) → 401. Deleted old `mumuni` key (token 15cbca18…), generated fresh (alias `mumuni`, 7 models: syslog-auto, qwen3.6-27B-code, gemma-4-12b, strix-moe, gpu-dense, gpu-light, qwen3.6-35B-udq4). New key sk-OzuWsoX2… written to /root/.hermes/.env (Rule 3/13 fallback). Vault sync PENDING (needs machine identity). WAL #1300. | + +## LiteLLM Master Key (use sparingly — agents should NOT use it directly) + +- Master key: `sk-litellm-7f96080dd99b15c36bd4b333b58a6796` (in /opt/inference-harness/.env on CT116, Infisical project=infrastructure env=production secret=LITELLM_MASTER_KEY) +- Used for /key/generate, /key/delete, /key/list (GET), DB queries +- **Known violation:** Abiba's LITELLM_API_KEY IS the master key (should be a dedicated `abiba` virtual key). TODO: generate `abiba` virtual key and stop using master key directly. +- LiteLLM key DB: `harness-postgres` container on CT116, table `"LiteLLM_VerificationToken"` (columns: token, key_alias, key_name, created_at, expires). Query: `docker exec harness-postgres psql -U litellm -d litellm -t -c "SELECT key_alias, substr(token,1,16) FROM \"LiteLLM_VerificationToken\" ORDER BY created_at;"`