fix: restore per-host probe coverage + sweep residual retired names #82
@@ -88,7 +88,7 @@ prose run memory-audit-maintenance memory_threshold=90 verify_configs=true
|
|||||||
prose run hermes-config-template agent_name=syslog-devops default_model=claude-sonnet-4
|
prose run hermes-config-template agent_name=syslog-devops default_model=claude-sonnet-4
|
||||||
|
|
||||||
# Configure an agent with a different auxiliary model
|
# Configure an agent with a different auxiliary model
|
||||||
prose run hermes-config-template agent_name=syslog-code default_model=qwen3.6-27B-code auxiliary_model=gpu-vision
|
prose run hermes-config-template agent_name=syslog-code default_model=gpu-dense auxiliary_model=gpu-vision
|
||||||
```
|
```
|
||||||
|
|
||||||
### Option B: Manual Execution
|
### Option B: Manual Execution
|
||||||
|
|||||||
@@ -163,7 +163,7 @@ def audit(path):
|
|||||||
check(
|
check(
|
||||||
comp.get("max_context_window") == 131072,
|
comp.get("max_context_window") == 131072,
|
||||||
"Rule 9",
|
"Rule 9",
|
||||||
f"compression.max_context_window must be 131072 (got {comp.get('max_context_window')!r}) — matches 128K GPU capacity",
|
f"compression.max_context_window must be 131072 (got {comp.get('max_context_window')!r}) — syslog-auto pool floor (NVIDIA hosts 128K; Strix Halo 256K)",
|
||||||
)
|
)
|
||||||
|
|
||||||
# --- Rule 10: Default Model Must Be syslog-auto ---
|
# --- Rule 10: Default Model Must Be syslog-auto ---
|
||||||
|
|||||||
+10
-10
@@ -8,12 +8,12 @@ description: >
|
|||||||
UPDATED 2026-07-15: Stable role-based aliases introduced: strix-moe,
|
UPDATED 2026-07-15: Stable role-based aliases introduced: strix-moe,
|
||||||
gpu-dense, gpu-vision (gpu-light was superseded by gpu-vision on 2026-09-12).
|
gpu-dense, gpu-vision (gpu-light was superseded by gpu-vision on 2026-09-12).
|
||||||
These never change — only the underlying model does.
|
These never change — only the underlying model does.
|
||||||
Strix Halo: strix-moe → unsloth/Qwen3.6-35B-A3B-MTP (UD-Q4_K_M, 22GB).
|
Strix Halo: strix-moe → Carnice-Qwen3.6-MoE-35B-A3B-Q4_K_M.gguf (22GB, 256K ctx).
|
||||||
RTX 5070: gpu-vision — IQ4_NL + MTP draft (~122 tok/s, 2x faster).
|
RTX 5070: gpu-vision — IQ4_NL + MTP draft (~122 tok/s, 2x faster).
|
||||||
UPDATED 2026-07-17: Context reduced fleet-wide from 256K to 128K for stability.
|
UPDATED 2026-07-17: NVIDIA host context reduced from 256K to 128K for stability.
|
||||||
Strix Halo model swapped to qwen3.6-35B-udq4 (22GB, strix-moe alias).
|
Strix Halo model: Carnice-Qwen3.6-MoE-35B-A3B-Q4_K_M.gguf (alias strix-moe, 256K context).
|
||||||
Instability observed near 100K at 256K (now all GPUs at 128K). 128K is the stable ceiling.
|
Strix Halo runs 256K (n_ctx 262144, --kv-unified); RTX 3090 and RTX 5070 remain at 128K.
|
||||||
For larger context needs → fall back to external providers (deepseek).
|
For >128K on NVIDIA hosts → fall back to external providers (deepseek).
|
||||||
VRAM headroom improved: RTX 3090 ~70%, RTX 5070 ~65%.
|
VRAM headroom improved: RTX 3090 ~70%, RTX 5070 ~65%.
|
||||||
agent: abiba
|
agent: abiba
|
||||||
triggers:
|
triggers:
|
||||||
@@ -92,7 +92,7 @@ CT 116 `/opt/inference-harness/litellm_config.yaml`. Do not duplicate those valu
|
|||||||
contracts — read them there.
|
contracts — read them there.
|
||||||
|
|
||||||
**No backward compatibility**: Old model-specific names (qwen3.6-27B-code, qwen3.5-9b-it) are retired as of 2026-09-12
|
**No backward compatibility**: Old model-specific names (qwen3.6-27B-code, qwen3.5-9b-it) are retired as of 2026-09-12
|
||||||
but are deprecated for agent configs. Only the stable aliases survive model swaps. `gemma-4-12b`
|
and no longer resolve; do not use them in agent configs. Only the stable aliases survive model swaps. `gemma-4-12b`
|
||||||
is retired and returns 400 `Invalid model name`.
|
is retired and returns 400 `Invalid model name`.
|
||||||
|
|
||||||
## Routing Configuration (LiteLLM — July 2026)
|
## Routing Configuration (LiteLLM — July 2026)
|
||||||
@@ -190,7 +190,7 @@ If no SSH access, send Zulip DM via abiba-bot with vault update instructions.
|
|||||||
| `/root/scripts/gpu-saturation-watchdog.py` | pi (.24) | Auto-restart stuck llama-server |
|
| `/root/scripts/gpu-saturation-watchdog.py` | pi (.24) | Auto-restart stuck llama-server |
|
||||||
| `/root/dashboard/gpu-fleet.html` | pi (.24) | Live HTML dashboard |
|
| `/root/dashboard/gpu-fleet.html` | pi (.24) | Live HTML dashboard |
|
||||||
| `/etc/systemd/system/llama-server.service` | .8, .110 | llama-server daemons (Nvidia GPUs) |
|
| `/etc/systemd/system/llama-server.service` | .8, .110 | llama-server daemons (Nvidia GPUs) |
|
||||||
| `/etc/systemd/system/strix-server.service` | .15 (amdpve) | llama-server daemon (Vulkan, Strix Halo) running unsloth/Qwen3.6-35B-A3B-MTP-GGUF. Note: `llama-server.service` and `llama-server@.service` are **masked** on .15 to prevent port 8080 collisions. |
|
| `/etc/systemd/system/strix-server.service` | .15 (amdpve) | llama-server daemon (Vulkan, Strix Halo) running Carnice-Qwen3.6-MoE-35B-A3B-Q4_K_M.gguf (256K ctx). Note: `llama-server.service` and `llama-server@.service` are **masked** on .15 to prevent port 8080 collisions. |
|
||||||
|
|
||||||
## Prometheus & Grafana
|
## Prometheus & Grafana
|
||||||
|
|
||||||
@@ -247,8 +247,8 @@ All agent configs MUST use stable role-based aliases, never model-specific names
|
|||||||
When the underlying model is swapped, only the LiteLLM config changes — agent configs are untouched.
|
When the underlying model is swapped, only the LiteLLM config changes — agent configs are untouched.
|
||||||
|
|
||||||
### Context Windows
|
### Context Windows
|
||||||
- RTX 3090: **128K** (reduced from 256K 2026-07-17) | RTX 5070: **128K** (reduced from 256K) | Strix Halo: **128K**
|
- RTX 3090: **128K** (reduced from 256K 2026-07-17) | RTX 5070: **128K** (reduced from 256K) | Strix Halo: **256K** (2026-09-12)
|
||||||
- **All agents**: 128K ceiling — stable margin. For >128K workloads, use external providers (deepseek)
|
- **Agents via `syslog-auto`**: 128K ceiling — the pool's safe floor (NVIDIA hosts are 128K). For >128K workloads, use external providers (deepseek)
|
||||||
- Compression threshold 0.65 (audit Rule 9): fires at ~85K (~43K headroom before 128K ceiling)
|
- Compression threshold 0.65 (audit Rule 9): fires at ~85K (~43K headroom before 128K ceiling)
|
||||||
- **Pi agents (Abiba)**: `compaction.reserveTokens: 52739` (≈60% of 128K)
|
- **Pi agents (Abiba)**: `compaction.reserveTokens: 52739` (≈60% of 128K)
|
||||||
- Mumuni compression model alias: `syslog-auto`
|
- Mumuni compression model alias: `syslog-auto`
|
||||||
@@ -266,7 +266,7 @@ Mumuni (kagentz CT105, 192.168.68.14 — migrated from CT100 2026-08-29) is the
|
|||||||
| `aux.vision.model` | `gpu-vision` | Vision tasks (RTX 5070) |
|
| `aux.vision.model` | `gpu-vision` | Vision tasks (RTX 5070) |
|
||||||
| `aux.web_extract.model` | `gpu-vision` | Web extraction |
|
| `aux.web_extract.model` | `gpu-vision` | Web extraction |
|
||||||
| `delegation.model` | `gpu-dense` | Sub-agent reasoning (RTX 3090) |
|
| `delegation.model` | `gpu-dense` | Sub-agent reasoning (RTX 3090) |
|
||||||
| `context.max_context_window` | 131072 (128K) | Reduced from 256K 2026-07-17 — stable 128K ceiling |
|
| `context.max_context_window` | 131072 (128K) | Conservative `syslog-auto` pool floor (NVIDIA hosts 128K; Strix Halo 256K) |
|
||||||
| `compression.threshold` | 0.65 | Rule 9: triggers at ~85K for a 128K window |
|
| `compression.threshold` | 0.65 | Rule 9: triggers at ~85K for a 128K window |
|
||||||
| `compression.target_ratio` | 0.3 | Compresses to ~38K |
|
| `compression.target_ratio` | 0.3 | Compresses to ~38K |
|
||||||
| `compression.protect_last_n` | 40 | Preserves last 40 messages |
|
| `compression.protect_last_n` | 40 | Preserves last 40 messages |
|
||||||
|
|||||||
@@ -52,7 +52,7 @@ depends_on:
|
|||||||
|
|
||||||
Key notes:
|
Key notes:
|
||||||
- All models use direct GPU routing via LiteLLM (`api_key: not-needed`). Router (port 9000) was decommissioned 2026-09-11 and is NOT in the inference path.
|
- All models use direct GPU routing via LiteLLM (`api_key: not-needed`). Router (port 9000) was decommissioned 2026-09-11 and is NOT in the inference path.
|
||||||
- Stable aliases (gpu-dense, gpu-vision, strix-moe) from gpu-fleet are the canonical names for agent configs. Model-specific names still work but are deprecated. The retired names `gpu-light` and `gemma-4-12b` were superseded by `gpu-vision` on 2026-09-12 and no longer resolve (400 `Invalid model name`).
|
- Stable aliases (gpu-dense, gpu-vision, strix-moe) from gpu-fleet are the canonical names for agent configs — use these, not model-specific names. The retired names `gpu-light` and `gemma-4-12b` were superseded by `gpu-vision` on 2026-09-12 and no longer resolve (400 `Invalid model name`).
|
||||||
- The RTX 5070 is the fastest endpoint per token — `gpu-vision` is its canonical alias. Route vision/web/light work there first. The RTX 5070 model is multimodal (image+text).
|
- The RTX 5070 is the fastest endpoint per token — `gpu-vision` is its canonical alias. Route vision/web/light work there first. The RTX 5070 model is multimodal (image+text).
|
||||||
- Strix Halo is 62.9 tok/s (89% of 70.5 baseline) — below optimal but stable. Check for competing workloads.
|
- Strix Halo is 62.9 tok/s (89% of 70.5 baseline) — below optimal but stable. Check for competing workloads.
|
||||||
- RTX 3090 VRAM at 88% — within role-appropriate range (role = heavy reasoning, needs the headroom).
|
- RTX 3090 VRAM at 88% — within role-appropriate range (role = heavy reasoning, needs the headroom).
|
||||||
@@ -142,7 +142,7 @@ Key notes:
|
|||||||
- **Escalate**: If Tier 2 triggers and temp still rising after 5 min → possible hardware failure
|
- **Escalate**: If Tier 2 triggers and temp still rising after 5 min → possible hardware failure
|
||||||
|
|
||||||
### Rule 9: Context Window Optimization
|
### Rule 9: Context Window Optimization
|
||||||
- **Detect**: Benchmark tok/s vs baseline for each GPU at current context (all 128K)
|
- **Detect**: Benchmark tok/s vs baseline for each GPU at its current context
|
||||||
- RTX 3090 (128K ctx, ThinkingCap): baseline 74.8 tok/s — currently at 74.9 (100%)
|
- RTX 3090 (128K ctx, ThinkingCap): baseline 74.8 tok/s — currently at 74.9 (100%)
|
||||||
- RTX 5070 (128K ctx, HauhauCS QAT): baseline 165.2 tok/s — currently at 169.6 (103%)
|
- RTX 5070 (128K ctx, HauhauCS QAT): baseline 165.2 tok/s — currently at 169.6 (103%)
|
||||||
- Strix Halo (256K ctx, Carnice-Qwen3.6-MoE-35B-A3B-Q4_K_M.gguf): baseline 70.5 tok/s — currently at 62.9 (89%)
|
- Strix Halo (256K ctx, Carnice-Qwen3.6-MoE-35B-A3B-Q4_K_M.gguf): baseline 70.5 tok/s — currently at 62.9 (89%)
|
||||||
|
|||||||
@@ -5,7 +5,7 @@ version: 1.0.0
|
|||||||
description: >
|
description: >
|
||||||
Canonical known-good baseline for all Syslog Hermes agents. Captures the exact
|
Canonical known-good baseline for all Syslog Hermes agents. Captures the exact
|
||||||
configuration state, keys, workarounds, and audit procedure. When an agent's
|
configuration state, keys, workarounds, and audit procedure. When an agent's
|
||||||
configuration goes sideways, restore from this baseline. Last verified 2026-07-16. All GPUs 128K context (reduced from 256K for stability Jul 2026) (RTX 3090 .8, RTX 5070 .110, Strix Halo .15). Parallel 1 fleet-wide (Strix Halo handles compression solo).
|
configuration goes sideways, restore from this baseline. Last verified 2026-07-16. RTX 3090/5070 at 128K (reduced from 256K for stability Jul 2026); Strix Halo at 256K (2026-09-12) (RTX 3090 .8, RTX 5070 .110, Strix Halo .15). Parallel 1 fleet-wide (Strix Halo handles compression solo).
|
||||||
author: Abiba (pi agent)
|
author: Abiba (pi agent)
|
||||||
---
|
---
|
||||||
|
|
||||||
|
|||||||
@@ -97,7 +97,7 @@ model:
|
|||||||
base_url: http://192.168.68.116/litellm/v1 # Rule 5 (2026-08-09): canonical authenticated path; /v1 also OK
|
base_url: http://192.168.68.116/litellm/v1 # Rule 5 (2026-08-09): canonical authenticated path; /v1 also OK
|
||||||
api_key_env: LITELLM_API_KEY # Injected via infisical run -- wrapper
|
api_key_env: LITELLM_API_KEY # Injected via infisical run -- wrapper
|
||||||
max_tokens: 4096 # ⚠️ CRITICAL: Prevents unbounded generation
|
max_tokens: 4096 # ⚠️ CRITICAL: Prevents unbounded generation
|
||||||
context_length: 131072 # For syslog-auto (all GPUs at 128K for stability).
|
context_length: 131072 # Conservative floor for syslog-auto (NVIDIA hosts 128K; Strix Halo 256K).
|
||||||
# ⚠️ MANDATORY: Hermes probes unknown models from 256K
|
# ⚠️ MANDATORY: Hermes probes unknown models from 256K
|
||||||
# and falls back to 256K when /v1/models lacks a context
|
# and falls back to 256K when /v1/models lacks a context
|
||||||
# field (llama-server does). Without this override, agents
|
# field (llama-server does). Without this override, agents
|
||||||
@@ -133,7 +133,7 @@ compression:
|
|||||||
enabled: true
|
enabled: true
|
||||||
model: syslog-auto # ⚠️ Must match auxiliary.compression.model. Stable alias (gpu-fleet § Stable Role-Based Aliases). NOT ornith-1.0-35b (LiteLLM does not serve that name).
|
model: syslog-auto # ⚠️ Must match auxiliary.compression.model. Stable alias (gpu-fleet § Stable Role-Based Aliases). NOT ornith-1.0-35b (LiteLLM does not serve that name).
|
||||||
provider: harness
|
provider: harness
|
||||||
max_context_window: 131072 # MUST match actual GPU capacity. All 3 GPUs are 128K (Jul 17).
|
max_context_window: 131072 # MUST stay at the syslog-auto pool floor: NVIDIA hosts are 128K, Strix Halo 256K (2026-09-12).
|
||||||
threshold: 0.65 # Fires at ~170K for 262K window, ~85K for 128K
|
threshold: 0.65 # Fires at ~170K for 262K window, ~85K for 128K
|
||||||
target_ratio: 0.30
|
target_ratio: 0.30
|
||||||
protect_last_n: 40
|
protect_last_n: 40
|
||||||
@@ -253,7 +253,7 @@ The following MUST be identical across ALL profiles:
|
|||||||
|
|
||||||
### Rule 7: Auxiliary Model Consistency (UPDATED 2026-07-16)
|
### Rule 7: Auxiliary Model Consistency (UPDATED 2026-07-16)
|
||||||
- Vision and web_extract use `gpu-vision` (RTX 5070 — 12GB, vision-optimized)
|
- Vision and web_extract use `gpu-vision` (RTX 5070 — 12GB, vision-optimized)
|
||||||
- Compression uses `syslog-auto` — the Strix Halo weighted pool (64GB, 128K ctx, compression-optimized); do NOT pin `compression.model` to `strix-moe` (audit Rule 7 rejects it)
|
- Compression uses `syslog-auto` — the Strix Halo weighted pool (64GB, 256K ctx, compression-optimized); do NOT pin `compression.model` to `strix-moe` (audit Rule 7 rejects it)
|
||||||
- **`ornith-1.0-35b` is NOT a valid compression model name** — LiteLLM does not serve it
|
- **`ornith-1.0-35b` is NOT a valid compression model name** — LiteLLM does not serve it
|
||||||
(do not restate the served model list here — CT 116 `/opt/inference-harness/litellm_config.yaml`
|
(do not restate the served model list here — CT 116 `/opt/inference-harness/litellm_config.yaml`
|
||||||
is the single source of truth for models, aliases, weights and fallbacks). Old configs with `ornith-1.0-35b` cause 403/model-not-found on compression calls.
|
is the single source of truth for models, aliases, weights and fallbacks). Old configs with `ornith-1.0-35b` cause 403/model-not-found on compression calls.
|
||||||
@@ -267,15 +267,15 @@ The following MUST be identical across ALL profiles:
|
|||||||
- `api_key_env: LITELLM_API_KEY`
|
- `api_key_env: LITELLM_API_KEY`
|
||||||
- **Do NOT use `syslog-auto` for `vision`/`web_extract`** — it routes unpredictably; compression is the deliberate exception (see the OPERATIONAL DECISION above)
|
- **Do NOT use `syslog-auto` for `vision`/`web_extract`** — it routes unpredictably; compression is the deliberate exception (see the OPERATIONAL DECISION above)
|
||||||
- **Compression on Strix Halo**: The strix-moe alias routes to Strix Halo
|
- **Compression on Strix Halo**: The strix-moe alias routes to Strix Halo
|
||||||
(64GB UMA, 128K context) — the designated compression GPU. This frees the
|
(64GB UMA, 256K context) — the designated compression GPU. This frees the
|
||||||
RTX 5070 for vision and web search, and the RTX 3090 for heavy reasoning.
|
RTX 5070 for vision and web search, and the RTX 3090 for heavy reasoning.
|
||||||
- The `compression:` block's `model` MUST match `auxiliary: compression: model`
|
- The `compression:` block's `model` MUST match `auxiliary: compression: model`
|
||||||
- The `compression: max_context_window: 131072` MUST match actual GPU capacity (128K)
|
- The `compression: max_context_window: 131072` MUST stay at the syslog-auto pool floor (NVIDIA hosts 128K; Strix Halo 256K)
|
||||||
|
|
||||||
### Rule 8: GPU Workload Distribution (UPDATED 2026-07-16)
|
### Rule 8: GPU Workload Distribution (UPDATED 2026-07-16)
|
||||||
- **RTX 3090 (24GB, 128K ctx, gpu-dense)**: Heavy reasoning, code gen, long conversations
|
- **RTX 3090 (24GB, 128K ctx, gpu-dense)**: Heavy reasoning, code gen, long conversations
|
||||||
- **RTX 5070 (12GB, 128K ctx, gpu-vision)**: Vision, web search, quick tasks, web_extract (IQ4_NL+MTP, ~65% VRAM at 128K)
|
- **RTX 5070 (12GB, 128K ctx, gpu-vision)**: Vision, web search, quick tasks, web_extract (IQ4_NL+MTP, ~65% VRAM at 128K)
|
||||||
- **Strix Halo (64GB, 128K ctx, syslog-auto)**: Context compression, summarization, long docs
|
- **Strix Halo (64GB, 256K ctx, syslog-auto)**: Context compression, summarization, long docs
|
||||||
- Agent profiles MUST route auxiliary tasks to the correct GPU:
|
- Agent profiles MUST route auxiliary tasks to the correct GPU:
|
||||||
- `auxiliary.vision.model: gpu-vision` (RTX 5070)
|
- `auxiliary.vision.model: gpu-vision` (RTX 5070)
|
||||||
- `auxiliary.web_extract.model: gpu-vision` (RTX 5070)
|
- `auxiliary.web_extract.model: gpu-vision` (RTX 5070)
|
||||||
@@ -284,14 +284,14 @@ The following MUST be identical across ALL profiles:
|
|||||||
- For 128K context window: `threshold: 0.65` (fires at ~85K tokens)
|
- For 128K context window: `threshold: 0.65` (fires at ~85K tokens)
|
||||||
- Do NOT use `threshold: 0.25` — this fires at 65K, causing premature context loss
|
- Do NOT use `threshold: 0.25` — this fires at 65K, causing premature context loss
|
||||||
- Do NOT use `threshold: 0.80` — this delays until ~105K, leaving only 23K margin
|
- Do NOT use `threshold: 0.80` — this delays until ~105K, leaving only 23K margin
|
||||||
- `max_context_window: 131072` MUST match the model's actual capacity (128K)
|
- `max_context_window: 131072` MUST stay at the pool floor (NVIDIA hosts 128K; Strix Halo 256K)
|
||||||
- See `devops-hermes-compression` skill for full reference
|
- See `devops-hermes-compression` skill for full reference
|
||||||
|
|
||||||
### Rule 9: Compression Threshold for 128K Models
|
### Rule 9: Compression Threshold for 128K Models
|
||||||
- For 128K context window: `threshold: 0.65` (fires at ~85K tokens)
|
- For 128K context window: `threshold: 0.65` (fires at ~85K tokens)
|
||||||
- Do NOT use `threshold: 0.25` — this fires at 65K, causing premature context loss
|
- Do NOT use `threshold: 0.25` — this fires at 65K, causing premature context loss
|
||||||
- Do NOT use `threshold: 0.80` — this delays until ~105K, leaving only 23K margin
|
- Do NOT use `threshold: 0.80` — this delays until ~105K, leaving only 23K margin
|
||||||
- `max_context_window: 131072` MUST match the model's actual capacity (all GPUs = 128K)
|
- `max_context_window: 131072` MUST stay at the pool floor (NVIDIA hosts 128K; Strix Halo 256K)
|
||||||
- See `devops-hermes-compression` skill for full reference
|
- See `devops-hermes-compression` skill for full reference
|
||||||
|
|
||||||
### Rule 10: Default Model Must Be `syslog-auto` (All Agents)
|
### Rule 10: Default Model Must Be `syslog-auto` (All Agents)
|
||||||
@@ -322,7 +322,8 @@ When an agent shows "context issues" (premature compression, 401s, 504s, DeepSee
|
|||||||
verify ALL FOUR of these against the live config. They are the only root causes found in production:
|
verify ALL FOUR of these against the live config. They are the only root causes found in production:
|
||||||
|
|
||||||
1. **max_context_window correct?** — BOTH `compression.max_context_window` AND `context.max_context_window`
|
1. **max_context_window correct?** — BOTH `compression.max_context_window` AND `context.max_context_window`
|
||||||
MUST be `131072` (all GPUs are 128K). A value of `262144` causes instability near 100K and must NOT be used.
|
MUST be `131072` (the syslog-auto pool floor: NVIDIA hosts are 128K; Strix Halo is 256K).
|
||||||
|
A `262144` client window can route to a 128K NVIDIA host and fail, so it must NOT be used.
|
||||||
~83K instead of ~170K. Check: `grep -n max_context_window ~/.hermes/config.yaml`
|
~83K instead of ~170K. Check: `grep -n max_context_window ~/.hermes/config.yaml`
|
||||||
2. **base_url uses authenticated path?** — `custom_providers[0].base_url`, `delegation.base_url`,
|
2. **base_url uses authenticated path?** — `custom_providers[0].base_url`, `delegation.base_url`,
|
||||||
and ALL `auxiliary.*.base_url` MUST be `http://192.168.68.116/litellm/v1` (Rule 5, canonical)
|
and ALL `auxiliary.*.base_url` MUST be `http://192.168.68.116/litellm/v1` (Rule 5, canonical)
|
||||||
|
|||||||
@@ -4,7 +4,7 @@ kind: responsibility
|
|||||||
description: >
|
description: >
|
||||||
Optimizes the full Syslog inference stack — LiteLLM routing weights, GPU model
|
Optimizes the full Syslog inference stack — LiteLLM routing weights, GPU model
|
||||||
assignments, agent context management, and prompt caching — to reduce response
|
assignments, agent context management, and prompt caching — to reduce response
|
||||||
times to sub-15s average. All GPUs now at 128K context (stable ceiling).
|
times to sub-15s average. NVIDIA GPUs at 128K context; Strix Halo at 256K (2026-09-12).
|
||||||
id: 067NC6KP02RG60S50M40E30928
|
id: 067NC6KP02RG60S50M40E30928
|
||||||
---
|
---
|
||||||
|
|
||||||
@@ -63,7 +63,7 @@ prefill time at 532 tok/s. Fix context first, routing second.
|
|||||||
- **Cache everything repeated**: system prompts, skill docs, AGENTS.md — these
|
- **Cache everything repeated**: system prompts, skill docs, AGENTS.md — these
|
||||||
never change between turns. Single-digit cache hit rate is unacceptable.
|
never change between turns. Single-digit cache hit rate is unacceptable.
|
||||||
- **Lower context ceiling**: 128K window is the stable ceiling for agent conversations.
|
- **Lower context ceiling**: 128K window is the stable ceiling for agent conversations.
|
||||||
GPUs reduced from 256K to 128K (2026-07-17). 128K window should compact at 85K (0.65 threshold). For larger contexts, route to external providers.
|
GPUs reduced from 256K to 128K (2026-07-17) for the NVIDIA hosts; Strix Halo runs 256K (2026-09-12). 128K window should compact at 85K (0.65 threshold). For larger contexts, route to external providers.
|
||||||
|
|
||||||
### Shape
|
### Shape
|
||||||
|
|
||||||
|
|||||||
@@ -223,7 +223,7 @@ description: >
|
|||||||
**Prometheus targets**:
|
**Prometheus targets**:
|
||||||
- 192.168.68.8:9400 (RTX 3090 — qwen)
|
- 192.168.68.8:9400 (RTX 3090 — qwen)
|
||||||
- 192.168.68.110:9400 (RTX 5070 — gpu-vision)
|
- 192.168.68.110:9400 (RTX 5070 — gpu-vision)
|
||||||
- 192.168.68.15:9400 (Strix Halo — qwen3.6-35B-udq4)
|
- 192.168.68.15:9400 (Strix Halo — strix-moe)
|
||||||
- harness-litellm:4000 (LiteLLM health)
|
- harness-litellm:4000 (LiteLLM health)
|
||||||
|
|
||||||
### Ecosystem C: Netbird (72.61.0.17 — Hostinger srv1079750.hstgr.cloud)
|
### Ecosystem C: Netbird (72.61.0.17 — Hostinger srv1079750.hstgr.cloud)
|
||||||
|
|||||||
@@ -44,7 +44,7 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2)
|
|||||||
**What changed (v3.2.0 → v4.0.0 — 2026-07-08)**:
|
**What changed (v3.2.0 → v4.0.0 — 2026-07-08)**:
|
||||||
- Router REMOVED from request path — LiteLLM proxies directly to GPU
|
- Router REMOVED from request path — LiteLLM proxies directly to GPU
|
||||||
- All GPUs at parallel 2 (was parallel 1)
|
- All GPUs at parallel 2 (was parallel 1)
|
||||||
- NVIDIA context reduced 256K→128K to free VRAM — now the stable ceiling across all GPUs (2026-07-17)
|
- NVIDIA context reduced 256K→128K to free VRAM — the stable NVIDIA ceiling (2026-07-17); Strix Halo runs 256K (2026-09-12)
|
||||||
- Timeouts and fallback chains are config state — read them from CT 116 `/opt/inference-harness/litellm_config.yaml`; they are not duplicated here.
|
- Timeouts and fallback chains are config state — read them from CT 116 `/opt/inference-harness/litellm_config.yaml`; they are not duplicated here.
|
||||||
|
|
||||||
## Parameters
|
## Parameters
|
||||||
@@ -69,7 +69,8 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2)
|
|||||||
- Network access to public_url, auth_host, and gpu_dashboard_url
|
- Network access to public_url, auth_host, and gpu_dashboard_url
|
||||||
- LiteLLM master key for admin endpoints only (`/key/list`, `/key/generate`, `/key/info`)
|
- LiteLLM master key for admin endpoints only (`/key/list`, `/key/generate`, `/key/info`)
|
||||||
- The dedicated `monitor` agent key on CT 116 at `/etc/litellm-monitor.env` (root-only 0600) for
|
- The dedicated `monitor` agent key on CT 116 at `/etc/litellm-monitor.env` (root-only 0600) for
|
||||||
model inference checks — the master key must never be used for inference
|
model inference checks, scoped for every alias step 7 probes (`gpu-dense`, `gpu-vision`,
|
||||||
|
`strix-moe`) — the master key must never be used for inference
|
||||||
|
|
||||||
## GPU Fleet Topology
|
## GPU Fleet Topology
|
||||||
|
|
||||||
@@ -144,12 +145,16 @@ contracts — read them there. Do not re-add retired names (`gemma-4-12b`, `gpu-
|
|||||||
7. **Check model inference via LiteLLM** — Test one model on each GPU host. The health
|
7. **Check model inference via LiteLLM** — Test one model on each GPU host. The health
|
||||||
check runs on the **backend edge**, not the public edge, so these paths carry the
|
check runs on the **backend edge**, not the public edge, so these paths carry the
|
||||||
`/litellm/` prefix:
|
`/litellm/` prefix:
|
||||||
- POST http://{{backend_host}}/litellm/v1/chat/completions model=gpu-vision → expect 200 (RTX 5070, .110)
|
- POST http://{{backend_host}}/litellm/v1/chat/completions model=gpu-dense → expect 200 (RTX 3090, .8)
|
||||||
- POST http://{{backend_host}}/litellm/v1/chat/completions model=gpu-vision → expect 200 (RTX 5070, .110)
|
- POST http://{{backend_host}}/litellm/v1/chat/completions model=gpu-vision → expect 200 (RTX 5070, .110)
|
||||||
- POST http://{{backend_host}}/litellm/v1/chat/completions model=strix-moe → expect 200 (Strix Halo, .15)
|
- POST http://{{backend_host}}/litellm/v1/chat/completions model=strix-moe → expect 200 (Strix Halo, .15)
|
||||||
- Auth uses the dedicated `monitor` agent key, read on CT 116 from
|
- Auth uses the dedicated `monitor` agent key, read on CT 116 from
|
||||||
`/etc/litellm-monitor.env` (root-only 0600). Do NOT use the master key for inference —
|
`/etc/litellm-monitor.env` (root-only 0600). Do NOT use the master key for inference —
|
||||||
the master key is for admin endpoints only (`/key/list`, `/key/generate`, `/key/info`).
|
the master key is for admin endpoints only (`/key/list`, `/key/generate`, `/key/info`).
|
||||||
|
- KEY SCOPE: the `monitor` key MUST be scoped for all three probed aliases (`gpu-dense`,
|
||||||
|
`gpu-vision`, `strix-moe`), otherwise the probe returns 403 and the host is not covered.
|
||||||
|
If a probe returns 403, widen the monitor key's model list on CT 116 (add the missing
|
||||||
|
alias) and re-run — never drop the host from the probe to make the check pass.
|
||||||
- `gemma-4-12b` was retired and returns 400 `Invalid model name` — do not re-add it to
|
- `gemma-4-12b` was retired and returns 400 `Invalid model name` — do not re-add it to
|
||||||
this list. The RTX 5070 host now serves `gpu-vision`.
|
this list. The RTX 5070 host now serves `gpu-vision`.
|
||||||
- `/v1/models` is **key-scoped**: a model is only visible to keys allowed to use it, so
|
- `/v1/models` is **key-scoped**: a model is only visible to keys allowed to use it, so
|
||||||
|
|||||||
@@ -90,8 +90,8 @@ raw data never provided.
|
|||||||
|
|
||||||
| Worker | Model | Toolsets | Role | Use When |
|
| Worker | Model | Toolsets | Role | Use When |
|
||||||
|--------|-------|----------|------|----------|
|
|--------|-------|----------|------|----------|
|
||||||
| `syslog-code` | qwen3.6-27B-code | terminal, file, web, memory, skills | Code patches, automation, scripts | Writing/modifying code, creating scripts, debugging, reading/writing files |
|
| `syslog-code` | gpu-dense | terminal, file, web, memory, skills | Code patches, automation, scripts | Writing/modifying code, creating scripts, debugging, reading/writing files |
|
||||||
| `syslog-devops` | qwen3.6-27B-code | terminal, file, web, memory, skills | Infrastructure, DB, bridge, Proxmox | Server ops, SSH, Docker, Proxmox, DB queries, hardware checks |
|
| `syslog-devops` | gpu-dense | terminal, file, web, memory, skills | Infrastructure, DB, bridge, Proxmox | Server ops, SSH, Docker, Proxmox, DB queries, hardware checks |
|
||||||
| `syslog-email` | strix-moe | terminal, file, web, memory, skills | Email automation, mail operations | Sending/receiving email, inbox management, SMTP operations |
|
| `syslog-email` | strix-moe | terminal, file, web, memory, skills | Email automation, mail operations | Sending/receiving email, inbox management, SMTP operations |
|
||||||
| `syslog-research` | strix-moe | terminal, file, web, memory, skills, **browser** | Analysis, classification, data processing | Web research, browser tasks, data analysis, classification, reading docs |
|
| `syslog-research` | strix-moe | terminal, file, web, memory, skills, **browser** | Analysis, classification, data processing | Web research, browser tasks, data analysis, classification, reading docs |
|
||||||
| `syslog-review` | strix-moe | terminal, file, web, memory, skills | Verification, QA, audit validation | **ALWAYS** verify worker output before delivery — especially for infra changes, code builds, and research findings |
|
| `syslog-review` | strix-moe | terminal, file, web, memory, skills | Verification, QA, audit validation | **ALWAYS** verify worker output before delivery — especially for infra changes, code builds, and research findings |
|
||||||
|
|||||||
@@ -87,7 +87,7 @@ agent: abiba
|
|||||||
| storepve | 192.168.68.6 | PVE |
|
| storepve | 192.168.68.6 | PVE |
|
||||||
| acerpve | 192.168.68.9 | PVE (hosts llm-gpu qemu/101) |
|
| acerpve | 192.168.68.9 | PVE (hosts llm-gpu qemu/101) |
|
||||||
| minipve | 192.168.68.12 | PVE |
|
| minipve | 192.168.68.12 | PVE |
|
||||||
| amdpve | 192.168.68.15 | PVE + Strix Halo LLM (qwen3.6-35B-udq4, strix-moe) |
|
| amdpve | 192.168.68.15 | PVE + Strix Halo LLM (strix-moe) |
|
||||||
|
|
||||||
## Operations
|
## Operations
|
||||||
|
|
||||||
|
|||||||
Reference in New Issue
Block a user