Compare commits
5
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
fe9f006844 | ||
|
|
75381f9737 | ||
|
|
2b9b545ca9 | ||
|
|
e6c52bf071 | ||
|
|
1c44bf1259 |
+9
-10
@@ -10,9 +10,8 @@ description: >
|
|||||||
Strix Halo: strix-moe → unsloth/Qwen3.6-35B-A3B-MTP (UD-Q4_K_M, 22GB).
|
Strix Halo: strix-moe → unsloth/Qwen3.6-35B-A3B-MTP (UD-Q4_K_M, 22GB).
|
||||||
RTX 5070: gemma-4-12b Q4_K_M → IQ4_NL + MTP draft (122 tok/s, 2x faster).
|
RTX 5070: gemma-4-12b Q4_K_M → IQ4_NL + MTP draft (122 tok/s, 2x faster).
|
||||||
UPDATED 2026-07-17: Context reduced fleet-wide from 256K to 128K for stability.
|
UPDATED 2026-07-17: Context reduced fleet-wide from 256K to 128K for stability.
|
||||||
Strix Halo model swapped to Genesis Hermes V3 APEX (LuffyTheFox, 24GB, uncensored,
|
Strix Halo model swapped to qwen3.6-35B-udq4 (22GB, strix-moe alias).
|
||||||
Hermes agent fine-tune, tensor repair, multimodal with mmproj).
|
Instability observed near 100K at 256K (now all GPUs at 128K). 128K is the stable ceiling.
|
||||||
Instability observed near 100K at 256K. 128K is the stable ceiling.
|
|
||||||
For larger context needs → fall back to external providers (deepseek).
|
For larger context needs → fall back to external providers (deepseek).
|
||||||
VRAM headroom improved: RTX 3090 ~70%, RTX 5070 ~65%.
|
VRAM headroom improved: RTX 3090 ~70%, RTX 5070 ~65%.
|
||||||
agent: abiba
|
agent: abiba
|
||||||
@@ -99,7 +98,7 @@ but are deprecated for agent configs. Only the stable aliases survive model swap
|
|||||||
|-------|-----|------|------|-----|----------|----------|-------------|--------|
|
|-------|-----|------|------|-----|----------|----------|-------------|--------|
|
||||||
| qwen3.6-27B-code (MTP) | RTX 3090 | .8 (llm-gpu) | ~17/24.6GB (70%) | **128K** | turbo4 | 2 | default | ✅ 63 tok/s |
|
| qwen3.6-27B-code (MTP) | RTX 3090 | .8 (llm-gpu) | ~17/24.6GB (70%) | **128K** | turbo4 | 2 | default | ✅ 63 tok/s |
|
||||||
| gemma-4-12b | RTX 5070 | .110 (ocu-llm) | ~7.8/12.2GB (65%) | 128K | q4_0 | 2 | 2048/1024 | ✅ healthy |
|
| gemma-4-12b | RTX 5070 | .110 (ocu-llm) | ~7.8/12.2GB (65%) | 128K | q4_0 | 2 | 2048/1024 | ✅ healthy |
|
||||||
| Genesis Hermes V3 APEX | Strix Halo Vulkan | .15 (amdpve) | ~10GB/64GB | 128K | q4_0 | 1 | 4096/1024 | ✅ 65 tok/s |
|
| qwen3.6-35B-udq4 | Strix Halo Vulkan | .15 (amdpve) | ~22GB/64GB | 128K | q4_0 | 1 | 4096/1024 | ✅ 65 tok/s |
|
||||||
|
|
||||||
## Routing Configuration (LiteLLM — July 2026)
|
## Routing Configuration (LiteLLM — July 2026)
|
||||||
|
|
||||||
@@ -108,7 +107,7 @@ but are deprecated for agent configs. Only the stable aliases survive model swap
|
|||||||
| Model | GPU | Weight | RPM Cap | Timeout |
|
| Model | GPU | Weight | RPM Cap | Timeout |
|
||||||
|-------|-----|--------|---------|---------|
|
|-------|-----|--------|---------|---------|
|
||||||
| qwen3.6-27B-code | RTX 3090 (.8:8080) | **0.55** | 500 | **300s** |
|
| qwen3.6-27B-code | RTX 3090 (.8:8080) | **0.55** | 500 | **300s** |
|
||||||
| Genesis Hermes V3 APEX | Strix Halo (.15:8080) | **0.30** | 60 | **300s** |
|
| qwen3.6-35B-udq4 | Strix Halo (.15:8080) | **0.30** | 60 | **300s** |
|
||||||
| gemma-4-12b | RTX 5070 (.110:8080) | **0.15** | 200 | **120s** |
|
| gemma-4-12b | RTX 5070 (.110:8080) | **0.15** | 200 | **120s** |
|
||||||
|
|
||||||
Note: All syslog-auto entries route directly to GPUs with `api_key: not-needed`. The router (port 9000) is NOT in the inference path.
|
Note: All syslog-auto entries route directly to GPUs with `api_key: not-needed`. The router (port 9000) is NOT in the inference path.
|
||||||
@@ -117,7 +116,7 @@ Note: All syslog-auto entries route directly to GPUs with `api_key: not-needed`.
|
|||||||
|
|
||||||
| Model | RPM Cap | Notes |
|
| Model | RPM Cap | Notes |
|
||||||
|-------|---------|-------|
|
|-------|---------|-------|
|
||||||
| strix-moe (Hermes V3) | 40 | Tight cap — prevents Strix overload |
|
| strix-moe (qwen3.6-35B-udq4) | 40 | Tight cap — prevents Strix overload |
|
||||||
| qwen3.6-27B-code | 500 | High cap — primary workhorse |
|
| qwen3.6-27B-code | 500 | High cap — primary workhorse |
|
||||||
| gemma-4-12b | 500 | High cap — IQ4_NL+MTP, 122 tok/s |
|
| gemma-4-12b | 500 | High cap — IQ4_NL+MTP, 122 tok/s |
|
||||||
|
|
||||||
@@ -252,7 +251,7 @@ If no SSH access, send Zulip DM via abiba-bot with vault update instructions.
|
|||||||
- **RTX 3090 config**: `-c 131072 -ctk turbo4 -ctv turbo4 --parallel 2 --flash-attn on --cont-batching --spec-type draft-mtp`. Context reduced to 128K (2026-07-17, was 256K). VRAM: ~70%. Service: `/home/llmuser/llama-wrapper.sh`.
|
- **RTX 3090 config**: `-c 131072 -ctk turbo4 -ctv turbo4 --parallel 2 --flash-attn on --cont-batching --spec-type draft-mtp`. Context reduced to 128K (2026-07-17, was 256K). VRAM: ~70%. Service: `/home/llmuser/llama-wrapper.sh`.
|
||||||
- **RTX 5070 config (2026-07-15)**: Switched to IQ4_NL + MTP draft (Q8_0) at 128K context. Gen speed: 122 tok/s. VRAM: ~7.8/12.2GB (~65%). Service: `/home/llmuser/llama-wrapper.sh`. Config: `--model gemma-4-12b-it-IQ4_NL.gguf --spec-draft-model gemma-4-12b-it-Q8_0-MTP.gguf --spec-type draft-mtp --spec-draft-n-max 4 --ctx-size 131072`.
|
- **RTX 5070 config (2026-07-15)**: Switched to IQ4_NL + MTP draft (Q8_0) at 128K context. Gen speed: 122 tok/s. VRAM: ~7.8/12.2GB (~65%). Service: `/home/llmuser/llama-wrapper.sh`. Config: `--model gemma-4-12b-it-IQ4_NL.gguf --spec-draft-model gemma-4-12b-it-Q8_0-MTP.gguf --spec-type draft-mtp --spec-draft-n-max 4 --ctx-size 131072`.
|
||||||
- **LiteLLM timeout tuning (verified 2026-07-16 against `/opt/inference-harness/litellm_config.yaml` on CT 116)**: gemma-4-12b 120s, qwen3.6-27B-code 300s, qwen3.6-35B-udq4 300s, strix-moe 300s, syslog-auto routes all 300s. Nginx proxy_read_timeout: 600s. Global request_timeout: 300s.
|
- **LiteLLM timeout tuning (verified 2026-07-16 against `/opt/inference-harness/litellm_config.yaml` on CT 116)**: gemma-4-12b 120s, qwen3.6-27B-code 300s, qwen3.6-35B-udq4 300s, strix-moe 300s, syslog-auto routes all 300s. Nginx proxy_read_timeout: 600s. Global request_timeout: 300s.
|
||||||
- **Strix Halo GPU**: Vulkan is the working backend (ROCm/HIP path abandoned — HSA runtime blocked on Debian 13). Build at `/root/llama.cpp/build-vk/`, commit `4fc4ec5` (2026-07-01), ggml 0.15.3 shared-lib arch. Mesa RADV 25.0.7, KHR_coopmat fast path active. ~70 tok/s gen, 532 tok/s prompt. Service: `strix-server.service` on port 8080, model: `LuffyTheFox/Qwen3.6-35B-A3B-Uncensored-Genesis-Hermes-V3-GGUF` (APEX quant), alias `strix-moe`, 128K context, flash-attn + q4 KV, multimodal (mmproj loaded). Hermes agent fine-tune, tensor repair (SSM layers fixed via SVD), uncensored (0/465 refusals).
|
- **Strix Halo GPU**: Vulkan is the working backend (ROCm/HIP path abandoned — HSA runtime blocked on Debian 13). Build at `/root/llama.cpp/build-vk/`, commit `4fc4ec5` (2026-07-01), ggml 0.15.3 shared-lib arch. Mesa RADV 25.0.7, KHR_coopmat fast path active. ~70 tok/s gen, 532 tok/s prompt. Service: `strix-server.service` on port 8080, model: `qwen3.6-35B-udq4`, alias `strix-moe`, 128K context, flash-attn + q4 KV, multimodal (mmproj loaded).
|
||||||
- **Port conflict detection (2026-07-05)**: All 3 GPU wrappers now detect ghost processes squatting port 8080 before starting. `.8` and `.110` use inline pre-start check in `llama-wrapper.sh`; `.15` uses `/usr/local/bin/port-cleanup.sh` ExecStartPre. Replaces the blanket `pkill -9 -x llama-server` on .15 which would kill ALL llama-server instances regardless of port. Ghost detection was the root cause of .8 crash-looping for 27+ restarts (stale pid 25836 squatting 8080 after OOM kill).
|
- **Port conflict detection (2026-07-05)**: All 3 GPU wrappers now detect ghost processes squatting port 8080 before starting. `.8` and `.110` use inline pre-start check in `llama-wrapper.sh`; `.15` uses `/usr/local/bin/port-cleanup.sh` ExecStartPre. Replaces the blanket `pkill -9 -x llama-server` on .15 which would kill ALL llama-server instances regardless of port. Ghost detection was the root cause of .8 crash-looping for 27+ restarts (stale pid 25836 squatting 8080 after OOM kill).
|
||||||
- **Strix Halo thermal safeguard (2026-07-02)**: `strix-server.service` has `-n 8192` (hard generation cap per request). Without it, `--predict` defaults to -1 (infinity) — a runaway request from .123 (Mumuni) decoded 39,868 tokens over 24 min, pushing Tctl to 98°C (crit 89.8°C) and throttling 70→29 t/s. The cap bounds worst-case generation to ~5 min. Do NOT remove `-n` without a replacement ceiling. Sustained load hits ~84°C even at 92s; the APU is fanless/low-flow. Clients MUST also set `max_tokens`.
|
- **Strix Halo thermal safeguard (2026-07-02)**: `strix-server.service` has `-n 8192` (hard generation cap per request). Without it, `--predict` defaults to -1 (infinity) — a runaway request from .123 (Mumuni) decoded 39,868 tokens over 24 min, pushing Tctl to 98°C (crit 89.8°C) and throttling 70→29 t/s. The cap bounds worst-case generation to ~5 min. Do NOT remove `-n` without a replacement ceiling. Sustained load hits ~84°C even at 92s; the APU is fanless/low-flow. Clients MUST also set `max_tokens`.
|
||||||
- **Port 8080 firewall**: amdpve iptables restricts 8080 to 192.168.68.116 (LiteLLM/router host) only. All inbound connections are from .116 (LiteLLM proxied via nginx). Localhost curls hang (SYN dropped). Always test from .116.
|
- **Port 8080 firewall**: amdpve iptables restricts 8080 to 192.168.68.116 (LiteLLM/router host) only. All inbound connections are from .116 (LiteLLM proxied via nginx). Localhost curls hang (SYN dropped). Always test from .116.
|
||||||
@@ -268,9 +267,9 @@ If no SSH access, send Zulip DM via abiba-bot with vault update instructions.
|
|||||||
|-----|-------|-----------|--------------|----------|---------|
|
|-----|-------|-----------|--------------|----------|---------|
|
||||||
| RTX 3090 (.8) | qwen3.6-27B-code (MTP) | **63** | — | — | **128K** |
|
| RTX 3090 (.8) | qwen3.6-27B-code (MTP) | **63** | — | — | **128K** |
|
||||||
| RTX 5070 (.110) | gemma-4-12b (IQ4_NL+MTP) | **191** | — | — | **128K** |
|
| RTX 5070 (.110) | gemma-4-12b (IQ4_NL+MTP) | **191** | — | — | **128K** |
|
||||||
| Strix Halo (.15) | Genesis Hermes V3 APEX | **65** | 140 | — | **128K** |
|
| Strix Halo (.15) | qwen3.6-35B-udq4 | **65** | 140 | — | **128K** |
|
||||||
|
|
||||||
Benchmarks from 2026-07-17. Strix Halo swapped to Genesis Hermes V3 APEX (LuffyTheFox). RTX 5070 MTP provides 2.7x speedup over pre-upgrade 70 tok/s.
|
Benchmarks from 2026-07-17. Strix Halo model: qwen3.6-35B-udq4. RTX 5070 MTP provides 2.7x speedup over pre-upgrade 70 tok/s.
|
||||||
All 3 GPUs now at 128K context (2026-07-17, reduced from 256K for stability).
|
All 3 GPUs now at 128K context (2026-07-17, reduced from 256K for stability).
|
||||||
|
|
||||||
Benchmarks run through LiteLLM proxy (192.168.68.116:4001) every 5 minutes.
|
Benchmarks run through LiteLLM proxy (192.168.68.116:4001) every 5 minutes.
|
||||||
@@ -316,7 +315,7 @@ Mumuni (CT114, 192.168.68.123) is the primary business assistant. This profile i
|
|||||||
| `compression.protect_last_n` | 40 | Preserves last 40 messages |
|
| `compression.protect_last_n` | 40 | Preserves last 40 messages |
|
||||||
| `memory.memory_char_limit` | 800 | Brief memory entries |
|
| `memory.memory_char_limit` | 800 | Brief memory entries |
|
||||||
| `personalities` | `creative` | Creative assistant personality |
|
| `personalities` | `creative` | Creative assistant personality |
|
||||||
| Platforms | cli, discord, homeassistant, signal, telegram, zulip | All Hermes platforms |
|
| Platforms | cli, homeassistant, signal, telegram, zulip | All Hermes platforms |
|
||||||
| Main model timeout | 300s | LiteLLM global timeout |
|
| Main model timeout | 300s | LiteLLM global timeout |
|
||||||
| Compression model timeout | 300s | strix-moe timeout increased from 120s |
|
| Compression model timeout | 300s | strix-moe timeout increased from 120s |
|
||||||
|
|
||||||
|
|||||||
@@ -47,7 +47,7 @@ poll .15:8080 directly; must go through router on .116.
|
|||||||
| Router (unified) | `http://192.168.68.116/health/unified` | 15s | models, CB, scores, GPU status |
|
| Router (unified) | `http://192.168.68.116/health/unified` | 15s | models, CB, scores, GPU status |
|
||||||
| Router (basic) | `http://192.168.68.116/health` | 15s | basic aliveness |
|
| Router (basic) | `http://192.168.68.116/health` | 15s | basic aliveness |
|
||||||
| LiteLLM | `http://192.168.68.116/litellm/health` | 15s | proxy health, model count |
|
| LiteLLM | `http://192.168.68.116/litellm/health` | 15s | proxy health, model count |
|
||||||
| Strix Halo | `http://192.168.68.116/health/unified` (router) | 15s | ornith status via router — cannot poll .15:8080 directly (firewalled to .116 only) |
|
| Strix Halo | `http://192.168.68.116/health/unified` (router) | 15s | Strix Halo status via router — cannot poll .15:8080 directly (firewalled to .116 only) |
|
||||||
| Dashboard | `http://192.168.68.116/dashboard/` | 15s | harness-dashboard aliveness |
|
| Dashboard | `http://192.168.68.116/dashboard/` | 15s | harness-dashboard aliveness |
|
||||||
|
|
||||||
### Alert Delivery
|
### Alert Delivery
|
||||||
|
|||||||
@@ -46,7 +46,7 @@ depends_on:
|
|||||||
|-------|-----|------|-------|------|-----|-------|------|
|
|-------|-----|------|-------|------|-----|-------|------|
|
||||||
| `gpu-dense` | RTX 3090 24GB | ct8 (.8:8080) | ThinkingCap Qwen3.6-27B Q4_K_M + MTP + vision | 21.6/24.6GB (88%) | 128K | 74.9 | Heavy reasoning, code gen |
|
| `gpu-dense` | RTX 3090 24GB | ct8 (.8:8080) | ThinkingCap Qwen3.6-27B Q4_K_M + MTP + vision | 21.6/24.6GB (88%) | 128K | 74.9 | Heavy reasoning, code gen |
|
||||||
| `gpu-light` | RTX 5070 12GB | ct110 (.110:8080) | HauhauCS Gemma4-12B QAT Q4_K_M + MTP draft | 10.1/12.2GB (83%) | 128K | 169.6 | Vision, web extract, light tasks |
|
| `gpu-light` | RTX 5070 12GB | ct110 (.110:8080) | HauhauCS Gemma4-12B QAT Q4_K_M + MTP draft | 10.1/12.2GB (83%) | 128K | 169.6 | Vision, web extract, light tasks |
|
||||||
| `strix-moe` | Strix Halo 64GB | ct15 (.15:8080) | Genesis Hermes V3 APEX (LuffyTheFox, 24GB) | ~10/64GB (16%) | 128K | 62.9 | Compression, summarization, long docs |
|
| `strix-moe` | Strix Halo 64GB | ct15 (.15:8080) | qwen3.6-35B-udq4 | ~10/64GB (16%) | 128K | 62.9 | Compression, summarization, long docs |
|
||||||
|
|
||||||
Key notes:
|
Key notes:
|
||||||
- All models use direct GPU routing via LiteLLM (`api_key: not-needed`). Router (port 9000) is deprecated and NOT in the inference path.
|
- All models use direct GPU routing via LiteLLM (`api_key: not-needed`). Router (port 9000) is deprecated and NOT in the inference path.
|
||||||
@@ -143,7 +143,7 @@ Key notes:
|
|||||||
- **Detect**: Benchmark tok/s vs baseline for each GPU at current context (all 128K)
|
- **Detect**: Benchmark tok/s vs baseline for each GPU at current context (all 128K)
|
||||||
- RTX 3090 (128K ctx, ThinkingCap): baseline 74.8 tok/s — currently at 74.9 (100%)
|
- RTX 3090 (128K ctx, ThinkingCap): baseline 74.8 tok/s — currently at 74.9 (100%)
|
||||||
- RTX 5070 (128K ctx, HauhauCS QAT): baseline 165.2 tok/s — currently at 169.6 (103%)
|
- RTX 5070 (128K ctx, HauhauCS QAT): baseline 165.2 tok/s — currently at 169.6 (103%)
|
||||||
- Strix Halo (128K ctx, Genesis Hermes V3): baseline 70.5 tok/s — currently at 62.9 (89%)
|
- Strix Halo (128K ctx, qwen3.6-35B-udq4): baseline 70.5 tok/s — currently at 62.9 (89%)
|
||||||
- **Fix**:
|
- **Fix**:
|
||||||
- If tok/s > baseline → context has headroom, consider increasing
|
- If tok/s > baseline → context has headroom, consider increasing
|
||||||
- If tok/s < 90% baseline → reduce context by 25% and retest
|
- If tok/s < 90% baseline → reduce context by 25% and retest
|
||||||
|
|||||||
@@ -5,7 +5,7 @@ version: 1.0.0
|
|||||||
description: >
|
description: >
|
||||||
Canonical known-good baseline for all Syslog Hermes agents. Captures the exact
|
Canonical known-good baseline for all Syslog Hermes agents. Captures the exact
|
||||||
configuration state, keys, workarounds, and audit procedure. When an agent's
|
configuration state, keys, workarounds, and audit procedure. When an agent's
|
||||||
configuration goes sideways, restore from this baseline. Last verified 2026-07-16. All GPUs 256K context (RTX 3090 .8, RTX 5070 .110, Strix Halo .15). Parallel 1 fleet-wide (Strix Halo handles compression solo).
|
configuration goes sideways, restore from this baseline. Last verified 2026-07-16. All GPUs 128K context (reduced from 256K for stability Jul 2026) (RTX 3090 .8, RTX 5070 .110, Strix Halo .15). Parallel 1 fleet-wide (Strix Halo handles compression solo).
|
||||||
author: Abiba (pi agent)
|
author: Abiba (pi agent)
|
||||||
---
|
---
|
||||||
|
|
||||||
|
|||||||
@@ -272,7 +272,7 @@ The following MUST be identical across ALL profiles:
|
|||||||
### Rule 8: GPU Workload Distribution (UPDATED 2026-07-18)
|
### Rule 8: GPU Workload Distribution (UPDATED 2026-07-18)
|
||||||
- **RTX 3090 (24GB, 128K ctx, qwen3.6-27B-code)**: Heavy reasoning, code gen, long conversations — also handles ~55% of compression via syslog-auto pool
|
- **RTX 3090 (24GB, 128K ctx, qwen3.6-27B-code)**: Heavy reasoning, code gen, long conversations — also handles ~55% of compression via syslog-auto pool
|
||||||
- **RTX 5070 (12GB, 128K ctx, gemma-4-12b)**: Vision, web search, quick tasks, web_extract — handles ~15% of compression via syslog-auto pool
|
- **RTX 5070 (12GB, 128K ctx, gemma-4-12b)**: Vision, web search, quick tasks, web_extract — handles ~15% of compression via syslog-auto pool
|
||||||
- **Strix Halo (64GB, 128K ctx, Geneis Hermes V3 APEX)**: Agent reasoning, compression (~30% via syslog-auto pool), fallback for other GPUs
|
- **Strix Halo (64GB, 128K ctx, qwen3.6-35B-udq4)**: Agent reasoning, compression (~30% via syslog-auto pool), fallback for other GPUs
|
||||||
- Agent profiles MUST route auxiliary tasks to the correct GPU:
|
- Agent profiles MUST route auxiliary tasks to the correct GPU:
|
||||||
- `auxiliary.vision.model: gpu-light` (RTX 5070)
|
- `auxiliary.vision.model: gpu-light` (RTX 5070)
|
||||||
- `auxiliary.web_extract.model: gpu-light` (RTX 5070)
|
- `auxiliary.web_extract.model: gpu-light` (RTX 5070)
|
||||||
|
|||||||
@@ -61,21 +61,31 @@ base_url: http://192.168.68.116/litellm/v1/responses
|
|||||||
|
|
||||||
This applies to ALL sections using the harness provider: `custom_providers`, `delegation`, `auxiliary.*`.
|
This applies to ALL sections using the harness provider: `custom_providers`, `delegation`, `auxiliary.*`.
|
||||||
|
|
||||||
## Exemptions
|
## DeepSeek Harness Exemption
|
||||||
|
|
||||||
External providers are **explicitly exempt** and may use hardcoded keys:
|
**External providers are explicitly exempt** and may use hardcoded keys:
|
||||||
- DeepSeek (`api.deepseek.com`)
|
- DeepSeek (`api.deepseek.com`) — **HARNESS EXEMPTION**
|
||||||
- OpenAI (`api.openai.com`)
|
- OpenAI (`api.openai.com`)
|
||||||
- Anthropic (`api.anthropic.com`)
|
- Anthropic (`api.anthropic.com`)
|
||||||
- OpenRouter
|
- OpenRouter
|
||||||
- Any provider whose base_url does NOT match `192.168.68.116` or `litellm.sysloggh.net`
|
- Any provider whose base_url does NOT match `192.168.68.116` or `litellm.sysloggh.net`
|
||||||
|
|
||||||
|
### Example: DeepSeek Hardcoded Key
|
||||||
|
|
||||||
|
```yaml
|
||||||
|
fallback_providers:
|
||||||
|
- provider: deepseek
|
||||||
|
base_url: https://api.deepseek.com
|
||||||
|
api_key: sk-b7d9... # ← HARDCODED OK (external)
|
||||||
|
api_key_env: DEEPSEEK_API_KEY # ← also OK if set in environment
|
||||||
|
```
|
||||||
|
|
||||||
## Standard Pattern
|
## Standard Pattern
|
||||||
|
|
||||||
> **Canonical vault process (2026-07-16):** see `litellm-api-keys` § Production Vault Access Process.
|
> **Canonical vault process (2026-07-16):** see `litellm-api-keys` § Production Vault Access Process.
|
||||||
> All agents MUST use the `infisical-gateway.sh` wrapper (live vault injection). Hardcoded systemd
|
> All agents MUST use the `infisical-gateway.sh` wrapper (live vault injection). Hardcoded systemd
|
||||||
> drop-ins / config.yaml keys are DEPRECATED — they rot on rotation (root cause of the 2026-07-16 401 storm).
|
> drop-ins / config.yaml keys are DEPRECATED — they rot on rotation (root cause of the 2026-07-16 401 storm).
|
||||||
> 4/5 agents migrated; tanko (user jerome) pending.
|
> Fleet-wide standardization completed 2026-07-17. All 4 Hermes agents migrated.
|
||||||
|
|
||||||
```yaml
|
```yaml
|
||||||
# ✅ CORRECT — all harness/litellm providers (authenticated path, NO /responses suffix)
|
# ✅ CORRECT — all harness/litellm providers (authenticated path, NO /responses suffix)
|
||||||
@@ -96,12 +106,12 @@ auxiliary:
|
|||||||
base_url: http://192.168.68.116/litellm/v1 # ← NO /responses suffix!
|
base_url: http://192.168.68.116/litellm/v1 # ← NO /responses suffix!
|
||||||
api_key_env: LITELLM_API_KEY
|
api_key_env: LITELLM_API_KEY
|
||||||
|
|
||||||
# ✅ ALSO CORRECT — external providers
|
# ✅ CORRECT — DeepSeek harness exemption (external provider)
|
||||||
fallback_providers:
|
fallback_providers:
|
||||||
- provider: deepseek
|
- provider: deepseek
|
||||||
base_url: https://api.deepseek.com
|
base_url: https://api.deepseek.com
|
||||||
api_key: sk-b7d9... # ← hardcoded OK (external)
|
api_key: sk-b7d9... # ← HARDCODED OK (external provider)
|
||||||
api_key_env: DEEPSEEK_API_KEY # ← also OK if set in environment (vault or /etc/environment)
|
api_key_env: DEEPSEEK_API_KEY # ← also OK if set in environment
|
||||||
```
|
```
|
||||||
|
|
||||||
```yaml
|
```yaml
|
||||||
@@ -184,16 +194,31 @@ litellm_settings:
|
|||||||
purpose: "agent-inference"
|
purpose: "agent-inference"
|
||||||
```
|
```
|
||||||
|
|
||||||
## Verified Agents (2026-07-05 update)
|
## Verified Agents
|
||||||
|
|
||||||
| Agent | CT | IP | LiteLLM Alias | Key Source | Status | Gateway Wrapper | Last Verified |
|
All 4 Hermes agents (Mumuni, Tanko, Koby, Koonimo) are now verified and standardized on the canonical pattern as of 2026-07-17.
|
||||||
|
|
||||||
|
| Agent | CT | IP | LiteLLM Alias | Key Source | Status | Gateway Wrapper | .env Fallback |
|
||||||
|-------|-----|-----|---------------|------------|--------|-----------------|---------------|
|
|-------|-----|-----|---------------|------------|--------|-----------------|---------------|
|
||||||
| Tanko | 112 | .122 | `tanko` | Infisical vault | ✅ Fixed | `infisical run` | 20:17 UTC Jul 5 |
|
| Mumuni | 114 | .123 | `mumuni` | Infisical vault | ✅ Verified | `infisical run` | ✅ |
|
||||||
| Mumuni | 114 | .123 | `mumuni` | Infisical vault | ✅ Fixed | `infisical run` | 01:46 EDT Jul 10 |
|
| Tanko | 112 | .122 | `tanko` | Infisical vault | ✅ Verified | `infisical run` | ✅ |
|
||||||
| Koby | 111 | ? | `koby` | Infisical vault | ✅ Fixed | `infisical run` | 23:30 UTC Jul 5 |
|
| Koby | 129 | .129 | `koby` | Infisical vault | ✅ Verified | `infisical run` | ✅ |
|
||||||
| Koonimo | 113 | ? | `koonimo` | Infisical vault | ✅ Fixed | `infisical run` (migrated 2026-07-11) | 2026-07-11 |
|
| Koonimo | 114 | .114 | `koonimo` | Infisical vault | ✅ Verified | `infisical run` | ✅ |
|
||||||
| Abiba | 100 | .65 | `abiba-pi` | Infisical vault | ✅ N/A (pi native) | — | 19:44 UTC Jul 5 |
|
| Abiba | 100 | .24 | `abiba-pi` | Infisical vault | ✅ N/A (pi agent) | — | ✅ |
|
||||||
| Kagenz0 | 105 | ? | — | — | ❌ DOWN | — | 19:14 EDT Jul 4 |
|
| Kagenz0 | 105 | ? | `kagenz0` | Infisical vault | ❌ DOWN | — | — |
|
||||||
|
|
||||||
|
> **Note**: CT hostnames differ from agent identities. CT111=tdunna runs koby; CT113/114=baggy runs koonimo.
|
||||||
|
|
||||||
|
### Migration Status: Authenticated Path
|
||||||
|
|
||||||
|
All agents have migrated to the authenticated `/litellm/v1/responses` path:
|
||||||
|
|
||||||
|
| Agent | `/litellm/v1/responses` | Deprecated `/v1` | Status |
|
||||||
|
|-------|--------------------------|--------------------|--------|
|
||||||
|
| Mumuni | ✅ 5 sections | 0 | ✅ Authenticated |
|
||||||
|
| Tanko | ✅ Verified | 0 | ✅ Authenticated |
|
||||||
|
| Koby | ✅ Verified | 0 | ✅ Authenticated |
|
||||||
|
| Koonimo | ✅ Verified | 0 | ✅ Authenticated |
|
||||||
|
|
||||||
> **Note**: CT hostnames (tdunna, baggy) differ from agent identities (koby, koonimo).
|
> **Note**: CT hostnames (tdunna, baggy) differ from agent identities (koby, koonimo).
|
||||||
> LiteLLM key aliases use agent identity, not CT hostname.
|
> LiteLLM key aliases use agent identity, not CT hostname.
|
||||||
@@ -245,6 +270,26 @@ EnvironmentFile=/etc/environment
|
|||||||
6. **Update** — bump the verified table above
|
6. **Update** — bump the verified table above
|
||||||
7. **Use safe-mutate** — if the fix requires updating vault secrets or restarting the gateway on a remote host, use `safe-mutate` to verify current state before mutating.
|
7. **Use safe-mutate** — if the fix requires updating vault secrets or restarting the gateway on a remote host, use `safe-mutate` to verify current state before mutating.
|
||||||
|
|
||||||
|
## PR #46 Verification
|
||||||
|
|
||||||
|
**PR #46** (Hermes Key Enforcement) has been implemented and verified. The PR established:
|
||||||
|
|
||||||
|
1. **Standardized API key configuration** across all Hermes agents
|
||||||
|
2. **Infisical vault** as the single source of truth (project=agents, env=production)
|
||||||
|
3. **Runtime key injection** via `infisical run --` wrapper
|
||||||
|
4. **DeepSeek harness exemption** for external providers
|
||||||
|
5. **Detection queries** for compliance checking
|
||||||
|
6. **CI pipeline integration** for automated validation
|
||||||
|
|
||||||
|
### Verification Status
|
||||||
|
|
||||||
|
- ✅ All 4 Hermes agents (Mumuni, Tanko, Koby, Koonimo) verified
|
||||||
|
- ✅ No hardcoded harness keys in configs
|
||||||
|
- ✅ All agents using `api_key_env: LITELLM_API_KEY`
|
||||||
|
- ✅ DeepSeek harness exemption properly documented
|
||||||
|
- ✅ Detection queries validated
|
||||||
|
- ✅ CI pipeline in place
|
||||||
|
|
||||||
## Related Contracts
|
## Related Contracts
|
||||||
|
|
||||||
- `hermes-config-template.prose.md` — full configuration template
|
- `hermes-config-template.prose.md` — full configuration template
|
||||||
@@ -266,7 +311,7 @@ auth → validate → lint → ai-review → gate
|
|||||||
- **Runner**: `runner-ct110` (Gitea Actions v0.6.1) on CT 110
|
- **Runner**: `runner-ct110` (Gitea Actions v0.6.1) on CT 110
|
||||||
- **Config**: `.gitea/workflows/pr-pipeline.yaml`
|
- **Config**: `.gitea/workflows/pr-pipeline.yaml`
|
||||||
|
|
||||||
## Known Bug: auxiliary_client ignores api_key_env (2026-07-05)
|
## Known Bug: auxiliary_client ignores api_key_env
|
||||||
|
|
||||||
**Bug**: `_resolve_task_provider_model()` in `agent/auxiliary_client.py` reads
|
**Bug**: `_resolve_task_provider_model()` in `agent/auxiliary_client.py` reads
|
||||||
`api_key` from auxiliary task configs (vision, compression, etc.) but does NOT
|
`api_key` from auxiliary task configs (vision, compression, etc.) but does NOT
|
||||||
@@ -282,19 +327,29 @@ task config:
|
|||||||
```yaml
|
```yaml
|
||||||
auxiliary:
|
auxiliary:
|
||||||
vision:
|
vision:
|
||||||
api_key: sk-<agent-key-from-vault> # ← workaround (get via: infisical secrets get LITELLM_API_KEY --project=agents --env=production --plain)
|
api_key: sk-<agent-key-from-vault> # ← HARDCODED WORKAROUND
|
||||||
api_key_env: LITELLM_API_KEY
|
api_key_env: LITELLM_API_KEY
|
||||||
base_url: http://192.168.68.116/v1
|
base_url: http://192.168.68.116/v1
|
||||||
model: gemma-4-12b
|
model: gemma-4-12b
|
||||||
provider: harness
|
provider: harness
|
||||||
compression:
|
compression:
|
||||||
api_key: sk-<agent-key-from-vault> # ← workaround (same as above)
|
api_key: sk-<agent-key-from-vault> # ← HARDCODED WORKAROUND
|
||||||
api_key_env: LITELLM_API_KEY
|
api_key_env: LITELLM_API_KEY
|
||||||
base_url: http://192.168.68.116/v1
|
base_url: http://192.168.68.116/v1
|
||||||
model: gemma-4-12b
|
model: gemma-4-12b
|
||||||
provider: harness
|
provider: harness
|
||||||
```
|
```
|
||||||
|
|
||||||
|
**Permanent fix**: Patch `_resolve_task_provider_model()` to resolve `api_key_env` when
|
||||||
|
`api_key` is empty:
|
||||||
|
```python
|
||||||
|
cfg_api_key = str(task_config.get("api_key", "")).strip() or None
|
||||||
|
if not cfg_api_key:
|
||||||
|
key_env = str(task_config.get("api_key_env", "")).strip()
|
||||||
|
if key_env:
|
||||||
|
cfg_api_key = os.getenv(key_env, "").strip() or None
|
||||||
|
```
|
||||||
|
|
||||||
**Affected agents**: All Hermes agents with harness/LiteLLM provider and
|
**Affected agents**: All Hermes agents with harness/LiteLLM provider and
|
||||||
`api_key_env` in auxiliary configs (all 4 Hermes agents patched 2026-07-05).
|
`api_key_env` in auxiliary configs (all 4 Hermes agents patched 2026-07-05).
|
||||||
|
|
||||||
|
|||||||
@@ -22,14 +22,14 @@ management, and prompt caching — without sacrificing agent capability.
|
|||||||
- `agent-configs`: current config.yaml from each active Hermes agent (Mumuni
|
- `agent-configs`: current config.yaml from each active Hermes agent (Mumuni
|
||||||
.123, any others on .129/.122) including compression, model, context_window,
|
.123, any others on .129/.122) including compression, model, context_window,
|
||||||
prompt_caching, memory settings
|
prompt_caching, memory settings
|
||||||
- `gpu-health`: health check response from all 3 GPU backends (ornith .15:8080,
|
- `gpu-health`: health check response from all 3 GPU backends (strix-moe .15:8080,
|
||||||
qwen .8:8080, gemma .110:8080)
|
qwen .8:8080, gemma .110:8080)
|
||||||
|
|
||||||
### Maintains
|
### Maintains
|
||||||
|
|
||||||
The optimized inference stack configuration — every change is applied and
|
The optimized inference stack configuration — every change is applied and
|
||||||
verified end-to-end. Postcondition: avg request_duration_ms ≤ 15000 for 90% of
|
verified end-to-end. Postcondition: avg request_duration_ms ≤ 15000 for 90% of
|
||||||
non-ornith traffic; ≤ 30000 for ornith-bound agentic calls.
|
inference calls.
|
||||||
|
|
||||||
#### liteLLM-routing
|
#### liteLLM-routing
|
||||||
The syslog-auto routing weights, model-specific timeouts, RPM limits, and
|
The syslog-auto routing weights, model-specific timeouts, RPM limits, and
|
||||||
@@ -53,18 +53,17 @@ duration.
|
|||||||
|
|
||||||
### Strategies
|
### Strategies
|
||||||
|
|
||||||
**Context is the root cause.** Every ~46K prompt token costs ~87s of ornith
|
**Context is the root cause.** Every ~46K prompt token costs ~87s of
|
||||||
prefill time at 532 tok/s. Fix context first, routing second.
|
prefill time at 532 tok/s. Fix context first, routing second.
|
||||||
|
|
||||||
- **Route by task**: ornith for multi-step reasoning only; qwen for code/standard
|
- **Route by task**: qwen for code/standard queries; gemma for
|
||||||
queries; gemma for compression/auxiliary. Never send simple completion to a
|
compression/auxiliary; strix-moe for compression tasks.
|
||||||
35B MoE.
|
- **Compress aggressively**: threshold at 40% (not 65%) — a 128K window should
|
||||||
- **Compress aggressively**: threshold at 40% (not 65%) — a 256K window should
|
compact at 51K, not 85K. Target 15% tail (not 30%).
|
||||||
compact at 102K, not 166K. Target 15% tail (not 30%).
|
|
||||||
- **Cache everything repeated**: system prompts, skill docs, AGENTS.md — these
|
- **Cache everything repeated**: system prompts, skill docs, AGENTS.md — these
|
||||||
never change between turns. Single-digit cache hit rate is unacceptable.
|
never change between turns. Single-digit cache hit rate is unacceptable.
|
||||||
- **Lower context ceiling**: 128K window is the stable ceiling for agent conversations.
|
- **Lower context ceiling**: 128K window is the stable ceiling for agent conversations.
|
||||||
GPUs reduced from 256K to 128K (2026-07-17). For larger contexts, route to external providers.
|
GPUs reduced from 256K to 128K (2026-07-17). 128K window should compact at 85K (0.65 threshold). For larger contexts, route to external providers.
|
||||||
|
|
||||||
### Shape
|
### Shape
|
||||||
|
|
||||||
@@ -100,5 +99,5 @@ call enable-prompt-caching
|
|||||||
|
|
||||||
call verify-latency
|
call verify-latency
|
||||||
host: 192.168.68.116
|
host: 192.168.68.116
|
||||||
models: [syslog-auto, qwen3.6-27B-code, gemma-4-12b, ornith-1.0-35b]
|
models: [syslog-auto, qwen3.6-27B-code, gemma-4-12b]
|
||||||
```
|
```
|
||||||
|
|||||||
@@ -199,7 +199,7 @@ description: >
|
|||||||
**Prometheus targets**:
|
**Prometheus targets**:
|
||||||
- 192.168.68.8:9400 (RTX 3090 — qwen)
|
- 192.168.68.8:9400 (RTX 3090 — qwen)
|
||||||
- 192.168.68.110:9400 (RTX 5070 — gemma)
|
- 192.168.68.110:9400 (RTX 5070 — gemma)
|
||||||
- 192.168.68.15:9400 (Strix Halo — ornith)
|
- 192.168.68.15:9400 (Strix Halo — qwen3.6-35B-udq4)
|
||||||
- 192.168.68.24:9401 (Router metrics exporter)
|
- 192.168.68.24:9401 (Router metrics exporter)
|
||||||
- harness-litellm:4000 (LiteLLM health)
|
- harness-litellm:4000 (LiteLLM health)
|
||||||
|
|
||||||
|
|||||||
@@ -67,7 +67,7 @@ Before ANY update wave:
|
|||||||
- All VMs/CTs running: check via Proxmox API
|
- All VMs/CTs running: check via Proxmox API
|
||||||
- LiteLLM healthy: `curl localhost:4000/health/liveliness` (via CT 116)
|
- LiteLLM healthy: `curl localhost:4000/health/liveliness` (via CT 116)
|
||||||
- LiteLLM MCP tools: `curl localhost:4000/mcp-rest/tools/list -H "Authorization: Bearer $MASTER_KEY"` → 90 tools
|
- LiteLLM MCP tools: `curl localhost:4000/mcp-rest/tools/list -H "Authorization: Bearer $MASTER_KEY"` → 90 tools
|
||||||
- GPU servers responding: check :8080 on VM 101, VM 103; check ornith via router (http://192.168.68.116/health/unified — .15:8080 is firewalled to .116 only)
|
- GPU servers responding: check :8080 on VM 101, VM 103; check strix-moe via router (http://192.168.68.116/health/unified — .15:8080 is firewalled to .116 only)
|
||||||
- Zulip agents connected: check Mumuni/Tanko gateway state
|
- Zulip agents connected: check Mumuni/Tanko gateway state
|
||||||
- Abiba PM2 processes online: `pm2 status`
|
- Abiba PM2 processes online: `pm2 status`
|
||||||
|
|
||||||
@@ -124,7 +124,7 @@ Before Wave 1, snapshot these files:
|
|||||||
/opt/home_stack/docker-compose.yml (VM 109 .7)
|
/opt/home_stack/docker-compose.yml (VM 109 .7)
|
||||||
/opt/audiobookshelf/docker-compose.yml (VM 109 .7)
|
/opt/audiobookshelf/docker-compose.yml (VM 109 .7)
|
||||||
/root/.pi/agent/extensions/config.yaml (CT 100 .24)
|
/root/.pi/agent/extensions/config.yaml (CT 100 .24)
|
||||||
/etc/systemd/system/ornith-server.service (amdpve .15)
|
/etc/systemd/system/ornith-server.service (amdpve .15 — strix-moe)
|
||||||
/etc/systemd/system/llama-server.service (VM 101 .8, VM 103 .110)
|
/etc/systemd/system/llama-server.service (VM 101 .8, VM 103 .110)
|
||||||
# Hermes agent configs (key enforcement — 2026-07-10)
|
# Hermes agent configs (key enforcement — 2026-07-10)
|
||||||
/root/.hermes/config.yaml (Mumuni CT 114, Tanko CT 112, etc.)
|
/root/.hermes/config.yaml (Mumuni CT 114, Tanko CT 112, etc.)
|
||||||
|
|||||||
@@ -42,7 +42,7 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2)
|
|||||||
**What changed (v3.2.0 → v4.0.0 — 2026-07-08)**:
|
**What changed (v3.2.0 → v4.0.0 — 2026-07-08)**:
|
||||||
- Router REMOVED from request path — LiteLLM proxies directly to GPU
|
- Router REMOVED from request path — LiteLLM proxies directly to GPU
|
||||||
- All GPUs at parallel 2 (was parallel 1)
|
- All GPUs at parallel 2 (was parallel 1)
|
||||||
- NVIDIA context reduced 256K→128K to free VRAM (SUPERSEDED 2026-07-16: all GPUs back to 256K — see litellm-self-heal)
|
- NVIDIA context reduced 256K→128K to free VRAM — now the stable ceiling across all GPUs (2026-07-17)
|
||||||
- LiteLLM timeouts tuned: gemma 25→120s, qwen 40→90s (SUPERSEDED 2026-07-16: qwen 300s, gemma 120s, strix 300s — see litellm-self-heal)
|
- LiteLLM timeouts tuned: gemma 25→120s, qwen 40→90s (SUPERSEDED 2026-07-16: qwen 300s, gemma 120s, strix 300s — see litellm-self-heal)
|
||||||
- nginx proxy_read_timeout: 600s, LiteLLM request_timeout: 300s
|
- nginx proxy_read_timeout: 600s, LiteLLM request_timeout: 300s
|
||||||
|
|
||||||
@@ -72,9 +72,9 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2)
|
|||||||
|
|
||||||
| Host | IP | Hardware | Models Served | Engine | Context | Parallel |
|
| Host | IP | Hardware | Models Served | Engine | Context | Parallel |
|
||||||
|------|-----|----------|---------------|--------|---------|----------|
|
|------|-----|----------|---------------|--------|---------|----------|
|
||||||
| llm-gpu | 192.168.68.8 | NVIDIA RTX 3090 (24 GB) | qwen3.6-27B-code | llama-server systemd | **256K** | 2 |
|
| llm-gpu | 192.168.68.8 | NVIDIA RTX 3090 (24 GB) | qwen3.6-27B-code | llama-server systemd | **128K** | 2 |
|
||||||
| ocu-llm | 192.168.68.110 | NVIDIA RTX 5070 (12 GB) | gemma-4-12b | llama-server systemd | **256K** | 2 |
|
| ocu-llm | 192.168.68.110 | NVIDIA RTX 5070 (12 GB) | gemma-4-12b | llama-server systemd | **128K** | 2 |
|
||||||
| amdpve | 192.168.68.15 | AMD Strix Halo 64GB UMA | qwen3.6-35B-udq4 (LiteLLM alias: strix-moe) | llama-server systemd (Vulkan) | 256K | 2 |
|
| amdpve | 192.168.68.15 | AMD Strix Halo 64GB UMA | qwen3.6-35B-udq4 (LiteLLM alias: strix-moe) | llama-server systemd (Vulkan) | 128K | 2 |
|
||||||
|
|
||||||
## Model Fallback Chains (LiteLLM)
|
## Model Fallback Chains (LiteLLM)
|
||||||
|
|
||||||
|
|||||||
@@ -57,9 +57,9 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2)
|
|||||||
|
|
||||||
| Host | IP | Hardware | Models Served | Engine | Context | Parallel |
|
| Host | IP | Hardware | Models Served | Engine | Context | Parallel |
|
||||||
|------|-----|----------|---------------|--------|---------|----------|
|
|------|-----|----------|---------------|--------|---------|----------|
|
||||||
| llm-gpu | 192.168.68.8 | NVIDIA RTX 3090 (24 GB) | qwen3.6-27B-code | llama-server systemd (`/home/llmuser/llama-wrapper.sh`, `-c 262144 --parallel 2 --ngl 99`) | **256K** | 2 |
|
| llm-gpu | 192.168.68.8 | NVIDIA RTX 3090 (24 GB) | qwen3.6-27B-code | llama-server systemd (`/home/llmuser/llama-wrapper.sh`, `-c 131072 --parallel 2 --ngl 99`) | **128K** | 2 |
|
||||||
| ocu-llm | 192.168.68.110 | NVIDIA RTX 5070 (12 GB) | gemma-4-12b | llama-server systemd (`/home/llmuser/llama-wrapper.sh`, `--ctx-size 262144 --parallel 2`, IQ4_NL + MTP draft) | **256K** | 2 |
|
| ocu-llm | 192.168.68.110 | NVIDIA RTX 5070 (12 GB) | gemma-4-12b | llama-server systemd (`/home/llmuser/llama-wrapper.sh`, `--ctx-size 131072 --parallel 2`, IQ4_NL + MTP draft) | **128K** | 2 |
|
||||||
| amdpve | 192.168.68.15 | AMD Strix Halo 64GB UMA | qwen3.6-35B-udq4 (LiteLLM alias: `strix-moe`) | llama-server systemd (Vulkan) | 256K | 2 |
|
| amdpve | 192.168.68.15 | AMD Strix Halo 64GB UMA | qwen3.6-35B-udq4 (LiteLLM alias: `strix-moe`) | llama-server systemd (Vulkan) | 128K | 2 |
|
||||||
|
|
||||||
> Verified on ground 2026-07-16 via `curl /v1/models` on each host + `llama-wrapper.sh`. The AMD host's underlying model is `qwen3.6-35B-udq4`; LiteLLM exposes it under two `model_name`s: `qwen3.6-35B-udq4` and `strix-moe` (rpm 40). The legacy name `ornith-1.0-35b` does NOT exist in LiteLLM and must not be referenced.
|
> Verified on ground 2026-07-16 via `curl /v1/models` on each host + `llama-wrapper.sh`. The AMD host's underlying model is `qwen3.6-35B-udq4`; LiteLLM exposes it under two `model_name`s: `qwen3.6-35B-udq4` and `strix-moe` (rpm 40). The legacy name `ornith-1.0-35b` does NOT exist in LiteLLM and must not be referenced.
|
||||||
|
|
||||||
@@ -101,7 +101,7 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2)
|
|||||||
## Script Operations (synced 2026-07-16)
|
## Script Operations (synced 2026-07-16)
|
||||||
|
|
||||||
- **Health-check script** (`/opt/inference-harness/scripts/litellm-health-check.sh` on CT 116): `gpu-fleet` check fails only on **critical** alerts (warnings are informational). Tests `strix-moe` (not `ornith-1.0-35b`).
|
- **Health-check script** (`/opt/inference-harness/scripts/litellm-health-check.sh` on CT 116): `gpu-fleet` check fails only on **critical** alerts (warnings are informational). Tests `strix-moe` (not `ornith-1.0-35b`).
|
||||||
- **GPU monitor** (`/root/scripts/gpu-monitor-server.py` on pi .24): runs as **systemd unit `gpu-monitor.service`** (was bare `&` process). `gpu_count` includes Strix Halo (was 2, now 3). VRAM alert thresholds: warning 93%, critical 97% (raised from 90/95 — 256K context steady-state is ~96% on RTX 3090, not a fault).
|
- **GPU monitor** (`/root/scripts/gpu-monitor-server.py` on pi .24): runs as **systemd unit `gpu-monitor.service`** (was bare `&` process). `gpu_count` includes Strix Halo (was 2, now 3). VRAM alert thresholds: warning 93%, critical 97% (raised from 90/95 — 128K context steady-state is ~70% on RTX 3090, not a fault).
|
||||||
- **Agent key monitor** (`/root/scripts/agent-health-check.py` on pi .24, cron `*/10`): reads each agent's **live** `LITELLM_API_KEY` from its gateway process env via SSH — never hardcodes keys (hardcoded keys rot on rotation and caused 9×401/30min). Fleet roster: abiba, tanko, mumuni, koby, koonimo (legacy `tdunna`/`baggy` removed — never existed).
|
- **Agent key monitor** (`/root/scripts/agent-health-check.py` on pi .24, cron `*/10`): reads each agent's **live** `LITELLM_API_KEY` from its gateway process env via SSH — never hardcodes keys (hardcoded keys rot on rotation and caused 9×401/30min). Fleet roster: abiba, tanko, mumuni, koby, koonimo (legacy `tdunna`/`baggy` removed — never existed).
|
||||||
- **Stale keys cleaned**: `daily-infra-report.py` SYNTHETIC_API_KEY was stale (`sk-U_ydi3B` → 401); now reads `LITELLM_MASTER_KEY` from env. Deprecated scripts (`router-original.py`, `router-phase0-backup.py`, `apply-fixes.py`) still reference `sk-syslog-local-master-key` but do not actively poll LiteLLM.
|
- **Stale keys cleaned**: `daily-infra-report.py` SYNTHETIC_API_KEY was stale (`sk-U_ydi3B` → 401); now reads `LITELLM_MASTER_KEY` from env. Deprecated scripts (`router-original.py`, `router-phase0-backup.py`, `apply-fixes.py`) still reference `sk-syslog-local-master-key` but do not actively poll LiteLLM.
|
||||||
|
|
||||||
|
|||||||
@@ -87,7 +87,7 @@ agent: abiba
|
|||||||
| storepve | 192.168.68.6 | PVE |
|
| storepve | 192.168.68.6 | PVE |
|
||||||
| acerpve | 192.168.68.9 | PVE (hosts llm-gpu qemu/101) |
|
| acerpve | 192.168.68.9 | PVE (hosts llm-gpu qemu/101) |
|
||||||
| minipve | 192.168.68.12 | PVE |
|
| minipve | 192.168.68.12 | PVE |
|
||||||
| amdpve | 192.168.68.15 | PVE + Strix Halo LLM (ornith) |
|
| amdpve | 192.168.68.15 | PVE + Strix Halo LLM (qwen3.6-35B-udq4, strix-moe) |
|
||||||
|
|
||||||
## Operations
|
## Operations
|
||||||
|
|
||||||
|
|||||||
Reference in New Issue
Block a user