docs(litellm-health): single source of truth, gpu-vision alias, litellm-health as live owner #79

Merged
abiba-bot merged 8 commits from fix/litellm-health-drift-20260912 into master 2026-09-12 16:05:28 +00:00
8 changed files with 132 additions and 174 deletions
+3 -3
View File
@@ -88,7 +88,7 @@ prose run memory-audit-maintenance memory_threshold=90 verify_configs=true
prose run hermes-config-template agent_name=syslog-devops default_model=claude-sonnet-4 prose run hermes-config-template agent_name=syslog-devops default_model=claude-sonnet-4
# Configure an agent with a different auxiliary model # Configure an agent with a different auxiliary model
prose run hermes-config-template agent_name=syslog-code default_model=qwen3.6-27B-code auxiliary_model=gemma-4-12b prose run hermes-config-template agent_name=syslog-code default_model=qwen3.6-27B-code auxiliary_model=gpu-vision
``` ```
### Option B: Manual Execution ### Option B: Manual Execution
@@ -116,7 +116,7 @@ Run on trigger or schedule. Maintain persistent world-model state across runs.
| `zulip-health` | Zulip | Checks Zulip connectivity, message flow, and bot responsiveness. | | `zulip-health` | Zulip | Checks Zulip connectivity, message flow, and bot responsiveness. |
| `zulip-mention-reliability` | Zulip | Diagnoses and fixes @mention detection issues in Zulip. | | `zulip-mention-reliability` | Zulip | Diagnoses and fixes @mention detection issues in Zulip. |
| `zulip-approval-fix` | Zulip | Fixes broken /approve and /deny slash commands for Hermes agents. | | `zulip-approval-fix` | Zulip | Fixes broken /approve and /deny slash commands for Hermes agents. |
| `litellm-self-heal` | LiteLLM | Consolidated health check + self-healing for the full nginx → LiteLLM → GPU chain. Verifies 8 containers, 3 GPUs, model inference, and agent keys. Applies 9 remediation rules. (litellm-health merged into this contract 2026-07-09.) | | `litellm-self-heal` | LiteLLM | Applies remediation rules for LiteLLM stack failures detected by `litellm-health` (full nginx → LiteLLM → GPU chain). 9 remediation rules. |
| `gpu-fleet` | GPU | Manages the GPU inference fleet: model deployment, registration, health checks, LiteLLM sync. | | `gpu-fleet` | GPU | Manages the GPU inference fleet: model deployment, registration, health checks, LiteLLM sync. |
| `gpu-monitor` | GPU | Comprehensive GPU fleet monitor — polls sidecars, router, LiteLLM every 15s, renders SSE dashboard. | | `gpu-monitor` | GPU | Comprehensive GPU fleet monitor — polls sidecars, router, LiteLLM every 15s, renders SSE dashboard. |
| `proxmox-monitor` | Infra | Proxmox cluster + Docker monitoring via the existing Grafana/Prometheus stack on CT 116. | | `proxmox-monitor` | Infra | Proxmox cluster + Docker monitoring via the existing Grafana/Prometheus stack on CT 116. |
@@ -149,7 +149,7 @@ Called on-demand as single-render tools.
| Contract | Description | | Contract | Description |
|---|---| |---|---|
| `litellm-api-keys` | Manages LiteLLM API keys for agent identity. Create, rotate, verify, and list agent keys. References gpu-fleet for current key inventory. | | `litellm-api-keys` | Manages LiteLLM API keys for agent identity. Create, rotate, verify, and list agent keys. References gpu-fleet for current key inventory. |
| `litellm-health` | ⚠️ **DEPRECATED** — consolidated into `litellm-self-heal` (2026-07-09). Retained for reference only. | | `litellm-health` | LiteLLM health check: public vs backend surfaces, CT 116 containers, GPU fleet, one model per GPU host, and agent keys. Owner of the probes; `litellm-self-heal` owns remediation. |
| `infrastructure-monitoring` | Target-state for Prometheus + GPU exporters + Grafana. Core stack deployed, GPU exporters NOT live. | | `infrastructure-monitoring` | Target-state for Prometheus + GPU exporters + Grafana. Core stack deployed, GPU exporters NOT live. |
| `stirling-pdf-agent-access` | Documents the Stirling-PDF API access pattern for agents — global API key, 12 operations, curl examples. Agents use the `stirling-pdf-api` shared skill for templates. | | `stirling-pdf-agent-access` | Documents the Stirling-PDF API access pattern for agents — global API key, 12 operations, curl examples. Agents use the `stirling-pdf-api` shared skill for templates. |
| `hello-world` | Minimal test contract — verifies the OpenProse execution pipeline works. | | `hello-world` | Minimal test contract — verifies the OpenProse execution pipeline works. |
+2 -2
View File
@@ -693,8 +693,8 @@ contracts:
version: 1.0.0 version: 1.0.0
trigger: trigger:
type: scheduled type: scheduled
cadence: '*/10 * * * *' cadence: '5 3,7,11,15,19,23 * * *'
description: "Every 10 minutes \u2014 LiteLLM proxy health" description: "4-hourly staggered dispatch via fm-send (run contract litellm-health)"
cron_job_id: null cron_job_id: null
execution: execution:
agent: abiba agent: abiba
+6 -2
View File
@@ -2,6 +2,10 @@
Generated: 2026-07-13 20:59:18 ET Generated: 2026-07-13 20:59:18 ET
> **Point-in-time snapshot.** Schedules and cadences are authoritative in
> `contract-registry.yaml`; any schedule quoted below may be stale. Do not use
> this file as the source of truth for a contract's trigger.
--- ---
## hermes-key-enforcement ## hermes-key-enforcement
@@ -442,7 +446,7 @@ IMPORTANT: If the contract file does not exist in prose-contracts/main, report f
## litellm-health ## litellm-health
**Category:** monitoring | **Domain:** litellm | **Owner:** abiba | **Schedule:** */10 * * * * **Category:** monitoring | **Domain:** litellm | **Owner:** abiba | **Schedule:** see contract-registry.yaml (authoritative)
``` ```
Contract Enforcement: litellm-health Contract Enforcement: litellm-health
@@ -450,7 +454,7 @@ Contract Enforcement: litellm-health
Category: monitoring Category: monitoring
Domain: litellm Domain: litellm
Owner: abiba Owner: abiba
Schedule: Every 10 minutes — LiteLLM proxy health Schedule: see contract-registry.yaml (authoritative)
This is a monitoring contract. Execute the monitoring checks defined in the contract. Report any deviations from expected state. This is a monitoring contract. Execute the monitoring checks defined in the contract. Report any deviations from expected state.
+1 -1
View File
@@ -13,7 +13,7 @@ the description:
1. **What system does this contract touch?** Name the hosts, CTs, containers, 1. **What system does this contract touch?** Name the hosts, CTs, containers,
and services explicitly. "The inference fleet" is vague. "GPU .8 (RTX 3090, and services explicitly. "The inference fleet" is vague. "GPU .8 (RTX 3090,
qwen), .110 (RTX 5070, gemma), .15 (Strix Halo, strix-moe), and LiteLLM on CT qwen), .110 (RTX 5070, gpu-vision), .15 (Strix Halo, strix-moe), and LiteLLM on CT
116" is specific. 116" is specific.
2. **Who runs this contract, and when?** State the agent, the trigger (cron, 2. **Who runs this contract, and when?** State the agent, the trigger (cron,
+28 -56
View File
@@ -6,9 +6,10 @@ description: >
registration, health checks, LiteLLM sync, agent key management, GPU registration, health checks, LiteLLM sync, agent key management, GPU
saturation watchdog, Prometheus/Grafana monitoring, and self-healing. saturation watchdog, Prometheus/Grafana monitoring, and self-healing.
UPDATED 2026-07-15: Stable role-based aliases introduced: strix-moe, UPDATED 2026-07-15: Stable role-based aliases introduced: strix-moe,
gpu-dense, gpu-light. These never change — only the underlying model does. gpu-dense, gpu-vision (gpu-light was superseded by gpu-vision on 2026-09-12).
These never change — only the underlying model does.
Strix Halo: strix-moe → unsloth/Qwen3.6-35B-A3B-MTP (UD-Q4_K_M, 22GB). Strix Halo: strix-moe → unsloth/Qwen3.6-35B-A3B-MTP (UD-Q4_K_M, 22GB).
RTX 5070: gemma-4-12b Q4_K_M → IQ4_NL + MTP draft (122 tok/s, 2x faster). RTX 5070: gpu-vision — IQ4_NL + MTP draft (~122 tok/s, 2x faster).
UPDATED 2026-07-17: Context reduced fleet-wide from 256K to 128K for stability. UPDATED 2026-07-17: Context reduced fleet-wide from 256K to 128K for stability.
Strix Halo model swapped to qwen3.6-35B-udq4 (22GB, strix-moe alias). Strix Halo model swapped to qwen3.6-35B-udq4 (22GB, strix-moe alias).
Instability observed near 100K at 256K (now all GPUs at 128K). 128K is the stable ceiling. Instability observed near 100K at 256K (now all GPUs at 128K). 128K is the stable ceiling.
@@ -24,7 +25,7 @@ triggers:
## Maintains ## Maintains
- gpu_roster: { models: map, hosts: map } — Single source of truth for all GPU models - gpu_roster: { models: map, hosts: map } — GPU host/model roster; the authoritative alias/weight/fallback registry is CT 116 `litellm_config.yaml`
- router: { status: "healthy", roster_loaded: bool, models: array } - router: { status: "healthy", roster_loaded: bool, models: array }
- litellm: { status: "healthy", keys: array, models: array } - litellm: { status: "healthy", keys: array, models: array }
- agent_keys: { agent: api_key } — All agent API keys registered in LiteLLM DB - agent_keys: { agent: api_key } — All agent API keys registered in LiteLLM DB
@@ -79,54 +80,28 @@ triggers:
Agent configs, cron jobs, and workflows MUST use these aliases, never model-specific names. Agent configs, cron jobs, and workflows MUST use these aliases, never model-specific names.
When a model is swapped on a GPU, ONLY the infrastructure layer changes — agent configs are untouched. When a model is swapped on a GPU, ONLY the infrastructure layer changes — agent configs are untouched.
| Alias | GPU | Current Model | Will Route To | | Alias | Serves | Where | Kind |
|-------|-----|---------------|---------------| |-------|--------|-------|------|
| `strix-moe` | Strix Halo (.15) | qwen3.6-35B-udq4 | Whatever runs on Strix Halo | | `gpu-dense` | heavy reasoning | RTX 3090 (192.168.68.8) | direct alias |
| `gpu-dense` | RTX 3090 (.8) | Qwen3.8-27B-Uncensored-Q4_K_M | Whatever runs on RTX 3090 | | `gpu-vision` | vision / web extract / light tasks | RTX 5070 (192.168.68.110) | direct alias AND `syslog-auto` pool member |
| `gpu-light` | RTX 5070 (.110) | gemma-4-12b | Whatever runs on RTX 5070 | | `strix-moe` | compression (MoE) | Strix Halo (192.168.68.15) | direct alias |
| `syslog-auto` | balanced default | weighted pool across the three GPU hosts | pool router |
**Backward compatibility**: Old model-specific names (qwen3.6-27B-code, gemma-4-12b, qwen3.5-9b-it) still work Single source of truth for models, aliases, rpm caps, weights and fallback chains:
but are deprecated for agent configs. Only the stable aliases survive model swaps. CT 116 `/opt/inference-harness/litellm_config.yaml`. Do not duplicate those values in
contracts — read them there.
## Current Model Assignments (2026-07-15) **Backward compatibility**: Old model-specific names (qwen3.6-27B-code, qwen3.5-9b-it) still work
but are deprecated for agent configs. Only the stable aliases survive model swaps. `gemma-4-12b`
| Model | GPU | Host | VRAM | Ctx | KV Cache | Parallel | Batch/Ubatch | Status | is retired and returns 400 `Invalid model name`.
|-------|-----|------|------|-----|----------|----------|-------------|--------|
## Routing Configuration (LiteLLM — July 2026) ## Routing Configuration (LiteLLM — July 2026)
### syslog-auto Weighted Pool (Direct GPU — bypasses router) Model, alias, rpm/weight and fallback values are owned by CT 116
`/opt/inference-harness/litellm_config.yaml` (see § Stable Role-Based Aliases above).
| Model | GPU | Weight | RPM Cap | Timeout |
|-------|-----|--------|---------|---------|
| Qwen3.8-27B-Uncensored-Q4_K_M | RTX 3090 (.8:8080) | **0.55** | 500 | **300s** |
Note: All syslog-auto entries route directly to GPUs with `api_key: not-needed`. The router (port 9000) was decommissioned 2026-09-11 and is NOT in the inference path. Note: All syslog-auto entries route directly to GPUs with `api_key: not-needed`. The router (port 9000) was decommissioned 2026-09-11 and is NOT in the inference path.
### Direct Model Endpoints
| Model | RPM Cap | Notes |
|-------|---------|-------|
### Stable Aliases (for agent configs — never change)
| Alias | RPM Cap | Routes To | Purpose |
|-------|---------|-----------|---------|
| `strix-moe` | 40 | Strix Halo | Compression tasks (MoE models) |
| `gpu-dense` | 500 | RTX 3090 | Heavy reasoning |
| `gpu-light` | 500 | RTX 5070 | Vision, web extract, light tasks |
### Fallback Chains
- gemma → qwen
- qwen → gemma
- strix-moe → qwen → gemma
- syslog-auto → qwen → gemma → qwen3.6-35B-udq4
### Why Strix Halo RPM Is Capped
- Direct (strix-moe): 40 RPM (tight) — Strix Halo is shared with compression tasks
- Via syslog-auto: 60 RPM (moderate) — prevents flooding when multiple agents use syslog-auto simultaneously
- Combined max: ~100 RPM across both paths — Strix Halo can sustain this at 80°C
## Operations ## Operations
### add-model ### add-model
@@ -176,9 +151,7 @@ Show full fleet status: GPUs, models, VRAM, context windows, parallel slots, act
2. Check llama-server processes: `ps aux | grep llama-server` on all 3 hosts 2. Check llama-server processes: `ps aux | grep llama-server` on all 3 hosts
3. Check LiteLLM: `curl http://192.168.68.116/health` (expect "I'm alive!") 3. Check LiteLLM: `curl http://192.168.68.116/health` (expect "I'm alive!")
4. Check LiteLLM models: `curl -H "Authorization: Bearer $MASTER_KEY" http://192.168.68.116/v1/models` 4. Check LiteLLM models: `curl -H "Authorization: Bearer $MASTER_KEY" http://192.168.68.116/v1/models`
5. Check LiteLLM timeouts: `grep -n 'timeout:' /opt/inference-harness/litellm_config.yaml` 5. Check LiteLLM timeouts: `grep -n 'timeout:' /opt/inference-harness/litellm_config.yaml` — read the live values from the authority config; do not assert them from this contract.
- Qwen3.5-9B: 120s, qwen3.6-27B-code: 300s, Carnice-Qwen3.6-MoE-35B-A3B/strix-moe: 300s (strix-moe alias retained, legacy name qwen3.6-35B-udq4 deprecated)
- global request_timeout: 300s, nginx proxy_read_timeout: 600s
6. Check AMD metrics: `curl http://192.168.68.15:9400/metrics` (Radeon 8060S, util%, VRAM, temp, power) 6. Check AMD metrics: `curl http://192.168.68.15:9400/metrics` (Radeon 8060S, util%, VRAM, temp, power)
7. Check port conflicts: verify only one llama-server on :8080 per host 7. Check port conflicts: verify only one llama-server on :8080 per host
8. Verify agent keys: 9 keys in LiteLLM DB (`GET /key/list`) 8. Verify agent keys: 9 keys in LiteLLM DB (`GET /key/list`)
@@ -243,6 +216,7 @@ If no SSH access, send Zulip DM via abiba-bot with vault update instructions.
- **Alert migration**: All alerts now go to `#agent-hub` topics (`alerts-gpu`, `alerts-pm2`, `alerts-infra`) instead of DMs. Cross-agent visibility enabled. - **Alert migration**: All alerts now go to `#agent-hub` topics (`alerts-gpu`, `alerts-pm2`, `alerts-infra`) instead of DMs. Cross-agent visibility enabled.
- **tok/s benchmarks**: Measured every 5 min via LiteLLM proxy. Baselines tracked with 30%/50% degradation thresholds. - **tok/s benchmarks**: Measured every 5 min via LiteLLM proxy. Baselines tracked with 30%/50% degradation thresholds.
- **NetBird 502**: Tanko routes through NetBird for litellm.sysloggh.net. Use direct IP if NetBird down. - **NetBird 502**: Tanko routes through NetBird for litellm.sysloggh.net. Use direct IP if NetBird down.
- **Alias-retirement follow-up (2026-09-12)**: the agent-facing templates (`hermes-config-template.prose.md`, `hermes-agent-baseline.prose.md`), `litellm-api-keys.prose.md`, and the executable `audit-hermes-config.py` still reference the retired `gpu-light`/`gemma-4-12b` names, and `audit-hermes-config.py` Rule 8 currently fails a config whose vision/web_extract model is `gpu-vision`. These are tracked separately and are intentionally NOT updated in this change.
## GPU Inference Benchmarks (Current) ## GPU Inference Benchmarks (Current)
@@ -265,10 +239,10 @@ History stored at `/root/data/toks-history.json` with 7-day rolling window.
### Stable Aliases — CRITICAL ### Stable Aliases — CRITICAL
All agent configs MUST use stable role-based aliases, never model-specific names: All agent configs MUST use stable role-based aliases, never model-specific names:
- `compression.model: strix-moe` (NOT `qwen3.6-35B-udq4`) - `compression.model: strix-moe`
- `auxiliary.vision.model: gpu-light` (NOT `gemma-4-12b`) - `auxiliary.vision.model: gpu-vision`
- `delegation.model: gpu-dense` (NOT `qwen3.6-27B-code`) - `delegation.model: gpu-dense`
- `auxiliary.web_extract.model: gpu-light` - `auxiliary.web_extract.model: gpu-vision`
When the underlying model is swapped, only the LiteLLM config changes — agent configs are untouched. When the underlying model is swapped, only the LiteLLM config changes — agent configs are untouched.
@@ -277,7 +251,7 @@ When the underlying model is swapped, only the LiteLLM config changes — agent
- **All agents**: 128K ceiling — stable margin. For >128K workloads, use external providers (deepseek) - **All agents**: 128K ceiling — stable margin. For >128K workloads, use external providers (deepseek)
- Compression threshold 0.60: fires at ~77K (~51K headroom before 128K ceiling) - Compression threshold 0.60: fires at ~77K (~51K headroom before 128K ceiling)
- **Pi agents (Abiba)**: `compaction.reserveTokens: 52739` (≈60% of 128K) - **Pi agents (Abiba)**: `compaction.reserveTokens: 52739` (≈60% of 128K)
- Mumuni compression model alias: `strix-moe` with 300s timeout - Mumuni compression model alias: `strix-moe`
### Mumuni Agent Profile ### Mumuni Agent Profile
@@ -285,12 +259,12 @@ Mumuni (kagentz CT105, 192.168.68.14 — migrated from CT100 2026-08-29) is the
| Setting | Value | Notes | | Setting | Value | Notes |
|---------|-------|-------| |---------|-------|-------|
| `model.default` | `syslog-auto` | Weighted pool (55% qwen, 30% strix, 15% gemma) | | `model.default` | `syslog-auto` | Balanced default (pool router) |
| `model.provider` | `custom:litellm` | LiteLLM on CT116 | | `model.provider` | `custom:litellm` | LiteLLM on CT116 |
| `compression.model` | `strix-moe` | Stable alias — survives model swaps | | `compression.model` | `strix-moe` | Stable alias — survives model swaps |
| `aux.compression.model` | `strix-moe` | Compression auxiliary model | | `aux.compression.model` | `strix-moe` | Compression auxiliary model |
| `aux.vision.model` | `gpu-light` | Vision tasks (RTX 5070) | | `aux.vision.model` | `gpu-vision` | Vision tasks (RTX 5070) |
| `aux.web_extract.model` | `gpu-light` | Web extraction | | `aux.web_extract.model` | `gpu-vision` | Web extraction |
| `delegation.model` | `gpu-dense` | Sub-agent reasoning (RTX 3090) | | `delegation.model` | `gpu-dense` | Sub-agent reasoning (RTX 3090) |
| `context.max_context_window` | 131072 (128K) | Reduced from 256K 2026-07-17 — stable 128K ceiling | | `context.max_context_window` | 131072 (128K) | Reduced from 256K 2026-07-17 — stable 128K ceiling |
| `compression.threshold` | 0.60 | Triggers at ~77K (~60% of 128K) — optimized for 128K context | | `compression.threshold` | 0.60 | Triggers at ~77K (~60% of 128K) — optimized for 128K context |
@@ -299,8 +273,6 @@ Mumuni (kagentz CT105, 192.168.68.14 — migrated from CT100 2026-08-29) is the
| `memory.memory_char_limit` | 800 | Brief memory entries | | `memory.memory_char_limit` | 800 | Brief memory entries |
| `personalities` | `creative` | Creative assistant personality | | `personalities` | `creative` | Creative assistant personality |
| Platforms | cli, homeassistant, signal, telegram, zulip | All Hermes platforms | | Platforms | cli, homeassistant, signal, telegram, zulip | All Hermes platforms |
| Main model timeout | 300s | LiteLLM global timeout |
| Compression model timeout | 300s | strix-moe timeout increased from 120s |
### Agent Update Status (2026-07-15) ### Agent Update Status (2026-07-15)
+3 -2
View File
@@ -79,8 +79,9 @@ proxy queuing.
Send a real model name and use a real (probe-designated) key. Send a real model name and use a real (probe-designated) key.
- Probe timeout: 30s. A probe that takes longer than 30s IS the alert — - Probe timeout: 30s. A probe that takes longer than 30s IS the alert —
report "backend slow (>30s)" rather than hanging. report "backend slow (>30s)" rather than hanging.
- Probe cadence: at most hourly. The 6h litellm-health cron cadence is the - Probe cadence: at most hourly. The litellm-health cron cadence (authoritative
standard; sub-hourly synthetic traffic distorts latency baselines. trigger in contract-registry.yaml) is the standard; sub-hourly synthetic
traffic distorts latency baselines.
### 5. Batch/benchmark jobs — schedule away from 04:00-07:00 EDT and chunk ### 5. Batch/benchmark jobs — schedule away from 04:00-07:00 EDT and chunk
+52 -35
View File
@@ -1,14 +1,7 @@
--- ---
kind: function kind: function
name: litellm-health name: litellm-health
status: deprecated status: active
deprecated_on: 2026-07-09
replaced_by: litellm-self-heal.prose.md
note: >
Consolidated into litellm-self-heal.prose.md to eliminate duplication
of architecture diagrams, GPU topology, timeout tables, and container
lists. Health check is now § Health Check within litellm-self-heal.
This file is retained for reference only — use litellm-self-heal instead.
description: > description: >
Verifies the LiteLLM inference stack health. Current architecture (2026-07-09): Verifies the LiteLLM inference stack health. Current architecture (2026-07-09):
nginx:80 → LiteLLM:4000 → GPU(llama-server) via direct proxy. nginx:80 → LiteLLM:4000 → GPU(llama-server) via direct proxy.
@@ -52,8 +45,7 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2)
- Router REMOVED from request path — LiteLLM proxies directly to GPU - Router REMOVED from request path — LiteLLM proxies directly to GPU
- All GPUs at parallel 2 (was parallel 1) - All GPUs at parallel 2 (was parallel 1)
- NVIDIA context reduced 256K→128K to free VRAM — now the stable ceiling across all GPUs (2026-07-17) - NVIDIA context reduced 256K→128K to free VRAM — now the stable ceiling across all GPUs (2026-07-17)
- LiteLLM timeouts tuned: gemma 25→120s, qwen 40→90s (SUPERSEDED 2026-07-16: qwen 300s, gemma 120s, strix 300s — see litellm-self-heal) - Timeouts and fallback chains are config state — read them from CT 116 `/opt/inference-harness/litellm_config.yaml`; they are not duplicated here.
- nginx proxy_read_timeout: 600s, LiteLLM request_timeout: 300s
## Parameters ## Parameters
@@ -75,26 +67,27 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2)
- SSH key access to backend_host for container checks - SSH key access to backend_host for container checks
- Network access to public_url, auth_host, and gpu_dashboard_url - Network access to public_url, auth_host, and gpu_dashboard_url
- LiteLLM master key for key management endpoints - LiteLLM master key for admin endpoints only (`/key/list`, `/key/generate`, `/key/info`)
- The dedicated `monitor` agent key on CT 116 at `/etc/litellm-monitor.env` (root-only 0600) for
model inference checks — the master key must never be used for inference
## GPU Fleet Topology ## GPU Fleet Topology
| Host | IP | Hardware | Models Served | Engine | Context | Parallel | | Host | IP | Hardware | Role |
|------|-----|----------|---------------|--------|---------|----------| |------|-----|----------|------|
| llm-gpu | 192.168.68.8 | NVIDIA RTX 3090 (24 GB) | qwen3.6-27B-code | llama-server systemd | **128K** | 2 | | llm-gpu | 192.168.68.8 | NVIDIA RTX 3090 (24 GB) | heavy reasoning (`gpu-dense`) |
| ocu-llm | 192.168.68.110 | NVIDIA RTX 5070 (12 GB) | gemma-4-12b | llama-server systemd | **128K** | 2 | | ocu-llm | 192.168.68.110 | NVIDIA RTX 5070 (12 GB) | vision / web extract / light tasks (`gpu-vision`) |
| amdpve | 192.168.68.15 | AMD Strix Halo 64GB UMA | qwen3.6-35B-udq4 (LiteLLM alias: strix-moe) | llama-server systemd (Vulkan) | 128K | 2 | | amdpve | 192.168.68.15 | AMD Strix Halo 64GB UMA | compression (`strix-moe`) |
## Model Fallback Chains (LiteLLM) Single source of truth for models, aliases, rpm caps, weights and fallback chains:
CT 116 `/opt/inference-harness/litellm_config.yaml`. Do not duplicate those values in
contracts — read them there. Do not re-add retired names (`gemma-4-12b`, `gpu-light`,
`crew-auto`).
| Primary | Timeout | Fallback | Timeout | > Re-scope note (2026-09-12): the earlier plan to restate the live fallback chains and
|---------|---------|----------|---------| > per-model timeouts in this contract is intentionally superseded — that state is config,
| qwen3.6-27B-code | 300s | gemma-4-12b | 120s | > and this contract points at the CT 116 config instead. Step 7 likewise authenticates with
| gemma-4-12b | 120s | qwen3.6-27B-code | 300s | > the dedicated `monitor` key, not the master key, which is admin-only.
| qwen3.6-35B-udq4 / strix-moe | 300s | qwen → gemma | — |
| syslog-auto (balanced) | 300s | qwen → gemma | — |
> Global: request_timeout=300s, nginx proxy_read_timeout=600s
## Containers on CT 116 ## Containers on CT 116
@@ -112,10 +105,22 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2)
1. **Read parameters** — Use provided values or defaults 1. **Read parameters** — Use provided values or defaults
2. **Check public endpoints**: 2. **Check the end-user surfaces** — the public edge and the backend edge serve the SAME
- GET {{public_url}}/litellm/ui/ → expect 200 ("LiteLLM Dashboard") app under DIFFERENT paths. They are not interchangeable, so every probe below must name
- GET {{public_url}}/litellm/docs → expect 200 ("LiteLLM API - Swagger UI") the surface it targets. Never point a check at a path that only resolves on the other
- GET {{public_url}}/ui/ and {{public_url}}/docs → expect 301 → /litellm/... → 200 (one-hop redirect helpers, added 2026-09-11) surface.
**Public edge** — `{{public_url}}` (https://litellm.sysloggh.net) serves the app at the
ROOT; the `/litellm/` prefix does not exist there and 404s:
- GET {{public_url}}/ui/ → expect 200 ("LiteLLM Dashboard")
- GET {{public_url}}/docs → expect 200 ("LiteLLM API - Swagger UI")
- GET {{public_url}}/litellm/ui/ and {{public_url}}/litellm/docs → expect 404 (not served on this edge)
**Backend edge** — `http://{{backend_host}}` (port 80) serves the app UNDER `/litellm/`:
- GET http://{{backend_host}}/litellm/ui/ → expect 200 ("LiteLLM Dashboard")
- GET http://{{backend_host}}/litellm/docs → expect 200 ("LiteLLM API - Swagger UI")
- GET http://{{backend_host}}/ui/ and /docs → expect 301 → /litellm/... → 200 (one-hop redirect helpers,
added 2026-09-11)
3. **Check LiteLLM health (no-auth)**: 3. **Check LiteLLM health (no-auth)**:
- GET http://{{backend_host}}/litellm/health/liveliness → expect 200 - GET http://{{backend_host}}/litellm/health/liveliness → expect 200
@@ -136,14 +141,26 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2)
- Verify GPUs reporting status "healthy" - Verify GPUs reporting status "healthy"
- Check alerts array for active warnings/critical - Check alerts array for active warnings/critical
7. **Check model inference via LiteLLM** — Test each model: 7. **Check model inference via LiteLLM** — Test one model on each GPU host. The health
- POST /v1/chat/completions model=gemma-4-12b → expect 200 check runs on the **backend edge**, not the public edge, so these paths carry the
- POST /v1/chat/completions model=qwen3.6-27B-code → expect 200 `/litellm/` prefix:
- POST /v1/chat/completions model=strix-moe → expect 200 - POST http://{{backend_host}}/litellm/v1/chat/completions model=qwen3.6-27B-code → expect 200 (RTX 3090, .8)
- Use master key for auth - POST http://{{backend_host}}/litellm/v1/chat/completions model=gpu-vision → expect 200 (RTX 5070, .110)
- POST http://{{backend_host}}/litellm/v1/chat/completions model=strix-moe → expect 200 (Strix Halo, .15)
- Auth uses the dedicated `monitor` agent key, read on CT 116 from
`/etc/litellm-monitor.env` (root-only 0600). Do NOT use the master key for inference —
the master key is for admin endpoints only (`/key/list`, `/key/generate`, `/key/info`).
- `gemma-4-12b` was retired and returns 400 `Invalid model name` — do not re-add it to
this list. The RTX 5070 host now serves `gpu-vision`.
- `/v1/models` is **key-scoped**: a model is only visible to keys allowed to use it, so
the set depends on the key. Always state which key a model list was read with — a
snapshot without its key is not evidence. This probe uses the `monitor` key on the
backend surface (`http://{{backend_host}}/litellm/v1/models`). The authoritative model
registry is CT 116 `/opt/inference-harness/litellm_config.yaml`; read it there rather
than freezing a list here.
8. **Check agent keys**: 8. **Check agent keys**:
- GET /key/list with master key → verify all 6 agents have keys - GET http://{{backend_host}}/litellm/key/list with master key (admin endpoint) → verify all 6 agents have keys
9. **Check Grafana**: 9. **Check Grafana**:
- GET {{grafana_url}}/api/health → expect 200 - GET {{grafana_url}}/api/health → expect 200
+37 -73
View File
@@ -11,22 +11,22 @@ note: >
Script: `/opt/inference-harness/scripts/litellm-health-check.sh` on CT 116 (cron `0 */6 * * *`). Script: `/opt/inference-harness/scripts/litellm-health-check.sh` on CT 116 (cron `0 */6 * * *`).
Reports to /var/log/litellm/health-*.json and Gitea (SyslogSolution/health-logs). Reports to /var/log/litellm/health-*.json and Gitea (SyslogSolution/health-logs).
GPU monitoring integrated from gpu-monitor on .24:9100. GPU monitoring integrated from gpu-monitor on .24:9100.
Consolidated from litellm-health + litellm-self-heal on 2026-07-09 to eliminate Health probes are owned by litellm-health.prose.md (dispatched as
duplication of architecture diagrams, GPU topology, timeout tables, and container `run contract: litellm-health`). This contract owns remediation only — it does not
lists. Health check is now § Health Check within this contract. re-specify the probes.
Source of truth for GPU topology and keys: gpu-fleet.prose.md Source of truth for GPU topology and keys: gpu-fleet.prose.md
Last verified: 2026-07-12 Last verified: 2026-07-12
description: > description: >
LiteLLM inference stack health monitoring + self-healing. Verifies the full LiteLLM inference stack remediation. Applies remediation rules for failures detected
nginx → LiteLLM → GPU chain, 12 containers on CT 116, 3 GPU hosts, model by litellm-health.prose.md (nginx → LiteLLM → GPU chain, CT 116 containers, GPU hosts,
inference, and agent keys. Applies remediation rules for common failures. model inference, and agent keys).
Reports every action via Zulip DM and Gitea (SyslogSolution/health-logs). Reports every action via Zulip DM and Gitea (SyslogSolution/health-logs).
--- ---
--- ---
# LiteLLM Operations — Health Check + Self-Heal # LiteLLM Operations — Self-Heal (Remediation)
## Architecture (v4.0.0 — Direct: nginx → LiteLLM → GPU) ## Architecture (v4.0.0 — Direct: nginx → LiteLLM → GPU)
@@ -58,41 +58,44 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2)
## GPU Fleet Topology ## GPU Fleet Topology
| Host | IP | Hardware | Models Served | Engine | Context | Parallel | | Host | IP | Hardware | Role |
|------|-----|----------|---------------|--------|---------|----------| |------|-----|----------|------|
| llm-gpu | 192.168.68.8 | NVIDIA RTX 3090 (24 GB) | qwen3.6-27B-code | llama-server systemd (`/home/llmuser/llama-wrapper.sh`, `-c 131072 --parallel 2 --ngl 99`) | **128K** | 2 | | llm-gpu | 192.168.68.8 | NVIDIA RTX 3090 (24 GB) | heavy reasoning (`gpu-dense`) |
| ocu-llm | 192.168.68.110 | NVIDIA RTX 5070 (12 GB) | gemma-4-12b | llama-server systemd (`/home/llmuser/llama-wrapper.sh`, `--ctx-size 131072 --parallel 2`, IQ4_NL + MTP draft) | **128K** | 2 | | ocu-llm | 192.168.68.110 | NVIDIA RTX 5070 (12 GB) | vision / web extract / light tasks (`gpu-vision`) |
| amdpve | 192.168.68.15 | AMD Strix Halo 64GB UMA | qwen3.6-35B-udq4 (LiteLLM alias: `strix-moe`) | llama-server systemd (Vulkan) | 128K | 2 | | amdpve | 192.168.68.15 | AMD Strix Halo 64GB UMA | compression (`strix-moe`) |
> Verified on ground 2026-07-16 via `curl /v1/models` on each host + `llama-wrapper.sh`. The AMD host's underlying model is `qwen3.6-35B-udq4`; LiteLLM exposes it under two `model_name`s: `qwen3.6-35B-udq4` and `strix-moe` (rpm 40). The legacy name `ornith-1.0-35b` does NOT exist in LiteLLM and must not be referenced. > Verified on the ground 2026-07-16. The legacy name `ornith-1.0-35b` does NOT exist in LiteLLM and must not be referenced.
## LiteLLM Model Surface (ground truth — `/opt/inference-harness/litellm_config.yaml` on CT 116) ## LiteLLM Model Surface
`model_name`s served: `qwen3.6-27B-code`, `gemma-4-12b`, `qwen3.6-35B-udq4`, `strix-moe`, `gpu-dense`, `gpu-light`, `syslog-auto`, `crew-auto` (new 2026-08-20). Single source of truth for models, aliases, rpm caps, weights and fallback chains:
CT 116 `/opt/inference-harness/litellm_config.yaml`. Do not duplicate those values in
contracts — read them there. Do not re-add retired names (`gemma-4-12b`, `gpu-light`,
`crew-auto`).
### Context Cap Split (2026-08-20) ### Alias Surface
| Alias | Serves | Where | Kind |
|-------|--------|-------|------|
| `gpu-dense` | heavy reasoning | RTX 3090 (192.168.68.8) | direct alias |
| `gpu-vision` | vision / web extract / light tasks | RTX 5070 (192.168.68.110) | direct alias AND `syslog-auto` pool member |
| `strix-moe` | compression (MoE) | Strix Halo (192.168.68.15) | direct alias |
| `syslog-auto` | balanced default | weighted pool across the three GPU hosts | pool router |
### Context Cap Split (2026-08-20; crew cap RETIRED)
- **Abiba (firstmate)**: 128K uncapped — unlimited context for primary workloads - **Abiba (firstmate)**: 128K uncapped — unlimited context for primary workloads
- **Hermes agents** (mumuni, tanko, koby, koonimo): 128K uncapped - **Hermes agents** (mumuni, tanko, koby, koonimo): 128K uncapped
- **Crewmates** (ops, tune, verify, auth-keys, build): 64K capped — alias `crew-auto` enforces 64K limit - **Crewmates** (ops, tune, verify, auth-keys, build): the 64K cap was retired together with the `crew-auto` alias; NO context cap is currently in force.
Preferred implementation: uncap shared pool, add capped alias for crew-only. > The 64K crew cap was retired with `crew-auto` (2026-09-12). No limit is currently in force; reinstating one would need per-key model limits as a separate, deliberately-scoped change.
- `syslog-auto` is a weighted router model: qwen3.6-27B-code (0.55, rpm 500) + qwen3.6-35B-udq4 (0.30, rpm 60) + gemma-4-12b (0.15, rpm 200). - Key scoping: `/v1/models` is key-scoped, so the set a caller sees must be read with a named key rather than assumed — a monitor key, an agent key, and the master key can each return a different set. Agents should use the stable aliases (`strix-moe`, `gpu-vision`, `gpu-dense`) rather than raw model names, so model swaps don't break them.
- `gpu-dense` / `gpu-light` are high-rpm aliases (rpm 500) onto qwen3.6-27B-code / gemma-4-12b respectively.
- Key scoping: agent keys are restricted to `['syslog-auto','qwen3.6-27B-code','gemma-4-12b','strix-moe','gpu-dense','gpu-light']`. As of 2026-07-16 the `baggy`/`koby`/`mumuni`/`abiba-pi` keys ALSO include `qwen3.6-35B-udq4`; `abiba-pi` additionally includes `deepseek-v4-pro` (cloud fallback). `kagenz0`/`koonimo`/`pi-agents-unified` have the standard 6 only. Agents should still use the stable alias `strix-moe` (not the raw `qwen3.6-35B-udq4`) so model swaps don't break them.
- **NEVER use `litellm_proxy_master_key` (the `sk-litellm-...` master key) for inference.** It is for admin endpoints only (`/key/list`, `/key/generate`, `/key/info`). All inference — agent traffic, health-check model tests, monitor scripts — uses agent-specific keys. The health-check script's model tests use a dedicated `monitor` agent key stored at `/etc/litellm-monitor.env` on CT 116 (root-only, `chmod 600`); `/key/list` is the only call that legitimately uses `$MASTER_KEY`. - **NEVER use `litellm_proxy_master_key` (the `sk-litellm-...` master key) for inference.** It is for admin endpoints only (`/key/list`, `/key/generate`, `/key/info`). All inference — agent traffic, health-check model tests, monitor scripts — uses agent-specific keys. The health-check script's model tests use a dedicated `monitor` agent key stored at `/etc/litellm-monitor.env` on CT 116 (root-only, `chmod 600`); `/key/list` is the only call that legitimately uses `$MASTER_KEY`.
## Model Fallback Chains (LiteLLM) > Re-scope note (2026-09-12): the fallback-chain and per-model timeout tables were removed
> by the single-source-of-truth re-scope; read those values from the CT 116 config named
| Primary | Timeout | Fallback | Timeout | > above rather than from this contract.
|---------|---------|----------|---------|
| qwen3.6-27B-code | 300s | gemma-4-12b | 120s |
| gemma-4-12b | 120s | qwen3.6-27B-code | 300s |
| qwen3.6-35B-udq4 / strix-moe | 300s | qwen → gemma | — |
| syslog-auto (balanced) | 300s | qwen → gemma | — |
> Global: request_timeout=300s, nginx proxy_read_timeout=600s
## Containers on CT 116 ## Containers on CT 116
@@ -135,46 +138,7 @@ Preferred implementation: uncap shared pool, add capped alias for crew-only.
## Health Check ## Health Check
Run this first on every cycle. Results feed into remediation rules below. Health probes are owned by `litellm-health.prose.md` (dispatched as `run contract: litellm-health`). This contract owns remediation only — it does not re-specify the probes.
### 1. Check public endpoints
- GET {{public_url}}/litellm/ui/ → expect 200 ("LiteLLM Dashboard") # served directly by nginx
- GET {{public_url}}/litellm/docs → expect 200 ("LiteLLM API - Swagger UI") # served directly by nginx
- GET {{public_url}}/ui/ and {{public_url}}/docs → expect 301 → /litellm/... → 200 (one-hop redirect helpers, added 2026-09-11)
### 2. Check LiteLLM health (no-auth)
- GET http://{{backend_host}}/litellm/health/liveliness → expect 200
### 3. Check backend container health
- SSH to {{backend_host}} → `docker ps` → verify 12 containers healthy (11 harness + trove-agent-docker, added 2026-09-11)
- Critical: harness-litellm, harness-nginx, harness-postgres
- Monitoring: harness-redis, harness-dashboard, harness-grafana, harness-prometheus,
harness-alertmanager, harness-zulip-bridge, harness-docker-stats, harness-pve-exporter,
trove-agent-docker (ghcr.io/techdox/trove-agent-docker, added 2026-09-11)
- Decommissioned 2026-09-11: harness-router (container, image and config removed)
### 4. Check GPU fleet health (via fleet dashboard)
- GET {{gpu_dashboard_url}}/gpu-data → expect 200 with GPU metrics JSON
- Verify GPUs reporting status "healthy"
- Check alerts array for active warnings/critical
### 5. Check model inference via LiteLLM — test each model
- POST /v1/chat/completions model=gemma-4-12b → expect 200
- POST /v1/chat/completions model=qwen3.6-27B-code → expect 200
- POST /v1/chat/completions model=strix-moe → expect 200
- Use master key for auth
### 6. Check agent keys
- GET /key/list with master key → verify all 6 agents have keys
### 7. Check Grafana
- GET {{grafana_url}}/api/health → expect 200
### 8. Compile overall status
Determine overall_status from individual check results:
- "healthy" — all checks pass
- "degraded" — 1-2 non-critical checks fail
- "down" — critical checks fail
--- ---
--- ---