docs(litellm-health): single source of truth, gpu-vision alias, litellm-health as live owner #79
@@ -88,7 +88,7 @@ prose run memory-audit-maintenance memory_threshold=90 verify_configs=true
|
||||
prose run hermes-config-template agent_name=syslog-devops default_model=claude-sonnet-4
|
||||
|
||||
# Configure an agent with a different auxiliary model
|
||||
prose run hermes-config-template agent_name=syslog-code default_model=qwen3.6-27B-code auxiliary_model=gemma-4-12b
|
||||
prose run hermes-config-template agent_name=syslog-code default_model=qwen3.6-27B-code auxiliary_model=gpu-vision
|
||||
```
|
||||
|
||||
### Option B: Manual Execution
|
||||
@@ -116,7 +116,7 @@ Run on trigger or schedule. Maintain persistent world-model state across runs.
|
||||
| `zulip-health` | Zulip | Checks Zulip connectivity, message flow, and bot responsiveness. |
|
||||
| `zulip-mention-reliability` | Zulip | Diagnoses and fixes @mention detection issues in Zulip. |
|
||||
| `zulip-approval-fix` | Zulip | Fixes broken /approve and /deny slash commands for Hermes agents. |
|
||||
| `litellm-self-heal` | LiteLLM | Consolidated health check + self-healing for the full nginx → LiteLLM → GPU chain. Verifies 8 containers, 3 GPUs, model inference, and agent keys. Applies 9 remediation rules. (litellm-health merged into this contract 2026-07-09.) |
|
||||
| `litellm-self-heal` | LiteLLM | Applies remediation rules for LiteLLM stack failures detected by `litellm-health` (full nginx → LiteLLM → GPU chain). 9 remediation rules. |
|
||||
| `gpu-fleet` | GPU | Manages the GPU inference fleet: model deployment, registration, health checks, LiteLLM sync. |
|
||||
| `gpu-monitor` | GPU | Comprehensive GPU fleet monitor — polls sidecars, router, LiteLLM every 15s, renders SSE dashboard. |
|
||||
| `proxmox-monitor` | Infra | Proxmox cluster + Docker monitoring via the existing Grafana/Prometheus stack on CT 116. |
|
||||
@@ -149,7 +149,7 @@ Called on-demand as single-render tools.
|
||||
| Contract | Description |
|
||||
|---|---|
|
||||
| `litellm-api-keys` | Manages LiteLLM API keys for agent identity. Create, rotate, verify, and list agent keys. References gpu-fleet for current key inventory. |
|
||||
| `litellm-health` | ⚠️ **DEPRECATED** — consolidated into `litellm-self-heal` (2026-07-09). Retained for reference only. |
|
||||
| `litellm-health` | LiteLLM health check: public vs backend surfaces, CT 116 containers, GPU fleet, one model per GPU host, and agent keys. Owner of the probes; `litellm-self-heal` owns remediation. |
|
||||
| `infrastructure-monitoring` | Target-state for Prometheus + GPU exporters + Grafana. Core stack deployed, GPU exporters NOT live. |
|
||||
| `stirling-pdf-agent-access` | Documents the Stirling-PDF API access pattern for agents — global API key, 12 operations, curl examples. Agents use the `stirling-pdf-api` shared skill for templates. |
|
||||
| `hello-world` | Minimal test contract — verifies the OpenProse execution pipeline works. |
|
||||
|
||||
@@ -693,8 +693,8 @@ contracts:
|
||||
version: 1.0.0
|
||||
trigger:
|
||||
type: scheduled
|
||||
cadence: '*/10 * * * *'
|
||||
description: "Every 10 minutes \u2014 LiteLLM proxy health"
|
||||
cadence: '5 3,7,11,15,19,23 * * *'
|
||||
description: "4-hourly staggered dispatch via fm-send (run contract litellm-health)"
|
||||
cron_job_id: null
|
||||
execution:
|
||||
agent: abiba
|
||||
|
||||
@@ -2,6 +2,10 @@
|
||||
|
||||
Generated: 2026-07-13 20:59:18 ET
|
||||
|
||||
> **Point-in-time snapshot.** Schedules and cadences are authoritative in
|
||||
> `contract-registry.yaml`; any schedule quoted below may be stale. Do not use
|
||||
> this file as the source of truth for a contract's trigger.
|
||||
|
||||
---
|
||||
|
||||
## hermes-key-enforcement
|
||||
@@ -442,7 +446,7 @@ IMPORTANT: If the contract file does not exist in prose-contracts/main, report f
|
||||
|
||||
## litellm-health
|
||||
|
||||
**Category:** monitoring | **Domain:** litellm | **Owner:** abiba | **Schedule:** */10 * * * *
|
||||
**Category:** monitoring | **Domain:** litellm | **Owner:** abiba | **Schedule:** see contract-registry.yaml (authoritative)
|
||||
|
||||
```
|
||||
Contract Enforcement: litellm-health
|
||||
@@ -450,7 +454,7 @@ Contract Enforcement: litellm-health
|
||||
Category: monitoring
|
||||
Domain: litellm
|
||||
Owner: abiba
|
||||
Schedule: Every 10 minutes — LiteLLM proxy health
|
||||
Schedule: see contract-registry.yaml (authoritative)
|
||||
|
||||
This is a monitoring contract. Execute the monitoring checks defined in the contract. Report any deviations from expected state.
|
||||
|
||||
|
||||
@@ -13,7 +13,7 @@ the description:
|
||||
|
||||
1. **What system does this contract touch?** Name the hosts, CTs, containers,
|
||||
and services explicitly. "The inference fleet" is vague. "GPU .8 (RTX 3090,
|
||||
qwen), .110 (RTX 5070, gemma), .15 (Strix Halo, strix-moe), and LiteLLM on CT
|
||||
qwen), .110 (RTX 5070, gpu-vision), .15 (Strix Halo, strix-moe), and LiteLLM on CT
|
||||
116" is specific.
|
||||
|
||||
2. **Who runs this contract, and when?** State the agent, the trigger (cron,
|
||||
|
||||
+28
-56
@@ -6,9 +6,10 @@ description: >
|
||||
registration, health checks, LiteLLM sync, agent key management, GPU
|
||||
saturation watchdog, Prometheus/Grafana monitoring, and self-healing.
|
||||
UPDATED 2026-07-15: Stable role-based aliases introduced: strix-moe,
|
||||
gpu-dense, gpu-light. These never change — only the underlying model does.
|
||||
gpu-dense, gpu-vision (gpu-light was superseded by gpu-vision on 2026-09-12).
|
||||
These never change — only the underlying model does.
|
||||
Strix Halo: strix-moe → unsloth/Qwen3.6-35B-A3B-MTP (UD-Q4_K_M, 22GB).
|
||||
RTX 5070: gemma-4-12b Q4_K_M → IQ4_NL + MTP draft (122 tok/s, 2x faster).
|
||||
RTX 5070: gpu-vision — IQ4_NL + MTP draft (~122 tok/s, 2x faster).
|
||||
UPDATED 2026-07-17: Context reduced fleet-wide from 256K to 128K for stability.
|
||||
Strix Halo model swapped to qwen3.6-35B-udq4 (22GB, strix-moe alias).
|
||||
Instability observed near 100K at 256K (now all GPUs at 128K). 128K is the stable ceiling.
|
||||
@@ -24,7 +25,7 @@ triggers:
|
||||
|
||||
## Maintains
|
||||
|
||||
- gpu_roster: { models: map, hosts: map } — Single source of truth for all GPU models
|
||||
- gpu_roster: { models: map, hosts: map } — GPU host/model roster; the authoritative alias/weight/fallback registry is CT 116 `litellm_config.yaml`
|
||||
- router: { status: "healthy", roster_loaded: bool, models: array }
|
||||
- litellm: { status: "healthy", keys: array, models: array }
|
||||
- agent_keys: { agent: api_key } — All agent API keys registered in LiteLLM DB
|
||||
@@ -79,54 +80,28 @@ triggers:
|
||||
Agent configs, cron jobs, and workflows MUST use these aliases, never model-specific names.
|
||||
When a model is swapped on a GPU, ONLY the infrastructure layer changes — agent configs are untouched.
|
||||
|
||||
| Alias | GPU | Current Model | Will Route To |
|
||||
|-------|-----|---------------|---------------|
|
||||
| `strix-moe` | Strix Halo (.15) | qwen3.6-35B-udq4 | Whatever runs on Strix Halo |
|
||||
| `gpu-dense` | RTX 3090 (.8) | Qwen3.8-27B-Uncensored-Q4_K_M | Whatever runs on RTX 3090 |
|
||||
| `gpu-light` | RTX 5070 (.110) | gemma-4-12b | Whatever runs on RTX 5070 |
|
||||
| Alias | Serves | Where | Kind |
|
||||
|-------|--------|-------|------|
|
||||
| `gpu-dense` | heavy reasoning | RTX 3090 (192.168.68.8) | direct alias |
|
||||
| `gpu-vision` | vision / web extract / light tasks | RTX 5070 (192.168.68.110) | direct alias AND `syslog-auto` pool member |
|
||||
| `strix-moe` | compression (MoE) | Strix Halo (192.168.68.15) | direct alias |
|
||||
| `syslog-auto` | balanced default | weighted pool across the three GPU hosts | pool router |
|
||||
|
||||
**Backward compatibility**: Old model-specific names (qwen3.6-27B-code, gemma-4-12b, qwen3.5-9b-it) still work
|
||||
but are deprecated for agent configs. Only the stable aliases survive model swaps.
|
||||
Single source of truth for models, aliases, rpm caps, weights and fallback chains:
|
||||
CT 116 `/opt/inference-harness/litellm_config.yaml`. Do not duplicate those values in
|
||||
contracts — read them there.
|
||||
|
||||
## Current Model Assignments (2026-07-15)
|
||||
|
||||
| Model | GPU | Host | VRAM | Ctx | KV Cache | Parallel | Batch/Ubatch | Status |
|
||||
|-------|-----|------|------|-----|----------|----------|-------------|--------|
|
||||
**Backward compatibility**: Old model-specific names (qwen3.6-27B-code, qwen3.5-9b-it) still work
|
||||
but are deprecated for agent configs. Only the stable aliases survive model swaps. `gemma-4-12b`
|
||||
is retired and returns 400 `Invalid model name`.
|
||||
|
||||
## Routing Configuration (LiteLLM — July 2026)
|
||||
|
||||
### syslog-auto Weighted Pool (Direct GPU — bypasses router)
|
||||
|
||||
| Model | GPU | Weight | RPM Cap | Timeout |
|
||||
|-------|-----|--------|---------|---------|
|
||||
| Qwen3.8-27B-Uncensored-Q4_K_M | RTX 3090 (.8:8080) | **0.55** | 500 | **300s** |
|
||||
Model, alias, rpm/weight and fallback values are owned by CT 116
|
||||
`/opt/inference-harness/litellm_config.yaml` (see § Stable Role-Based Aliases above).
|
||||
|
||||
Note: All syslog-auto entries route directly to GPUs with `api_key: not-needed`. The router (port 9000) was decommissioned 2026-09-11 and is NOT in the inference path.
|
||||
|
||||
### Direct Model Endpoints
|
||||
|
||||
| Model | RPM Cap | Notes |
|
||||
|-------|---------|-------|
|
||||
|
||||
### Stable Aliases (for agent configs — never change)
|
||||
|
||||
| Alias | RPM Cap | Routes To | Purpose |
|
||||
|-------|---------|-----------|---------|
|
||||
| `strix-moe` | 40 | Strix Halo | Compression tasks (MoE models) |
|
||||
| `gpu-dense` | 500 | RTX 3090 | Heavy reasoning |
|
||||
| `gpu-light` | 500 | RTX 5070 | Vision, web extract, light tasks |
|
||||
|
||||
### Fallback Chains
|
||||
- gemma → qwen
|
||||
- qwen → gemma
|
||||
- strix-moe → qwen → gemma
|
||||
- syslog-auto → qwen → gemma → qwen3.6-35B-udq4
|
||||
|
||||
### Why Strix Halo RPM Is Capped
|
||||
- Direct (strix-moe): 40 RPM (tight) — Strix Halo is shared with compression tasks
|
||||
- Via syslog-auto: 60 RPM (moderate) — prevents flooding when multiple agents use syslog-auto simultaneously
|
||||
- Combined max: ~100 RPM across both paths — Strix Halo can sustain this at 80°C
|
||||
|
||||
## Operations
|
||||
|
||||
### add-model
|
||||
@@ -176,9 +151,7 @@ Show full fleet status: GPUs, models, VRAM, context windows, parallel slots, act
|
||||
2. Check llama-server processes: `ps aux | grep llama-server` on all 3 hosts
|
||||
3. Check LiteLLM: `curl http://192.168.68.116/health` (expect "I'm alive!")
|
||||
4. Check LiteLLM models: `curl -H "Authorization: Bearer $MASTER_KEY" http://192.168.68.116/v1/models`
|
||||
5. Check LiteLLM timeouts: `grep -n 'timeout:' /opt/inference-harness/litellm_config.yaml`
|
||||
- Qwen3.5-9B: 120s, qwen3.6-27B-code: 300s, Carnice-Qwen3.6-MoE-35B-A3B/strix-moe: 300s (strix-moe alias retained, legacy name qwen3.6-35B-udq4 deprecated)
|
||||
- global request_timeout: 300s, nginx proxy_read_timeout: 600s
|
||||
5. Check LiteLLM timeouts: `grep -n 'timeout:' /opt/inference-harness/litellm_config.yaml` — read the live values from the authority config; do not assert them from this contract.
|
||||
6. Check AMD metrics: `curl http://192.168.68.15:9400/metrics` (Radeon 8060S, util%, VRAM, temp, power)
|
||||
7. Check port conflicts: verify only one llama-server on :8080 per host
|
||||
8. Verify agent keys: 9 keys in LiteLLM DB (`GET /key/list`)
|
||||
@@ -243,6 +216,7 @@ If no SSH access, send Zulip DM via abiba-bot with vault update instructions.
|
||||
- **Alert migration**: All alerts now go to `#agent-hub` topics (`alerts-gpu`, `alerts-pm2`, `alerts-infra`) instead of DMs. Cross-agent visibility enabled.
|
||||
- **tok/s benchmarks**: Measured every 5 min via LiteLLM proxy. Baselines tracked with 30%/50% degradation thresholds.
|
||||
- **NetBird 502**: Tanko routes through NetBird for litellm.sysloggh.net. Use direct IP if NetBird down.
|
||||
- **Alias-retirement follow-up (2026-09-12)**: the agent-facing templates (`hermes-config-template.prose.md`, `hermes-agent-baseline.prose.md`), `litellm-api-keys.prose.md`, and the executable `audit-hermes-config.py` still reference the retired `gpu-light`/`gemma-4-12b` names, and `audit-hermes-config.py` Rule 8 currently fails a config whose vision/web_extract model is `gpu-vision`. These are tracked separately and are intentionally NOT updated in this change.
|
||||
|
||||
## GPU Inference Benchmarks (Current)
|
||||
|
||||
@@ -265,10 +239,10 @@ History stored at `/root/data/toks-history.json` with 7-day rolling window.
|
||||
### Stable Aliases — CRITICAL
|
||||
|
||||
All agent configs MUST use stable role-based aliases, never model-specific names:
|
||||
- `compression.model: strix-moe` (NOT `qwen3.6-35B-udq4`)
|
||||
- `auxiliary.vision.model: gpu-light` (NOT `gemma-4-12b`)
|
||||
- `delegation.model: gpu-dense` (NOT `qwen3.6-27B-code`)
|
||||
- `auxiliary.web_extract.model: gpu-light`
|
||||
- `compression.model: strix-moe`
|
||||
- `auxiliary.vision.model: gpu-vision`
|
||||
- `delegation.model: gpu-dense`
|
||||
- `auxiliary.web_extract.model: gpu-vision`
|
||||
|
||||
When the underlying model is swapped, only the LiteLLM config changes — agent configs are untouched.
|
||||
|
||||
@@ -277,7 +251,7 @@ When the underlying model is swapped, only the LiteLLM config changes — agent
|
||||
- **All agents**: 128K ceiling — stable margin. For >128K workloads, use external providers (deepseek)
|
||||
- Compression threshold 0.60: fires at ~77K (~51K headroom before 128K ceiling)
|
||||
- **Pi agents (Abiba)**: `compaction.reserveTokens: 52739` (≈60% of 128K)
|
||||
- Mumuni compression model alias: `strix-moe` with 300s timeout
|
||||
- Mumuni compression model alias: `strix-moe`
|
||||
|
||||
### Mumuni Agent Profile
|
||||
|
||||
@@ -285,12 +259,12 @@ Mumuni (kagentz CT105, 192.168.68.14 — migrated from CT100 2026-08-29) is the
|
||||
|
||||
| Setting | Value | Notes |
|
||||
|---------|-------|-------|
|
||||
| `model.default` | `syslog-auto` | Weighted pool (55% qwen, 30% strix, 15% gemma) |
|
||||
| `model.default` | `syslog-auto` | Balanced default (pool router) |
|
||||
| `model.provider` | `custom:litellm` | LiteLLM on CT116 |
|
||||
| `compression.model` | `strix-moe` | Stable alias — survives model swaps |
|
||||
| `aux.compression.model` | `strix-moe` | Compression auxiliary model |
|
||||
| `aux.vision.model` | `gpu-light` | Vision tasks (RTX 5070) |
|
||||
| `aux.web_extract.model` | `gpu-light` | Web extraction |
|
||||
| `aux.vision.model` | `gpu-vision` | Vision tasks (RTX 5070) |
|
||||
| `aux.web_extract.model` | `gpu-vision` | Web extraction |
|
||||
| `delegation.model` | `gpu-dense` | Sub-agent reasoning (RTX 3090) |
|
||||
| `context.max_context_window` | 131072 (128K) | Reduced from 256K 2026-07-17 — stable 128K ceiling |
|
||||
| `compression.threshold` | 0.60 | Triggers at ~77K (~60% of 128K) — optimized for 128K context |
|
||||
@@ -299,8 +273,6 @@ Mumuni (kagentz CT105, 192.168.68.14 — migrated from CT100 2026-08-29) is the
|
||||
| `memory.memory_char_limit` | 800 | Brief memory entries |
|
||||
| `personalities` | `creative` | Creative assistant personality |
|
||||
| Platforms | cli, homeassistant, signal, telegram, zulip | All Hermes platforms |
|
||||
| Main model timeout | 300s | LiteLLM global timeout |
|
||||
| Compression model timeout | 300s | strix-moe timeout increased from 120s |
|
||||
|
||||
### Agent Update Status (2026-07-15)
|
||||
|
||||
|
||||
@@ -79,8 +79,9 @@ proxy queuing.
|
||||
Send a real model name and use a real (probe-designated) key.
|
||||
- Probe timeout: 30s. A probe that takes longer than 30s IS the alert —
|
||||
report "backend slow (>30s)" rather than hanging.
|
||||
- Probe cadence: at most hourly. The 6h litellm-health cron cadence is the
|
||||
standard; sub-hourly synthetic traffic distorts latency baselines.
|
||||
- Probe cadence: at most hourly. The litellm-health cron cadence (authoritative
|
||||
trigger in contract-registry.yaml) is the standard; sub-hourly synthetic
|
||||
traffic distorts latency baselines.
|
||||
|
||||
### 5. Batch/benchmark jobs — schedule away from 04:00-07:00 EDT and chunk
|
||||
|
||||
|
||||
+52
-35
@@ -1,14 +1,7 @@
|
||||
---
|
||||
kind: function
|
||||
name: litellm-health
|
||||
status: deprecated
|
||||
deprecated_on: 2026-07-09
|
||||
replaced_by: litellm-self-heal.prose.md
|
||||
note: >
|
||||
Consolidated into litellm-self-heal.prose.md to eliminate duplication
|
||||
of architecture diagrams, GPU topology, timeout tables, and container
|
||||
lists. Health check is now § Health Check within litellm-self-heal.
|
||||
This file is retained for reference only — use litellm-self-heal instead.
|
||||
status: active
|
||||
description: >
|
||||
Verifies the LiteLLM inference stack health. Current architecture (2026-07-09):
|
||||
nginx:80 → LiteLLM:4000 → GPU(llama-server) via direct proxy.
|
||||
@@ -52,8 +45,7 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2)
|
||||
- Router REMOVED from request path — LiteLLM proxies directly to GPU
|
||||
- All GPUs at parallel 2 (was parallel 1)
|
||||
- NVIDIA context reduced 256K→128K to free VRAM — now the stable ceiling across all GPUs (2026-07-17)
|
||||
- LiteLLM timeouts tuned: gemma 25→120s, qwen 40→90s (SUPERSEDED 2026-07-16: qwen 300s, gemma 120s, strix 300s — see litellm-self-heal)
|
||||
- nginx proxy_read_timeout: 600s, LiteLLM request_timeout: 300s
|
||||
- Timeouts and fallback chains are config state — read them from CT 116 `/opt/inference-harness/litellm_config.yaml`; they are not duplicated here.
|
||||
|
||||
## Parameters
|
||||
|
||||
@@ -75,26 +67,27 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2)
|
||||
|
||||
- SSH key access to backend_host for container checks
|
||||
- Network access to public_url, auth_host, and gpu_dashboard_url
|
||||
- LiteLLM master key for key management endpoints
|
||||
- LiteLLM master key for admin endpoints only (`/key/list`, `/key/generate`, `/key/info`)
|
||||
- The dedicated `monitor` agent key on CT 116 at `/etc/litellm-monitor.env` (root-only 0600) for
|
||||
model inference checks — the master key must never be used for inference
|
||||
|
||||
## GPU Fleet Topology
|
||||
|
||||
| Host | IP | Hardware | Models Served | Engine | Context | Parallel |
|
||||
|------|-----|----------|---------------|--------|---------|----------|
|
||||
| llm-gpu | 192.168.68.8 | NVIDIA RTX 3090 (24 GB) | qwen3.6-27B-code | llama-server systemd | **128K** | 2 |
|
||||
| ocu-llm | 192.168.68.110 | NVIDIA RTX 5070 (12 GB) | gemma-4-12b | llama-server systemd | **128K** | 2 |
|
||||
| amdpve | 192.168.68.15 | AMD Strix Halo 64GB UMA | qwen3.6-35B-udq4 (LiteLLM alias: strix-moe) | llama-server systemd (Vulkan) | 128K | 2 |
|
||||
| Host | IP | Hardware | Role |
|
||||
|------|-----|----------|------|
|
||||
| llm-gpu | 192.168.68.8 | NVIDIA RTX 3090 (24 GB) | heavy reasoning (`gpu-dense`) |
|
||||
| ocu-llm | 192.168.68.110 | NVIDIA RTX 5070 (12 GB) | vision / web extract / light tasks (`gpu-vision`) |
|
||||
| amdpve | 192.168.68.15 | AMD Strix Halo 64GB UMA | compression (`strix-moe`) |
|
||||
|
||||
## Model Fallback Chains (LiteLLM)
|
||||
Single source of truth for models, aliases, rpm caps, weights and fallback chains:
|
||||
CT 116 `/opt/inference-harness/litellm_config.yaml`. Do not duplicate those values in
|
||||
contracts — read them there. Do not re-add retired names (`gemma-4-12b`, `gpu-light`,
|
||||
`crew-auto`).
|
||||
|
||||
| Primary | Timeout | Fallback | Timeout |
|
||||
|---------|---------|----------|---------|
|
||||
| qwen3.6-27B-code | 300s | gemma-4-12b | 120s |
|
||||
| gemma-4-12b | 120s | qwen3.6-27B-code | 300s |
|
||||
| qwen3.6-35B-udq4 / strix-moe | 300s | qwen → gemma | — |
|
||||
| syslog-auto (balanced) | 300s | qwen → gemma | — |
|
||||
|
||||
> Global: request_timeout=300s, nginx proxy_read_timeout=600s
|
||||
> Re-scope note (2026-09-12): the earlier plan to restate the live fallback chains and
|
||||
> per-model timeouts in this contract is intentionally superseded — that state is config,
|
||||
> and this contract points at the CT 116 config instead. Step 7 likewise authenticates with
|
||||
> the dedicated `monitor` key, not the master key, which is admin-only.
|
||||
|
||||
## Containers on CT 116
|
||||
|
||||
@@ -112,10 +105,22 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2)
|
||||
|
||||
1. **Read parameters** — Use provided values or defaults
|
||||
|
||||
2. **Check public endpoints**:
|
||||
- GET {{public_url}}/litellm/ui/ → expect 200 ("LiteLLM Dashboard")
|
||||
- GET {{public_url}}/litellm/docs → expect 200 ("LiteLLM API - Swagger UI")
|
||||
- GET {{public_url}}/ui/ and {{public_url}}/docs → expect 301 → /litellm/... → 200 (one-hop redirect helpers, added 2026-09-11)
|
||||
2. **Check the end-user surfaces** — the public edge and the backend edge serve the SAME
|
||||
app under DIFFERENT paths. They are not interchangeable, so every probe below must name
|
||||
the surface it targets. Never point a check at a path that only resolves on the other
|
||||
surface.
|
||||
|
||||
**Public edge** — `{{public_url}}` (https://litellm.sysloggh.net) serves the app at the
|
||||
ROOT; the `/litellm/` prefix does not exist there and 404s:
|
||||
- GET {{public_url}}/ui/ → expect 200 ("LiteLLM Dashboard")
|
||||
- GET {{public_url}}/docs → expect 200 ("LiteLLM API - Swagger UI")
|
||||
- GET {{public_url}}/litellm/ui/ and {{public_url}}/litellm/docs → expect 404 (not served on this edge)
|
||||
|
||||
**Backend edge** — `http://{{backend_host}}` (port 80) serves the app UNDER `/litellm/`:
|
||||
- GET http://{{backend_host}}/litellm/ui/ → expect 200 ("LiteLLM Dashboard")
|
||||
- GET http://{{backend_host}}/litellm/docs → expect 200 ("LiteLLM API - Swagger UI")
|
||||
- GET http://{{backend_host}}/ui/ and /docs → expect 301 → /litellm/... → 200 (one-hop redirect helpers,
|
||||
added 2026-09-11)
|
||||
|
||||
3. **Check LiteLLM health (no-auth)**:
|
||||
- GET http://{{backend_host}}/litellm/health/liveliness → expect 200
|
||||
@@ -136,14 +141,26 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2)
|
||||
- Verify GPUs reporting status "healthy"
|
||||
- Check alerts array for active warnings/critical
|
||||
|
||||
7. **Check model inference via LiteLLM** — Test each model:
|
||||
- POST /v1/chat/completions model=gemma-4-12b → expect 200
|
||||
- POST /v1/chat/completions model=qwen3.6-27B-code → expect 200
|
||||
- POST /v1/chat/completions model=strix-moe → expect 200
|
||||
- Use master key for auth
|
||||
7. **Check model inference via LiteLLM** — Test one model on each GPU host. The health
|
||||
check runs on the **backend edge**, not the public edge, so these paths carry the
|
||||
`/litellm/` prefix:
|
||||
- POST http://{{backend_host}}/litellm/v1/chat/completions model=qwen3.6-27B-code → expect 200 (RTX 3090, .8)
|
||||
- POST http://{{backend_host}}/litellm/v1/chat/completions model=gpu-vision → expect 200 (RTX 5070, .110)
|
||||
- POST http://{{backend_host}}/litellm/v1/chat/completions model=strix-moe → expect 200 (Strix Halo, .15)
|
||||
- Auth uses the dedicated `monitor` agent key, read on CT 116 from
|
||||
`/etc/litellm-monitor.env` (root-only 0600). Do NOT use the master key for inference —
|
||||
the master key is for admin endpoints only (`/key/list`, `/key/generate`, `/key/info`).
|
||||
- `gemma-4-12b` was retired and returns 400 `Invalid model name` — do not re-add it to
|
||||
this list. The RTX 5070 host now serves `gpu-vision`.
|
||||
- `/v1/models` is **key-scoped**: a model is only visible to keys allowed to use it, so
|
||||
the set depends on the key. Always state which key a model list was read with — a
|
||||
snapshot without its key is not evidence. This probe uses the `monitor` key on the
|
||||
backend surface (`http://{{backend_host}}/litellm/v1/models`). The authoritative model
|
||||
registry is CT 116 `/opt/inference-harness/litellm_config.yaml`; read it there rather
|
||||
than freezing a list here.
|
||||
|
||||
8. **Check agent keys**:
|
||||
- GET /key/list with master key → verify all 6 agents have keys
|
||||
- GET http://{{backend_host}}/litellm/key/list with master key (admin endpoint) → verify all 6 agents have keys
|
||||
|
||||
9. **Check Grafana**:
|
||||
- GET {{grafana_url}}/api/health → expect 200
|
||||
|
||||
+37
-73
@@ -11,22 +11,22 @@ note: >
|
||||
Script: `/opt/inference-harness/scripts/litellm-health-check.sh` on CT 116 (cron `0 */6 * * *`).
|
||||
Reports to /var/log/litellm/health-*.json and Gitea (SyslogSolution/health-logs).
|
||||
GPU monitoring integrated from gpu-monitor on .24:9100.
|
||||
|
||||
Consolidated from litellm-health + litellm-self-heal on 2026-07-09 to eliminate
|
||||
duplication of architecture diagrams, GPU topology, timeout tables, and container
|
||||
lists. Health check is now § Health Check within this contract.
|
||||
|
||||
|
||||
Health probes are owned by litellm-health.prose.md (dispatched as
|
||||
`run contract: litellm-health`). This contract owns remediation only — it does not
|
||||
re-specify the probes.
|
||||
|
||||
Source of truth for GPU topology and keys: gpu-fleet.prose.md
|
||||
Last verified: 2026-07-12
|
||||
description: >
|
||||
LiteLLM inference stack health monitoring + self-healing. Verifies the full
|
||||
nginx → LiteLLM → GPU chain, 12 containers on CT 116, 3 GPU hosts, model
|
||||
inference, and agent keys. Applies remediation rules for common failures.
|
||||
LiteLLM inference stack remediation. Applies remediation rules for failures detected
|
||||
by litellm-health.prose.md (nginx → LiteLLM → GPU chain, CT 116 containers, GPU hosts,
|
||||
model inference, and agent keys).
|
||||
Reports every action via Zulip DM and Gitea (SyslogSolution/health-logs).
|
||||
---
|
||||
---
|
||||
|
||||
# LiteLLM Operations — Health Check + Self-Heal
|
||||
# LiteLLM Operations — Self-Heal (Remediation)
|
||||
|
||||
## Architecture (v4.0.0 — Direct: nginx → LiteLLM → GPU)
|
||||
|
||||
@@ -58,41 +58,44 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2)
|
||||
|
||||
## GPU Fleet Topology
|
||||
|
||||
| Host | IP | Hardware | Models Served | Engine | Context | Parallel |
|
||||
|------|-----|----------|---------------|--------|---------|----------|
|
||||
| llm-gpu | 192.168.68.8 | NVIDIA RTX 3090 (24 GB) | qwen3.6-27B-code | llama-server systemd (`/home/llmuser/llama-wrapper.sh`, `-c 131072 --parallel 2 --ngl 99`) | **128K** | 2 |
|
||||
| ocu-llm | 192.168.68.110 | NVIDIA RTX 5070 (12 GB) | gemma-4-12b | llama-server systemd (`/home/llmuser/llama-wrapper.sh`, `--ctx-size 131072 --parallel 2`, IQ4_NL + MTP draft) | **128K** | 2 |
|
||||
| amdpve | 192.168.68.15 | AMD Strix Halo 64GB UMA | qwen3.6-35B-udq4 (LiteLLM alias: `strix-moe`) | llama-server systemd (Vulkan) | 128K | 2 |
|
||||
| Host | IP | Hardware | Role |
|
||||
|------|-----|----------|------|
|
||||
| llm-gpu | 192.168.68.8 | NVIDIA RTX 3090 (24 GB) | heavy reasoning (`gpu-dense`) |
|
||||
| ocu-llm | 192.168.68.110 | NVIDIA RTX 5070 (12 GB) | vision / web extract / light tasks (`gpu-vision`) |
|
||||
| amdpve | 192.168.68.15 | AMD Strix Halo 64GB UMA | compression (`strix-moe`) |
|
||||
|
||||
> Verified on ground 2026-07-16 via `curl /v1/models` on each host + `llama-wrapper.sh`. The AMD host's underlying model is `qwen3.6-35B-udq4`; LiteLLM exposes it under two `model_name`s: `qwen3.6-35B-udq4` and `strix-moe` (rpm 40). The legacy name `ornith-1.0-35b` does NOT exist in LiteLLM and must not be referenced.
|
||||
> Verified on the ground 2026-07-16. The legacy name `ornith-1.0-35b` does NOT exist in LiteLLM and must not be referenced.
|
||||
|
||||
## LiteLLM Model Surface (ground truth — `/opt/inference-harness/litellm_config.yaml` on CT 116)
|
||||
## LiteLLM Model Surface
|
||||
|
||||
`model_name`s served: `qwen3.6-27B-code`, `gemma-4-12b`, `qwen3.6-35B-udq4`, `strix-moe`, `gpu-dense`, `gpu-light`, `syslog-auto`, `crew-auto` (new 2026-08-20).
|
||||
Single source of truth for models, aliases, rpm caps, weights and fallback chains:
|
||||
CT 116 `/opt/inference-harness/litellm_config.yaml`. Do not duplicate those values in
|
||||
contracts — read them there. Do not re-add retired names (`gemma-4-12b`, `gpu-light`,
|
||||
`crew-auto`).
|
||||
|
||||
### Context Cap Split (2026-08-20)
|
||||
### Alias Surface
|
||||
|
||||
| Alias | Serves | Where | Kind |
|
||||
|-------|--------|-------|------|
|
||||
| `gpu-dense` | heavy reasoning | RTX 3090 (192.168.68.8) | direct alias |
|
||||
| `gpu-vision` | vision / web extract / light tasks | RTX 5070 (192.168.68.110) | direct alias AND `syslog-auto` pool member |
|
||||
| `strix-moe` | compression (MoE) | Strix Halo (192.168.68.15) | direct alias |
|
||||
| `syslog-auto` | balanced default | weighted pool across the three GPU hosts | pool router |
|
||||
|
||||
### Context Cap Split (2026-08-20; crew cap RETIRED)
|
||||
|
||||
- **Abiba (firstmate)**: 128K uncapped — unlimited context for primary workloads
|
||||
- **Hermes agents** (mumuni, tanko, koby, koonimo): 128K uncapped
|
||||
- **Crewmates** (ops, tune, verify, auth-keys, build): 64K capped — alias `crew-auto` enforces 64K limit
|
||||
- **Crewmates** (ops, tune, verify, auth-keys, build): the 64K cap was retired together with the `crew-auto` alias; NO context cap is currently in force.
|
||||
|
||||
Preferred implementation: uncap shared pool, add capped alias for crew-only.
|
||||
> The 64K crew cap was retired with `crew-auto` (2026-09-12). No limit is currently in force; reinstating one would need per-key model limits as a separate, deliberately-scoped change.
|
||||
|
||||
- `syslog-auto` is a weighted router model: qwen3.6-27B-code (0.55, rpm 500) + qwen3.6-35B-udq4 (0.30, rpm 60) + gemma-4-12b (0.15, rpm 200).
|
||||
- `gpu-dense` / `gpu-light` are high-rpm aliases (rpm 500) onto qwen3.6-27B-code / gemma-4-12b respectively.
|
||||
- Key scoping: agent keys are restricted to `['syslog-auto','qwen3.6-27B-code','gemma-4-12b','strix-moe','gpu-dense','gpu-light']`. As of 2026-07-16 the `baggy`/`koby`/`mumuni`/`abiba-pi` keys ALSO include `qwen3.6-35B-udq4`; `abiba-pi` additionally includes `deepseek-v4-pro` (cloud fallback). `kagenz0`/`koonimo`/`pi-agents-unified` have the standard 6 only. Agents should still use the stable alias `strix-moe` (not the raw `qwen3.6-35B-udq4`) so model swaps don't break them.
|
||||
- Key scoping: `/v1/models` is key-scoped, so the set a caller sees must be read with a named key rather than assumed — a monitor key, an agent key, and the master key can each return a different set. Agents should use the stable aliases (`strix-moe`, `gpu-vision`, `gpu-dense`) rather than raw model names, so model swaps don't break them.
|
||||
- **NEVER use `litellm_proxy_master_key` (the `sk-litellm-...` master key) for inference.** It is for admin endpoints only (`/key/list`, `/key/generate`, `/key/info`). All inference — agent traffic, health-check model tests, monitor scripts — uses agent-specific keys. The health-check script's model tests use a dedicated `monitor` agent key stored at `/etc/litellm-monitor.env` on CT 116 (root-only, `chmod 600`); `/key/list` is the only call that legitimately uses `$MASTER_KEY`.
|
||||
|
||||
## Model Fallback Chains (LiteLLM)
|
||||
|
||||
| Primary | Timeout | Fallback | Timeout |
|
||||
|---------|---------|----------|---------|
|
||||
| qwen3.6-27B-code | 300s | gemma-4-12b | 120s |
|
||||
| gemma-4-12b | 120s | qwen3.6-27B-code | 300s |
|
||||
| qwen3.6-35B-udq4 / strix-moe | 300s | qwen → gemma | — |
|
||||
| syslog-auto (balanced) | 300s | qwen → gemma | — |
|
||||
|
||||
> Global: request_timeout=300s, nginx proxy_read_timeout=600s
|
||||
> Re-scope note (2026-09-12): the fallback-chain and per-model timeout tables were removed
|
||||
> by the single-source-of-truth re-scope; read those values from the CT 116 config named
|
||||
> above rather than from this contract.
|
||||
|
||||
## Containers on CT 116
|
||||
|
||||
@@ -135,46 +138,7 @@ Preferred implementation: uncap shared pool, add capped alias for crew-only.
|
||||
|
||||
## Health Check
|
||||
|
||||
Run this first on every cycle. Results feed into remediation rules below.
|
||||
|
||||
### 1. Check public endpoints
|
||||
- GET {{public_url}}/litellm/ui/ → expect 200 ("LiteLLM Dashboard") # served directly by nginx
|
||||
- GET {{public_url}}/litellm/docs → expect 200 ("LiteLLM API - Swagger UI") # served directly by nginx
|
||||
- GET {{public_url}}/ui/ and {{public_url}}/docs → expect 301 → /litellm/... → 200 (one-hop redirect helpers, added 2026-09-11)
|
||||
|
||||
### 2. Check LiteLLM health (no-auth)
|
||||
- GET http://{{backend_host}}/litellm/health/liveliness → expect 200
|
||||
|
||||
### 3. Check backend container health
|
||||
- SSH to {{backend_host}} → `docker ps` → verify 12 containers healthy (11 harness + trove-agent-docker, added 2026-09-11)
|
||||
- Critical: harness-litellm, harness-nginx, harness-postgres
|
||||
- Monitoring: harness-redis, harness-dashboard, harness-grafana, harness-prometheus,
|
||||
harness-alertmanager, harness-zulip-bridge, harness-docker-stats, harness-pve-exporter,
|
||||
trove-agent-docker (ghcr.io/techdox/trove-agent-docker, added 2026-09-11)
|
||||
- Decommissioned 2026-09-11: harness-router (container, image and config removed)
|
||||
|
||||
### 4. Check GPU fleet health (via fleet dashboard)
|
||||
- GET {{gpu_dashboard_url}}/gpu-data → expect 200 with GPU metrics JSON
|
||||
- Verify GPUs reporting status "healthy"
|
||||
- Check alerts array for active warnings/critical
|
||||
|
||||
### 5. Check model inference via LiteLLM — test each model
|
||||
- POST /v1/chat/completions model=gemma-4-12b → expect 200
|
||||
- POST /v1/chat/completions model=qwen3.6-27B-code → expect 200
|
||||
- POST /v1/chat/completions model=strix-moe → expect 200
|
||||
- Use master key for auth
|
||||
|
||||
### 6. Check agent keys
|
||||
- GET /key/list with master key → verify all 6 agents have keys
|
||||
|
||||
### 7. Check Grafana
|
||||
- GET {{grafana_url}}/api/health → expect 200
|
||||
|
||||
### 8. Compile overall status
|
||||
Determine overall_status from individual check results:
|
||||
- "healthy" — all checks pass
|
||||
- "degraded" — 1-2 non-critical checks fail
|
||||
- "down" — critical checks fail
|
||||
Health probes are owned by `litellm-health.prose.md` (dispatched as `run contract: litellm-health`). This contract owns remediation only — it does not re-specify the probes.
|
||||
|
||||
---
|
||||
---
|
||||
|
||||
Reference in New Issue
Block a user