fix(router): decommission legacy GPU router across contracts and fleet scripts #74

Merged
kagentz-bot merged 2 commits from fix/decommission-router-20260911 into master 2026-09-11 19:19:14 +00:00
11 changed files with 96 additions and 98 deletions
+29 -33
View File
@@ -19,7 +19,7 @@ triggers:
- on model add/remove - on model add/remove
- on GPU health degradation - on GPU health degradation
- on agent key rotation - on agent key rotation
- on router restart (roster must be loaded) - on harness container restart (LiteLLM reloads its model list)
--- ---
## Maintains ## Maintains
@@ -36,46 +36,42 @@ triggers:
- prometheus: { status: "running", targets: 5 } — Scrapes GPU :9400 exporters + LiteLLM - prometheus: { status: "running", targets: 5 } — Scrapes GPU :9400 exporters + LiteLLM
- port_conflict_detection: { status: "active" } — All 3 GPU wrappers detect ghost processes before binding - port_conflict_detection: { status: "active" } — All 3 GPU wrappers detect ghost processes before binding
## Fleet Topology (Current — July 2026) ## Fleet Topology (Current — 2026-09-11, router decommissioned)
``` ```
┌──────────────────────────────────────────────────────────────────┐ ┌──────────────────────────────────────────────────────────────────┐
│ CT 116 (192.168.68.116) — Inference Harness Host │ │ CT 116 (192.168.68.116) — Inference Harness Host │
│ │ │ │
│ nginx:80 (entrypoint) │ │ nginx:80 (entrypoint) │
│ ├─ /v1/* → harness-litellm:4000 (API requests) │ │ ├─ /v1/* → harness-litellm:4000 (API requests) │
│ ├─ /admin/* → harness-litellm:4000 (admin endpoints) │ │ ├─ /admin/* → harness-litellm:4000 (admin endpoints) │
│ ├─ /dashboard/ → harness-dashboard:3000 (harness UI) │ │ ├─ /dashboard/ → harness-dashboard:3000 (harness UI) │
│ ├─ /litellm/* → harness-litellm:4000 (LiteLLM UI + API) │ │ ├─ /litellm/* → harness-litellm:4000 (LiteLLM UI + API) │
│ ├─ /health/* → harness-litellm:4000 (health probes) │ │ ├─ /health/* → harness-litellm:4000 (health probes) │
│ └─ /gpu/* → 192.168.68.24:9100 (fleet monitor) │ │ └─ /gpu/* → 192.168.68.24:9100 (fleet monitor) │
│ │ │ │
│ Containers: │ │ Containers: │
│ ┌──────────┐ ┌──────────┐ ┌──────────┐ ┌──────────┐ │ │ ┌──────────┐ ┌──────────┐ ┌──────────┐ │
│ │ LiteLLM │ │ Router │ │Dashboard │ │ Grafana │ │ │ │ LiteLLM │ │Dashboard │ │ Grafana │ │
│ │ :4000 │ │ :9000 │ │ :3000 │ │ :3000 │ │ │ │ :4000 │ │ :3000 │ │ :3001 │ │
│ │ keys+sync│ │deprecated│ │ harness │ │ Prometheus│ │ │ │ keys+sync│ │ harness │ │Prometheus│ │
│ │ fallback │ │not in │ │ UI │ │ data src │ │ │ │ fallback │ │ UI │ │ data src │ │
│ └──────────┘ └───┬──────┘ └──────────┘ └──────────┘ │ │ └──────────┘ └──────────┘ └──────────┘ │
│ │ │ │ │
│ ┌──────────┐ ┌──────────┐ ┌──────────┐ │ │ ┌──────────┐ ┌──────────┐ ┌──────────┐ │
│ │PostgreSQL│ │ Redis │ │Prometheus│ │ │ │PostgreSQL│ │ Redis │ │Prometheus│ │
│ │ :5432 │ │ :6379 │ │ :9090 │ │ │ │ :5432 │ │ :6379 │ │ :9090 │ │
│ └──────────┘ └──────────┘ └──────────┘ │ │ └──────────┘ └──────────┘ └──────────┘ │
└────────────────────┼─────────────────────────────────────────────┘ └───────┼──────────────────────────────────────────────────────────┘
│ │
┌───────────────┼───────────────┬──────────────────┐ ┌─┴─────────────┬───────────────┬───────────────┐
│ │ │ │ │ │ │ │
┌────▼─────┐ ┌──────▼──────┐ ┌────▼──────┐ ┌───────▼──────┐ ┌─────▼──────┐ ┌─────▼──────┐ ┌─────▼──────┐ ┌─────▼──────┐
│ CT 8 │ │ CT 110 │ │ CT 15 │ │ pi (.24) │ │ CT 8 │ │ CT 110 │ │ CT 15 │ │ pi (.24) │
│ RTX 3090 │ │ RTX 5070 │ │ Strix Halo│ │ GPU Monitor │ │ RTX 3090 │ │ RTX 5070 │ │ Strix Halo │ │ GPU Monitor│
│ 24GB │ │ 12GB │ │ 64GB UMA │ │ :9100 │ │ 24GB │ │ 12GB │ │ 64GB UMA │ │ :9100 │
│ 128K ctx │ │ 128K ctx │ │ 128K ctx │ │ Watchdog │ │ :8080 │ │ :9400 exp │ │ :9400 exp │ │ :9401 │
│ qwen3.6 │ │ Qwen3.5-9B │ │ qwen3.5 │ │ Prometheus │ └────────────┘ └────────────┘ └────────────┘ └────────────┘
│ 27B-code │ │ :8080 │ │ -35B-udq4 │ │ exporter │
│ :8080 │ │ :9400 (exp) │ │ :9400(exp)│ │ :9401 │
│ :9400 │ └─────────────┘ └───────────┘ └──────────────┘
└──────────┘
``` ```
## Stable Role-Based Aliases (Introduced 2026-07-15) ## Stable Role-Based Aliases (Introduced 2026-07-15)
@@ -105,7 +101,7 @@ but are deprecated for agent configs. Only the stable aliases survive model swap
|-------|-----|--------|---------|---------| |-------|-----|--------|---------|---------|
| Qwen3.8-27B-Uncensored-Q4_K_M | RTX 3090 (.8:8080) | **0.55** | 500 | **300s** | | Qwen3.8-27B-Uncensored-Q4_K_M | RTX 3090 (.8:8080) | **0.55** | 500 | **300s** |
Note: All syslog-auto entries route directly to GPUs with `api_key: not-needed`. The router (port 9000) is NOT in the inference path. Note: All syslog-auto entries route directly to GPUs with `api_key: not-needed`. The router (port 9000) was decommissioned 2026-09-11 and is NOT in the inference path.
### Direct Model Endpoints ### Direct Model Endpoints
@@ -153,7 +149,7 @@ Note: All syslog-auto entries route directly to GPUs with `api_key: not-needed`.
6. Cleanup model files (optional) 6. Cleanup model files (optional)
### heal ### heal
1. Check all GPUs via router internal `:9000/health/unified` 1. Check all GPUs via gpu-monitor `{{gpu_dashboard_url}}/gpu-data`
2. Check LiteLLM health via nginx `:80/litellm/health/liveliness` 2. Check LiteLLM health via nginx `:80/litellm/health/liveliness`
3. Reset stuck circuit breakers if idle (Redis) 3. Reset stuck circuit breakers if idle (Redis)
4. Restart dead llama-server instances via SSH 4. Restart dead llama-server instances via SSH
@@ -161,7 +157,7 @@ Note: All syslog-auto entries route directly to GPUs with `api_key: not-needed`.
6. Verify GPU monitor server is running on pi (:9100) 6. Verify GPU monitor server is running on pi (:9100)
7. Verify watchdog is running on pi 7. Verify watchdog is running on pi
8. Restart router if roster not loaded (check logs for STARTUP ROSTER) 8. Restart router if roster not loaded (check logs for STARTUP ROSTER)
9. Reload roster via `POST :9000/admin/roster/reload` if available 9. Router roster reload — REMOVED 2026-09-11 (router decommissioned; LiteLLM fallbacks handle routing)
### sync-keys ### sync-keys
1. List all agent keys in LiteLLM DB via `GET /key/list` 1. List all agent keys in LiteLLM DB via `GET /key/list`
@@ -258,7 +254,7 @@ If no SSH access, send Zulip DM via abiba-bot with vault update instructions.
Benchmarks from 2026-07-17. Strix Halo model: qwen3.6-35B-udq4. RTX 5070 MTP provides 2.7x speedup over pre-upgrade 70 tok/s. Benchmarks from 2026-07-17. Strix Halo model: qwen3.6-35B-udq4. RTX 5070 MTP provides 2.7x speedup over pre-upgrade 70 tok/s.
All 3 GPUs now at 128K context (2026-07-17, reduced from 256K for stability). All 3 GPUs now at 128K context (2026-07-17, reduced from 256K for stability).
Benchmarks run through LiteLLM proxy (192.168.68.116:4001) every 5 minutes. Benchmarks run through LiteLLM proxy (192.168.68.116:4000) every 5 minutes.
Degradation alerts fire at 30% (warning) and 50% (critical) below baseline. Degradation alerts fire at 30% (warning) and 50% (critical) below baseline.
History stored at `/root/data/toks-history.json` with 7-day rolling window. History stored at `/root/data/toks-history.json` with 7-day rolling window.
+17 -17
View File
@@ -2,7 +2,7 @@
kind: responsibility kind: responsibility
name: gpu-monitor name: gpu-monitor
description: > description: >
Comprehensive GPU fleet monitor — polls every subsystem (sidecars, router, Comprehensive GPU fleet monitor — polls every subsystem (sidecars,
LiteLLM, Strix Halo, dashboard) every 15s, renders a live HTML dashboard, LiteLLM, Strix Halo, dashboard) every 15s, renders a live HTML dashboard,
checks alert thresholds, and exposes a JSON API for downstream consumers. checks alert thresholds, and exposes a JSON API for downstream consumers.
agent: abiba agent: abiba
@@ -46,19 +46,19 @@ and `http://192.168.68.110/health` answer `000`. Never use a bare-port-80 probe
as a GPU liveness signal: on 2026-09-09 that produced three false as a GPU liveness signal: on 2026-09-09 that produced three false
`DEGRADED — GPU-rtx3090 000, GPU-rtx5070 000` rounds while `DEGRADED — GPU-rtx3090 000, GPU-rtx5070 000` rounds while
`http://192.168.68.8:8080/health` and `http://192.168.68.110:8080/health` `http://192.168.68.8:8080/health` and `http://192.168.68.110:8080/health`
answered `200`. Port 80 is valid only on the router (.116), never on a GPU host. answered `200`. Port 80 is valid only on the harness host (.116), never on a GPU host.
### Subsystems Polled ### Subsystems Polled
| Subsystem | Endpoint | Frequency | Metrics | | Subsystem | Endpoint | Frequency | Metrics |
|-----------|----------|-----------|---------| |-----------|----------|-----------|---------|
| GPU Status (all, via router) | `http://192.168.68.116/health/unified` | 15s | models, CB, scores, GPU status (router probes each GPU /health directly) — `301` → `/gpu/gpu-data` is **alive** | | GPU Status (all, via fleet API) | `http://192.168.68.116/gpu/gpu-data` | 15s | models, CB, scores, GPU status from gpu-monitor on .24:9100 |
| GPU .8 (RTX 3090) health | `http://192.168.68.8:8080/health` | 15s | direct liveness fallback — **:8080 ONLY, never bare port 80** | | GPU .8 (RTX 3090) health | `http://192.168.68.8:8080/health` | 15s | direct liveness fallback — **:8080 ONLY, never bare port 80** |
| GPU .110 (RTX 5070) health | `http://192.168.68.110:8080/health` | 15s | direct liveness fallback — **:8080 ONLY, never bare port 80** | | GPU .110 (RTX 5070) health | `http://192.168.68.110:8080/health` | 15s | direct liveness fallback — **:8080 ONLY, never bare port 80** |
| Router (unified) | `http://192.168.68.116/health/unified` | 15s | models, CB, scores, GPU status (`301` → `/gpu/gpu-data` = alive) | | Fleet (unified) | `http://192.168.68.116/health/unified` | 15s | nginx `301` → `/gpu/gpu-data` served by gpu-monitor = alive (router decommissioned 2026-09-11) |
| Router (basic) | `http://192.168.68.116/health` | 15s | basic aliveness | | Harness (basic) | `http://192.168.68.116/health` | 15s | nginx → LiteLLM `/health/liveliness` |
| LiteLLM | `http://192.168.68.116/litellm/health` | 15s | proxy health, model count | | LiteLLM | `http://192.168.68.116/litellm/health` | 15s | proxy health, model count |
| Strix Halo | `http://192.168.68.116/health/unified` (router) | 15s | Strix Halo status via router — cannot poll .15:8080 directly (firewalled to .116 only) | | Strix Halo | `http://192.168.68.116/health/unified` (nginx → fleet API) | 15s | Strix Halo status via gpu-monitor — cannot poll .15:8080 directly (firewalled to .116 only) |
| Dashboard | `http://192.168.68.116/dashboard/` | 15s | harness-dashboard aliveness | | Dashboard | `http://192.168.68.116/dashboard/` | 15s | harness-dashboard aliveness |
### Alert Delivery ### Alert Delivery
@@ -75,8 +75,8 @@ This replaces the previous DM-only delivery. All agents on the mesh can see and
### Liveness rule (scoped) ### Liveness rule (scoped)
The any-HTTP-response rule applies ONLY to redirect/auth-gated liveness The any-HTTP-response rule applies ONLY to redirect/auth-gated liveness
endpoints, where any HTTP answer proves a listener is up. Applied here: the endpoints, where any HTTP answer proves a listener is up. Applied here: nginx's
router's `/health/unified` answers `301 Moved Permanently` → `/gpu/gpu-data` `/health/unified` answers `301 Moved Permanently` → `/gpu/gpu-data`
(the same payload) and LiteLLM's `/litellm/health` answers `301` → (the same payload) and LiteLLM's `/litellm/health` answers `301` →
`/litellm/health/liveliness`. For those endpoints a probe is **ALIVE** on `/litellm/health/liveliness`. For those endpoints a probe is **ALIVE** on
**ANY** HTTP status — `3xx` redirects and `401`/`403` auth challenges included — **ANY** HTTP status — `3xx` redirects and `401`/`403` auth challenges included —
@@ -84,8 +84,8 @@ and **DOWN = connection refused (`000`) or timeout only**. Same scoped rule as
zulip-health (Tanko) and infrastructure-monitoring. zulip-health (Tanko) and infrastructure-monitoring.
Probes whose success condition is specifically a bare `200` are NOT covered by Probes whose success condition is specifically a bare `200` are NOT covered by
the any-HTTP rule. On those — the GPU `:8080/health` endpoints, the router the any-HTTP rule. On those — the GPU `:8080/health` endpoints, nginx `/health`,
`/health`, and the dashboard — an unexpected status (`401`/`403`, `5xx`, or and the dashboard — an unexpected status (`401`/`403`, `5xx`, or
anything other than the expected `200`) is an **ALERT**, not "alive". anything other than the expected `200`) is an **ALERT**, not "alive".
| Metric | Warning | Critical | | Metric | Warning | Critical |
@@ -93,7 +93,7 @@ anything other than the expected `200`) is an **ALERT**, not "alive".
| GPU Temp | >80°C | >90°C | | GPU Temp | >80°C | >90°C |
| VRAM Usage | >90% | >95% | | VRAM Usage | >90% | >95% |
| GPU Util | >95% | >98% | | GPU Util | >95% | >98% |
| Sidecar Unreachable | — | info (sidecars not deployed — use router /health/unified) | | Sidecar Unreachable | — | info (sidecars not deployed — use gpu-monitor /gpu-data) |
| Model Down | — | critical (circuit breaker open) | | Model Down | — | critical (circuit breaker open) |
### JSON API Response Schema (/gpu-data) ### JSON API Response Schema (/gpu-data)
@@ -174,13 +174,13 @@ Open `http://localhost:9100/` in browser — Live HTML dashboard
pkill -f gpu-monitor-server.py pkill -f gpu-monitor-server.py
python3 /root/scripts/gpu-monitor-server.py & python3 /root/scripts/gpu-monitor-server.py &
``` ```
Or via PM2: `pm2 restart gpu-monitor` Managed by systemd (verified 2026-09-11): `systemctl restart gpu-monitor`
### check-router ### check-fleet
The router health is accessed through nginx on port 80 on the **router** Fleet health is accessed through nginx on port 80 on the harness host
(.116) — NOT port 9000 directly, and NOT bare port 80 on a GPU host. (.116) — NOT port 9000 (router decommissioned 2026-09-11), and NOT bare
`curl http://192.168.68.116/health/unified` — Router unified health via nginx proxy; answers `301` → `/gpu/gpu-data` (same payload) = alive port 80 on a GPU host.
`curl http://192.168.68.116:9000/health/unified` — ❌ WILL FAIL (port bound to 127.0.0.1 only) `curl http://192.168.68.116/health/unified` — nginx answers `301` → `/gpu/gpu-data` (fleet monitor payload) = alive
`curl http://192.168.68.8/health` — ❌ NEVER USE (GPU host, no port-80 listener → false `000`/DEGRADED) `curl http://192.168.68.8/health` — ❌ NEVER USE (GPU host, no port-80 listener → false `000`/DEGRADED)
## Configuration Files ## Configuration Files
+5 -5
View File
@@ -9,7 +9,7 @@ description: >
(v2.1.0) with active remediation rules, Prometheus metrics consumption, (v2.1.0) with active remediation rules, Prometheus metrics consumption,
VRAM trend analysis, and predictive alerting. VRAM trend analysis, and predictive alerting.
UPDATED 2026-07-18: Model assignments synced to 2026-07-17 swaps. UPDATED 2026-07-18: Model assignments synced to 2026-07-17 swaps.
Router (port 9000) references replaced with direct GPU routing. Router (port 9000) DECOMMISSIONED 2026-09-11; references replaced with direct GPU routing.
Benchmark baselines refreshed to live values. Benchmark baselines refreshed to live values.
Prometheus exporters removed — not deployed; fall back to direct sidecar probes. Prometheus exporters removed — not deployed; fall back to direct sidecar probes.
Stable role-based aliases (strix-moe, gpu-dense, gpu-light) from gpu-fleet. Stable role-based aliases (strix-moe, gpu-dense, gpu-light) from gpu-fleet.
@@ -51,7 +51,7 @@ depends_on:
| `strix-moe` | Strix Halo 64GB | ct15 (.15:8080) | qwen3.6-35B-udq4 | ~10/64GB (16%) | 128K | 62.9 | Compression, summarization, long docs | | `strix-moe` | Strix Halo 64GB | ct15 (.15:8080) | qwen3.6-35B-udq4 | ~10/64GB (16%) | 128K | 62.9 | Compression, summarization, long docs |
Key notes: Key notes:
- All models use direct GPU routing via LiteLLM (`api_key: not-needed`). Router (port 9000) is deprecated and NOT in the inference path. - All models use direct GPU routing via LiteLLM (`api_key: not-needed`). Router (port 9000) was decommissioned 2026-09-11 and is NOT in the inference path.
- Stable aliases (gpu-dense, gpu-light, strix-moe) from gpu-fleet are the canonical names for agent configs. Model-specific names still work but are deprecated. - Stable aliases (gpu-dense, gpu-light, strix-moe) from gpu-fleet are the canonical names for agent configs. Model-specific names still work but are deprecated.
- RTX 5070 tok/s is ~145 for Qwen3.5-9B — gpu-light is the fastest endpoint. Route vision/web/light work there first. NOTE: Qwen3.5-9B is multimodal (image+text), NOT text-only like gemma-4-12b was. - RTX 5070 tok/s is ~145 for Qwen3.5-9B — gpu-light is the fastest endpoint. Route vision/web/light work there first. NOTE: Qwen3.5-9B is multimodal (image+text), NOT text-only like gemma-4-12b was.
- Strix Halo is 62.9 tok/s (89% of 70.5 baseline) — below optimal but stable. Check for competing workloads. - Strix Halo is 62.9 tok/s (89% of 70.5 baseline) — below optimal but stable. Check for competing workloads.
@@ -104,10 +104,10 @@ Key notes:
### Rule 5: Circuit Breaker Stuck Open ### Rule 5: Circuit Breaker Stuck Open
- **Detect**: Circuit breaker open >10 minutes with GPU reporting healthy - **Detect**: Circuit breaker open >10 minutes with GPU reporting healthy
- **Note**: Router (port 9000) is deprecated. If circuit breakers are reported by gpu-monitor, they come from LiteLLM's internal tracking, not the old router. - **Note**: Router (port 9000) was decommissioned 2026-09-11. If circuit breakers are reported by gpu-monitor, they come from LiteLLM's internal tracking, not the old router.
- **Fix**: - **Fix**:
1. Verify GPU /health returns 200 on direct port (:8080) 1. Verify GPU /health returns 200 on direct port (:8080)
2. If GPU healthy, alert but do NOT reset via router API (deprecated) 2. If GPU healthy, alert but do NOT reset via router API (decommissioned 2026-09-11)
3. Check LiteLLM health directly: http://192.168.68.116/litellm/health/liveliness 3. Check LiteLLM health directly: http://192.168.68.116/litellm/health/liveliness
4. Restart LiteLLM container on CT 116 if circuit breakers are stuck 4. Restart LiteLLM container on CT 116 if circuit breakers are stuck
- **Verify**: LiteLLM returns healthy, circuit breaker clears within 60s - **Verify**: LiteLLM returns healthy, circuit breaker clears within 60s
@@ -276,7 +276,7 @@ Pushed to `SyslogSolution/health-logs/gpu/{run_id}.json` — versioned, searchab
2. **Model restart**: ✅ Only if >50% failure rate over 60s + 30s grace period. Not on single stuck request. 2. **Model restart**: ✅ Only if >50% failure rate over 60s + 30s grace period. Not on single stuck request.
3. **Strix direct access**: ✅ Open firewall .15:8080 → .24 for direct health probe + restart. 3. **Strix direct access**: ✅ Open firewall .15:8080 → .24 for direct health probe + restart.
4. **VRAM thresholds**: Tiered — **300MB/h** (RTX 3090), **300MB/h** (RTX 5070), 200MB/h (Strix). Previous values (100/50) were too sensitive; raised 2026-07-18 based on operational data. 4. **VRAM thresholds**: Tiered — **300MB/h** (RTX 3090), **300MB/h** (RTX 5070), 200MB/h (Strix). Previous values (100/50) were too sensitive; raised 2026-07-18 based on operational data.
5. **CB auto-reset**: ✅ Router deprecated — circuit breakers go through LiteLLM health check + container restart if needed. No per-GPU auto-reset. 5. **CB auto-reset**: ✅ Router decommissioned 2026-09-11 — circuit breakers go through LiteLLM health check + container restart if needed. No per-GPU auto-reset.
6. **Benchmark baseline**: Rolling 30-day average, recalculated weekly. Current baselines live in gpu-monitor. 6. **Benchmark baseline**: Rolling 30-day average, recalculated weekly. Current baselines live in gpu-monitor.
7. **Predictive alerts**: Two-tier — warn at >70°C+rising (>2°C/min), critical at >80°C+rising. 7. **Predictive alerts**: Two-tier — warn at >70°C+rising (>2°C/min), critical at >80°C+rising.
8. **Prometheus**: ❌ Not deployed. Use direct sidecar probes (:8080/health) and gpu-monitor API. Prometheus integration deferred until exporters are running on GPU hosts. 8. **Prometheus**: ❌ Not deployed. Use direct sidecar probes (:8080/health) and gpu-monitor API. Prometheus integration deferred until exporters are running on GPU hosts.
+1 -1
View File
@@ -39,7 +39,7 @@ runtime via `infisical run --` wrapper. Plaintext keys removed from this baselin
``` ```
Infisical vault → infisical run -- hermes gateway → LITELLM_API_KEY (runtime) Infisical vault → infisical run -- hermes gateway → LITELLM_API_KEY (runtime)
↓ ↓
Agent (systemd) → LITELLM_API_KEY → LiteLLM (:116/v1) → Router (:9000) → GPU (llama-server) Agent (systemd) → LITELLM_API_KEY → LiteLLM (:116/v1) → GPU (llama-server)
└── Key DB (Postgres) └── Key DB (Postgres)
``` ```
+3 -7
View File
@@ -199,15 +199,14 @@ description: >
### Ecosystem B: CT 116 syslog-api (192.168.68.116) ### Ecosystem B: CT 116 syslog-api (192.168.68.116)
8 containers in inference-harness stack: 11 containers in inference-harness stack (verified live 2026-09-11; LiteLLM upgraded 1.90.0-rc.1 -> 1.99.1):
| Container | Image | Port | Role | | Container | Image | Port | Role |
|-----------|-------|------|------| |-----------|-------|------|------|
| harness-litellm | berriai/litellm:1.90.0-rc.1 | :4000→:4001 | API proxy, key mgmt, fallbacks | | harness-litellm | ghcr.io/docker.litellm.ai/berriai/litellm:1.99.1 | :4000→:4000 | API proxy, key mgmt, fallbacks |
| harness-router | inference-harness-router | :9000 (127.0.0.1) | GPU routing, slot booking, CB |
| harness-nginx | nginx:alpine | :80 | Entrypoint, /v1→LiteLLM, /dashboard/ | | harness-nginx | nginx:alpine | :80 | Entrypoint, /v1→LiteLLM, /dashboard/ |
| harness-postgres | postgres:16-alpine | :5432 | LiteLLM DB (keys, spend, config) | | harness-postgres | postgres:16-alpine | :5432 | LiteLLM DB (keys, spend, config) |
| harness-redis | redis:7-alpine | :6379 | Router slots, circuit breakers | | harness-redis | redis:7-alpine | :6379 | LiteLLM cache + rate-limit state |
| harness-dashboard | inference-harness-dashboard | :3000 | SyslogAI Harness UI | | harness-dashboard | inference-harness-dashboard | :3000 | SyslogAI Harness UI |
| harness-grafana | grafana/grafana | :3000→:3001 (direct LAN, not behind nginx) | GPU + Proxmox + Docker dashboards | | harness-grafana | grafana/grafana | :3000→:3001 (direct LAN, not behind nginx) | GPU + Proxmox + Docker dashboards |
| harness-prometheus | prom/prometheus | :9090 | Metrics scraper, 6 jobs | | harness-prometheus | prom/prometheus | :9090 | Metrics scraper, 6 jobs |
@@ -588,9 +587,6 @@ curl -s -H "Authorization: Bearer $(infisical secrets get LITELLM_MASTER_KEY --p
# Storage check # Storage check
ssh root@192.168.68.7 "df -h /media/storage /media/mediastore" ssh root@192.168.68.7 "df -h /media/storage /media/mediastore"
# Router roster reload (if needed)
curl -s -X POST http://192.168.68.116:9000/admin/roster/reload \
-H "Authorization: Bearer sk-admin-ee09fffd04978b61a1569ac670c68814"
# Restart stuck GPU (saturation watchdog alternative) # Restart stuck GPU (saturation watchdog alternative)
ssh root@192.168.68.8 "systemctl restart llama-server" ssh root@192.168.68.8 "systemctl restart llama-server"
+2 -2
View File
@@ -173,7 +173,7 @@ curl -s http://192.168.68.116:3001/api/health | jq '{status, version}'
# Expected: {"status":"ok","version":"..."} # Expected: {"status":"ok","version":"..."}
# LiteLLM metrics (Prometheus endpoint) # LiteLLM metrics (Prometheus endpoint)
curl -s http://192.168.68.116:4001/metrics | head -20 curl -s http://192.168.68.116:4000/metrics | head -20
# Expected: Prometheus-formatted metrics output # Expected: Prometheus-formatted metrics output
``` ```
@@ -242,5 +242,5 @@ curl -s http://192.168.68.116:9090/api/v1/targets
curl -s http://192.168.68.116:3001/api/health curl -s http://192.168.68.116:3001/api/health
# LiteLLM metrics (already live) # LiteLLM metrics (already live)
curl -s http://192.168.68.116:4001/metrics | head -20 curl -s http://192.168.68.116:4000/metrics | head -20
``` ```
+1 -1
View File
@@ -206,7 +206,7 @@ mcp_servers:
| Key | MCP Access | | Key | MCP Access |
|-----|-----------| |-----|-----------|
| Master key | ✅ Full — 90 tools (vault-injected) | | Master key | ✅ Full — 90 tools (vault-injected) |
| Agent keys (mumuni, tanko, etc.) | ❌ Per-key grants not supported in v1.90.0-rc.1 | | Agent keys (mumuni, tanko, etc.) | ❌ Per-key grants not supported in v1.99.1 |
### Known Limitations ### Known Limitations
- Per-key MCP server grants not functional — only master key has access - Per-key MCP server grants not functional — only master key has access
+12 -13
View File
@@ -12,8 +12,8 @@ note: >
description: > description: >
Verifies the LiteLLM inference stack health. Current architecture (2026-07-09): Verifies the LiteLLM inference stack health. Current architecture (2026-07-09):
nginx:80 → LiteLLM:4000 → GPU(llama-server) via direct proxy. nginx:80 → LiteLLM:4000 → GPU(llama-server) via direct proxy.
Router (harness-router :9000) is DEPRECATED — container still runs but Router (harness-router :9000) was DECOMMISSIONED 2026-09-11 (container,
is not in the request path. GPU monitoring via Prometheus/Grafana and image and config removed). GPU monitoring via Prometheus/Grafana and
fleet dashboard (gpu-monitor :9100). fleet dashboard (gpu-monitor :9100).
Designed as a reusable contract for any Syslog agent. Designed as a reusable contract for any Syslog agent.
@@ -42,8 +42,8 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2)
│ │
Grafana :3001 Grafana :3001
harness-router :9000 — DEPRECATED, container still runs but harness-router :9000 — DECOMMISSIONED 2026-09-11 (container, image
NOT in request path. nginx routes /v1 → LiteLLM directly. and config removed). nginx routes /v1 → LiteLLM directly.
Router slot booking + circuit breakers replaced by Router slot booking + circuit breakers replaced by
LiteLLM native fallbacks + timeouts. LiteLLM native fallbacks + timeouts.
``` ```
@@ -100,8 +100,7 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2)
| Container | Image | Port | Health Check | | Container | Image | Port | Health Check |
|-----------|-------|------|-------------| |-----------|-------|------|-------------|
| harness-litellm | berriai/litellm:1.90.0-rc.1 | :4000→:4001 | /health/liveliness | | harness-litellm | ghcr.io/docker.litellm.ai/berriai/litellm:1.99.1 | :4000→:4000 | /health/liveliness |
| harness-router | inference-harness-router | :9000 (127.0.0.1) | /health (DEPRECATED — not in path) |
| harness-nginx | nginx:alpine | :80 | HTTP 200 on /health | | harness-nginx | nginx:alpine | :80 | HTTP 200 on /health |
| harness-postgres | postgres:16-alpine | :5432 | pg_isready | | harness-postgres | postgres:16-alpine | :5432 | pg_isready |
| harness-redis | redis:7-alpine | :6379 | PING | | harness-redis | redis:7-alpine | :6379 | PING |
@@ -121,14 +120,14 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2)
- GET http://{{backend_host}}/litellm/health/liveliness → expect 200 - GET http://{{backend_host}}/litellm/health/liveliness → expect 200
4. **Check backend container health**: 4. **Check backend container health**:
- SSH to {{backend_host}} → `docker ps` → verify 8 containers healthy - SSH to {{backend_host}} → `docker ps` → verify 11 containers healthy
- Critical: harness-litellm, harness-router, harness-nginx, harness-postgres - Critical: harness-litellm, harness-nginx, harness-postgres
- Monitoring: harness-redis, harness-dashboard, harness-grafana, harness-prometheus - Monitoring: harness-redis, harness-dashboard, harness-grafana, harness-prometheus,
harness-alertmanager, harness-zulip-bridge, harness-docker-stats, harness-pve-exporter
5. **Check router roster loaded**: 5. **Check GPU fleet health via gpu-monitor** (router decommissioned 2026-09-11):
- GET http://{{backend_host}}:9000/health → expect 200 - GET {{gpu_dashboard_url}}/gpu-data → expect 200 with GPU metrics JSON
- GET http://{{backend_host}}:9000/health/unified → expect 3 models - nginx `/health/unified` is now a `301` redirect to `/gpu/gpu-data` (same payload)
- If router returns "all GPUs saturated" but GPUs idle: roster not loaded → reload
6. **Check GPU fleet health** (via fleet dashboard): 6. **Check GPU fleet health** (via fleet dashboard):
- GET {{gpu_dashboard_url}}/gpu-data → expect 200 with GPU metrics JSON - GET {{gpu_dashboard_url}}/gpu-data → expect 200 with GPU metrics JSON
+11 -10
View File
@@ -20,7 +20,7 @@ note: >
Last verified: 2026-07-12 Last verified: 2026-07-12
description: > description: >
LiteLLM inference stack health monitoring + self-healing. Verifies the full LiteLLM inference stack health monitoring + self-healing. Verifies the full
nginx → LiteLLM → GPU chain, 8 containers on CT 116, 3 GPU hosts, model nginx → LiteLLM → GPU chain, 11 containers on CT 116, 3 GPU hosts, model
inference, and agent keys. Applies remediation rules for common failures. inference, and agent keys. Applies remediation rules for common failures.
Reports every action via Zulip DM and Gitea (SyslogSolution/health-logs). Reports every action via Zulip DM and Gitea (SyslogSolution/health-logs).
--- ---
@@ -41,8 +41,8 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2)
│ │
Grafana :3001 Grafana :3001
harness-router :9000 — DEPRECATED, container still runs but harness-router :9000 — DECOMMISSIONED 2026-09-11 (container, image
NOT in request path. nginx routes /v1 → LiteLLM directly. and config removed). nginx routes /v1 → LiteLLM directly.
Router slot booking + circuit breakers replaced by Router slot booking + circuit breakers replaced by
LiteLLM native fallbacks + timeouts. LiteLLM native fallbacks + timeouts.
``` ```
@@ -98,8 +98,7 @@ Preferred implementation: uncap shared pool, add capped alias for crew-only.
| Container | Image | Port | Health Check | | Container | Image | Port | Health Check |
|-----------|-------|------|-------------| |-----------|-------|------|-------------|
| harness-litellm | berriai/litellm:1.90.0-rc.1 | :4000→:4001 | /health/liveliness | | harness-litellm | ghcr.io/docker.litellm.ai/berriai/litellm:1.99.1 | :4000→:4000 | /health/liveliness |
| harness-router | inference-harness-router | :9000 (127.0.0.1) | /health (DEPRECATED — not in path) |
| harness-nginx | nginx:alpine | :80 | HTTP 200 on /health | | harness-nginx | nginx:alpine | :80 | HTTP 200 on /health |
| harness-postgres | postgres:16-alpine | :5432 | pg_isready | | harness-postgres | postgres:16-alpine | :5432 | pg_isready |
| harness-redis | redis:7-alpine | :6379 | PING | | harness-redis | redis:7-alpine | :6379 | PING |
@@ -146,10 +145,11 @@ Run this first on every cycle. Results feed into remediation rules below.
- GET http://{{backend_host}}/litellm/health/liveliness → expect 200 - GET http://{{backend_host}}/litellm/health/liveliness → expect 200
### 3. Check backend container health ### 3. Check backend container health
- SSH to {{backend_host}} → `docker ps` → verify 10 containers healthy - SSH to {{backend_host}} → `docker ps` → verify 11 containers healthy
- Critical: harness-litellm, harness-nginx, harness-postgres - Critical: harness-litellm, harness-nginx, harness-postgres
- Monitoring: harness-redis, harness-dashboard, harness-grafana, harness-prometheus, harness-docker-stats, harness-pve-exporter - Monitoring: harness-redis, harness-dashboard, harness-grafana, harness-prometheus,
- Deprecated but running: harness-router (not in path, reference only) harness-alertmanager, harness-zulip-bridge, harness-docker-stats, harness-pve-exporter
- Decommissioned 2026-09-11: harness-router (container, image and config removed)
### 4. Check GPU fleet health (via fleet dashboard) ### 4. Check GPU fleet health (via fleet dashboard)
- GET {{gpu_dashboard_url}}/gpu-data → expect 200 with GPU metrics JSON - GET {{gpu_dashboard_url}}/gpu-data → expect 200 with GPU metrics JSON
@@ -215,8 +215,9 @@ Fix → generate keys in LiteLLM via /key/generate → update /etc/environment o
Escalate → if SSH access unavailable, send Zulip DM Escalate → if SSH access unavailable, send Zulip DM
### Rule 9: Stale Active Counter in Redis — DEPRECATED ### Rule 9: Stale Active Counter in Redis — DEPRECATED
Router no longer in path so Redis active counters are unused. Rule retained Router no longer in path so router active-slot counters are unused. Rule retained
for reference but inactive. If Redis issues occur, check harness-redis container. for reference but inactive. `harness-redis` now serves only LiteLLM cache and
rate-limit state; check the container if cache errors appear.
--- ---
--- ---
+9 -5
View File
@@ -213,10 +213,11 @@ def collect():
# ── LiteLLM Specific Checks (from litellm-health prose contract) ── # ── LiteLLM Specific Checks (from litellm-health prose contract) ──
report["litellm"] = {"checks": []} report["litellm"] = {"checks": []}
# Check 1: LiteLLM aggregate health endpoint (router binds to 127.0.0.1, check via SSH) # Check 1: Fleet health via nginx /health/unified (301 -> /gpu/gpu-data served by gpu-monitor on .24:9100).
health_unified = ssh(LITELLM_BACKEND, "curl -sf http://127.0.0.1:9000/health/unified -o /dev/null -w '%{http_code}' 2>/dev/null") # Router (:9000) was decommissioned 2026-09-11; probing it was a guaranteed daily failure.
health_unified = ssh(LITELLM_BACKEND, "curl -s -o /dev/null -w '%{http_code}' --max-time 8 http://127.0.0.1/health/unified 2>/dev/null")
report["litellm"]["health_unified"] = health_unified or "000" report["litellm"]["health_unified"] = health_unified or "000"
report["litellm"]["checks"].append({"name": "unified-health", "status": "pass" if health_unified == "200" else "fail", "code": health_unified or "000"}) report["litellm"]["checks"].append({"name": "fleet-health-via-nginx", "status": "pass" if health_unified in ("200", "301") else "fail", "code": health_unified or "000"})
# Check 2: Nginx-proxied internal endpoints # Check 2: Nginx-proxied internal endpoints
for path, name in [("/litellm/ui/", "nginx-ui"), ("/litellm/docs", "nginx-docs")]: for path, name in [("/litellm/ui/", "nginx-ui"), ("/litellm/docs", "nginx-docs")]:
@@ -224,8 +225,11 @@ def collect():
report["litellm"]["checks"].append({"name": name, "status": "pass" if code == "200" else "fail", "code": code}) report["litellm"]["checks"].append({"name": name, "status": "pass" if code == "200" else "fail", "code": code})
# Check 3: Docker container health for LiteLLM stack # Check 3: Docker container health for LiteLLM stack
expected_containers = ["harness-litellm", "harness-nginx", "harness-router", expected_containers = ["harness-litellm", "harness-nginx", "harness-postgres",
"harness-postgres", "harness-redis", "harness-dashboard"] "harness-redis", "harness-dashboard", "harness-grafana",
"harness-prometheus", "harness-alertmanager",
"harness-zulip-bridge", "harness-docker-stats",
"harness-pve-exporter"]
actual_names = [c["name"] for c in containers2] actual_names = [c["name"] for c in containers2]
report["litellm"]["expected_containers"] = expected_containers report["litellm"]["expected_containers"] = expected_containers
report["litellm"]["missing_containers"] = [e for e in expected_containers if e not in actual_names] report["litellm"]["missing_containers"] = [e for e in expected_containers if e not in actual_names]
+6 -4
View File
@@ -66,11 +66,13 @@ The infrastructure-control.prose.md contract is the canonical reference for the
2. NO .19 IP — Zulip is CT 117 on storepve. 2. NO .19 IP — Zulip is CT 117 on storepve.
3. NO CT 122/123 — Tanko=CT 112, Mumuni=CT 100 (inside Abiba). No CT 114 anywhere. 3. NO CT 122/123 — Tanko=CT 112, Mumuni=CT 100 (inside Abiba). No CT 114 anywhere.
4. Strix Halo :8080 is FIREWALLED to .116 only — cannot be probed from abiba (.24). 4. Strix Halo :8080 is FIREWALLED to .116 only — cannot be probed from abiba (.24).
5. abiba-zulip PM2 process is DECOMMISSIONED (2026-07-04) — abiba uses Telegram only. 5. abiba-zulip PM2 process is ONLINE (verified 2026-09-11) — this rule was stale.
**Docker on CT 116 (8 containers):** **Docker on CT 116 (11 containers, verified 2026-09-11):**
harness-litellm, harness-router, harness-nginx, harness-postgres, harness-litellm, harness-nginx, harness-postgres, harness-redis,
harness-redis, harness-dashboard, harness-grafana, harness-prometheus harness-dashboard, harness-grafana, harness-prometheus, harness-alertmanager,
harness-zulip-bridge, harness-docker-stats, harness-pve-exporter
(harness-router was decommissioned 2026-09-11)
## DIFF TO REVIEW ## DIFF TO REVIEW