fix(router): decommission legacy GPU router across contracts and fleet scripts #74
+29
-33
@@ -19,7 +19,7 @@ triggers:
|
|||||||
- on model add/remove
|
- on model add/remove
|
||||||
- on GPU health degradation
|
- on GPU health degradation
|
||||||
- on agent key rotation
|
- on agent key rotation
|
||||||
- on router restart (roster must be loaded)
|
- on harness container restart (LiteLLM reloads its model list)
|
||||||
---
|
---
|
||||||
|
|
||||||
## Maintains
|
## Maintains
|
||||||
@@ -36,46 +36,42 @@ triggers:
|
|||||||
- prometheus: { status: "running", targets: 5 } — Scrapes GPU :9400 exporters + LiteLLM
|
- prometheus: { status: "running", targets: 5 } — Scrapes GPU :9400 exporters + LiteLLM
|
||||||
- port_conflict_detection: { status: "active" } — All 3 GPU wrappers detect ghost processes before binding
|
- port_conflict_detection: { status: "active" } — All 3 GPU wrappers detect ghost processes before binding
|
||||||
|
|
||||||
## Fleet Topology (Current — July 2026)
|
## Fleet Topology (Current — 2026-09-11, router decommissioned)
|
||||||
|
|
||||||
```
|
```
|
||||||
┌──────────────────────────────────────────────────────────────────┐
|
┌──────────────────────────────────────────────────────────────────┐
|
||||||
│ CT 116 (192.168.68.116) — Inference Harness Host │
|
│ CT 116 (192.168.68.116) — Inference Harness Host │
|
||||||
│ │
|
│ │
|
||||||
│ nginx:80 (entrypoint) │
|
│ nginx:80 (entrypoint) │
|
||||||
│ ├─ /v1/* → harness-litellm:4000 (API requests) │
|
│ ├─ /v1/* → harness-litellm:4000 (API requests) │
|
||||||
│ ├─ /admin/* → harness-litellm:4000 (admin endpoints) │
|
│ ├─ /admin/* → harness-litellm:4000 (admin endpoints) │
|
||||||
│ ├─ /dashboard/ → harness-dashboard:3000 (harness UI) │
|
│ ├─ /dashboard/ → harness-dashboard:3000 (harness UI) │
|
||||||
│ ├─ /litellm/* → harness-litellm:4000 (LiteLLM UI + API) │
|
│ ├─ /litellm/* → harness-litellm:4000 (LiteLLM UI + API) │
|
||||||
│ ├─ /health/* → harness-litellm:4000 (health probes) │
|
│ ├─ /health/* → harness-litellm:4000 (health probes) │
|
||||||
│ └─ /gpu/* → 192.168.68.24:9100 (fleet monitor) │
|
│ └─ /gpu/* → 192.168.68.24:9100 (fleet monitor) │
|
||||||
│ │
|
│ │
|
||||||
│ Containers: │
|
│ Containers: │
|
||||||
│ ┌──────────┐ ┌──────────┐ ┌──────────┐ ┌──────────┐ │
|
│ ┌──────────┐ ┌──────────┐ ┌──────────┐ │
|
||||||
│ │ LiteLLM │ │ Router │ │Dashboard │ │ Grafana │ │
|
│ │ LiteLLM │ │Dashboard │ │ Grafana │ │
|
||||||
│ │ :4000 │ │ :9000 │ │ :3000 │ │ :3000 │ │
|
│ │ :4000 │ │ :3000 │ │ :3001 │ │
|
||||||
│ │ keys+sync│ │deprecated│ │ harness │ │ Prometheus│ │
|
│ │ keys+sync│ │ harness │ │Prometheus│ │
|
||||||
│ │ fallback │ │not in │ │ UI │ │ data src │ │
|
│ │ fallback │ │ UI │ │ data src │ │
|
||||||
│ └──────────┘ └───┬──────┘ └──────────┘ └──────────┘ │
|
│ └──────────┘ └──────────┘ └──────────┘ │
|
||||||
│ │ │
|
│ │
|
||||||
│ ┌──────────┐ ┌──────────┐ ┌──────────┐ │
|
│ ┌──────────┐ ┌──────────┐ ┌──────────┐ │
|
||||||
│ │PostgreSQL│ │ Redis │ │Prometheus│ │
|
│ │PostgreSQL│ │ Redis │ │Prometheus│ │
|
||||||
│ │ :5432 │ │ :6379 │ │ :9090 │ │
|
│ │ :5432 │ │ :6379 │ │ :9090 │ │
|
||||||
│ └──────────┘ └──────────┘ └──────────┘ │
|
│ └──────────┘ └──────────┘ └──────────┘ │
|
||||||
└────────────────────┼─────────────────────────────────────────────┘
|
└───────┼──────────────────────────────────────────────────────────┘
|
||||||
│
|
│
|
||||||
┌───────────────┼───────────────┬──────────────────┐
|
┌─┴─────────────┬───────────────┬───────────────┐
|
||||||
│ │ │ │
|
│ │ │ │
|
||||||
┌────▼─────┐ ┌──────▼──────┐ ┌────▼──────┐ ┌───────▼──────┐
|
┌─────▼──────┐ ┌─────▼──────┐ ┌─────▼──────┐ ┌─────▼──────┐
|
||||||
│ CT 8 │ │ CT 110 │ │ CT 15 │ │ pi (.24) │
|
│ CT 8 │ │ CT 110 │ │ CT 15 │ │ pi (.24) │
|
||||||
│ RTX 3090 │ │ RTX 5070 │ │ Strix Halo│ │ GPU Monitor │
|
│ RTX 3090 │ │ RTX 5070 │ │ Strix Halo │ │ GPU Monitor│
|
||||||
│ 24GB │ │ 12GB │ │ 64GB UMA │ │ :9100 │
|
│ 24GB │ │ 12GB │ │ 64GB UMA │ │ :9100 │
|
||||||
│ 128K ctx │ │ 128K ctx │ │ 128K ctx │ │ Watchdog │
|
│ :8080 │ │ :9400 exp │ │ :9400 exp │ │ :9401 │
|
||||||
│ qwen3.6 │ │ Qwen3.5-9B │ │ qwen3.5 │ │ Prometheus │
|
└────────────┘ └────────────┘ └────────────┘ └────────────┘
|
||||||
│ 27B-code │ │ :8080 │ │ -35B-udq4 │ │ exporter │
|
|
||||||
│ :8080 │ │ :9400 (exp) │ │ :9400(exp)│ │ :9401 │
|
|
||||||
│ :9400 │ └─────────────┘ └───────────┘ └──────────────┘
|
|
||||||
└──────────┘
|
|
||||||
```
|
```
|
||||||
|
|
||||||
## Stable Role-Based Aliases (Introduced 2026-07-15)
|
## Stable Role-Based Aliases (Introduced 2026-07-15)
|
||||||
@@ -105,7 +101,7 @@ but are deprecated for agent configs. Only the stable aliases survive model swap
|
|||||||
|-------|-----|--------|---------|---------|
|
|-------|-----|--------|---------|---------|
|
||||||
| Qwen3.8-27B-Uncensored-Q4_K_M | RTX 3090 (.8:8080) | **0.55** | 500 | **300s** |
|
| Qwen3.8-27B-Uncensored-Q4_K_M | RTX 3090 (.8:8080) | **0.55** | 500 | **300s** |
|
||||||
|
|
||||||
Note: All syslog-auto entries route directly to GPUs with `api_key: not-needed`. The router (port 9000) is NOT in the inference path.
|
Note: All syslog-auto entries route directly to GPUs with `api_key: not-needed`. The router (port 9000) was decommissioned 2026-09-11 and is NOT in the inference path.
|
||||||
|
|
||||||
### Direct Model Endpoints
|
### Direct Model Endpoints
|
||||||
|
|
||||||
@@ -153,7 +149,7 @@ Note: All syslog-auto entries route directly to GPUs with `api_key: not-needed`.
|
|||||||
6. Cleanup model files (optional)
|
6. Cleanup model files (optional)
|
||||||
|
|
||||||
### heal
|
### heal
|
||||||
1. Check all GPUs via router internal `:9000/health/unified`
|
1. Check all GPUs via gpu-monitor `{{gpu_dashboard_url}}/gpu-data`
|
||||||
2. Check LiteLLM health via nginx `:80/litellm/health/liveliness`
|
2. Check LiteLLM health via nginx `:80/litellm/health/liveliness`
|
||||||
3. Reset stuck circuit breakers if idle (Redis)
|
3. Reset stuck circuit breakers if idle (Redis)
|
||||||
4. Restart dead llama-server instances via SSH
|
4. Restart dead llama-server instances via SSH
|
||||||
@@ -161,7 +157,7 @@ Note: All syslog-auto entries route directly to GPUs with `api_key: not-needed`.
|
|||||||
6. Verify GPU monitor server is running on pi (:9100)
|
6. Verify GPU monitor server is running on pi (:9100)
|
||||||
7. Verify watchdog is running on pi
|
7. Verify watchdog is running on pi
|
||||||
8. Restart router if roster not loaded (check logs for STARTUP ROSTER)
|
8. Restart router if roster not loaded (check logs for STARTUP ROSTER)
|
||||||
9. Reload roster via `POST :9000/admin/roster/reload` if available
|
9. Router roster reload — REMOVED 2026-09-11 (router decommissioned; LiteLLM fallbacks handle routing)
|
||||||
|
|
||||||
### sync-keys
|
### sync-keys
|
||||||
1. List all agent keys in LiteLLM DB via `GET /key/list`
|
1. List all agent keys in LiteLLM DB via `GET /key/list`
|
||||||
@@ -258,7 +254,7 @@ If no SSH access, send Zulip DM via abiba-bot with vault update instructions.
|
|||||||
Benchmarks from 2026-07-17. Strix Halo model: qwen3.6-35B-udq4. RTX 5070 MTP provides 2.7x speedup over pre-upgrade 70 tok/s.
|
Benchmarks from 2026-07-17. Strix Halo model: qwen3.6-35B-udq4. RTX 5070 MTP provides 2.7x speedup over pre-upgrade 70 tok/s.
|
||||||
All 3 GPUs now at 128K context (2026-07-17, reduced from 256K for stability).
|
All 3 GPUs now at 128K context (2026-07-17, reduced from 256K for stability).
|
||||||
|
|
||||||
Benchmarks run through LiteLLM proxy (192.168.68.116:4001) every 5 minutes.
|
Benchmarks run through LiteLLM proxy (192.168.68.116:4000) every 5 minutes.
|
||||||
Degradation alerts fire at 30% (warning) and 50% (critical) below baseline.
|
Degradation alerts fire at 30% (warning) and 50% (critical) below baseline.
|
||||||
History stored at `/root/data/toks-history.json` with 7-day rolling window.
|
History stored at `/root/data/toks-history.json` with 7-day rolling window.
|
||||||
|
|
||||||
|
|||||||
+17
-17
@@ -2,7 +2,7 @@
|
|||||||
kind: responsibility
|
kind: responsibility
|
||||||
name: gpu-monitor
|
name: gpu-monitor
|
||||||
description: >
|
description: >
|
||||||
Comprehensive GPU fleet monitor — polls every subsystem (sidecars, router,
|
Comprehensive GPU fleet monitor — polls every subsystem (sidecars,
|
||||||
LiteLLM, Strix Halo, dashboard) every 15s, renders a live HTML dashboard,
|
LiteLLM, Strix Halo, dashboard) every 15s, renders a live HTML dashboard,
|
||||||
checks alert thresholds, and exposes a JSON API for downstream consumers.
|
checks alert thresholds, and exposes a JSON API for downstream consumers.
|
||||||
agent: abiba
|
agent: abiba
|
||||||
@@ -46,19 +46,19 @@ and `http://192.168.68.110/health` answer `000`. Never use a bare-port-80 probe
|
|||||||
as a GPU liveness signal: on 2026-09-09 that produced three false
|
as a GPU liveness signal: on 2026-09-09 that produced three false
|
||||||
`DEGRADED — GPU-rtx3090 000, GPU-rtx5070 000` rounds while
|
`DEGRADED — GPU-rtx3090 000, GPU-rtx5070 000` rounds while
|
||||||
`http://192.168.68.8:8080/health` and `http://192.168.68.110:8080/health`
|
`http://192.168.68.8:8080/health` and `http://192.168.68.110:8080/health`
|
||||||
answered `200`. Port 80 is valid only on the router (.116), never on a GPU host.
|
answered `200`. Port 80 is valid only on the harness host (.116), never on a GPU host.
|
||||||
|
|
||||||
### Subsystems Polled
|
### Subsystems Polled
|
||||||
|
|
||||||
| Subsystem | Endpoint | Frequency | Metrics |
|
| Subsystem | Endpoint | Frequency | Metrics |
|
||||||
|-----------|----------|-----------|---------|
|
|-----------|----------|-----------|---------|
|
||||||
| GPU Status (all, via router) | `http://192.168.68.116/health/unified` | 15s | models, CB, scores, GPU status (router probes each GPU /health directly) — `301` → `/gpu/gpu-data` is **alive** |
|
| GPU Status (all, via fleet API) | `http://192.168.68.116/gpu/gpu-data` | 15s | models, CB, scores, GPU status from gpu-monitor on .24:9100 |
|
||||||
| GPU .8 (RTX 3090) health | `http://192.168.68.8:8080/health` | 15s | direct liveness fallback — **:8080 ONLY, never bare port 80** |
|
| GPU .8 (RTX 3090) health | `http://192.168.68.8:8080/health` | 15s | direct liveness fallback — **:8080 ONLY, never bare port 80** |
|
||||||
| GPU .110 (RTX 5070) health | `http://192.168.68.110:8080/health` | 15s | direct liveness fallback — **:8080 ONLY, never bare port 80** |
|
| GPU .110 (RTX 5070) health | `http://192.168.68.110:8080/health` | 15s | direct liveness fallback — **:8080 ONLY, never bare port 80** |
|
||||||
| Router (unified) | `http://192.168.68.116/health/unified` | 15s | models, CB, scores, GPU status (`301` → `/gpu/gpu-data` = alive) |
|
| Fleet (unified) | `http://192.168.68.116/health/unified` | 15s | nginx `301` → `/gpu/gpu-data` served by gpu-monitor = alive (router decommissioned 2026-09-11) |
|
||||||
| Router (basic) | `http://192.168.68.116/health` | 15s | basic aliveness |
|
| Harness (basic) | `http://192.168.68.116/health` | 15s | nginx → LiteLLM `/health/liveliness` |
|
||||||
| LiteLLM | `http://192.168.68.116/litellm/health` | 15s | proxy health, model count |
|
| LiteLLM | `http://192.168.68.116/litellm/health` | 15s | proxy health, model count |
|
||||||
| Strix Halo | `http://192.168.68.116/health/unified` (router) | 15s | Strix Halo status via router — cannot poll .15:8080 directly (firewalled to .116 only) |
|
| Strix Halo | `http://192.168.68.116/health/unified` (nginx → fleet API) | 15s | Strix Halo status via gpu-monitor — cannot poll .15:8080 directly (firewalled to .116 only) |
|
||||||
| Dashboard | `http://192.168.68.116/dashboard/` | 15s | harness-dashboard aliveness |
|
| Dashboard | `http://192.168.68.116/dashboard/` | 15s | harness-dashboard aliveness |
|
||||||
|
|
||||||
### Alert Delivery
|
### Alert Delivery
|
||||||
@@ -75,8 +75,8 @@ This replaces the previous DM-only delivery. All agents on the mesh can see and
|
|||||||
### Liveness rule (scoped)
|
### Liveness rule (scoped)
|
||||||
|
|
||||||
The any-HTTP-response rule applies ONLY to redirect/auth-gated liveness
|
The any-HTTP-response rule applies ONLY to redirect/auth-gated liveness
|
||||||
endpoints, where any HTTP answer proves a listener is up. Applied here: the
|
endpoints, where any HTTP answer proves a listener is up. Applied here: nginx's
|
||||||
router's `/health/unified` answers `301 Moved Permanently` → `/gpu/gpu-data`
|
`/health/unified` answers `301 Moved Permanently` → `/gpu/gpu-data`
|
||||||
(the same payload) and LiteLLM's `/litellm/health` answers `301` →
|
(the same payload) and LiteLLM's `/litellm/health` answers `301` →
|
||||||
`/litellm/health/liveliness`. For those endpoints a probe is **ALIVE** on
|
`/litellm/health/liveliness`. For those endpoints a probe is **ALIVE** on
|
||||||
**ANY** HTTP status — `3xx` redirects and `401`/`403` auth challenges included —
|
**ANY** HTTP status — `3xx` redirects and `401`/`403` auth challenges included —
|
||||||
@@ -84,8 +84,8 @@ and **DOWN = connection refused (`000`) or timeout only**. Same scoped rule as
|
|||||||
zulip-health (Tanko) and infrastructure-monitoring.
|
zulip-health (Tanko) and infrastructure-monitoring.
|
||||||
|
|
||||||
Probes whose success condition is specifically a bare `200` are NOT covered by
|
Probes whose success condition is specifically a bare `200` are NOT covered by
|
||||||
the any-HTTP rule. On those — the GPU `:8080/health` endpoints, the router
|
the any-HTTP rule. On those — the GPU `:8080/health` endpoints, nginx `/health`,
|
||||||
`/health`, and the dashboard — an unexpected status (`401`/`403`, `5xx`, or
|
and the dashboard — an unexpected status (`401`/`403`, `5xx`, or
|
||||||
anything other than the expected `200`) is an **ALERT**, not "alive".
|
anything other than the expected `200`) is an **ALERT**, not "alive".
|
||||||
|
|
||||||
| Metric | Warning | Critical |
|
| Metric | Warning | Critical |
|
||||||
@@ -93,7 +93,7 @@ anything other than the expected `200`) is an **ALERT**, not "alive".
|
|||||||
| GPU Temp | >80°C | >90°C |
|
| GPU Temp | >80°C | >90°C |
|
||||||
| VRAM Usage | >90% | >95% |
|
| VRAM Usage | >90% | >95% |
|
||||||
| GPU Util | >95% | >98% |
|
| GPU Util | >95% | >98% |
|
||||||
| Sidecar Unreachable | — | info (sidecars not deployed — use router /health/unified) |
|
| Sidecar Unreachable | — | info (sidecars not deployed — use gpu-monitor /gpu-data) |
|
||||||
| Model Down | — | critical (circuit breaker open) |
|
| Model Down | — | critical (circuit breaker open) |
|
||||||
|
|
||||||
### JSON API Response Schema (/gpu-data)
|
### JSON API Response Schema (/gpu-data)
|
||||||
@@ -174,13 +174,13 @@ Open `http://localhost:9100/` in browser — Live HTML dashboard
|
|||||||
pkill -f gpu-monitor-server.py
|
pkill -f gpu-monitor-server.py
|
||||||
python3 /root/scripts/gpu-monitor-server.py &
|
python3 /root/scripts/gpu-monitor-server.py &
|
||||||
```
|
```
|
||||||
Or via PM2: `pm2 restart gpu-monitor`
|
Managed by systemd (verified 2026-09-11): `systemctl restart gpu-monitor`
|
||||||
|
|
||||||
### check-router
|
### check-fleet
|
||||||
The router health is accessed through nginx on port 80 on the **router**
|
Fleet health is accessed through nginx on port 80 on the harness host
|
||||||
(.116) — NOT port 9000 directly, and NOT bare port 80 on a GPU host.
|
(.116) — NOT port 9000 (router decommissioned 2026-09-11), and NOT bare
|
||||||
`curl http://192.168.68.116/health/unified` — Router unified health via nginx proxy; answers `301` → `/gpu/gpu-data` (same payload) = alive
|
port 80 on a GPU host.
|
||||||
`curl http://192.168.68.116:9000/health/unified` — ❌ WILL FAIL (port bound to 127.0.0.1 only)
|
`curl http://192.168.68.116/health/unified` — nginx answers `301` → `/gpu/gpu-data` (fleet monitor payload) = alive
|
||||||
`curl http://192.168.68.8/health` — ❌ NEVER USE (GPU host, no port-80 listener → false `000`/DEGRADED)
|
`curl http://192.168.68.8/health` — ❌ NEVER USE (GPU host, no port-80 listener → false `000`/DEGRADED)
|
||||||
|
|
||||||
## Configuration Files
|
## Configuration Files
|
||||||
|
|||||||
@@ -9,7 +9,7 @@ description: >
|
|||||||
(v2.1.0) with active remediation rules, Prometheus metrics consumption,
|
(v2.1.0) with active remediation rules, Prometheus metrics consumption,
|
||||||
VRAM trend analysis, and predictive alerting.
|
VRAM trend analysis, and predictive alerting.
|
||||||
UPDATED 2026-07-18: Model assignments synced to 2026-07-17 swaps.
|
UPDATED 2026-07-18: Model assignments synced to 2026-07-17 swaps.
|
||||||
Router (port 9000) references replaced with direct GPU routing.
|
Router (port 9000) DECOMMISSIONED 2026-09-11; references replaced with direct GPU routing.
|
||||||
Benchmark baselines refreshed to live values.
|
Benchmark baselines refreshed to live values.
|
||||||
Prometheus exporters removed — not deployed; fall back to direct sidecar probes.
|
Prometheus exporters removed — not deployed; fall back to direct sidecar probes.
|
||||||
Stable role-based aliases (strix-moe, gpu-dense, gpu-light) from gpu-fleet.
|
Stable role-based aliases (strix-moe, gpu-dense, gpu-light) from gpu-fleet.
|
||||||
@@ -51,7 +51,7 @@ depends_on:
|
|||||||
| `strix-moe` | Strix Halo 64GB | ct15 (.15:8080) | qwen3.6-35B-udq4 | ~10/64GB (16%) | 128K | 62.9 | Compression, summarization, long docs |
|
| `strix-moe` | Strix Halo 64GB | ct15 (.15:8080) | qwen3.6-35B-udq4 | ~10/64GB (16%) | 128K | 62.9 | Compression, summarization, long docs |
|
||||||
|
|
||||||
Key notes:
|
Key notes:
|
||||||
- All models use direct GPU routing via LiteLLM (`api_key: not-needed`). Router (port 9000) is deprecated and NOT in the inference path.
|
- All models use direct GPU routing via LiteLLM (`api_key: not-needed`). Router (port 9000) was decommissioned 2026-09-11 and is NOT in the inference path.
|
||||||
- Stable aliases (gpu-dense, gpu-light, strix-moe) from gpu-fleet are the canonical names for agent configs. Model-specific names still work but are deprecated.
|
- Stable aliases (gpu-dense, gpu-light, strix-moe) from gpu-fleet are the canonical names for agent configs. Model-specific names still work but are deprecated.
|
||||||
- RTX 5070 tok/s is ~145 for Qwen3.5-9B — gpu-light is the fastest endpoint. Route vision/web/light work there first. NOTE: Qwen3.5-9B is multimodal (image+text), NOT text-only like gemma-4-12b was.
|
- RTX 5070 tok/s is ~145 for Qwen3.5-9B — gpu-light is the fastest endpoint. Route vision/web/light work there first. NOTE: Qwen3.5-9B is multimodal (image+text), NOT text-only like gemma-4-12b was.
|
||||||
- Strix Halo is 62.9 tok/s (89% of 70.5 baseline) — below optimal but stable. Check for competing workloads.
|
- Strix Halo is 62.9 tok/s (89% of 70.5 baseline) — below optimal but stable. Check for competing workloads.
|
||||||
@@ -104,10 +104,10 @@ Key notes:
|
|||||||
|
|
||||||
### Rule 5: Circuit Breaker Stuck Open
|
### Rule 5: Circuit Breaker Stuck Open
|
||||||
- **Detect**: Circuit breaker open >10 minutes with GPU reporting healthy
|
- **Detect**: Circuit breaker open >10 minutes with GPU reporting healthy
|
||||||
- **Note**: Router (port 9000) is deprecated. If circuit breakers are reported by gpu-monitor, they come from LiteLLM's internal tracking, not the old router.
|
- **Note**: Router (port 9000) was decommissioned 2026-09-11. If circuit breakers are reported by gpu-monitor, they come from LiteLLM's internal tracking, not the old router.
|
||||||
- **Fix**:
|
- **Fix**:
|
||||||
1. Verify GPU /health returns 200 on direct port (:8080)
|
1. Verify GPU /health returns 200 on direct port (:8080)
|
||||||
2. If GPU healthy, alert but do NOT reset via router API (deprecated)
|
2. If GPU healthy, alert but do NOT reset via router API (decommissioned 2026-09-11)
|
||||||
3. Check LiteLLM health directly: http://192.168.68.116/litellm/health/liveliness
|
3. Check LiteLLM health directly: http://192.168.68.116/litellm/health/liveliness
|
||||||
4. Restart LiteLLM container on CT 116 if circuit breakers are stuck
|
4. Restart LiteLLM container on CT 116 if circuit breakers are stuck
|
||||||
- **Verify**: LiteLLM returns healthy, circuit breaker clears within 60s
|
- **Verify**: LiteLLM returns healthy, circuit breaker clears within 60s
|
||||||
@@ -276,7 +276,7 @@ Pushed to `SyslogSolution/health-logs/gpu/{run_id}.json` — versioned, searchab
|
|||||||
2. **Model restart**: ✅ Only if >50% failure rate over 60s + 30s grace period. Not on single stuck request.
|
2. **Model restart**: ✅ Only if >50% failure rate over 60s + 30s grace period. Not on single stuck request.
|
||||||
3. **Strix direct access**: ✅ Open firewall .15:8080 → .24 for direct health probe + restart.
|
3. **Strix direct access**: ✅ Open firewall .15:8080 → .24 for direct health probe + restart.
|
||||||
4. **VRAM thresholds**: Tiered — **300MB/h** (RTX 3090), **300MB/h** (RTX 5070), 200MB/h (Strix). Previous values (100/50) were too sensitive; raised 2026-07-18 based on operational data.
|
4. **VRAM thresholds**: Tiered — **300MB/h** (RTX 3090), **300MB/h** (RTX 5070), 200MB/h (Strix). Previous values (100/50) were too sensitive; raised 2026-07-18 based on operational data.
|
||||||
5. **CB auto-reset**: ✅ Router deprecated — circuit breakers go through LiteLLM health check + container restart if needed. No per-GPU auto-reset.
|
5. **CB auto-reset**: ✅ Router decommissioned 2026-09-11 — circuit breakers go through LiteLLM health check + container restart if needed. No per-GPU auto-reset.
|
||||||
6. **Benchmark baseline**: Rolling 30-day average, recalculated weekly. Current baselines live in gpu-monitor.
|
6. **Benchmark baseline**: Rolling 30-day average, recalculated weekly. Current baselines live in gpu-monitor.
|
||||||
7. **Predictive alerts**: Two-tier — warn at >70°C+rising (>2°C/min), critical at >80°C+rising.
|
7. **Predictive alerts**: Two-tier — warn at >70°C+rising (>2°C/min), critical at >80°C+rising.
|
||||||
8. **Prometheus**: ❌ Not deployed. Use direct sidecar probes (:8080/health) and gpu-monitor API. Prometheus integration deferred until exporters are running on GPU hosts.
|
8. **Prometheus**: ❌ Not deployed. Use direct sidecar probes (:8080/health) and gpu-monitor API. Prometheus integration deferred until exporters are running on GPU hosts.
|
||||||
|
|||||||
@@ -39,7 +39,7 @@ runtime via `infisical run --` wrapper. Plaintext keys removed from this baselin
|
|||||||
```
|
```
|
||||||
Infisical vault → infisical run -- hermes gateway → LITELLM_API_KEY (runtime)
|
Infisical vault → infisical run -- hermes gateway → LITELLM_API_KEY (runtime)
|
||||||
↓
|
↓
|
||||||
Agent (systemd) → LITELLM_API_KEY → LiteLLM (:116/v1) → Router (:9000) → GPU (llama-server)
|
Agent (systemd) → LITELLM_API_KEY → LiteLLM (:116/v1) → GPU (llama-server)
|
||||||
└── Key DB (Postgres)
|
└── Key DB (Postgres)
|
||||||
```
|
```
|
||||||
|
|
||||||
|
|||||||
@@ -199,15 +199,14 @@ description: >
|
|||||||
|
|
||||||
### Ecosystem B: CT 116 syslog-api (192.168.68.116)
|
### Ecosystem B: CT 116 syslog-api (192.168.68.116)
|
||||||
|
|
||||||
8 containers in inference-harness stack:
|
11 containers in inference-harness stack (verified live 2026-09-11; LiteLLM upgraded 1.90.0-rc.1 -> 1.99.1):
|
||||||
|
|
||||||
| Container | Image | Port | Role |
|
| Container | Image | Port | Role |
|
||||||
|-----------|-------|------|------|
|
|-----------|-------|------|------|
|
||||||
| harness-litellm | berriai/litellm:1.90.0-rc.1 | :4000→:4001 | API proxy, key mgmt, fallbacks |
|
| harness-litellm | ghcr.io/docker.litellm.ai/berriai/litellm:1.99.1 | :4000→:4000 | API proxy, key mgmt, fallbacks |
|
||||||
| harness-router | inference-harness-router | :9000 (127.0.0.1) | GPU routing, slot booking, CB |
|
|
||||||
| harness-nginx | nginx:alpine | :80 | Entrypoint, /v1→LiteLLM, /dashboard/ |
|
| harness-nginx | nginx:alpine | :80 | Entrypoint, /v1→LiteLLM, /dashboard/ |
|
||||||
| harness-postgres | postgres:16-alpine | :5432 | LiteLLM DB (keys, spend, config) |
|
| harness-postgres | postgres:16-alpine | :5432 | LiteLLM DB (keys, spend, config) |
|
||||||
| harness-redis | redis:7-alpine | :6379 | Router slots, circuit breakers |
|
| harness-redis | redis:7-alpine | :6379 | LiteLLM cache + rate-limit state |
|
||||||
| harness-dashboard | inference-harness-dashboard | :3000 | SyslogAI Harness UI |
|
| harness-dashboard | inference-harness-dashboard | :3000 | SyslogAI Harness UI |
|
||||||
| harness-grafana | grafana/grafana | :3000→:3001 (direct LAN, not behind nginx) | GPU + Proxmox + Docker dashboards |
|
| harness-grafana | grafana/grafana | :3000→:3001 (direct LAN, not behind nginx) | GPU + Proxmox + Docker dashboards |
|
||||||
| harness-prometheus | prom/prometheus | :9090 | Metrics scraper, 6 jobs |
|
| harness-prometheus | prom/prometheus | :9090 | Metrics scraper, 6 jobs |
|
||||||
@@ -588,9 +587,6 @@ curl -s -H "Authorization: Bearer $(infisical secrets get LITELLM_MASTER_KEY --p
|
|||||||
# Storage check
|
# Storage check
|
||||||
ssh root@192.168.68.7 "df -h /media/storage /media/mediastore"
|
ssh root@192.168.68.7 "df -h /media/storage /media/mediastore"
|
||||||
|
|
||||||
# Router roster reload (if needed)
|
|
||||||
curl -s -X POST http://192.168.68.116:9000/admin/roster/reload \
|
|
||||||
-H "Authorization: Bearer sk-admin-ee09fffd04978b61a1569ac670c68814"
|
|
||||||
|
|
||||||
# Restart stuck GPU (saturation watchdog alternative)
|
# Restart stuck GPU (saturation watchdog alternative)
|
||||||
ssh root@192.168.68.8 "systemctl restart llama-server"
|
ssh root@192.168.68.8 "systemctl restart llama-server"
|
||||||
|
|||||||
@@ -173,7 +173,7 @@ curl -s http://192.168.68.116:3001/api/health | jq '{status, version}'
|
|||||||
# Expected: {"status":"ok","version":"..."}
|
# Expected: {"status":"ok","version":"..."}
|
||||||
|
|
||||||
# LiteLLM metrics (Prometheus endpoint)
|
# LiteLLM metrics (Prometheus endpoint)
|
||||||
curl -s http://192.168.68.116:4001/metrics | head -20
|
curl -s http://192.168.68.116:4000/metrics | head -20
|
||||||
# Expected: Prometheus-formatted metrics output
|
# Expected: Prometheus-formatted metrics output
|
||||||
```
|
```
|
||||||
|
|
||||||
@@ -242,5 +242,5 @@ curl -s http://192.168.68.116:9090/api/v1/targets
|
|||||||
curl -s http://192.168.68.116:3001/api/health
|
curl -s http://192.168.68.116:3001/api/health
|
||||||
|
|
||||||
# LiteLLM metrics (already live)
|
# LiteLLM metrics (already live)
|
||||||
curl -s http://192.168.68.116:4001/metrics | head -20
|
curl -s http://192.168.68.116:4000/metrics | head -20
|
||||||
```
|
```
|
||||||
|
|||||||
@@ -206,7 +206,7 @@ mcp_servers:
|
|||||||
| Key | MCP Access |
|
| Key | MCP Access |
|
||||||
|-----|-----------|
|
|-----|-----------|
|
||||||
| Master key | ✅ Full — 90 tools (vault-injected) |
|
| Master key | ✅ Full — 90 tools (vault-injected) |
|
||||||
| Agent keys (mumuni, tanko, etc.) | ❌ Per-key grants not supported in v1.90.0-rc.1 |
|
| Agent keys (mumuni, tanko, etc.) | ❌ Per-key grants not supported in v1.99.1 |
|
||||||
|
|
||||||
### Known Limitations
|
### Known Limitations
|
||||||
- Per-key MCP server grants not functional — only master key has access
|
- Per-key MCP server grants not functional — only master key has access
|
||||||
|
|||||||
+12
-13
@@ -12,8 +12,8 @@ note: >
|
|||||||
description: >
|
description: >
|
||||||
Verifies the LiteLLM inference stack health. Current architecture (2026-07-09):
|
Verifies the LiteLLM inference stack health. Current architecture (2026-07-09):
|
||||||
nginx:80 → LiteLLM:4000 → GPU(llama-server) via direct proxy.
|
nginx:80 → LiteLLM:4000 → GPU(llama-server) via direct proxy.
|
||||||
Router (harness-router :9000) is DEPRECATED — container still runs but
|
Router (harness-router :9000) was DECOMMISSIONED 2026-09-11 (container,
|
||||||
is not in the request path. GPU monitoring via Prometheus/Grafana and
|
image and config removed). GPU monitoring via Prometheus/Grafana and
|
||||||
fleet dashboard (gpu-monitor :9100).
|
fleet dashboard (gpu-monitor :9100).
|
||||||
Designed as a reusable contract for any Syslog agent.
|
Designed as a reusable contract for any Syslog agent.
|
||||||
|
|
||||||
@@ -42,8 +42,8 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2)
|
|||||||
│
|
│
|
||||||
Grafana :3001
|
Grafana :3001
|
||||||
|
|
||||||
harness-router :9000 — DEPRECATED, container still runs but
|
harness-router :9000 — DECOMMISSIONED 2026-09-11 (container, image
|
||||||
NOT in request path. nginx routes /v1 → LiteLLM directly.
|
and config removed). nginx routes /v1 → LiteLLM directly.
|
||||||
Router slot booking + circuit breakers replaced by
|
Router slot booking + circuit breakers replaced by
|
||||||
LiteLLM native fallbacks + timeouts.
|
LiteLLM native fallbacks + timeouts.
|
||||||
```
|
```
|
||||||
@@ -100,8 +100,7 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2)
|
|||||||
|
|
||||||
| Container | Image | Port | Health Check |
|
| Container | Image | Port | Health Check |
|
||||||
|-----------|-------|------|-------------|
|
|-----------|-------|------|-------------|
|
||||||
| harness-litellm | berriai/litellm:1.90.0-rc.1 | :4000→:4001 | /health/liveliness |
|
| harness-litellm | ghcr.io/docker.litellm.ai/berriai/litellm:1.99.1 | :4000→:4000 | /health/liveliness |
|
||||||
| harness-router | inference-harness-router | :9000 (127.0.0.1) | /health (DEPRECATED — not in path) |
|
|
||||||
| harness-nginx | nginx:alpine | :80 | HTTP 200 on /health |
|
| harness-nginx | nginx:alpine | :80 | HTTP 200 on /health |
|
||||||
| harness-postgres | postgres:16-alpine | :5432 | pg_isready |
|
| harness-postgres | postgres:16-alpine | :5432 | pg_isready |
|
||||||
| harness-redis | redis:7-alpine | :6379 | PING |
|
| harness-redis | redis:7-alpine | :6379 | PING |
|
||||||
@@ -121,14 +120,14 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2)
|
|||||||
- GET http://{{backend_host}}/litellm/health/liveliness → expect 200
|
- GET http://{{backend_host}}/litellm/health/liveliness → expect 200
|
||||||
|
|
||||||
4. **Check backend container health**:
|
4. **Check backend container health**:
|
||||||
- SSH to {{backend_host}} → `docker ps` → verify 8 containers healthy
|
- SSH to {{backend_host}} → `docker ps` → verify 11 containers healthy
|
||||||
- Critical: harness-litellm, harness-router, harness-nginx, harness-postgres
|
- Critical: harness-litellm, harness-nginx, harness-postgres
|
||||||
- Monitoring: harness-redis, harness-dashboard, harness-grafana, harness-prometheus
|
- Monitoring: harness-redis, harness-dashboard, harness-grafana, harness-prometheus,
|
||||||
|
harness-alertmanager, harness-zulip-bridge, harness-docker-stats, harness-pve-exporter
|
||||||
|
|
||||||
5. **Check router roster loaded**:
|
5. **Check GPU fleet health via gpu-monitor** (router decommissioned 2026-09-11):
|
||||||
- GET http://{{backend_host}}:9000/health → expect 200
|
- GET {{gpu_dashboard_url}}/gpu-data → expect 200 with GPU metrics JSON
|
||||||
- GET http://{{backend_host}}:9000/health/unified → expect 3 models
|
- nginx `/health/unified` is now a `301` redirect to `/gpu/gpu-data` (same payload)
|
||||||
- If router returns "all GPUs saturated" but GPUs idle: roster not loaded → reload
|
|
||||||
|
|
||||||
6. **Check GPU fleet health** (via fleet dashboard):
|
6. **Check GPU fleet health** (via fleet dashboard):
|
||||||
- GET {{gpu_dashboard_url}}/gpu-data → expect 200 with GPU metrics JSON
|
- GET {{gpu_dashboard_url}}/gpu-data → expect 200 with GPU metrics JSON
|
||||||
|
|||||||
+11
-10
@@ -20,7 +20,7 @@ note: >
|
|||||||
Last verified: 2026-07-12
|
Last verified: 2026-07-12
|
||||||
description: >
|
description: >
|
||||||
LiteLLM inference stack health monitoring + self-healing. Verifies the full
|
LiteLLM inference stack health monitoring + self-healing. Verifies the full
|
||||||
nginx → LiteLLM → GPU chain, 8 containers on CT 116, 3 GPU hosts, model
|
nginx → LiteLLM → GPU chain, 11 containers on CT 116, 3 GPU hosts, model
|
||||||
inference, and agent keys. Applies remediation rules for common failures.
|
inference, and agent keys. Applies remediation rules for common failures.
|
||||||
Reports every action via Zulip DM and Gitea (SyslogSolution/health-logs).
|
Reports every action via Zulip DM and Gitea (SyslogSolution/health-logs).
|
||||||
---
|
---
|
||||||
@@ -41,8 +41,8 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2)
|
|||||||
│
|
│
|
||||||
Grafana :3001
|
Grafana :3001
|
||||||
|
|
||||||
harness-router :9000 — DEPRECATED, container still runs but
|
harness-router :9000 — DECOMMISSIONED 2026-09-11 (container, image
|
||||||
NOT in request path. nginx routes /v1 → LiteLLM directly.
|
and config removed). nginx routes /v1 → LiteLLM directly.
|
||||||
Router slot booking + circuit breakers replaced by
|
Router slot booking + circuit breakers replaced by
|
||||||
LiteLLM native fallbacks + timeouts.
|
LiteLLM native fallbacks + timeouts.
|
||||||
```
|
```
|
||||||
@@ -98,8 +98,7 @@ Preferred implementation: uncap shared pool, add capped alias for crew-only.
|
|||||||
|
|
||||||
| Container | Image | Port | Health Check |
|
| Container | Image | Port | Health Check |
|
||||||
|-----------|-------|------|-------------|
|
|-----------|-------|------|-------------|
|
||||||
| harness-litellm | berriai/litellm:1.90.0-rc.1 | :4000→:4001 | /health/liveliness |
|
| harness-litellm | ghcr.io/docker.litellm.ai/berriai/litellm:1.99.1 | :4000→:4000 | /health/liveliness |
|
||||||
| harness-router | inference-harness-router | :9000 (127.0.0.1) | /health (DEPRECATED — not in path) |
|
|
||||||
| harness-nginx | nginx:alpine | :80 | HTTP 200 on /health |
|
| harness-nginx | nginx:alpine | :80 | HTTP 200 on /health |
|
||||||
| harness-postgres | postgres:16-alpine | :5432 | pg_isready |
|
| harness-postgres | postgres:16-alpine | :5432 | pg_isready |
|
||||||
| harness-redis | redis:7-alpine | :6379 | PING |
|
| harness-redis | redis:7-alpine | :6379 | PING |
|
||||||
@@ -146,10 +145,11 @@ Run this first on every cycle. Results feed into remediation rules below.
|
|||||||
- GET http://{{backend_host}}/litellm/health/liveliness → expect 200
|
- GET http://{{backend_host}}/litellm/health/liveliness → expect 200
|
||||||
|
|
||||||
### 3. Check backend container health
|
### 3. Check backend container health
|
||||||
- SSH to {{backend_host}} → `docker ps` → verify 10 containers healthy
|
- SSH to {{backend_host}} → `docker ps` → verify 11 containers healthy
|
||||||
- Critical: harness-litellm, harness-nginx, harness-postgres
|
- Critical: harness-litellm, harness-nginx, harness-postgres
|
||||||
- Monitoring: harness-redis, harness-dashboard, harness-grafana, harness-prometheus, harness-docker-stats, harness-pve-exporter
|
- Monitoring: harness-redis, harness-dashboard, harness-grafana, harness-prometheus,
|
||||||
- Deprecated but running: harness-router (not in path, reference only)
|
harness-alertmanager, harness-zulip-bridge, harness-docker-stats, harness-pve-exporter
|
||||||
|
- Decommissioned 2026-09-11: harness-router (container, image and config removed)
|
||||||
|
|
||||||
### 4. Check GPU fleet health (via fleet dashboard)
|
### 4. Check GPU fleet health (via fleet dashboard)
|
||||||
- GET {{gpu_dashboard_url}}/gpu-data → expect 200 with GPU metrics JSON
|
- GET {{gpu_dashboard_url}}/gpu-data → expect 200 with GPU metrics JSON
|
||||||
@@ -215,8 +215,9 @@ Fix → generate keys in LiteLLM via /key/generate → update /etc/environment o
|
|||||||
Escalate → if SSH access unavailable, send Zulip DM
|
Escalate → if SSH access unavailable, send Zulip DM
|
||||||
|
|
||||||
### Rule 9: Stale Active Counter in Redis — DEPRECATED
|
### Rule 9: Stale Active Counter in Redis — DEPRECATED
|
||||||
Router no longer in path so Redis active counters are unused. Rule retained
|
Router no longer in path so router active-slot counters are unused. Rule retained
|
||||||
for reference but inactive. If Redis issues occur, check harness-redis container.
|
for reference but inactive. `harness-redis` now serves only LiteLLM cache and
|
||||||
|
rate-limit state; check the container if cache errors appear.
|
||||||
|
|
||||||
---
|
---
|
||||||
---
|
---
|
||||||
|
|||||||
@@ -213,10 +213,11 @@ def collect():
|
|||||||
# ── LiteLLM Specific Checks (from litellm-health prose contract) ──
|
# ── LiteLLM Specific Checks (from litellm-health prose contract) ──
|
||||||
report["litellm"] = {"checks": []}
|
report["litellm"] = {"checks": []}
|
||||||
|
|
||||||
# Check 1: LiteLLM aggregate health endpoint (router binds to 127.0.0.1, check via SSH)
|
# Check 1: Fleet health via nginx /health/unified (301 -> /gpu/gpu-data served by gpu-monitor on .24:9100).
|
||||||
health_unified = ssh(LITELLM_BACKEND, "curl -sf http://127.0.0.1:9000/health/unified -o /dev/null -w '%{http_code}' 2>/dev/null")
|
# Router (:9000) was decommissioned 2026-09-11; probing it was a guaranteed daily failure.
|
||||||
|
health_unified = ssh(LITELLM_BACKEND, "curl -s -o /dev/null -w '%{http_code}' --max-time 8 http://127.0.0.1/health/unified 2>/dev/null")
|
||||||
report["litellm"]["health_unified"] = health_unified or "000"
|
report["litellm"]["health_unified"] = health_unified or "000"
|
||||||
report["litellm"]["checks"].append({"name": "unified-health", "status": "pass" if health_unified == "200" else "fail", "code": health_unified or "000"})
|
report["litellm"]["checks"].append({"name": "fleet-health-via-nginx", "status": "pass" if health_unified in ("200", "301") else "fail", "code": health_unified or "000"})
|
||||||
|
|
||||||
# Check 2: Nginx-proxied internal endpoints
|
# Check 2: Nginx-proxied internal endpoints
|
||||||
for path, name in [("/litellm/ui/", "nginx-ui"), ("/litellm/docs", "nginx-docs")]:
|
for path, name in [("/litellm/ui/", "nginx-ui"), ("/litellm/docs", "nginx-docs")]:
|
||||||
@@ -224,8 +225,11 @@ def collect():
|
|||||||
report["litellm"]["checks"].append({"name": name, "status": "pass" if code == "200" else "fail", "code": code})
|
report["litellm"]["checks"].append({"name": name, "status": "pass" if code == "200" else "fail", "code": code})
|
||||||
|
|
||||||
# Check 3: Docker container health for LiteLLM stack
|
# Check 3: Docker container health for LiteLLM stack
|
||||||
expected_containers = ["harness-litellm", "harness-nginx", "harness-router",
|
expected_containers = ["harness-litellm", "harness-nginx", "harness-postgres",
|
||||||
"harness-postgres", "harness-redis", "harness-dashboard"]
|
"harness-redis", "harness-dashboard", "harness-grafana",
|
||||||
|
"harness-prometheus", "harness-alertmanager",
|
||||||
|
"harness-zulip-bridge", "harness-docker-stats",
|
||||||
|
"harness-pve-exporter"]
|
||||||
actual_names = [c["name"] for c in containers2]
|
actual_names = [c["name"] for c in containers2]
|
||||||
report["litellm"]["expected_containers"] = expected_containers
|
report["litellm"]["expected_containers"] = expected_containers
|
||||||
report["litellm"]["missing_containers"] = [e for e in expected_containers if e not in actual_names]
|
report["litellm"]["missing_containers"] = [e for e in expected_containers if e not in actual_names]
|
||||||
|
|||||||
@@ -66,11 +66,13 @@ The infrastructure-control.prose.md contract is the canonical reference for the
|
|||||||
2. NO .19 IP — Zulip is CT 117 on storepve.
|
2. NO .19 IP — Zulip is CT 117 on storepve.
|
||||||
3. NO CT 122/123 — Tanko=CT 112, Mumuni=CT 100 (inside Abiba). No CT 114 anywhere.
|
3. NO CT 122/123 — Tanko=CT 112, Mumuni=CT 100 (inside Abiba). No CT 114 anywhere.
|
||||||
4. Strix Halo :8080 is FIREWALLED to .116 only — cannot be probed from abiba (.24).
|
4. Strix Halo :8080 is FIREWALLED to .116 only — cannot be probed from abiba (.24).
|
||||||
5. abiba-zulip PM2 process is DECOMMISSIONED (2026-07-04) — abiba uses Telegram only.
|
5. abiba-zulip PM2 process is ONLINE (verified 2026-09-11) — this rule was stale.
|
||||||
|
|
||||||
**Docker on CT 116 (8 containers):**
|
**Docker on CT 116 (11 containers, verified 2026-09-11):**
|
||||||
harness-litellm, harness-router, harness-nginx, harness-postgres,
|
harness-litellm, harness-nginx, harness-postgres, harness-redis,
|
||||||
harness-redis, harness-dashboard, harness-grafana, harness-prometheus
|
harness-dashboard, harness-grafana, harness-prometheus, harness-alertmanager,
|
||||||
|
harness-zulip-bridge, harness-docker-stats, harness-pve-exporter
|
||||||
|
(harness-router was decommissioned 2026-09-11)
|
||||||
|
|
||||||
## DIFF TO REVIEW
|
## DIFF TO REVIEW
|
||||||
|
|
||||||
|
|||||||
Reference in New Issue
Block a user