fix(router): decommission legacy GPU router across contracts and fleet scripts #74

Merged
kagentz-bot merged 2 commits from fix/decommission-router-20260911 into master 2026-09-11 19:19:14 +00:00
11 changed files with 96 additions and 98 deletions
+29 -33
View File
@@ -19,7 +19,7 @@ triggers:
- on model add/remove
- on GPU health degradation
- on agent key rotation
- on router restart (roster must be loaded)
- on harness container restart (LiteLLM reloads its model list)
---
## Maintains
@@ -36,46 +36,42 @@ triggers:
- prometheus: { status: "running", targets: 5 } — Scrapes GPU :9400 exporters + LiteLLM
- port_conflict_detection: { status: "active" } — All 3 GPU wrappers detect ghost processes before binding
## Fleet Topology (Current — July 2026)
## Fleet Topology (Current — 2026-09-11, router decommissioned)
```
┌──────────────────────────────────────────────────────────────────┐
│ CT 116 (192.168.68.116) — Inference Harness Host │
│ │
│ nginx:80 (entrypoint) │
│ ├─ /v1/* → harness-litellm:4000 (API requests) │
│ ├─ /v1/* → harness-litellm:4000 (API requests) │
│ ├─ /admin/* → harness-litellm:4000 (admin endpoints) │
│ ├─ /dashboard/ → harness-dashboard:3000 (harness UI) │
│ ├─ /litellm/* → harness-litellm:4000 (LiteLLM UI + API) │
│ ├─ /litellm/* → harness-litellm:4000 (LiteLLM UI + API) │
│ ├─ /health/* → harness-litellm:4000 (health probes) │
│ └─ /gpu/* → 192.168.68.24:9100 (fleet monitor) │
│ │
│ Containers: │
│ ┌──────────┐ ┌──────────┐ ┌──────────┐ ┌──────────┐ │
│ │ LiteLLM │ │ Router │ │Dashboard │ │ Grafana │ │
│ │ :4000 │ │ :9000 │ │ :3000 │ │ :3000 │ │
│ │ keys+sync│ │deprecated│ │ harness │ │ Prometheus│ │
│ │ fallback │ │not in │ │ UI │ │ data src │ │
│ └──────────┘ └───┬──────┘ └──────────┘ └──────────┘ │
│ │ │
│ ┌──────────┐ ┌──────────┐ ┌──────────┐ │
│ │PostgreSQL│ │ Redis │ │Prometheus│ │
│ │ :5432 │ │ :6379 │ │ :9090 │ │
│ └──────────┘ └──────────┘ └──────────┘ │
└────────────────────┼─────────────────────────────────────────────┘
│
┌───────────────┼───────────────┬──────────────────┐
│ │ │ │
┌────▼─────┐ ┌──────▼──────┐ ┌────▼──────┐ ┌───────▼──────┐
│ CT 8 │ │ CT 110 │ │ CT 15 │ │ pi (.24) │
│ RTX 3090 │ │ RTX 5070 │ │ Strix Halo│ │ GPU Monitor │
│ 24GB │ │ 12GB │ │ 64GB UMA │ │ :9100 │
│ 128K ctx │ │ 128K ctx │ │ 128K ctx │ │ Watchdog │
│ qwen3.6 │ │ Qwen3.5-9B │ │ qwen3.5 │ │ Prometheus │
│ 27B-code │ │ :8080 │ │ -35B-udq4 │ │ exporter │
│ :8080 │ │ :9400 (exp) │ │ :9400(exp)│ │ :9401 │
│ :9400 │ └─────────────┘ └───────────┘ └──────────────┘
└──────────┘
│ ┌──────────┐ ┌──────────┐ ┌──────────┐ │
│ │ LiteLLM │ │Dashboard │ │ Grafana │ │
│ │ :4000 │ │ :3000 │ │ :3001 │ │
│ │ keys+sync│ │ harness │ │Prometheus│ │
│ │ fallback │ │ UI │ │ data src │ │
│ └──────────┘ └──────────┘ └──────────┘ │
│ │
│ ┌──────────┐ ┌──────────┐ ┌──────────┐ │
│ │PostgreSQL│ │ Redis │ │Prometheus│ │
│ │ :5432 │ │ :6379 │ │ :9090 │ │
│ └──────────┘ └──────────┘ └──────────┘ │
└───────┼──────────────────────────────────────────────────────────┘
│
┌─┴─────────────┬───────────────┬───────────────┐
│ │ │ │
┌─────▼──────┐ ┌─────▼──────┐ ┌─────▼──────┐ ┌─────▼──────┐
│ CT 8 │ │ CT 110 │ │ CT 15 │ │ pi (.24) │
│ RTX 3090 │ │ RTX 5070 │ │ Strix Halo │ │ GPU Monitor│
│ 24GB │ │ 12GB │ │ 64GB UMA │ │ :9100 │
│ :8080 │ │ :9400 exp │ │ :9400 exp │ │ :9401 │
└────────────┘ └────────────┘ └────────────┘ └────────────┘
```
## Stable Role-Based Aliases (Introduced 2026-07-15)
@@ -105,7 +101,7 @@ but are deprecated for agent configs. Only the stable aliases survive model swap
|-------|-----|--------|---------|---------|
| Qwen3.8-27B-Uncensored-Q4_K_M | RTX 3090 (.8:8080) | **0.55** | 500 | **300s** |
Note: All syslog-auto entries route directly to GPUs with `api_key: not-needed`. The router (port 9000) is NOT in the inference path.
Note: All syslog-auto entries route directly to GPUs with `api_key: not-needed`. The router (port 9000) was decommissioned 2026-09-11 and is NOT in the inference path.
### Direct Model Endpoints
@@ -153,7 +149,7 @@ Note: All syslog-auto entries route directly to GPUs with `api_key: not-needed`.
6. Cleanup model files (optional)
### heal
1. Check all GPUs via router internal `:9000/health/unified`
1. Check all GPUs via gpu-monitor `{{gpu_dashboard_url}}/gpu-data`
2. Check LiteLLM health via nginx `:80/litellm/health/liveliness`
3. Reset stuck circuit breakers if idle (Redis)
4. Restart dead llama-server instances via SSH
@@ -161,7 +157,7 @@ Note: All syslog-auto entries route directly to GPUs with `api_key: not-needed`.
6. Verify GPU monitor server is running on pi (:9100)
7. Verify watchdog is running on pi
8. Restart router if roster not loaded (check logs for STARTUP ROSTER)
9. Reload roster via `POST :9000/admin/roster/reload` if available
9. Router roster reload — REMOVED 2026-09-11 (router decommissioned; LiteLLM fallbacks handle routing)
### sync-keys
1. List all agent keys in LiteLLM DB via `GET /key/list`
@@ -258,7 +254,7 @@ If no SSH access, send Zulip DM via abiba-bot with vault update instructions.
Benchmarks from 2026-07-17. Strix Halo model: qwen3.6-35B-udq4. RTX 5070 MTP provides 2.7x speedup over pre-upgrade 70 tok/s.
All 3 GPUs now at 128K context (2026-07-17, reduced from 256K for stability).
Benchmarks run through LiteLLM proxy (192.168.68.116:4001) every 5 minutes.
Benchmarks run through LiteLLM proxy (192.168.68.116:4000) every 5 minutes.
Degradation alerts fire at 30% (warning) and 50% (critical) below baseline.
History stored at `/root/data/toks-history.json` with 7-day rolling window.
+17 -17
View File
@@ -2,7 +2,7 @@
kind: responsibility
name: gpu-monitor
description: >
Comprehensive GPU fleet monitor — polls every subsystem (sidecars, router,
Comprehensive GPU fleet monitor — polls every subsystem (sidecars,
LiteLLM, Strix Halo, dashboard) every 15s, renders a live HTML dashboard,
checks alert thresholds, and exposes a JSON API for downstream consumers.
agent: abiba
@@ -46,19 +46,19 @@ and `http://192.168.68.110/health` answer `000`. Never use a bare-port-80 probe
as a GPU liveness signal: on 2026-09-09 that produced three false
`DEGRADED — GPU-rtx3090 000, GPU-rtx5070 000` rounds while
`http://192.168.68.8:8080/health` and `http://192.168.68.110:8080/health`
answered `200`. Port 80 is valid only on the router (.116), never on a GPU host.
answered `200`. Port 80 is valid only on the harness host (.116), never on a GPU host.
### Subsystems Polled
| Subsystem | Endpoint | Frequency | Metrics |
|-----------|----------|-----------|---------|
| GPU Status (all, via router) | `http://192.168.68.116/health/unified` | 15s | models, CB, scores, GPU status (router probes each GPU /health directly) — `301` → `/gpu/gpu-data` is **alive** |
| GPU Status (all, via fleet API) | `http://192.168.68.116/gpu/gpu-data` | 15s | models, CB, scores, GPU status from gpu-monitor on .24:9100 |
| GPU .8 (RTX 3090) health | `http://192.168.68.8:8080/health` | 15s | direct liveness fallback — **:8080 ONLY, never bare port 80** |
| GPU .110 (RTX 5070) health | `http://192.168.68.110:8080/health` | 15s | direct liveness fallback — **:8080 ONLY, never bare port 80** |
| Router (unified) | `http://192.168.68.116/health/unified` | 15s | models, CB, scores, GPU status (`301` → `/gpu/gpu-data` = alive) |
| Router (basic) | `http://192.168.68.116/health` | 15s | basic aliveness |
| Fleet (unified) | `http://192.168.68.116/health/unified` | 15s | nginx `301` → `/gpu/gpu-data` served by gpu-monitor = alive (router decommissioned 2026-09-11) |
| Harness (basic) | `http://192.168.68.116/health` | 15s | nginx → LiteLLM `/health/liveliness` |
| LiteLLM | `http://192.168.68.116/litellm/health` | 15s | proxy health, model count |
| Strix Halo | `http://192.168.68.116/health/unified` (router) | 15s | Strix Halo status via router — cannot poll .15:8080 directly (firewalled to .116 only) |
| Strix Halo | `http://192.168.68.116/health/unified` (nginx → fleet API) | 15s | Strix Halo status via gpu-monitor — cannot poll .15:8080 directly (firewalled to .116 only) |
| Dashboard | `http://192.168.68.116/dashboard/` | 15s | harness-dashboard aliveness |
### Alert Delivery
@@ -75,8 +75,8 @@ This replaces the previous DM-only delivery. All agents on the mesh can see and
### Liveness rule (scoped)
The any-HTTP-response rule applies ONLY to redirect/auth-gated liveness
endpoints, where any HTTP answer proves a listener is up. Applied here: the
router's `/health/unified` answers `301 Moved Permanently` → `/gpu/gpu-data`
endpoints, where any HTTP answer proves a listener is up. Applied here: nginx's
`/health/unified` answers `301 Moved Permanently` → `/gpu/gpu-data`
(the same payload) and LiteLLM's `/litellm/health` answers `301` →
`/litellm/health/liveliness`. For those endpoints a probe is **ALIVE** on
**ANY** HTTP status — `3xx` redirects and `401`/`403` auth challenges included —
@@ -84,8 +84,8 @@ and **DOWN = connection refused (`000`) or timeout only**. Same scoped rule as
zulip-health (Tanko) and infrastructure-monitoring.
Probes whose success condition is specifically a bare `200` are NOT covered by
the any-HTTP rule. On those — the GPU `:8080/health` endpoints, the router
`/health`, and the dashboard — an unexpected status (`401`/`403`, `5xx`, or
the any-HTTP rule. On those — the GPU `:8080/health` endpoints, nginx `/health`,
and the dashboard — an unexpected status (`401`/`403`, `5xx`, or
anything other than the expected `200`) is an **ALERT**, not "alive".
| Metric | Warning | Critical |
@@ -93,7 +93,7 @@ anything other than the expected `200`) is an **ALERT**, not "alive".
| GPU Temp | >80°C | >90°C |
| VRAM Usage | >90% | >95% |
| GPU Util | >95% | >98% |
| Sidecar Unreachable | — | info (sidecars not deployed — use router /health/unified) |
| Sidecar Unreachable | — | info (sidecars not deployed — use gpu-monitor /gpu-data) |
| Model Down | — | critical (circuit breaker open) |
### JSON API Response Schema (/gpu-data)
@@ -174,13 +174,13 @@ Open `http://localhost:9100/` in browser — Live HTML dashboard
pkill -f gpu-monitor-server.py
python3 /root/scripts/gpu-monitor-server.py &
```
Or via PM2: `pm2 restart gpu-monitor`
Managed by systemd (verified 2026-09-11): `systemctl restart gpu-monitor`
### check-router
The router health is accessed through nginx on port 80 on the **router**
(.116) — NOT port 9000 directly, and NOT bare port 80 on a GPU host.
`curl http://192.168.68.116/health/unified` — Router unified health via nginx proxy; answers `301` → `/gpu/gpu-data` (same payload) = alive
`curl http://192.168.68.116:9000/health/unified` — ❌ WILL FAIL (port bound to 127.0.0.1 only)
### check-fleet
Fleet health is accessed through nginx on port 80 on the harness host
(.116) — NOT port 9000 (router decommissioned 2026-09-11), and NOT bare
port 80 on a GPU host.
`curl http://192.168.68.116/health/unified` — nginx answers `301` → `/gpu/gpu-data` (fleet monitor payload) = alive
`curl http://192.168.68.8/health` — ❌ NEVER USE (GPU host, no port-80 listener → false `000`/DEGRADED)
## Configuration Files
+5 -5
View File
@@ -9,7 +9,7 @@ description: >
(v2.1.0) with active remediation rules, Prometheus metrics consumption,
VRAM trend analysis, and predictive alerting.
UPDATED 2026-07-18: Model assignments synced to 2026-07-17 swaps.
Router (port 9000) references replaced with direct GPU routing.
Router (port 9000) DECOMMISSIONED 2026-09-11; references replaced with direct GPU routing.
Benchmark baselines refreshed to live values.
Prometheus exporters removed — not deployed; fall back to direct sidecar probes.
Stable role-based aliases (strix-moe, gpu-dense, gpu-light) from gpu-fleet.
@@ -51,7 +51,7 @@ depends_on:
| `strix-moe` | Strix Halo 64GB | ct15 (.15:8080) | qwen3.6-35B-udq4 | ~10/64GB (16%) | 128K | 62.9 | Compression, summarization, long docs |
Key notes:
- All models use direct GPU routing via LiteLLM (`api_key: not-needed`). Router (port 9000) is deprecated and NOT in the inference path.
- All models use direct GPU routing via LiteLLM (`api_key: not-needed`). Router (port 9000) was decommissioned 2026-09-11 and is NOT in the inference path.
- Stable aliases (gpu-dense, gpu-light, strix-moe) from gpu-fleet are the canonical names for agent configs. Model-specific names still work but are deprecated.
- RTX 5070 tok/s is ~145 for Qwen3.5-9B — gpu-light is the fastest endpoint. Route vision/web/light work there first. NOTE: Qwen3.5-9B is multimodal (image+text), NOT text-only like gemma-4-12b was.
- Strix Halo is 62.9 tok/s (89% of 70.5 baseline) — below optimal but stable. Check for competing workloads.
@@ -104,10 +104,10 @@ Key notes:
### Rule 5: Circuit Breaker Stuck Open
- **Detect**: Circuit breaker open >10 minutes with GPU reporting healthy
- **Note**: Router (port 9000) is deprecated. If circuit breakers are reported by gpu-monitor, they come from LiteLLM's internal tracking, not the old router.
- **Note**: Router (port 9000) was decommissioned 2026-09-11. If circuit breakers are reported by gpu-monitor, they come from LiteLLM's internal tracking, not the old router.
- **Fix**:
1. Verify GPU /health returns 200 on direct port (:8080)
2. If GPU healthy, alert but do NOT reset via router API (deprecated)
2. If GPU healthy, alert but do NOT reset via router API (decommissioned 2026-09-11)
3. Check LiteLLM health directly: http://192.168.68.116/litellm/health/liveliness
4. Restart LiteLLM container on CT 116 if circuit breakers are stuck
- **Verify**: LiteLLM returns healthy, circuit breaker clears within 60s
@@ -276,7 +276,7 @@ Pushed to `SyslogSolution/health-logs/gpu/{run_id}.json` — versioned, searchab
2. **Model restart**: ✅ Only if >50% failure rate over 60s + 30s grace period. Not on single stuck request.
3. **Strix direct access**: ✅ Open firewall .15:8080 → .24 for direct health probe + restart.
4. **VRAM thresholds**: Tiered — **300MB/h** (RTX 3090), **300MB/h** (RTX 5070), 200MB/h (Strix). Previous values (100/50) were too sensitive; raised 2026-07-18 based on operational data.
5. **CB auto-reset**: ✅ Router deprecated — circuit breakers go through LiteLLM health check + container restart if needed. No per-GPU auto-reset.
5. **CB auto-reset**: ✅ Router decommissioned 2026-09-11 — circuit breakers go through LiteLLM health check + container restart if needed. No per-GPU auto-reset.
6. **Benchmark baseline**: Rolling 30-day average, recalculated weekly. Current baselines live in gpu-monitor.
7. **Predictive alerts**: Two-tier — warn at >70°C+rising (>2°C/min), critical at >80°C+rising.
8. **Prometheus**: ❌ Not deployed. Use direct sidecar probes (:8080/health) and gpu-monitor API. Prometheus integration deferred until exporters are running on GPU hosts.
+1 -1
View File
@@ -39,7 +39,7 @@ runtime via `infisical run --` wrapper. Plaintext keys removed from this baselin
```
Infisical vault → infisical run -- hermes gateway → LITELLM_API_KEY (runtime)
↓
Agent (systemd) → LITELLM_API_KEY → LiteLLM (:116/v1) → Router (:9000) → GPU (llama-server)
Agent (systemd) → LITELLM_API_KEY → LiteLLM (:116/v1) → GPU (llama-server)
└── Key DB (Postgres)
```
+3 -7
View File
@@ -199,15 +199,14 @@ description: >
### Ecosystem B: CT 116 syslog-api (192.168.68.116)
8 containers in inference-harness stack:
11 containers in inference-harness stack (verified live 2026-09-11; LiteLLM upgraded 1.90.0-rc.1 -> 1.99.1):
| Container | Image | Port | Role |
|-----------|-------|------|------|
| harness-litellm | berriai/litellm:1.90.0-rc.1 | :4000→:4001 | API proxy, key mgmt, fallbacks |
| harness-router | inference-harness-router | :9000 (127.0.0.1) | GPU routing, slot booking, CB |
| harness-litellm | ghcr.io/docker.litellm.ai/berriai/litellm:1.99.1 | :4000→:4000 | API proxy, key mgmt, fallbacks |
| harness-nginx | nginx:alpine | :80 | Entrypoint, /v1→LiteLLM, /dashboard/ |
| harness-postgres | postgres:16-alpine | :5432 | LiteLLM DB (keys, spend, config) |
| harness-redis | redis:7-alpine | :6379 | Router slots, circuit breakers |
| harness-redis | redis:7-alpine | :6379 | LiteLLM cache + rate-limit state |
| harness-dashboard | inference-harness-dashboard | :3000 | SyslogAI Harness UI |
| harness-grafana | grafana/grafana | :3000→:3001 (direct LAN, not behind nginx) | GPU + Proxmox + Docker dashboards |
| harness-prometheus | prom/prometheus | :9090 | Metrics scraper, 6 jobs |
@@ -588,9 +587,6 @@ curl -s -H "Authorization: Bearer $(infisical secrets get LITELLM_MASTER_KEY --p
# Storage check
ssh root@192.168.68.7 "df -h /media/storage /media/mediastore"
# Router roster reload (if needed)
curl -s -X POST http://192.168.68.116:9000/admin/roster/reload \
-H "Authorization: Bearer sk-admin-ee09fffd04978b61a1569ac670c68814"
# Restart stuck GPU (saturation watchdog alternative)
ssh root@192.168.68.8 "systemctl restart llama-server"
+2 -2
View File
@@ -173,7 +173,7 @@ curl -s http://192.168.68.116:3001/api/health | jq '{status, version}'
# Expected: {"status":"ok","version":"..."}
# LiteLLM metrics (Prometheus endpoint)
curl -s http://192.168.68.116:4001/metrics | head -20
curl -s http://192.168.68.116:4000/metrics | head -20
# Expected: Prometheus-formatted metrics output
```
@@ -242,5 +242,5 @@ curl -s http://192.168.68.116:9090/api/v1/targets
curl -s http://192.168.68.116:3001/api/health
# LiteLLM metrics (already live)
curl -s http://192.168.68.116:4001/metrics | head -20
curl -s http://192.168.68.116:4000/metrics | head -20
```
+1 -1
View File
@@ -206,7 +206,7 @@ mcp_servers:
| Key | MCP Access |
|-----|-----------|
| Master key | ✅ Full — 90 tools (vault-injected) |
| Agent keys (mumuni, tanko, etc.) | ❌ Per-key grants not supported in v1.90.0-rc.1 |
| Agent keys (mumuni, tanko, etc.) | ❌ Per-key grants not supported in v1.99.1 |
### Known Limitations
- Per-key MCP server grants not functional — only master key has access
+12 -13
View File
@@ -12,8 +12,8 @@ note: >
description: >
Verifies the LiteLLM inference stack health. Current architecture (2026-07-09):
nginx:80 → LiteLLM:4000 → GPU(llama-server) via direct proxy.
Router (harness-router :9000) is DEPRECATED — container still runs but
is not in the request path. GPU monitoring via Prometheus/Grafana and
Router (harness-router :9000) was DECOMMISSIONED 2026-09-11 (container,
image and config removed). GPU monitoring via Prometheus/Grafana and
fleet dashboard (gpu-monitor :9100).
Designed as a reusable contract for any Syslog agent.
@@ -42,8 +42,8 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2)
│
Grafana :3001
harness-router :9000 — DEPRECATED, container still runs but
NOT in request path. nginx routes /v1 → LiteLLM directly.
harness-router :9000 — DECOMMISSIONED 2026-09-11 (container, image
and config removed). nginx routes /v1 → LiteLLM directly.
Router slot booking + circuit breakers replaced by
LiteLLM native fallbacks + timeouts.
```
@@ -100,8 +100,7 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2)
| Container | Image | Port | Health Check |
|-----------|-------|------|-------------|
| harness-litellm | berriai/litellm:1.90.0-rc.1 | :4000→:4001 | /health/liveliness |
| harness-router | inference-harness-router | :9000 (127.0.0.1) | /health (DEPRECATED — not in path) |
| harness-litellm | ghcr.io/docker.litellm.ai/berriai/litellm:1.99.1 | :4000→:4000 | /health/liveliness |
| harness-nginx | nginx:alpine | :80 | HTTP 200 on /health |
| harness-postgres | postgres:16-alpine | :5432 | pg_isready |
| harness-redis | redis:7-alpine | :6379 | PING |
@@ -121,14 +120,14 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2)
- GET http://{{backend_host}}/litellm/health/liveliness → expect 200
4. **Check backend container health**:
- SSH to {{backend_host}} → `docker ps` → verify 8 containers healthy
- Critical: harness-litellm, harness-router, harness-nginx, harness-postgres
- Monitoring: harness-redis, harness-dashboard, harness-grafana, harness-prometheus
- SSH to {{backend_host}} → `docker ps` → verify 11 containers healthy
- Critical: harness-litellm, harness-nginx, harness-postgres
- Monitoring: harness-redis, harness-dashboard, harness-grafana, harness-prometheus,
harness-alertmanager, harness-zulip-bridge, harness-docker-stats, harness-pve-exporter
5. **Check router roster loaded**:
- GET http://{{backend_host}}:9000/health → expect 200
- GET http://{{backend_host}}:9000/health/unified → expect 3 models
- If router returns "all GPUs saturated" but GPUs idle: roster not loaded → reload
5. **Check GPU fleet health via gpu-monitor** (router decommissioned 2026-09-11):
- GET {{gpu_dashboard_url}}/gpu-data → expect 200 with GPU metrics JSON
- nginx `/health/unified` is now a `301` redirect to `/gpu/gpu-data` (same payload)
6. **Check GPU fleet health** (via fleet dashboard):
- GET {{gpu_dashboard_url}}/gpu-data → expect 200 with GPU metrics JSON
+11 -10
View File
@@ -20,7 +20,7 @@ note: >
Last verified: 2026-07-12
description: >
LiteLLM inference stack health monitoring + self-healing. Verifies the full
nginx → LiteLLM → GPU chain, 8 containers on CT 116, 3 GPU hosts, model
nginx → LiteLLM → GPU chain, 11 containers on CT 116, 3 GPU hosts, model
inference, and agent keys. Applies remediation rules for common failures.
Reports every action via Zulip DM and Gitea (SyslogSolution/health-logs).
---
@@ -41,8 +41,8 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2)
│
Grafana :3001
harness-router :9000 — DEPRECATED, container still runs but
NOT in request path. nginx routes /v1 → LiteLLM directly.
harness-router :9000 — DECOMMISSIONED 2026-09-11 (container, image
and config removed). nginx routes /v1 → LiteLLM directly.
Router slot booking + circuit breakers replaced by
LiteLLM native fallbacks + timeouts.
```
@@ -98,8 +98,7 @@ Preferred implementation: uncap shared pool, add capped alias for crew-only.
| Container | Image | Port | Health Check |
|-----------|-------|------|-------------|
| harness-litellm | berriai/litellm:1.90.0-rc.1 | :4000→:4001 | /health/liveliness |
| harness-router | inference-harness-router | :9000 (127.0.0.1) | /health (DEPRECATED — not in path) |
| harness-litellm | ghcr.io/docker.litellm.ai/berriai/litellm:1.99.1 | :4000→:4000 | /health/liveliness |
| harness-nginx | nginx:alpine | :80 | HTTP 200 on /health |
| harness-postgres | postgres:16-alpine | :5432 | pg_isready |
| harness-redis | redis:7-alpine | :6379 | PING |
@@ -146,10 +145,11 @@ Run this first on every cycle. Results feed into remediation rules below.
- GET http://{{backend_host}}/litellm/health/liveliness → expect 200
### 3. Check backend container health
- SSH to {{backend_host}} → `docker ps` → verify 10 containers healthy
- SSH to {{backend_host}} → `docker ps` → verify 11 containers healthy
- Critical: harness-litellm, harness-nginx, harness-postgres
- Monitoring: harness-redis, harness-dashboard, harness-grafana, harness-prometheus, harness-docker-stats, harness-pve-exporter
- Deprecated but running: harness-router (not in path, reference only)
- Monitoring: harness-redis, harness-dashboard, harness-grafana, harness-prometheus,
harness-alertmanager, harness-zulip-bridge, harness-docker-stats, harness-pve-exporter
- Decommissioned 2026-09-11: harness-router (container, image and config removed)
### 4. Check GPU fleet health (via fleet dashboard)
- GET {{gpu_dashboard_url}}/gpu-data → expect 200 with GPU metrics JSON
@@ -215,8 +215,9 @@ Fix → generate keys in LiteLLM via /key/generate → update /etc/environment o
Escalate → if SSH access unavailable, send Zulip DM
### Rule 9: Stale Active Counter in Redis — DEPRECATED
Router no longer in path so Redis active counters are unused. Rule retained
for reference but inactive. If Redis issues occur, check harness-redis container.
Router no longer in path so router active-slot counters are unused. Rule retained
for reference but inactive. `harness-redis` now serves only LiteLLM cache and
rate-limit state; check the container if cache errors appear.
---
---
+9 -5
View File
@@ -213,10 +213,11 @@ def collect():
# ── LiteLLM Specific Checks (from litellm-health prose contract) ──
report["litellm"] = {"checks": []}
# Check 1: LiteLLM aggregate health endpoint (router binds to 127.0.0.1, check via SSH)
health_unified = ssh(LITELLM_BACKEND, "curl -sf http://127.0.0.1:9000/health/unified -o /dev/null -w '%{http_code}' 2>/dev/null")
# Check 1: Fleet health via nginx /health/unified (301 -> /gpu/gpu-data served by gpu-monitor on .24:9100).
# Router (:9000) was decommissioned 2026-09-11; probing it was a guaranteed daily failure.
health_unified = ssh(LITELLM_BACKEND, "curl -s -o /dev/null -w '%{http_code}' --max-time 8 http://127.0.0.1/health/unified 2>/dev/null")
report["litellm"]["health_unified"] = health_unified or "000"
report["litellm"]["checks"].append({"name": "unified-health", "status": "pass" if health_unified == "200" else "fail", "code": health_unified or "000"})
report["litellm"]["checks"].append({"name": "fleet-health-via-nginx", "status": "pass" if health_unified in ("200", "301") else "fail", "code": health_unified or "000"})
# Check 2: Nginx-proxied internal endpoints
for path, name in [("/litellm/ui/", "nginx-ui"), ("/litellm/docs", "nginx-docs")]:
@@ -224,8 +225,11 @@ def collect():
report["litellm"]["checks"].append({"name": name, "status": "pass" if code == "200" else "fail", "code": code})
# Check 3: Docker container health for LiteLLM stack
expected_containers = ["harness-litellm", "harness-nginx", "harness-router",
"harness-postgres", "harness-redis", "harness-dashboard"]
expected_containers = ["harness-litellm", "harness-nginx", "harness-postgres",
"harness-redis", "harness-dashboard", "harness-grafana",
"harness-prometheus", "harness-alertmanager",
"harness-zulip-bridge", "harness-docker-stats",
"harness-pve-exporter"]
actual_names = [c["name"] for c in containers2]
report["litellm"]["expected_containers"] = expected_containers
report["litellm"]["missing_containers"] = [e for e in expected_containers if e not in actual_names]
+6 -4
View File
@@ -66,11 +66,13 @@ The infrastructure-control.prose.md contract is the canonical reference for the
2. NO .19 IP — Zulip is CT 117 on storepve.
3. NO CT 122/123 — Tanko=CT 112, Mumuni=CT 100 (inside Abiba). No CT 114 anywhere.
4. Strix Halo :8080 is FIREWALLED to .116 only — cannot be probed from abiba (.24).
5. abiba-zulip PM2 process is DECOMMISSIONED (2026-07-04) — abiba uses Telegram only.
5. abiba-zulip PM2 process is ONLINE (verified 2026-09-11) — this rule was stale.
**Docker on CT 116 (8 containers):**
harness-litellm, harness-router, harness-nginx, harness-postgres,
harness-redis, harness-dashboard, harness-grafana, harness-prometheus
**Docker on CT 116 (11 containers, verified 2026-09-11):**
harness-litellm, harness-nginx, harness-postgres, harness-redis,
harness-dashboard, harness-grafana, harness-prometheus, harness-alertmanager,
harness-zulip-bridge, harness-docker-stats, harness-pve-exporter
(harness-router was decommissioned 2026-09-11)
## DIFF TO REVIEW