diff --git a/gpu-fleet.prose.md b/gpu-fleet.prose.md index f3b4aab..e539ee4 100644 --- a/gpu-fleet.prose.md +++ b/gpu-fleet.prose.md @@ -19,7 +19,7 @@ triggers: - on model add/remove - on GPU health degradation - on agent key rotation - - on router restart (roster must be loaded) + - on harness container restart (LiteLLM reloads its model list) --- ## Maintains @@ -36,46 +36,42 @@ triggers: - prometheus: { status: "running", targets: 5 } — Scrapes GPU :9400 exporters + LiteLLM - port_conflict_detection: { status: "active" } — All 3 GPU wrappers detect ghost processes before binding -## Fleet Topology (Current — July 2026) +## Fleet Topology (Current — 2026-09-11, router decommissioned) ``` ┌──────────────────────────────────────────────────────────────────┐ │ CT 116 (192.168.68.116) — Inference Harness Host │ │ │ │ nginx:80 (entrypoint) │ -│ ├─ /v1/* → harness-litellm:4000 (API requests) │ +│ ├─ /v1/* → harness-litellm:4000 (API requests) │ │ ├─ /admin/* → harness-litellm:4000 (admin endpoints) │ │ ├─ /dashboard/ → harness-dashboard:3000 (harness UI) │ -│ ├─ /litellm/* → harness-litellm:4000 (LiteLLM UI + API) │ +│ ├─ /litellm/* → harness-litellm:4000 (LiteLLM UI + API) │ │ ├─ /health/* → harness-litellm:4000 (health probes) │ │ └─ /gpu/* → 192.168.68.24:9100 (fleet monitor) │ │ │ │ Containers: │ -│ ┌──────────┐ ┌──────────┐ ┌──────────┐ ┌──────────┐ │ -│ │ LiteLLM │ │ Router │ │Dashboard │ │ Grafana │ │ -│ │ :4000 │ │ :9000 │ │ :3000 │ │ :3000 │ │ -│ │ keys+sync│ │deprecated│ │ harness │ │ Prometheus│ │ -│ │ fallback │ │not in │ │ UI │ │ data src │ │ -│ └──────────┘ └───┬──────┘ └──────────┘ └──────────┘ │ -│ │ │ -│ ┌──────────┐ ┌──────────┐ ┌──────────┐ │ -│ │PostgreSQL│ │ Redis │ │Prometheus│ │ -│ │ :5432 │ │ :6379 │ │ :9090 │ │ -│ └──────────┘ └──────────┘ └──────────┘ │ -└────────────────────┼─────────────────────────────────────────────┘ - │ - ┌───────────────┼───────────────┬──────────────────┐ - │ │ │ │ -┌────▼─────┐ ┌──────▼──────┐ ┌────▼──────┐ ┌───────▼──────┐ -│ CT 8 │ │ CT 110 │ │ CT 15 │ │ pi (.24) │ -│ RTX 3090 │ │ RTX 5070 │ │ Strix Halo│ │ GPU Monitor │ -│ 24GB │ │ 12GB │ │ 64GB UMA │ │ :9100 │ -│ 128K ctx │ │ 128K ctx │ │ 128K ctx │ │ Watchdog │ -│ qwen3.6 │ │ Qwen3.5-9B │ │ qwen3.5 │ │ Prometheus │ -│ 27B-code │ │ :8080 │ │ -35B-udq4 │ │ exporter │ -│ :8080 │ │ :9400 (exp) │ │ :9400(exp)│ │ :9401 │ -│ :9400 │ └─────────────┘ └───────────┘ └──────────────┘ -└──────────┘ +│ ┌──────────┐ ┌──────────┐ ┌──────────┐ │ +│ │ LiteLLM │ │Dashboard │ │ Grafana │ │ +│ │ :4000 │ │ :3000 │ │ :3001 │ │ +│ │ keys+sync│ │ harness │ │Prometheus│ │ +│ │ fallback │ │ UI │ │ data src │ │ +│ └──────────┘ └──────────┘ └──────────┘ │ +│ │ +│ ┌──────────┐ ┌──────────┐ ┌──────────┐ │ +│ │PostgreSQL│ │ Redis │ │Prometheus│ │ +│ │ :5432 │ │ :6379 │ │ :9090 │ │ +│ └──────────┘ └──────────┘ └──────────┘ │ +└───────┼──────────────────────────────────────────────────────────┘ + │ + ┌─┴─────────────┬───────────────┬───────────────┐ + │ │ │ │ +┌─────▼──────┐ ┌─────▼──────┐ ┌─────▼──────┐ ┌─────▼──────┐ +│ CT 8 │ │ CT 110 │ │ CT 15 │ │ pi (.24) │ +│ RTX 3090 │ │ RTX 5070 │ │ Strix Halo │ │ GPU Monitor│ +│ 24GB │ │ 12GB │ │ 64GB UMA │ │ :9100 │ +│ :8080 │ │ :9400 exp │ │ :9400 exp │ │ :9401 │ +└────────────┘ └────────────┘ └────────────┘ └────────────┘ ``` ## Stable Role-Based Aliases (Introduced 2026-07-15) @@ -105,7 +101,7 @@ but are deprecated for agent configs. Only the stable aliases survive model swap |-------|-----|--------|---------|---------| | Qwen3.8-27B-Uncensored-Q4_K_M | RTX 3090 (.8:8080) | **0.55** | 500 | **300s** | -Note: All syslog-auto entries route directly to GPUs with `api_key: not-needed`. The router (port 9000) is NOT in the inference path. +Note: All syslog-auto entries route directly to GPUs with `api_key: not-needed`. The router (port 9000) was decommissioned 2026-09-11 and is NOT in the inference path. ### Direct Model Endpoints @@ -153,7 +149,7 @@ Note: All syslog-auto entries route directly to GPUs with `api_key: not-needed`. 6. Cleanup model files (optional) ### heal -1. Check all GPUs via router internal `:9000/health/unified` +1. Check all GPUs via gpu-monitor `{{gpu_dashboard_url}}/gpu-data` 2. Check LiteLLM health via nginx `:80/litellm/health/liveliness` 3. Reset stuck circuit breakers if idle (Redis) 4. Restart dead llama-server instances via SSH @@ -161,7 +157,7 @@ Note: All syslog-auto entries route directly to GPUs with `api_key: not-needed`. 6. Verify GPU monitor server is running on pi (:9100) 7. Verify watchdog is running on pi 8. Restart router if roster not loaded (check logs for STARTUP ROSTER) -9. Reload roster via `POST :9000/admin/roster/reload` if available +9. Router roster reload — REMOVED 2026-09-11 (router decommissioned; LiteLLM fallbacks handle routing) ### sync-keys 1. List all agent keys in LiteLLM DB via `GET /key/list` @@ -258,7 +254,7 @@ If no SSH access, send Zulip DM via abiba-bot with vault update instructions. Benchmarks from 2026-07-17. Strix Halo model: qwen3.6-35B-udq4. RTX 5070 MTP provides 2.7x speedup over pre-upgrade 70 tok/s. All 3 GPUs now at 128K context (2026-07-17, reduced from 256K for stability). -Benchmarks run through LiteLLM proxy (192.168.68.116:4001) every 5 minutes. +Benchmarks run through LiteLLM proxy (192.168.68.116:4000) every 5 minutes. Degradation alerts fire at 30% (warning) and 50% (critical) below baseline. History stored at `/root/data/toks-history.json` with 7-day rolling window. diff --git a/gpu-monitor.prose.md b/gpu-monitor.prose.md index 881bac4..a6808e1 100644 --- a/gpu-monitor.prose.md +++ b/gpu-monitor.prose.md @@ -2,7 +2,7 @@ kind: responsibility name: gpu-monitor description: > - Comprehensive GPU fleet monitor — polls every subsystem (sidecars, router, + Comprehensive GPU fleet monitor — polls every subsystem (sidecars, LiteLLM, Strix Halo, dashboard) every 15s, renders a live HTML dashboard, checks alert thresholds, and exposes a JSON API for downstream consumers. agent: abiba @@ -46,19 +46,19 @@ and `http://192.168.68.110/health` answer `000`. Never use a bare-port-80 probe as a GPU liveness signal: on 2026-09-09 that produced three false `DEGRADED — GPU-rtx3090 000, GPU-rtx5070 000` rounds while `http://192.168.68.8:8080/health` and `http://192.168.68.110:8080/health` -answered `200`. Port 80 is valid only on the router (.116), never on a GPU host. +answered `200`. Port 80 is valid only on the harness host (.116), never on a GPU host. ### Subsystems Polled | Subsystem | Endpoint | Frequency | Metrics | |-----------|----------|-----------|---------| -| GPU Status (all, via router) | `http://192.168.68.116/health/unified` | 15s | models, CB, scores, GPU status (router probes each GPU /health directly) — `301` → `/gpu/gpu-data` is **alive** | +| GPU Status (all, via fleet API) | `http://192.168.68.116/gpu/gpu-data` | 15s | models, CB, scores, GPU status from gpu-monitor on .24:9100 | | GPU .8 (RTX 3090) health | `http://192.168.68.8:8080/health` | 15s | direct liveness fallback — **:8080 ONLY, never bare port 80** | | GPU .110 (RTX 5070) health | `http://192.168.68.110:8080/health` | 15s | direct liveness fallback — **:8080 ONLY, never bare port 80** | -| Router (unified) | `http://192.168.68.116/health/unified` | 15s | models, CB, scores, GPU status (`301` → `/gpu/gpu-data` = alive) | -| Router (basic) | `http://192.168.68.116/health` | 15s | basic aliveness | +| Fleet (unified) | `http://192.168.68.116/health/unified` | 15s | nginx `301` → `/gpu/gpu-data` served by gpu-monitor = alive (router decommissioned 2026-09-11) | +| Harness (basic) | `http://192.168.68.116/health` | 15s | nginx → LiteLLM `/health/liveliness` | | LiteLLM | `http://192.168.68.116/litellm/health` | 15s | proxy health, model count | -| Strix Halo | `http://192.168.68.116/health/unified` (router) | 15s | Strix Halo status via router — cannot poll .15:8080 directly (firewalled to .116 only) | +| Strix Halo | `http://192.168.68.116/health/unified` (nginx → fleet API) | 15s | Strix Halo status via gpu-monitor — cannot poll .15:8080 directly (firewalled to .116 only) | | Dashboard | `http://192.168.68.116/dashboard/` | 15s | harness-dashboard aliveness | ### Alert Delivery @@ -75,8 +75,8 @@ This replaces the previous DM-only delivery. All agents on the mesh can see and ### Liveness rule (scoped) The any-HTTP-response rule applies ONLY to redirect/auth-gated liveness -endpoints, where any HTTP answer proves a listener is up. Applied here: the -router's `/health/unified` answers `301 Moved Permanently` → `/gpu/gpu-data` +endpoints, where any HTTP answer proves a listener is up. Applied here: nginx's +`/health/unified` answers `301 Moved Permanently` → `/gpu/gpu-data` (the same payload) and LiteLLM's `/litellm/health` answers `301` → `/litellm/health/liveliness`. For those endpoints a probe is **ALIVE** on **ANY** HTTP status — `3xx` redirects and `401`/`403` auth challenges included — @@ -84,8 +84,8 @@ and **DOWN = connection refused (`000`) or timeout only**. Same scoped rule as zulip-health (Tanko) and infrastructure-monitoring. Probes whose success condition is specifically a bare `200` are NOT covered by -the any-HTTP rule. On those — the GPU `:8080/health` endpoints, the router -`/health`, and the dashboard — an unexpected status (`401`/`403`, `5xx`, or +the any-HTTP rule. On those — the GPU `:8080/health` endpoints, nginx `/health`, +and the dashboard — an unexpected status (`401`/`403`, `5xx`, or anything other than the expected `200`) is an **ALERT**, not "alive". | Metric | Warning | Critical | @@ -93,7 +93,7 @@ anything other than the expected `200`) is an **ALERT**, not "alive". | GPU Temp | >80°C | >90°C | | VRAM Usage | >90% | >95% | | GPU Util | >95% | >98% | -| Sidecar Unreachable | — | info (sidecars not deployed — use router /health/unified) | +| Sidecar Unreachable | — | info (sidecars not deployed — use gpu-monitor /gpu-data) | | Model Down | — | critical (circuit breaker open) | ### JSON API Response Schema (/gpu-data) @@ -174,13 +174,13 @@ Open `http://localhost:9100/` in browser — Live HTML dashboard pkill -f gpu-monitor-server.py python3 /root/scripts/gpu-monitor-server.py & ``` -Or via PM2: `pm2 restart gpu-monitor` +Managed by systemd (verified 2026-09-11): `systemctl restart gpu-monitor` -### check-router -The router health is accessed through nginx on port 80 on the **router** -(.116) — NOT port 9000 directly, and NOT bare port 80 on a GPU host. -`curl http://192.168.68.116/health/unified` — Router unified health via nginx proxy; answers `301` → `/gpu/gpu-data` (same payload) = alive -`curl http://192.168.68.116:9000/health/unified` — ❌ WILL FAIL (port bound to 127.0.0.1 only) +### check-fleet +Fleet health is accessed through nginx on port 80 on the harness host +(.116) — NOT port 9000 (router decommissioned 2026-09-11), and NOT bare +port 80 on a GPU host. +`curl http://192.168.68.116/health/unified` — nginx answers `301` → `/gpu/gpu-data` (fleet monitor payload) = alive `curl http://192.168.68.8/health` — ❌ NEVER USE (GPU host, no port-80 listener → false `000`/DEGRADED) ## Configuration Files diff --git a/gpu-self-heal.prose.md b/gpu-self-heal.prose.md index 1161933..04ef8e4 100644 --- a/gpu-self-heal.prose.md +++ b/gpu-self-heal.prose.md @@ -9,7 +9,7 @@ description: > (v2.1.0) with active remediation rules, Prometheus metrics consumption, VRAM trend analysis, and predictive alerting. UPDATED 2026-07-18: Model assignments synced to 2026-07-17 swaps. - Router (port 9000) references replaced with direct GPU routing. + Router (port 9000) DECOMMISSIONED 2026-09-11; references replaced with direct GPU routing. Benchmark baselines refreshed to live values. Prometheus exporters removed — not deployed; fall back to direct sidecar probes. Stable role-based aliases (strix-moe, gpu-dense, gpu-light) from gpu-fleet. @@ -51,7 +51,7 @@ depends_on: | `strix-moe` | Strix Halo 64GB | ct15 (.15:8080) | qwen3.6-35B-udq4 | ~10/64GB (16%) | 128K | 62.9 | Compression, summarization, long docs | Key notes: -- All models use direct GPU routing via LiteLLM (`api_key: not-needed`). Router (port 9000) is deprecated and NOT in the inference path. +- All models use direct GPU routing via LiteLLM (`api_key: not-needed`). Router (port 9000) was decommissioned 2026-09-11 and is NOT in the inference path. - Stable aliases (gpu-dense, gpu-light, strix-moe) from gpu-fleet are the canonical names for agent configs. Model-specific names still work but are deprecated. - RTX 5070 tok/s is ~145 for Qwen3.5-9B — gpu-light is the fastest endpoint. Route vision/web/light work there first. NOTE: Qwen3.5-9B is multimodal (image+text), NOT text-only like gemma-4-12b was. - Strix Halo is 62.9 tok/s (89% of 70.5 baseline) — below optimal but stable. Check for competing workloads. @@ -104,10 +104,10 @@ Key notes: ### Rule 5: Circuit Breaker Stuck Open - **Detect**: Circuit breaker open >10 minutes with GPU reporting healthy -- **Note**: Router (port 9000) is deprecated. If circuit breakers are reported by gpu-monitor, they come from LiteLLM's internal tracking, not the old router. +- **Note**: Router (port 9000) was decommissioned 2026-09-11. If circuit breakers are reported by gpu-monitor, they come from LiteLLM's internal tracking, not the old router. - **Fix**: 1. Verify GPU /health returns 200 on direct port (:8080) - 2. If GPU healthy, alert but do NOT reset via router API (deprecated) + 2. If GPU healthy, alert but do NOT reset via router API (decommissioned 2026-09-11) 3. Check LiteLLM health directly: http://192.168.68.116/litellm/health/liveliness 4. Restart LiteLLM container on CT 116 if circuit breakers are stuck - **Verify**: LiteLLM returns healthy, circuit breaker clears within 60s @@ -276,7 +276,7 @@ Pushed to `SyslogSolution/health-logs/gpu/{run_id}.json` — versioned, searchab 2. **Model restart**: ✅ Only if >50% failure rate over 60s + 30s grace period. Not on single stuck request. 3. **Strix direct access**: ✅ Open firewall .15:8080 → .24 for direct health probe + restart. 4. **VRAM thresholds**: Tiered — **300MB/h** (RTX 3090), **300MB/h** (RTX 5070), 200MB/h (Strix). Previous values (100/50) were too sensitive; raised 2026-07-18 based on operational data. -5. **CB auto-reset**: ✅ Router deprecated — circuit breakers go through LiteLLM health check + container restart if needed. No per-GPU auto-reset. +5. **CB auto-reset**: ✅ Router decommissioned 2026-09-11 — circuit breakers go through LiteLLM health check + container restart if needed. No per-GPU auto-reset. 6. **Benchmark baseline**: Rolling 30-day average, recalculated weekly. Current baselines live in gpu-monitor. 7. **Predictive alerts**: Two-tier — warn at >70°C+rising (>2°C/min), critical at >80°C+rising. 8. **Prometheus**: ❌ Not deployed. Use direct sidecar probes (:8080/health) and gpu-monitor API. Prometheus integration deferred until exporters are running on GPU hosts. diff --git a/hermes-agent-baseline.prose.md b/hermes-agent-baseline.prose.md index 50d0ae3..5fcd5f1 100644 --- a/hermes-agent-baseline.prose.md +++ b/hermes-agent-baseline.prose.md @@ -39,7 +39,7 @@ runtime via `infisical run --` wrapper. Plaintext keys removed from this baselin ``` Infisical vault → infisical run -- hermes gateway → LITELLM_API_KEY (runtime) ↓ -Agent (systemd) → LITELLM_API_KEY → LiteLLM (:116/v1) → Router (:9000) → GPU (llama-server) +Agent (systemd) → LITELLM_API_KEY → LiteLLM (:116/v1) → GPU (llama-server) └── Key DB (Postgres) ``` diff --git a/infrastructure-control.prose.md b/infrastructure-control.prose.md index fe3e350..cbbdbc0 100644 --- a/infrastructure-control.prose.md +++ b/infrastructure-control.prose.md @@ -199,12 +199,11 @@ description: > ### Ecosystem B: CT 116 syslog-api (192.168.68.116) -8 containers in inference-harness stack: +11 containers in inference-harness stack (verified live 2026-09-11): | Container | Image | Port | Role | |-----------|-------|------|------| -| harness-litellm | berriai/litellm:1.90.0-rc.1 | :4000→:4001 | API proxy, key mgmt, fallbacks | -| harness-router | inference-harness-router | :9000 (127.0.0.1) | GPU routing, slot booking, CB | +| harness-litellm | ghcr.io/docker.litellm.ai/berriai/litellm:1.90.0-rc.1 | :4000→:4000 | API proxy, key mgmt, fallbacks | | harness-nginx | nginx:alpine | :80 | Entrypoint, /v1→LiteLLM, /dashboard/ | | harness-postgres | postgres:16-alpine | :5432 | LiteLLM DB (keys, spend, config) | | harness-redis | redis:7-alpine | :6379 | Router slots, circuit breakers | @@ -588,9 +587,6 @@ curl -s -H "Authorization: Bearer $(infisical secrets get LITELLM_MASTER_KEY --p # Storage check ssh root@192.168.68.7 "df -h /media/storage /media/mediastore" -# Router roster reload (if needed) -curl -s -X POST http://192.168.68.116:9000/admin/roster/reload \ - -H "Authorization: Bearer sk-admin-ee09fffd04978b61a1569ac670c68814" # Restart stuck GPU (saturation watchdog alternative) ssh root@192.168.68.8 "systemctl restart llama-server" diff --git a/infrastructure-monitoring.prose.md b/infrastructure-monitoring.prose.md index 2969ac6..c87c35a 100644 --- a/infrastructure-monitoring.prose.md +++ b/infrastructure-monitoring.prose.md @@ -173,7 +173,7 @@ curl -s http://192.168.68.116:3001/api/health | jq '{status, version}' # Expected: {"status":"ok","version":"..."} # LiteLLM metrics (Prometheus endpoint) -curl -s http://192.168.68.116:4001/metrics | head -20 +curl -s http://192.168.68.116:4000/metrics | head -20 # Expected: Prometheus-formatted metrics output ``` @@ -242,5 +242,5 @@ curl -s http://192.168.68.116:9090/api/v1/targets curl -s http://192.168.68.116:3001/api/health # LiteLLM metrics (already live) -curl -s http://192.168.68.116:4001/metrics | head -20 +curl -s http://192.168.68.116:4000/metrics | head -20 ``` diff --git a/litellm-health.prose.md b/litellm-health.prose.md index 2cd8792..dad97bf 100644 --- a/litellm-health.prose.md +++ b/litellm-health.prose.md @@ -12,8 +12,8 @@ note: > description: > Verifies the LiteLLM inference stack health. Current architecture (2026-07-09): nginx:80 → LiteLLM:4000 → GPU(llama-server) via direct proxy. - Router (harness-router :9000) is DEPRECATED — container still runs but - is not in the request path. GPU monitoring via Prometheus/Grafana and + Router (harness-router :9000) was DECOMMISSIONED 2026-09-11 (container, + image and config removed). GPU monitoring via Prometheus/Grafana and fleet dashboard (gpu-monitor :9100). Designed as a reusable contract for any Syslog agent. @@ -42,8 +42,8 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2) │ Grafana :3001 - harness-router :9000 — DEPRECATED, container still runs but - NOT in request path. nginx routes /v1 → LiteLLM directly. + harness-router :9000 — DECOMMISSIONED 2026-09-11 (container, image + and config removed). nginx routes /v1 → LiteLLM directly. Router slot booking + circuit breakers replaced by LiteLLM native fallbacks + timeouts. ``` @@ -100,8 +100,7 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2) | Container | Image | Port | Health Check | |-----------|-------|------|-------------| -| harness-litellm | berriai/litellm:1.90.0-rc.1 | :4000→:4001 | /health/liveliness | -| harness-router | inference-harness-router | :9000 (127.0.0.1) | /health (DEPRECATED — not in path) | +| harness-litellm | ghcr.io/docker.litellm.ai/berriai/litellm:1.90.0-rc.1 | :4000→:4000 | /health/liveliness | | harness-nginx | nginx:alpine | :80 | HTTP 200 on /health | | harness-postgres | postgres:16-alpine | :5432 | pg_isready | | harness-redis | redis:7-alpine | :6379 | PING | @@ -121,14 +120,14 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2) - GET http://{{backend_host}}/litellm/health/liveliness → expect 200 4. **Check backend container health**: - - SSH to {{backend_host}} → `docker ps` → verify 8 containers healthy - - Critical: harness-litellm, harness-router, harness-nginx, harness-postgres - - Monitoring: harness-redis, harness-dashboard, harness-grafana, harness-prometheus + - SSH to {{backend_host}} → `docker ps` → verify 11 containers healthy + - Critical: harness-litellm, harness-nginx, harness-postgres + - Monitoring: harness-redis, harness-dashboard, harness-grafana, harness-prometheus, + harness-alertmanager, harness-zulip-bridge, harness-docker-stats, harness-pve-exporter -5. **Check router roster loaded**: - - GET http://{{backend_host}}:9000/health → expect 200 - - GET http://{{backend_host}}:9000/health/unified → expect 3 models - - If router returns "all GPUs saturated" but GPUs idle: roster not loaded → reload +5. **Check GPU fleet health via gpu-monitor** (router decommissioned 2026-09-11): + - GET {{gpu_dashboard_url}}/gpu-data → expect 200 with GPU metrics JSON + - nginx `/health/unified` is now a `301` redirect to `/gpu/gpu-data` (same payload) 6. **Check GPU fleet health** (via fleet dashboard): - GET {{gpu_dashboard_url}}/gpu-data → expect 200 with GPU metrics JSON diff --git a/litellm-self-heal.prose.md b/litellm-self-heal.prose.md index 734c26a..70ee7d3 100644 --- a/litellm-self-heal.prose.md +++ b/litellm-self-heal.prose.md @@ -20,7 +20,7 @@ note: > Last verified: 2026-07-12 description: > LiteLLM inference stack health monitoring + self-healing. Verifies the full - nginx → LiteLLM → GPU chain, 8 containers on CT 116, 3 GPU hosts, model + nginx → LiteLLM → GPU chain, 11 containers on CT 116, 3 GPU hosts, model inference, and agent keys. Applies remediation rules for common failures. Reports every action via Zulip DM and Gitea (SyslogSolution/health-logs). --- @@ -41,8 +41,8 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2) │ Grafana :3001 - harness-router :9000 — DEPRECATED, container still runs but - NOT in request path. nginx routes /v1 → LiteLLM directly. + harness-router :9000 — DECOMMISSIONED 2026-09-11 (container, image + and config removed). nginx routes /v1 → LiteLLM directly. Router slot booking + circuit breakers replaced by LiteLLM native fallbacks + timeouts. ``` @@ -98,8 +98,7 @@ Preferred implementation: uncap shared pool, add capped alias for crew-only. | Container | Image | Port | Health Check | |-----------|-------|------|-------------| -| harness-litellm | berriai/litellm:1.90.0-rc.1 | :4000→:4001 | /health/liveliness | -| harness-router | inference-harness-router | :9000 (127.0.0.1) | /health (DEPRECATED — not in path) | +| harness-litellm | ghcr.io/docker.litellm.ai/berriai/litellm:1.90.0-rc.1 | :4000→:4000 | /health/liveliness | | harness-nginx | nginx:alpine | :80 | HTTP 200 on /health | | harness-postgres | postgres:16-alpine | :5432 | pg_isready | | harness-redis | redis:7-alpine | :6379 | PING | @@ -146,10 +145,11 @@ Run this first on every cycle. Results feed into remediation rules below. - GET http://{{backend_host}}/litellm/health/liveliness → expect 200 ### 3. Check backend container health -- SSH to {{backend_host}} → `docker ps` → verify 10 containers healthy +- SSH to {{backend_host}} → `docker ps` → verify 11 containers healthy - Critical: harness-litellm, harness-nginx, harness-postgres -- Monitoring: harness-redis, harness-dashboard, harness-grafana, harness-prometheus, harness-docker-stats, harness-pve-exporter -- Deprecated but running: harness-router (not in path, reference only) +- Monitoring: harness-redis, harness-dashboard, harness-grafana, harness-prometheus, + harness-alertmanager, harness-zulip-bridge, harness-docker-stats, harness-pve-exporter +- Decommissioned 2026-09-11: harness-router (container, image and config removed) ### 4. Check GPU fleet health (via fleet dashboard) - GET {{gpu_dashboard_url}}/gpu-data → expect 200 with GPU metrics JSON diff --git a/scripts/daily-infra-report.py b/scripts/daily-infra-report.py index 3386d4d..e1468b7 100755 --- a/scripts/daily-infra-report.py +++ b/scripts/daily-infra-report.py @@ -213,10 +213,11 @@ def collect(): # ── LiteLLM Specific Checks (from litellm-health prose contract) ── report["litellm"] = {"checks": []} - # Check 1: LiteLLM aggregate health endpoint (router binds to 127.0.0.1, check via SSH) - health_unified = ssh(LITELLM_BACKEND, "curl -sf http://127.0.0.1:9000/health/unified -o /dev/null -w '%{http_code}' 2>/dev/null") + # Check 1: Fleet health via nginx /health/unified (301 -> /gpu/gpu-data served by gpu-monitor on .24:9100). + # Router (:9000) was decommissioned 2026-09-11; probing it was a guaranteed daily failure. + health_unified = ssh(LITELLM_BACKEND, "curl -s -o /dev/null -w '%{http_code}' --max-time 8 http://127.0.0.1/health/unified 2>/dev/null") report["litellm"]["health_unified"] = health_unified or "000" - report["litellm"]["checks"].append({"name": "unified-health", "status": "pass" if health_unified == "200" else "fail", "code": health_unified or "000"}) + report["litellm"]["checks"].append({"name": "fleet-health-via-nginx", "status": "pass" if health_unified in ("200", "301") else "fail", "code": health_unified or "000"}) # Check 2: Nginx-proxied internal endpoints for path, name in [("/litellm/ui/", "nginx-ui"), ("/litellm/docs", "nginx-docs")]: @@ -224,8 +225,11 @@ def collect(): report["litellm"]["checks"].append({"name": name, "status": "pass" if code == "200" else "fail", "code": code}) # Check 3: Docker container health for LiteLLM stack - expected_containers = ["harness-litellm", "harness-nginx", "harness-router", - "harness-postgres", "harness-redis", "harness-dashboard"] + expected_containers = ["harness-litellm", "harness-nginx", "harness-postgres", + "harness-redis", "harness-dashboard", "harness-grafana", + "harness-prometheus", "harness-alertmanager", + "harness-zulip-bridge", "harness-docker-stats", + "harness-pve-exporter"] actual_names = [c["name"] for c in containers2] report["litellm"]["expected_containers"] = expected_containers report["litellm"]["missing_containers"] = [e for e in expected_containers if e not in actual_names] diff --git a/scripts/prose-ai-review.sh b/scripts/prose-ai-review.sh index eaf8a4c..cefc089 100755 --- a/scripts/prose-ai-review.sh +++ b/scripts/prose-ai-review.sh @@ -66,11 +66,13 @@ The infrastructure-control.prose.md contract is the canonical reference for the 2. NO .19 IP — Zulip is CT 117 on storepve. 3. NO CT 122/123 — Tanko=CT 112, Mumuni=CT 100 (inside Abiba). No CT 114 anywhere. 4. Strix Halo :8080 is FIREWALLED to .116 only — cannot be probed from abiba (.24). -5. abiba-zulip PM2 process is DECOMMISSIONED (2026-07-04) — abiba uses Telegram only. +5. abiba-zulip PM2 process is ONLINE (verified 2026-09-11) — this rule was stale. -**Docker on CT 116 (8 containers):** -harness-litellm, harness-router, harness-nginx, harness-postgres, -harness-redis, harness-dashboard, harness-grafana, harness-prometheus +**Docker on CT 116 (11 containers, verified 2026-09-11):** +harness-litellm, harness-nginx, harness-postgres, harness-redis, +harness-dashboard, harness-grafana, harness-prometheus, harness-alertmanager, +harness-zulip-bridge, harness-docker-stats, harness-pve-exporter +(harness-router was decommissioned 2026-09-11) ## DIFF TO REVIEW