fix(router): decommission legacy GPU router across contracts and fleet scripts
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s
Router container, image, and config are purged on CT 116 (verified: no container, no image, inference-harness-router:latest removed, :9000 free, 11 containers healthy, 7 models, live syslog-auto completion OK). Updates: - gpu-fleet.prose.md: topology diagram rebuilt without the router tier - gpu-monitor / gpu-self-heal / litellm-health / litellm-self-heal: router steps dropped - infrastructure-control.prose.md: container inventory + litellm row corrected - scripts/daily-infra-report.py, scripts/prose-ai-review.sh: host-scoped checks Verified: 40/40 anchors matched uniquely, 0 U+FFFD across changed files, daily-infra-report.py compiles, no stray harness-router references remain.
This commit is contained in:
+29
-33
@@ -19,7 +19,7 @@ triggers:
|
||||
- on model add/remove
|
||||
- on GPU health degradation
|
||||
- on agent key rotation
|
||||
- on router restart (roster must be loaded)
|
||||
- on harness container restart (LiteLLM reloads its model list)
|
||||
---
|
||||
|
||||
## Maintains
|
||||
@@ -36,46 +36,42 @@ triggers:
|
||||
- prometheus: { status: "running", targets: 5 } — Scrapes GPU :9400 exporters + LiteLLM
|
||||
- port_conflict_detection: { status: "active" } — All 3 GPU wrappers detect ghost processes before binding
|
||||
|
||||
## Fleet Topology (Current — July 2026)
|
||||
## Fleet Topology (Current — 2026-09-11, router decommissioned)
|
||||
|
||||
```
|
||||
┌──────────────────────────────────────────────────────────────────┐
|
||||
│ CT 116 (192.168.68.116) — Inference Harness Host │
|
||||
│ │
|
||||
│ nginx:80 (entrypoint) │
|
||||
│ ├─ /v1/* → harness-litellm:4000 (API requests) │
|
||||
│ ├─ /v1/* → harness-litellm:4000 (API requests) │
|
||||
│ ├─ /admin/* → harness-litellm:4000 (admin endpoints) │
|
||||
│ ├─ /dashboard/ → harness-dashboard:3000 (harness UI) │
|
||||
│ ├─ /litellm/* → harness-litellm:4000 (LiteLLM UI + API) │
|
||||
│ ├─ /litellm/* → harness-litellm:4000 (LiteLLM UI + API) │
|
||||
│ ├─ /health/* → harness-litellm:4000 (health probes) │
|
||||
│ └─ /gpu/* → 192.168.68.24:9100 (fleet monitor) │
|
||||
│ │
|
||||
│ Containers: │
|
||||
│ ┌──────────┐ ┌──────────┐ ┌──────────┐ ┌──────────┐ │
|
||||
│ │ LiteLLM │ │ Router │ │Dashboard │ │ Grafana │ │
|
||||
│ │ :4000 │ │ :9000 │ │ :3000 │ │ :3000 │ │
|
||||
│ │ keys+sync│ │deprecated│ │ harness │ │ Prometheus│ │
|
||||
│ │ fallback │ │not in │ │ UI │ │ data src │ │
|
||||
│ └──────────┘ └───┬──────┘ └──────────┘ └──────────┘ │
|
||||
│ │ │
|
||||
│ ┌──────────┐ ┌──────────┐ ┌──────────┐ │
|
||||
│ │PostgreSQL│ │ Redis │ │Prometheus│ │
|
||||
│ │ :5432 │ │ :6379 │ │ :9090 │ │
|
||||
│ └──────────┘ └──────────┘ └──────────┘ │
|
||||
└────────────────────┼─────────────────────────────────────────────┘
|
||||
│
|
||||
┌───────────────┼───────────────┬──────────────────┐
|
||||
│ │ │ │
|
||||
┌────▼─────┐ ┌──────▼──────┐ ┌────▼──────┐ ┌───────▼──────┐
|
||||
│ CT 8 │ │ CT 110 │ │ CT 15 │ │ pi (.24) │
|
||||
│ RTX 3090 │ │ RTX 5070 │ │ Strix Halo│ │ GPU Monitor │
|
||||
│ 24GB │ │ 12GB │ │ 64GB UMA │ │ :9100 │
|
||||
│ 128K ctx │ │ 128K ctx │ │ 128K ctx │ │ Watchdog │
|
||||
│ qwen3.6 │ │ Qwen3.5-9B │ │ qwen3.5 │ │ Prometheus │
|
||||
│ 27B-code │ │ :8080 │ │ -35B-udq4 │ │ exporter │
|
||||
│ :8080 │ │ :9400 (exp) │ │ :9400(exp)│ │ :9401 │
|
||||
│ :9400 │ └─────────────┘ └───────────┘ └──────────────┘
|
||||
└──────────┘
|
||||
│ ┌──────────┐ ┌──────────┐ ┌──────────┐ │
|
||||
│ │ LiteLLM │ │Dashboard │ │ Grafana │ │
|
||||
│ │ :4000 │ │ :3000 │ │ :3001 │ │
|
||||
│ │ keys+sync│ │ harness │ │Prometheus│ │
|
||||
│ │ fallback │ │ UI │ │ data src │ │
|
||||
│ └──────────┘ └──────────┘ └──────────┘ │
|
||||
│ │
|
||||
│ ┌──────────┐ ┌──────────┐ ┌──────────┐ │
|
||||
│ │PostgreSQL│ │ Redis │ │Prometheus│ │
|
||||
│ │ :5432 │ │ :6379 │ │ :9090 │ │
|
||||
│ └──────────┘ └──────────┘ └──────────┘ │
|
||||
└───────┼──────────────────────────────────────────────────────────┘
|
||||
│
|
||||
┌─┴─────────────┬───────────────┬───────────────┐
|
||||
│ │ │ │
|
||||
┌─────▼──────┐ ┌─────▼──────┐ ┌─────▼──────┐ ┌─────▼──────┐
|
||||
│ CT 8 │ │ CT 110 │ │ CT 15 │ │ pi (.24) │
|
||||
│ RTX 3090 │ │ RTX 5070 │ │ Strix Halo │ │ GPU Monitor│
|
||||
│ 24GB │ │ 12GB │ │ 64GB UMA │ │ :9100 │
|
||||
│ :8080 │ │ :9400 exp │ │ :9400 exp │ │ :9401 │
|
||||
└────────────┘ └────────────┘ └────────────┘ └────────────┘
|
||||
```
|
||||
|
||||
## Stable Role-Based Aliases (Introduced 2026-07-15)
|
||||
@@ -105,7 +101,7 @@ but are deprecated for agent configs. Only the stable aliases survive model swap
|
||||
|-------|-----|--------|---------|---------|
|
||||
| Qwen3.8-27B-Uncensored-Q4_K_M | RTX 3090 (.8:8080) | **0.55** | 500 | **300s** |
|
||||
|
||||
Note: All syslog-auto entries route directly to GPUs with `api_key: not-needed`. The router (port 9000) is NOT in the inference path.
|
||||
Note: All syslog-auto entries route directly to GPUs with `api_key: not-needed`. The router (port 9000) was decommissioned 2026-09-11 and is NOT in the inference path.
|
||||
|
||||
### Direct Model Endpoints
|
||||
|
||||
@@ -153,7 +149,7 @@ Note: All syslog-auto entries route directly to GPUs with `api_key: not-needed`.
|
||||
6. Cleanup model files (optional)
|
||||
|
||||
### heal
|
||||
1. Check all GPUs via router internal `:9000/health/unified`
|
||||
1. Check all GPUs via gpu-monitor `{{gpu_dashboard_url}}/gpu-data`
|
||||
2. Check LiteLLM health via nginx `:80/litellm/health/liveliness`
|
||||
3. Reset stuck circuit breakers if idle (Redis)
|
||||
4. Restart dead llama-server instances via SSH
|
||||
@@ -161,7 +157,7 @@ Note: All syslog-auto entries route directly to GPUs with `api_key: not-needed`.
|
||||
6. Verify GPU monitor server is running on pi (:9100)
|
||||
7. Verify watchdog is running on pi
|
||||
8. Restart router if roster not loaded (check logs for STARTUP ROSTER)
|
||||
9. Reload roster via `POST :9000/admin/roster/reload` if available
|
||||
9. Router roster reload — REMOVED 2026-09-11 (router decommissioned; LiteLLM fallbacks handle routing)
|
||||
|
||||
### sync-keys
|
||||
1. List all agent keys in LiteLLM DB via `GET /key/list`
|
||||
@@ -258,7 +254,7 @@ If no SSH access, send Zulip DM via abiba-bot with vault update instructions.
|
||||
Benchmarks from 2026-07-17. Strix Halo model: qwen3.6-35B-udq4. RTX 5070 MTP provides 2.7x speedup over pre-upgrade 70 tok/s.
|
||||
All 3 GPUs now at 128K context (2026-07-17, reduced from 256K for stability).
|
||||
|
||||
Benchmarks run through LiteLLM proxy (192.168.68.116:4001) every 5 minutes.
|
||||
Benchmarks run through LiteLLM proxy (192.168.68.116:4000) every 5 minutes.
|
||||
Degradation alerts fire at 30% (warning) and 50% (critical) below baseline.
|
||||
History stored at `/root/data/toks-history.json` with 7-day rolling window.
|
||||
|
||||
|
||||
Reference in New Issue
Block a user