fix(router): decommission legacy GPU router across contracts and fleet scripts
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s

Router container, image, and config are purged on CT 116 (verified: no
container, no image, inference-harness-router:latest removed, :9000 free,
11 containers healthy, 7 models, live syslog-auto completion OK). Updates:

- gpu-fleet.prose.md: topology diagram rebuilt without the router tier
- gpu-monitor / gpu-self-heal / litellm-health / litellm-self-heal: router steps dropped
- infrastructure-control.prose.md: container inventory + litellm row corrected
- scripts/daily-infra-report.py, scripts/prose-ai-review.sh: host-scoped checks

Verified: 40/40 anchors matched uniquely, 0 U+FFFD across changed files,
daily-infra-report.py compiles, no stray harness-router references remain.
This commit is contained in:
agent-zero
2026-09-11 14:20:18 -04:00
parent 64ccf65eaf
commit c26255f5ff
10 changed files with 91 additions and 94 deletions
+29 -33
View File
@@ -19,7 +19,7 @@ triggers:
- on model add/remove
- on GPU health degradation
- on agent key rotation
- on router restart (roster must be loaded)
- on harness container restart (LiteLLM reloads its model list)
---
## Maintains
@@ -36,46 +36,42 @@ triggers:
- prometheus: { status: "running", targets: 5 } — Scrapes GPU :9400 exporters + LiteLLM
- port_conflict_detection: { status: "active" } — All 3 GPU wrappers detect ghost processes before binding
## Fleet Topology (Current — July 2026)
## Fleet Topology (Current — 2026-09-11, router decommissioned)
```
┌──────────────────────────────────────────────────────────────────┐
│ CT 116 (192.168.68.116) — Inference Harness Host │
│ │
│ nginx:80 (entrypoint) │
│ ├─ /v1/* → harness-litellm:4000 (API requests) │
│ ├─ /v1/* → harness-litellm:4000 (API requests) │
│ ├─ /admin/* → harness-litellm:4000 (admin endpoints) │
│ ├─ /dashboard/ → harness-dashboard:3000 (harness UI) │
│ ├─ /litellm/* → harness-litellm:4000 (LiteLLM UI + API) │
│ ├─ /litellm/* → harness-litellm:4000 (LiteLLM UI + API) │
│ ├─ /health/* → harness-litellm:4000 (health probes) │
│ └─ /gpu/* → 192.168.68.24:9100 (fleet monitor) │
│ │
│ Containers: │
│ ┌──────────┐ ┌──────────┐ ┌──────────┐ ┌──────────┐ │
│ │ LiteLLM │ │ Router │ │Dashboard │ │ Grafana │ │
│ │ :4000 │ │ :9000 │ │ :3000 │ │ :3000 │ │
│ │ keys+sync│ │deprecated│ │ harness │ │ Prometheus│ │
│ │ fallback │ │not in │ │ UI │ │ data src │ │
│ └──────────┘ └───┬──────┘ └──────────┘ └──────────┘ │
│ │ │
│ ┌──────────┐ ┌──────────┐ ┌──────────┐ │
│ │PostgreSQL│ │ Redis │ │Prometheus│ │
│ │ :5432 │ │ :6379 │ │ :9090 │ │
│ └──────────┘ └──────────┘ └──────────┘ │
└────────────────────┼─────────────────────────────────────────────┘
│
┌───────────────┼───────────────┬──────────────────┐
│ │ │ │
┌────▼─────┐ ┌──────▼──────┐ ┌────▼──────┐ ┌───────▼──────┐
│ CT 8 │ │ CT 110 │ │ CT 15 │ │ pi (.24) │
│ RTX 3090 │ │ RTX 5070 │ │ Strix Halo│ │ GPU Monitor │
│ 24GB │ │ 12GB │ │ 64GB UMA │ │ :9100 │
│ 128K ctx │ │ 128K ctx │ │ 128K ctx │ │ Watchdog │
│ qwen3.6 │ │ Qwen3.5-9B │ │ qwen3.5 │ │ Prometheus │
│ 27B-code │ │ :8080 │ │ -35B-udq4 │ │ exporter │
│ :8080 │ │ :9400 (exp) │ │ :9400(exp)│ │ :9401 │
│ :9400 │ └─────────────┘ └───────────┘ └──────────────┘
└──────────┘
│ ┌──────────┐ ┌──────────┐ ┌──────────┐ │
│ │ LiteLLM │ │Dashboard │ │ Grafana │ │
│ │ :4000 │ │ :3000 │ │ :3001 │ │
│ │ keys+sync│ │ harness │ │Prometheus│ │
│ │ fallback │ │ UI │ │ data src │ │
│ └──────────┘ └──────────┘ └──────────┘ │
│ │
│ ┌──────────┐ ┌──────────┐ ┌──────────┐ │
│ │PostgreSQL│ │ Redis │ │Prometheus│ │
│ │ :5432 │ │ :6379 │ │ :9090 │ │
│ └──────────┘ └──────────┘ └──────────┘ │
└───────┼──────────────────────────────────────────────────────────┘
│
┌─┴─────────────┬───────────────┬───────────────┐
│ │ │ │
┌─────▼──────┐ ┌─────▼──────┐ ┌─────▼──────┐ ┌─────▼──────┐
│ CT 8 │ │ CT 110 │ │ CT 15 │ │ pi (.24) │
│ RTX 3090 │ │ RTX 5070 │ │ Strix Halo │ │ GPU Monitor│
│ 24GB │ │ 12GB │ │ 64GB UMA │ │ :9100 │
│ :8080 │ │ :9400 exp │ │ :9400 exp │ │ :9401 │
└────────────┘ └────────────┘ └────────────┘ └────────────┘
```
## Stable Role-Based Aliases (Introduced 2026-07-15)
@@ -105,7 +101,7 @@ but are deprecated for agent configs. Only the stable aliases survive model swap
|-------|-----|--------|---------|---------|
| Qwen3.8-27B-Uncensored-Q4_K_M | RTX 3090 (.8:8080) | **0.55** | 500 | **300s** |
Note: All syslog-auto entries route directly to GPUs with `api_key: not-needed`. The router (port 9000) is NOT in the inference path.
Note: All syslog-auto entries route directly to GPUs with `api_key: not-needed`. The router (port 9000) was decommissioned 2026-09-11 and is NOT in the inference path.
### Direct Model Endpoints
@@ -153,7 +149,7 @@ Note: All syslog-auto entries route directly to GPUs with `api_key: not-needed`.
6. Cleanup model files (optional)
### heal
1. Check all GPUs via router internal `:9000/health/unified`
1. Check all GPUs via gpu-monitor `{{gpu_dashboard_url}}/gpu-data`
2. Check LiteLLM health via nginx `:80/litellm/health/liveliness`
3. Reset stuck circuit breakers if idle (Redis)
4. Restart dead llama-server instances via SSH
@@ -161,7 +157,7 @@ Note: All syslog-auto entries route directly to GPUs with `api_key: not-needed`.
6. Verify GPU monitor server is running on pi (:9100)
7. Verify watchdog is running on pi
8. Restart router if roster not loaded (check logs for STARTUP ROSTER)
9. Reload roster via `POST :9000/admin/roster/reload` if available
9. Router roster reload — REMOVED 2026-09-11 (router decommissioned; LiteLLM fallbacks handle routing)
### sync-keys
1. List all agent keys in LiteLLM DB via `GET /key/list`
@@ -258,7 +254,7 @@ If no SSH access, send Zulip DM via abiba-bot with vault update instructions.
Benchmarks from 2026-07-17. Strix Halo model: qwen3.6-35B-udq4. RTX 5070 MTP provides 2.7x speedup over pre-upgrade 70 tok/s.
All 3 GPUs now at 128K context (2026-07-17, reduced from 256K for stability).
Benchmarks run through LiteLLM proxy (192.168.68.116:4001) every 5 minutes.
Benchmarks run through LiteLLM proxy (192.168.68.116:4000) every 5 minutes.
Degradation alerts fire at 30% (warning) and 50% (critical) below baseline.
History stored at `/root/data/toks-history.json` with 7-day rolling window.