fix(router): decommission legacy GPU router across contracts and fleet scripts
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s

Router container, image, and config are purged on CT 116 (verified: no
container, no image, inference-harness-router:latest removed, :9000 free,
11 containers healthy, 7 models, live syslog-auto completion OK). Updates:

- gpu-fleet.prose.md: topology diagram rebuilt without the router tier
- gpu-monitor / gpu-self-heal / litellm-health / litellm-self-heal: router steps dropped
- infrastructure-control.prose.md: container inventory + litellm row corrected
- scripts/daily-infra-report.py, scripts/prose-ai-review.sh: host-scoped checks

Verified: 40/40 anchors matched uniquely, 0 U+FFFD across changed files,
daily-infra-report.py compiles, no stray harness-router references remain.
This commit is contained in:
agent-zero
2026-09-11 14:20:18 -04:00
parent 64ccf65eaf
commit c26255f5ff
10 changed files with 91 additions and 94 deletions
+9 -5
View File
@@ -213,10 +213,11 @@ def collect():
# ── LiteLLM Specific Checks (from litellm-health prose contract) ──
report["litellm"] = {"checks": []}
# Check 1: LiteLLM aggregate health endpoint (router binds to 127.0.0.1, check via SSH)
health_unified = ssh(LITELLM_BACKEND, "curl -sf http://127.0.0.1:9000/health/unified -o /dev/null -w '%{http_code}' 2>/dev/null")
# Check 1: Fleet health via nginx /health/unified (301 -> /gpu/gpu-data served by gpu-monitor on .24:9100).
# Router (:9000) was decommissioned 2026-09-11; probing it was a guaranteed daily failure.
health_unified = ssh(LITELLM_BACKEND, "curl -s -o /dev/null -w '%{http_code}' --max-time 8 http://127.0.0.1/health/unified 2>/dev/null")
report["litellm"]["health_unified"] = health_unified or "000"
report["litellm"]["checks"].append({"name": "unified-health", "status": "pass" if health_unified == "200" else "fail", "code": health_unified or "000"})
report["litellm"]["checks"].append({"name": "fleet-health-via-nginx", "status": "pass" if health_unified in ("200", "301") else "fail", "code": health_unified or "000"})
# Check 2: Nginx-proxied internal endpoints
for path, name in [("/litellm/ui/", "nginx-ui"), ("/litellm/docs", "nginx-docs")]:
@@ -224,8 +225,11 @@ def collect():
report["litellm"]["checks"].append({"name": name, "status": "pass" if code == "200" else "fail", "code": code})
# Check 3: Docker container health for LiteLLM stack
expected_containers = ["harness-litellm", "harness-nginx", "harness-router",
"harness-postgres", "harness-redis", "harness-dashboard"]
expected_containers = ["harness-litellm", "harness-nginx", "harness-postgres",
"harness-redis", "harness-dashboard", "harness-grafana",
"harness-prometheus", "harness-alertmanager",
"harness-zulip-bridge", "harness-docker-stats",
"harness-pve-exporter"]
actual_names = [c["name"] for c in containers2]
report["litellm"]["expected_containers"] = expected_containers
report["litellm"]["missing_containers"] = [e for e in expected_containers if e not in actual_names]
+6 -4
View File
@@ -66,11 +66,13 @@ The infrastructure-control.prose.md contract is the canonical reference for the
2. NO .19 IP — Zulip is CT 117 on storepve.
3. NO CT 122/123 — Tanko=CT 112, Mumuni=CT 100 (inside Abiba). No CT 114 anywhere.
4. Strix Halo :8080 is FIREWALLED to .116 only — cannot be probed from abiba (.24).
5. abiba-zulip PM2 process is DECOMMISSIONED (2026-07-04) — abiba uses Telegram only.
5. abiba-zulip PM2 process is ONLINE (verified 2026-09-11) — this rule was stale.
**Docker on CT 116 (8 containers):**
harness-litellm, harness-router, harness-nginx, harness-postgres,
harness-redis, harness-dashboard, harness-grafana, harness-prometheus
**Docker on CT 116 (11 containers, verified 2026-09-11):**
harness-litellm, harness-nginx, harness-postgres, harness-redis,
harness-dashboard, harness-grafana, harness-prometheus, harness-alertmanager,
harness-zulip-bridge, harness-docker-stats, harness-pve-exporter
(harness-router was decommissioned 2026-09-11)
## DIFF TO REVIEW