fix(router): decommission legacy GPU router across contracts and fleet scripts
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s
Router container, image, and config are purged on CT 116 (verified: no container, no image, inference-harness-router:latest removed, :9000 free, 11 containers healthy, 7 models, live syslog-auto completion OK). Updates: - gpu-fleet.prose.md: topology diagram rebuilt without the router tier - gpu-monitor / gpu-self-heal / litellm-health / litellm-self-heal: router steps dropped - infrastructure-control.prose.md: container inventory + litellm row corrected - scripts/daily-infra-report.py, scripts/prose-ai-review.sh: host-scoped checks Verified: 40/40 anchors matched uniquely, 0 U+FFFD across changed files, daily-infra-report.py compiles, no stray harness-router references remain.
This commit is contained in:
@@ -213,10 +213,11 @@ def collect():
|
||||
# ── LiteLLM Specific Checks (from litellm-health prose contract) ──
|
||||
report["litellm"] = {"checks": []}
|
||||
|
||||
# Check 1: LiteLLM aggregate health endpoint (router binds to 127.0.0.1, check via SSH)
|
||||
health_unified = ssh(LITELLM_BACKEND, "curl -sf http://127.0.0.1:9000/health/unified -o /dev/null -w '%{http_code}' 2>/dev/null")
|
||||
# Check 1: Fleet health via nginx /health/unified (301 -> /gpu/gpu-data served by gpu-monitor on .24:9100).
|
||||
# Router (:9000) was decommissioned 2026-09-11; probing it was a guaranteed daily failure.
|
||||
health_unified = ssh(LITELLM_BACKEND, "curl -s -o /dev/null -w '%{http_code}' --max-time 8 http://127.0.0.1/health/unified 2>/dev/null")
|
||||
report["litellm"]["health_unified"] = health_unified or "000"
|
||||
report["litellm"]["checks"].append({"name": "unified-health", "status": "pass" if health_unified == "200" else "fail", "code": health_unified or "000"})
|
||||
report["litellm"]["checks"].append({"name": "fleet-health-via-nginx", "status": "pass" if health_unified in ("200", "301") else "fail", "code": health_unified or "000"})
|
||||
|
||||
# Check 2: Nginx-proxied internal endpoints
|
||||
for path, name in [("/litellm/ui/", "nginx-ui"), ("/litellm/docs", "nginx-docs")]:
|
||||
@@ -224,8 +225,11 @@ def collect():
|
||||
report["litellm"]["checks"].append({"name": name, "status": "pass" if code == "200" else "fail", "code": code})
|
||||
|
||||
# Check 3: Docker container health for LiteLLM stack
|
||||
expected_containers = ["harness-litellm", "harness-nginx", "harness-router",
|
||||
"harness-postgres", "harness-redis", "harness-dashboard"]
|
||||
expected_containers = ["harness-litellm", "harness-nginx", "harness-postgres",
|
||||
"harness-redis", "harness-dashboard", "harness-grafana",
|
||||
"harness-prometheus", "harness-alertmanager",
|
||||
"harness-zulip-bridge", "harness-docker-stats",
|
||||
"harness-pve-exporter"]
|
||||
actual_names = [c["name"] for c in containers2]
|
||||
report["litellm"]["expected_containers"] = expected_containers
|
||||
report["litellm"]["missing_containers"] = [e for e in expected_containers if e not in actual_names]
|
||||
|
||||
@@ -66,11 +66,13 @@ The infrastructure-control.prose.md contract is the canonical reference for the
|
||||
2. NO .19 IP — Zulip is CT 117 on storepve.
|
||||
3. NO CT 122/123 — Tanko=CT 112, Mumuni=CT 100 (inside Abiba). No CT 114 anywhere.
|
||||
4. Strix Halo :8080 is FIREWALLED to .116 only — cannot be probed from abiba (.24).
|
||||
5. abiba-zulip PM2 process is DECOMMISSIONED (2026-07-04) — abiba uses Telegram only.
|
||||
5. abiba-zulip PM2 process is ONLINE (verified 2026-09-11) — this rule was stale.
|
||||
|
||||
**Docker on CT 116 (8 containers):**
|
||||
harness-litellm, harness-router, harness-nginx, harness-postgres,
|
||||
harness-redis, harness-dashboard, harness-grafana, harness-prometheus
|
||||
**Docker on CT 116 (11 containers, verified 2026-09-11):**
|
||||
harness-litellm, harness-nginx, harness-postgres, harness-redis,
|
||||
harness-dashboard, harness-grafana, harness-prometheus, harness-alertmanager,
|
||||
harness-zulip-bridge, harness-docker-stats, harness-pve-exporter
|
||||
(harness-router was decommissioned 2026-09-11)
|
||||
|
||||
## DIFF TO REVIEW
|
||||
|
||||
|
||||
Reference in New Issue
Block a user