feat(litellm): 1.99.1 update coverage, trove agent, nginx /ui /docs fixes #75

Merged
kagentz-bot merged 1 commits from update/litellm-1991-trove-20260911 into master 2026-09-11 19:34:23 +00:00
5 changed files with 21 additions and 15 deletions
+3 -3
View File
@@ -1,5 +1,5 @@
registry_version: 0.1.0 registry_version: 0.1.0
last_updated: '2026-07-13T00:00:00Z' last_updated: '2026-09-11T00:00:00Z'
updated_by: mumuni updated_by: mumuni
categories: categories:
- compliance - compliance
@@ -750,7 +750,7 @@ contracts:
sensitivity: critical sensitivity: critical
status: active status: active
owner: abiba owner: abiba
version: 1.0.0 version: 1.1.0
trigger: trigger:
type: event_driven type: event_driven
description: Triggered by relay message from litellm-health or infrastructure-monitoring description: Triggered by relay message from litellm-health or infrastructure-monitoring
@@ -1365,7 +1365,7 @@ contracts:
sensitivity: high sensitivity: high
status: active status: active
owner: ops owner: ops
version: 1.0.0 version: 1.1.0
trigger: trigger:
type: scheduled type: scheduled
cadence: 0 2 * * 0 cadence: 0 2 * * 0
+2 -2
View File
@@ -199,7 +199,7 @@ description: >
### Ecosystem B: CT 116 syslog-api (192.168.68.116) ### Ecosystem B: CT 116 syslog-api (192.168.68.116)
11 containers in inference-harness stack (verified live 2026-09-11; LiteLLM upgraded 1.90.0-rc.1 -> 1.99.1): 12 containers on CT 116 — 11 in the inference-harness stack + trove-agent-docker (verified live 2026-09-11; LiteLLM upgraded 1.90.0-rc.1 -> 1.99.1; trove-agent-docker added 2026-09-11):
| Container | Image | Port | Role | | Container | Image | Port | Role |
|-----------|-------|------|------| |-----------|-------|------|------|
@@ -210,6 +210,7 @@ description: >
| harness-dashboard | inference-harness-dashboard | :3000 | SyslogAI Harness UI | | harness-dashboard | inference-harness-dashboard | :3000 | SyslogAI Harness UI |
| harness-grafana | grafana/grafana | :3000→:3001 (direct LAN, not behind nginx) | GPU + Proxmox + Docker dashboards | | harness-grafana | grafana/grafana | :3000→:3001 (direct LAN, not behind nginx) | GPU + Proxmox + Docker dashboards |
| harness-prometheus | prom/prometheus | :9090 | Metrics scraper, 6 jobs | | harness-prometheus | prom/prometheus | :9090 | Metrics scraper, 6 jobs |
| trove-agent-docker | ghcr.io/techdox/trove-agent-docker:latest | outbound agent (no port) | Trove host agent — service inventory + metrics, added 2026-09-11 |
**Nginx routing**: **Nginx routing**:
- `/v1/*` → harness-litellm:4000 (API) - `/v1/*` → harness-litellm:4000 (API)
@@ -223,7 +224,6 @@ description: >
- 192.168.68.8:9400 (RTX 3090 — qwen) - 192.168.68.8:9400 (RTX 3090 — qwen)
- 192.168.68.110:9400 (RTX 5070 — gemma) - 192.168.68.110:9400 (RTX 5070 — gemma)
- 192.168.68.15:9400 (Strix Halo — qwen3.6-35B-udq4) - 192.168.68.15:9400 (Strix Halo — qwen3.6-35B-udq4)
- 192.168.68.24:9401 (Router metrics exporter)
- harness-litellm:4000 (LiteLLM health) - harness-litellm:4000 (LiteLLM health)
### Ecosystem C: Netbird (72.61.0.17 — Hostinger srv1079750.hstgr.cloud) ### Ecosystem C: Netbird (72.61.0.17 — Hostinger srv1079750.hstgr.cloud)
+3 -1
View File
@@ -12,7 +12,7 @@ triggers:
- on "infra update" command - on "infra update" command
- weekly (Sunday 03:00 America/New_York) via Agent Zero scheduler task "weekly-fleet-docker-update" (qSOOVzsU) — implemented 2026-09-08 - weekly (Sunday 03:00 America/New_York) via Agent Zero scheduler task "weekly-fleet-docker-update" (qSOOVzsU) — implemented 2026-09-08
- on security advisory relay from Mumuni - on security advisory relay from Mumuni
version: 1.3.0 version: 1.4.0
--- ---
## Maintains ## Maintains
@@ -81,6 +81,7 @@ Before ANY update wave:
| VM 109 (.7) | Home stack (Pulse, Stirling PDF) — JDownloader moved to CT 118 LXC 2026-08-01 | `cd /opt/home_stack && docker compose pull && docker compose up -d` | | VM 109 (.7) | Home stack (Pulse, Stirling PDF) — JDownloader moved to CT 118 LXC 2026-08-01 | `cd /opt/home_stack && docker compose pull && docker compose up -d` |
| VM 109 (.7) | Audiobookshelf | `cd /opt/audiobookshelf && docker compose pull && docker compose up -d` | | VM 109 (.7) | Audiobookshelf | `cd /opt/audiobookshelf && docker compose pull && docker compose up -d` |
| CT 116 (.116) | Inference Harness (LiteLLM, Prometheus, Grafana) | `cd /opt/inference-harness && docker compose pull && docker compose up -d` | | CT 116 (.116) | Inference Harness (LiteLLM, Prometheus, Grafana) | `cd /opt/inference-harness && docker compose pull && docker compose up -d` |
| CT 116 (.116, via minipve) | Trove docker agent (trove-agent-docker) | `pct exec 116 -- bash -c 'cd /opt/trove-agent && docker compose pull && docker compose up -d'` |
| CT 117 (storepve) | Zulip | `pct exec 117 -- bash -c 'cd /opt/zulip && docker compose pull && docker compose up -d'` (from storepve; compose recreates on zulip_default network) | | CT 117 (storepve) | Zulip | `pct exec 117 -- bash -c 'cd /opt/zulip && docker compose pull && docker compose up -d'` (from storepve; compose recreates on zulip_default network) |
| CT 117 (storepve) | Jitsi | `pct exec 117 -- bash -c 'cd /opt/jitsi && docker compose pull && docker compose up -d'` (from storepve) | | CT 117 (storepve) | Jitsi | `pct exec 117 -- bash -c 'cd /opt/jitsi && docker compose pull && docker compose up -d'` (from storepve) |
| hwpve (.11) | Authentik (server, worker, postgres) | `ssh root@192.168.68.11 'cd /root && docker compose pull && docker compose up -d'` | | hwpve (.11) | Authentik (server, worker, postgres) | `ssh root@192.168.68.11 'cd /root && docker compose pull && docker compose up -d'` |
@@ -101,6 +102,7 @@ Before ANY update wave:
- harness-litellm cold start: allow 3-5 min after recreate — reports unhealthy and :4000 refuses connections while loading config/DB, then recovers to 200 on its own (verified 2026-09-08) - harness-litellm cold start: allow 3-5 min after recreate — reports unhealthy and :4000 refuses connections while loading config/DB, then recovers to 200 on its own (verified 2026-09-08)
- SearXNG test: `curl :8888` - SearXNG test: `curl :8888`
- Digest-pin sweep: `grep -rn '@sha256:' /opt/*/docker-compose.y*` on every host — digest-pinned images are INVISIBLE to `docker compose pull` (the pin re-pulls the same digest forever, so new releases never appear). Flag every pin in the run report and propose un-pinning to a floating tag with user approval before editing. Found 2026-09-10: audiobookshelf was digest-pinned at 2.34.0 (container created 2026-07-18) and silently missed by every sweep; dockhand stack was also pinned (stack removed 2026-09-10, unused). After un-pinning audiobookshelf to :latest it updated to 2.36.0 and verified HTTP 200. - Digest-pin sweep: `grep -rn '@sha256:' /opt/*/docker-compose.y*` on every host — digest-pinned images are INVISIBLE to `docker compose pull` (the pin re-pulls the same digest forever, so new releases never appear). Flag every pin in the run report and propose un-pinning to a floating tag with user approval before editing. Found 2026-09-10: audiobookshelf was digest-pinned at 2.34.0 (container created 2026-07-18) and silently missed by every sweep; dockhand stack was also pinned (stack removed 2026-09-10, unused). After un-pinning audiobookshelf to :latest it updated to 2.36.0 and verified HTTP 200.
- Version-pin awareness: a fixed version tag (e.g. `image: ...litellm:1.99.1`) is a no-op for `docker compose pull` just like a digest pin, so the stack silently stops advancing. CT 116 `harness-litellm` is INTENTIONALLY pinned to `1.99.1` (registry `main-stable`/`latest` currently resolve to `1.100.1`, sha256:a3715fa7 — a bleeding-edge jump explicitly declined 2026-09-11). Every run must look up the newest STABLE release tag for any version-pinned image, bump the pin deliberately with user approval, recreate, and re-verify. Never silently revert a pin to a floating tag.
## Wave 4: Proxmox Kernel Reboot ## Wave 4: Proxmox Kernel Reboot
+6 -4
View File
@@ -113,17 +113,19 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2)
1. **Read parameters** — Use provided values or defaults 1. **Read parameters** — Use provided values or defaults
2. **Check public endpoints**: 2. **Check public endpoints**:
- GET {{public_url}}/ui/ → expect 200 ("LiteLLM Dashboard") - GET {{public_url}}/litellm/ui/ → expect 200 ("LiteLLM Dashboard")
- GET {{public_url}}/docs → expect 200 ("LiteLLM API - Swagger UI") - GET {{public_url}}/litellm/docs → expect 200 ("LiteLLM API - Swagger UI")
- GET {{public_url}}/ui/ and {{public_url}}/docs → expect 301 → /litellm/... → 200 (one-hop redirect helpers, added 2026-09-11)
3. **Check LiteLLM health (no-auth)**: 3. **Check LiteLLM health (no-auth)**:
- GET http://{{backend_host}}/litellm/health/liveliness → expect 200 - GET http://{{backend_host}}/litellm/health/liveliness → expect 200
4. **Check backend container health**: 4. **Check backend container health**:
- SSH to {{backend_host}} → `docker ps` → verify 11 containers healthy - SSH to {{backend_host}} → `docker ps` → verify 12 containers healthy (11 harness + trove-agent-docker, added 2026-09-11)
- Critical: harness-litellm, harness-nginx, harness-postgres - Critical: harness-litellm, harness-nginx, harness-postgres
- Monitoring: harness-redis, harness-dashboard, harness-grafana, harness-prometheus, - Monitoring: harness-redis, harness-dashboard, harness-grafana, harness-prometheus,
harness-alertmanager, harness-zulip-bridge, harness-docker-stats, harness-pve-exporter harness-alertmanager, harness-zulip-bridge, harness-docker-stats, harness-pve-exporter,
trove-agent-docker (ghcr.io/techdox/trove-agent-docker, added 2026-09-11)
5. **Check GPU fleet health via gpu-monitor** (router decommissioned 2026-09-11): 5. **Check GPU fleet health via gpu-monitor** (router decommissioned 2026-09-11):
- GET {{gpu_dashboard_url}}/gpu-data → expect 200 with GPU metrics JSON - GET {{gpu_dashboard_url}}/gpu-data → expect 200 with GPU metrics JSON
+7 -5
View File
@@ -20,7 +20,7 @@ note: >
Last verified: 2026-07-12 Last verified: 2026-07-12
description: > description: >
LiteLLM inference stack health monitoring + self-healing. Verifies the full LiteLLM inference stack health monitoring + self-healing. Verifies the full
nginx → LiteLLM → GPU chain, 11 containers on CT 116, 3 GPU hosts, model nginx → LiteLLM → GPU chain, 12 containers on CT 116, 3 GPU hosts, model
inference, and agent keys. Applies remediation rules for common failures. inference, and agent keys. Applies remediation rules for common failures.
Reports every action via Zulip DM and Gitea (SyslogSolution/health-logs). Reports every action via Zulip DM and Gitea (SyslogSolution/health-logs).
--- ---
@@ -138,17 +138,19 @@ Preferred implementation: uncap shared pool, add capped alias for crew-only.
Run this first on every cycle. Results feed into remediation rules below. Run this first on every cycle. Results feed into remediation rules below.
### 1. Check public endpoints ### 1. Check public endpoints
- GET {{public_url}}/ui/ → expect 200 ("LiteLLM Dashboard") - GET {{public_url}}/litellm/ui/ → expect 200 ("LiteLLM Dashboard") # served directly by nginx
- GET {{public_url}}/docs → expect 200 ("LiteLLM API - Swagger UI") - GET {{public_url}}/litellm/docs → expect 200 ("LiteLLM API - Swagger UI") # served directly by nginx
- GET {{public_url}}/ui/ and {{public_url}}/docs → expect 301 → /litellm/... → 200 (one-hop redirect helpers, added 2026-09-11)
### 2. Check LiteLLM health (no-auth) ### 2. Check LiteLLM health (no-auth)
- GET http://{{backend_host}}/litellm/health/liveliness → expect 200 - GET http://{{backend_host}}/litellm/health/liveliness → expect 200
### 3. Check backend container health ### 3. Check backend container health
- SSH to {{backend_host}} → `docker ps` → verify 11 containers healthy - SSH to {{backend_host}} → `docker ps` → verify 12 containers healthy (11 harness + trove-agent-docker, added 2026-09-11)
- Critical: harness-litellm, harness-nginx, harness-postgres - Critical: harness-litellm, harness-nginx, harness-postgres
- Monitoring: harness-redis, harness-dashboard, harness-grafana, harness-prometheus, - Monitoring: harness-redis, harness-dashboard, harness-grafana, harness-prometheus,
harness-alertmanager, harness-zulip-bridge, harness-docker-stats, harness-pve-exporter harness-alertmanager, harness-zulip-bridge, harness-docker-stats, harness-pve-exporter,
trove-agent-docker (ghcr.io/techdox/trove-agent-docker, added 2026-09-11)
- Decommissioned 2026-09-11: harness-router (container, image and config removed) - Decommissioned 2026-09-11: harness-router (container, image and config removed)
### 4. Check GPU fleet health (via fleet dashboard) ### 4. Check GPU fleet health (via fleet dashboard)