diff --git a/contract-registry.yaml b/contract-registry.yaml index 49731af..7135893 100644 --- a/contract-registry.yaml +++ b/contract-registry.yaml @@ -1,5 +1,5 @@ registry_version: 0.1.0 -last_updated: '2026-07-13T00:00:00Z' +last_updated: '2026-09-11T00:00:00Z' updated_by: mumuni categories: - compliance @@ -750,7 +750,7 @@ contracts: sensitivity: critical status: active owner: abiba - version: 1.0.0 + version: 1.1.0 trigger: type: event_driven description: Triggered by relay message from litellm-health or infrastructure-monitoring @@ -1365,7 +1365,7 @@ contracts: sensitivity: high status: active owner: ops - version: 1.0.0 + version: 1.1.0 trigger: type: scheduled cadence: 0 2 * * 0 diff --git a/infrastructure-control.prose.md b/infrastructure-control.prose.md index 52c8281..b9b682b 100644 --- a/infrastructure-control.prose.md +++ b/infrastructure-control.prose.md @@ -199,7 +199,7 @@ description: > ### Ecosystem B: CT 116 syslog-api (192.168.68.116) -11 containers in inference-harness stack (verified live 2026-09-11; LiteLLM upgraded 1.90.0-rc.1 -> 1.99.1): +12 containers on CT 116 — 11 in the inference-harness stack + trove-agent-docker (verified live 2026-09-11; LiteLLM upgraded 1.90.0-rc.1 -> 1.99.1; trove-agent-docker added 2026-09-11): | Container | Image | Port | Role | |-----------|-------|------|------| @@ -210,6 +210,7 @@ description: > | harness-dashboard | inference-harness-dashboard | :3000 | SyslogAI Harness UI | | harness-grafana | grafana/grafana | :3000→:3001 (direct LAN, not behind nginx) | GPU + Proxmox + Docker dashboards | | harness-prometheus | prom/prometheus | :9090 | Metrics scraper, 6 jobs | +| trove-agent-docker | ghcr.io/techdox/trove-agent-docker:latest | outbound agent (no port) | Trove host agent — service inventory + metrics, added 2026-09-11 | **Nginx routing**: - `/v1/*` → harness-litellm:4000 (API) @@ -223,7 +224,6 @@ description: > - 192.168.68.8:9400 (RTX 3090 — qwen) - 192.168.68.110:9400 (RTX 5070 — gemma) - 192.168.68.15:9400 (Strix Halo — qwen3.6-35B-udq4) -- 192.168.68.24:9401 (Router metrics exporter) - harness-litellm:4000 (LiteLLM health) ### Ecosystem C: Netbird (72.61.0.17 — Hostinger srv1079750.hstgr.cloud) diff --git a/infrastructure-update.prose.md b/infrastructure-update.prose.md index f1411c6..228e78c 100644 --- a/infrastructure-update.prose.md +++ b/infrastructure-update.prose.md @@ -12,7 +12,7 @@ triggers: - on "infra update" command - weekly (Sunday 03:00 America/New_York) via Agent Zero scheduler task "weekly-fleet-docker-update" (qSOOVzsU) — implemented 2026-09-08 - on security advisory relay from Mumuni -version: 1.3.0 +version: 1.4.0 --- ## Maintains @@ -81,6 +81,7 @@ Before ANY update wave: | VM 109 (.7) | Home stack (Pulse, Stirling PDF) — JDownloader moved to CT 118 LXC 2026-08-01 | `cd /opt/home_stack && docker compose pull && docker compose up -d` | | VM 109 (.7) | Audiobookshelf | `cd /opt/audiobookshelf && docker compose pull && docker compose up -d` | | CT 116 (.116) | Inference Harness (LiteLLM, Prometheus, Grafana) | `cd /opt/inference-harness && docker compose pull && docker compose up -d` | +| CT 116 (.116, via minipve) | Trove docker agent (trove-agent-docker) | `pct exec 116 -- bash -c 'cd /opt/trove-agent && docker compose pull && docker compose up -d'` | | CT 117 (storepve) | Zulip | `pct exec 117 -- bash -c 'cd /opt/zulip && docker compose pull && docker compose up -d'` (from storepve; compose recreates on zulip_default network) | | CT 117 (storepve) | Jitsi | `pct exec 117 -- bash -c 'cd /opt/jitsi && docker compose pull && docker compose up -d'` (from storepve) | | hwpve (.11) | Authentik (server, worker, postgres) | `ssh root@192.168.68.11 'cd /root && docker compose pull && docker compose up -d'` | @@ -101,6 +102,7 @@ Before ANY update wave: - harness-litellm cold start: allow 3-5 min after recreate — reports unhealthy and :4000 refuses connections while loading config/DB, then recovers to 200 on its own (verified 2026-09-08) - SearXNG test: `curl :8888` - Digest-pin sweep: `grep -rn '@sha256:' /opt/*/docker-compose.y*` on every host — digest-pinned images are INVISIBLE to `docker compose pull` (the pin re-pulls the same digest forever, so new releases never appear). Flag every pin in the run report and propose un-pinning to a floating tag with user approval before editing. Found 2026-09-10: audiobookshelf was digest-pinned at 2.34.0 (container created 2026-07-18) and silently missed by every sweep; dockhand stack was also pinned (stack removed 2026-09-10, unused). After un-pinning audiobookshelf to :latest it updated to 2.36.0 and verified HTTP 200. +- Version-pin awareness: a fixed version tag (e.g. `image: ...litellm:1.99.1`) is a no-op for `docker compose pull` just like a digest pin, so the stack silently stops advancing. CT 116 `harness-litellm` is INTENTIONALLY pinned to `1.99.1` (registry `main-stable`/`latest` currently resolve to `1.100.1`, sha256:a3715fa7 — a bleeding-edge jump explicitly declined 2026-09-11). Every run must look up the newest STABLE release tag for any version-pinned image, bump the pin deliberately with user approval, recreate, and re-verify. Never silently revert a pin to a floating tag. ## Wave 4: Proxmox Kernel Reboot diff --git a/litellm-health.prose.md b/litellm-health.prose.md index 15555b3..ab502c8 100644 --- a/litellm-health.prose.md +++ b/litellm-health.prose.md @@ -113,17 +113,19 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2) 1. **Read parameters** — Use provided values or defaults 2. **Check public endpoints**: - - GET {{public_url}}/ui/ → expect 200 ("LiteLLM Dashboard") - - GET {{public_url}}/docs → expect 200 ("LiteLLM API - Swagger UI") + - GET {{public_url}}/litellm/ui/ → expect 200 ("LiteLLM Dashboard") + - GET {{public_url}}/litellm/docs → expect 200 ("LiteLLM API - Swagger UI") + - GET {{public_url}}/ui/ and {{public_url}}/docs → expect 301 → /litellm/... → 200 (one-hop redirect helpers, added 2026-09-11) 3. **Check LiteLLM health (no-auth)**: - GET http://{{backend_host}}/litellm/health/liveliness → expect 200 4. **Check backend container health**: - - SSH to {{backend_host}} → `docker ps` → verify 11 containers healthy + - SSH to {{backend_host}} → `docker ps` → verify 12 containers healthy (11 harness + trove-agent-docker, added 2026-09-11) - Critical: harness-litellm, harness-nginx, harness-postgres - Monitoring: harness-redis, harness-dashboard, harness-grafana, harness-prometheus, - harness-alertmanager, harness-zulip-bridge, harness-docker-stats, harness-pve-exporter + harness-alertmanager, harness-zulip-bridge, harness-docker-stats, harness-pve-exporter, + trove-agent-docker (ghcr.io/techdox/trove-agent-docker, added 2026-09-11) 5. **Check GPU fleet health via gpu-monitor** (router decommissioned 2026-09-11): - GET {{gpu_dashboard_url}}/gpu-data → expect 200 with GPU metrics JSON diff --git a/litellm-self-heal.prose.md b/litellm-self-heal.prose.md index 7524458..f961460 100644 --- a/litellm-self-heal.prose.md +++ b/litellm-self-heal.prose.md @@ -20,7 +20,7 @@ note: > Last verified: 2026-07-12 description: > LiteLLM inference stack health monitoring + self-healing. Verifies the full - nginx → LiteLLM → GPU chain, 11 containers on CT 116, 3 GPU hosts, model + nginx → LiteLLM → GPU chain, 12 containers on CT 116, 3 GPU hosts, model inference, and agent keys. Applies remediation rules for common failures. Reports every action via Zulip DM and Gitea (SyslogSolution/health-logs). --- @@ -138,17 +138,19 @@ Preferred implementation: uncap shared pool, add capped alias for crew-only. Run this first on every cycle. Results feed into remediation rules below. ### 1. Check public endpoints -- GET {{public_url}}/ui/ → expect 200 ("LiteLLM Dashboard") -- GET {{public_url}}/docs → expect 200 ("LiteLLM API - Swagger UI") +- GET {{public_url}}/litellm/ui/ → expect 200 ("LiteLLM Dashboard") # served directly by nginx +- GET {{public_url}}/litellm/docs → expect 200 ("LiteLLM API - Swagger UI") # served directly by nginx +- GET {{public_url}}/ui/ and {{public_url}}/docs → expect 301 → /litellm/... → 200 (one-hop redirect helpers, added 2026-09-11) ### 2. Check LiteLLM health (no-auth) - GET http://{{backend_host}}/litellm/health/liveliness → expect 200 ### 3. Check backend container health -- SSH to {{backend_host}} → `docker ps` → verify 11 containers healthy +- SSH to {{backend_host}} → `docker ps` → verify 12 containers healthy (11 harness + trove-agent-docker, added 2026-09-11) - Critical: harness-litellm, harness-nginx, harness-postgres - Monitoring: harness-redis, harness-dashboard, harness-grafana, harness-prometheus, - harness-alertmanager, harness-zulip-bridge, harness-docker-stats, harness-pve-exporter + harness-alertmanager, harness-zulip-bridge, harness-docker-stats, harness-pve-exporter, + trove-agent-docker (ghcr.io/techdox/trove-agent-docker, added 2026-09-11) - Decommissioned 2026-09-11: harness-router (container, image and config removed) ### 4. Check GPU fleet health (via fleet dashboard)