diff --git a/LITELLM-MIGRATION-PLAN.md b/LITELLM-MIGRATION-PLAN.md index 5e89e00..3c0d05e 100644 --- a/LITELLM-MIGRATION-PLAN.md +++ b/LITELLM-MIGRATION-PLAN.md @@ -37,7 +37,7 @@ 5-Tier Content-Based Routing GPU Slot Management (Redis) Agent Spread Prevention - GPU Health Scoring (40/30/30) + GPU Health Scoring Sidecar VRAM/Temp/Power Circuit Breaker Context Window Tracking @@ -76,6 +76,7 @@ | **Context Tracking** | Session-level token accumulation with compaction warnings in headers | | **SSE Streaming** | Real-time dashboard updates, per-model timeseries | | **Admin** | `/admin/keys`, `/admin/keys/generate`, `/admin/keys/revoke`, `/admin/keys/deprecation-summary` | +| **Strict Passthrough** | Explicit model requests go to that GPU exactly (no silent fallback). LiteLLM owns failover. | ### 1.2 GPU Backends @@ -89,12 +90,13 @@ CT 116 already has a LiteLLM container running (POC, 6 days uptime): -```\nharness-litellm | ghcr.io/berriai/litellm:main-stable | 127.0.0.1:80814000 +``` +harness-litellm | ghcr.io/berriai/litellm:main-stable | 127.0.0.1:8081->4000 harness-redis | redis:7-alpine | 127.0.0.1:6379 harness-router | inference-harness-router | 127.0.0.1:9000 harness-nginx | nginx:alpine | 0.0.0.0:80 harness-dashboard | inference-harness-dashboard | 127.0.0.1:3000 -```\ +``` - `/opt/litellm/` previous setup directory on CT 116 - Configured with Postgres, host networking, master key @@ -133,14 +135,14 @@ services: | Feature | Our Router | LiteLLM | Value Add | |---------|-----------|---------|-----------| | **Admin UI** | | Full dashboard at /ui | Non-technical users can manage keys, view spend | -| **Virtual Key Permissions** | (binary keytier) | Granular: per-model, per-team, budget caps | Fine-grained access control | +| **Virtual Key Permissions** | (binary key->tier) | Granular: per-model, per-team, budget caps | Fine-grained access control | | **Spend Tracking** | | Per-request $ cost with model-specific pricing | Billing, cost allocation, client invoicing | -| **Teams & Orgs** | | Multi-tenant: orgteamuser hierarchy | Segregate clients/projects | +| **Teams & Orgs** | | Multi-tenant: org->team->user hierarchy | Segregate clients/projects | | **SSO/OIDC** | | Google, GitHub, Microsoft, Okta, Keycloak | Enterprise auth integration | | **Budget Alerts** | | Per-key, per-user, per-team budget with webhooks | Prevent overspend | | **Usage Analytics** | (custom /metrics) | Built-in: daily trends, model breakdown, per-customer | Better visualization | | **100+ Provider Support** | (3 local GPUs) | OpenAI, Anthropic, Bedrock, Vertex, etc. | Future cloud model access | -| **Fallback Chains** | | Multi-provider: OpenAIAzureTogether | External model resilience | +| **Fallback Chains** | (silent rerouting) | Explicit multi-provider failover with per-model logging | Accurate per-model tracking, visible failover | | **RPM/TPM Weighted LB** | | Weighted load balancing across deployments | Fine-grained traffic shaping | --- @@ -150,7 +152,7 @@ services: | Feature | Why We Must Keep It | |---------|---------------------| | **Content-based 5-tier routing** | LiteLLM routes by model name only; we analyze prompt complexity, tokens, turns, and routing_hints | -| **GPU hardware health scoring** | LiteLLM doesn't monitor VRAM, temp, power our 40/30/30 scoring prevents routing to overheating GPUs | +| **GPU hardware health scoring** | LiteLLM doesn't monitor VRAM, temp, power our scoring prevents routing to overheating GPUs | | **GPU slot management** | LiteLLM doesn't know about llama.cpp --parallel limits; our Redis counters prevent overloading | | **Agent spread prevention** | Our `select_best_gpu()` spreads agents across GPUs to prevent hotspots; LiteLLM only does simple-shuffle | | **Cross-turn context tracking** | Session-level token accumulation with compaction warnings via X-Context-Warning headers | @@ -188,7 +190,9 @@ Agent Request X-Session-Id: 7. On response: log spend, update budgets - 8. Return response to agent + 8. If router returns 503 (GPU saturated): + consult fallback chain, retry next model + 9. Return response to agent @@ -197,9 +201,8 @@ Agent Request 1. Authenticate agent key (sk-syslog-...) 2. Hardware rate limit (per-tier RPM) - 3. Content-based tier routing - - Estimate tokens, detect system msg - - Count turns, check routing_hints + 3. Content-based tier routing (for syslog-auto) + OR strict passthrough (for explicit models) 4. GPU slot availability (Redis counter) 5. GPU health check (sidecar) 6. Agent spread logic (select_best_gpu) @@ -225,16 +228,15 @@ general_settings: store_model_in_db: true model_list: - # All three GPUs exposed as a single virtual "syslog-router" model - # LiteLLM passes through to our router, which handles actual GPU selection - - model_name: syslog-auto # Default auto-routing + # Content-based auto-routing (router picks GPU via 5-tier analysis) + - model_name: syslog-auto litellm_params: - model: openai/syslog-auto # Using OpenAI-compatible format + model: openai/syslog-auto api_base: http://router:9000/v1 api_key: os.environ/ROUTER_API_KEY - rpm: 600 # Cap total RPM across all GPUs + rpm: 600 - # Individual GPU pass-through (for explicit model requests) + # Individual GPU strict passthrough (exact GPU, no silent fallback) - model_name: qwen3.6-35B-A3B litellm_params: model: openai/qwen3.6-35B-A3B @@ -291,37 +293,69 @@ litellm_settings: router_settings: routing_strategy: "usage-based-routing" # For external models only - # Note: All local GPU routing is handled by custom router enable_loadbalancing_on_proxy: false # Disable LiteLLM's internal LB - allowed_fails: 0 # Set to 0 router's circuit breaker is authoritative + allowed_fails: 100 # Router returns 503 on saturated GPUs cooldown disabled + # Fallback chains: LiteLLM retries down the chain when router returns saturated + # This gives accurate per-model metrics because router no longer silently reroutes + fallbacks: + - qwen3.6-35B-A3B: ["qwen3.6-27B-code", "gemma-4-12b"] + - qwen3.6-27B-code: ["qwen3.6-35B-A3B", "gemma-4-12b"] + - gemma-4-12b: ["qwen3.6-27B-code", "qwen3.6-35B-A3B"] -# Cost tracking: map model names to per-token pricing -# These are passed through from our router's X-Usage-Tokens header +# Cost tracking: map model names to per-token pricing for spend tracking +litellm_settings: + model_cost: + qwen3.6-35B-A3B: + input_cost_per_token: 0.0 + output_cost_per_token: 0.0 + qwen3.6-27B-code: + input_cost_per_token: 0.0 + output_cost_per_token: 0.0 + gemma-4-12b: + input_cost_per_token: 0.0 + output_cost_per_token: 0.0 + # For internal cost allocation, set symbolic rates: + # e.g., MoE = $2/M tokens, Dense = $1/M tokens, VLM = $0.50/M tokens ``` -### 4.4 Router Modifications (Engine Room Updates) +### 4.4 Router Modifications -To accommodate LiteLLM, the `router-fixed.py` requires the following updates: +To accommodate LiteLLM, `router-fixed.py` requires the following updates: + +1. **Strict passthrough for explicit models** (DEPLOYED): + ```python + # In route(), the explicit model section changed from silent fallback to strict: + req = rd.get("model","auto") + if req != "auto": + # STRICT MODE: no silent fallback LiteLLM handles failover chains. + # This keeps per-model metrics accurate. Returns saturated if busy. + target = req if req in avail else avail[0] + if req not in avail: + return {"model": req, "reason": "explicit_unavailable", "saturated": True} + if is_gpu_busy(target): + return {"model": target, "reason": "explicit_saturated", "saturated": True} + return {"model": target, "reason": "explicit"} + ``` + +2. **New header passthrough**: Forward `X-LiteLLM-*` headers to GPU (transparent already works) + +3. **New endpoint for health passthrough**: `GET /v1/models` already works + +4. **Keep ALL routing logic**: No changes to `select_best_gpu()`, `check_gpu_health()`, slot management, etc. Content-based routing for `syslog-auto` is fully intact. -1. **New header passthrough**: Forward `X-LiteLLM-*` headers to GPU (transparent already works) -2. **New endpoint for health passthrough**: `GET /v1/models` already works -3. **Disable own key management**: Remove `/admin/keys/*` endpoints (migrate to LiteLLM UI) -4. **Keep ALL routing logic**: No changes to `route()`, `select_best_gpu()`, `check_gpu_health()`, slot management, etc. 5. **Add LiteLLM-compatible response**: Return `X-Usage-Tokens` header so LiteLLM can track token costs - -```python -# ADD to router-fixed.py chat() response: -resp.headers["X-Usage-Tokens"] = json.dumps({ - "prompt_tokens": prompt_tokens, - "completion_tokens": completion_tokens, - "model": model -}) -``` + ```python + resp.headers["X-Usage-Tokens"] = json.dumps({ + "prompt_tokens": prompt_tokens, + "completion_tokens": completion_tokens, + "model": model + }) + ``` ### 4.5 Router Logic Refinements **GPU Health Scoring (Updated):** -We are updating the scoring algorithm to include Power metrics (30% weight): +We are updating the scoring algorithm to include Power metrics: ```python def gpu_health_score(model): h = check_gpu_health(model, sidecar_timeout=1.5, gpu_timeout=1) @@ -333,7 +367,7 @@ def gpu_health_score(model): active = gpu_active_count(model) max_c = GPU_MAX_CONCURRENT.get(model, 1) load_pct = (active / max_c) * 100 if max_c > 0 else 0 - # Score: lower = better (more headroom, cooler, less power-constrained, less loaded) + # Score: lower = better score = (vram_pct * 0.3) + (max((temp_c - 30, 0) * 0.3) + (power_w * 0.2) + (load_pct * 0.2)) return score ``` @@ -349,174 +383,72 @@ def gpu_health_score(model): **Tasks:** 1. **Set up DNS split-horizon on CT 116** ```bash - # On CT 116 host (syslog-api) echo "192.168.68.11 auth.sysloggh.net" >> /etc/hosts ``` - Add to docker-compose.yml (see 7): - ```yaml - extra_hosts: - - "auth.sysloggh.net:192.168.68.11" - ``` -2. **Deploy Postgres container** alongside existing services: - ```bash - cd /opt/litellm - # Add postgres to docker-compose.yml - docker compose up -d postgres - ``` - *Note: Use a dedicated volume for `pgdata` to ensure LiteLLM database persistence.* +2. **Deploy Postgres container** alongside existing services 3. **Replace LiteLLM config** with production config.yaml (see 4.3) - All 4 models `http://router:9000/v1` - - Add guardrails (pre-call, post-call, content filter) - - Set `num_retries: 0` (router handles retry) - - Set `request_timeout: 600` + - Fallback chains for explicit models + - Guardrails (pre-call, post-call, content filter) + - `allowed_fails: 100` (router returns 503 on saturated) + - `num_retries: 0` (LiteLLM retries handled by fallback chains) 4. **Deploy custom_sso.py** for Authentik OIDC integration - - Mount to LiteLLM container volume - - Reference in config.yaml: `custom_ui_sso_sign_in_handler: custom_sso.custom_ui_sso_sign_in_handler` 5. **Restart LiteLLM container** with new config - ```bash - docker compose restart litellm - ``` -6. **Verify internal routing** LiteLLM Router pass-through works: - ```bash - curl -X POST http://127.0.0.1:4000/v1/chat/completions \ - -H "Authorization: Bearer *** \ - -H "Content-Type: application/json" \ - -d '{"model":"syslog-auto","messages":[{"role":"user","content":"test"}]}' - ``` +6. **Verify internal routing** **Verification Checklist:** - [ ] DNS resolution: `getent hosts auth.sysloggh.net` 192.168.68.11 - [ ] Postgres container healthy - [ ] LiteLLM `/health` returns 200 - [ ] LiteLLM Router pass-through returns valid chat completion -- [ ] GPU health metrics unaffected (watch `/metrics`) +- [ ] GPU health metrics unaffected +- [ ] Explicit model request returns saturated (not silently rerouted) when GPU busy ### Phase 1: Shadow Mode (Week 1) Zero Risk, Zero Downtime **Goal:** Deploy LiteLLM alongside existing router, test in shadow mode. **Agents continue using :9000 directly.** -```\nAgent LiteLLM (:4000) Router (:9000) GPU - (new, testing) (existing, unchanged) - -Agent can also directly hit :9000 as fallback (unchanged) -``` +**Tasks:** +1. Create virtual keys for test agents via LiteLLM UI +2. Verify pass-through works for all 4 models +3. Validate fallback chains: saturate MoE confirm LiteLLM retries Dense confirm VLM +4. Run 24-hour shadow: monitor LiteLLM spend logs vs router metrics +5. Verify GPU health metrics unaffected +6. Check guardrails not generating false positives + +### Phase 2: Cutover (Week 2) Gradual Agent Migration + +**Goal:** Move agents one-by-one to LiteLLM endpoint. **Tasks:** -1. **Create virtual keys for test agents** via LiteLLM UI or API: - ```bash - curl -X POST http://127.0.0.1:4000/key/generate \ - -H "Authorization: Bearer *** \ - -H "Content-Type: application/json" \ - -d '{"models":["syslog-auto"],"metadata":{"agent":"test"}}' - ``` - - Mirror existing API_KEYS in LiteLLM's key store - - Set per-key budgets (test with $100 cap) +1. Migrate API keys to LiteLLM virtual keys +2. Create teams: "Core Agents" (enterprise), "Dev Agents" (professional) +3. Update agent configs one at a time: `OPENAI_API_BASE` `:4000` +4. Test each agent individually +5. Enable SSO via Authentik + custom_sso.py +6. Keep router :9000 as emergency fallback for 48 hours -2. **Verify pass-through works** - ```bash - curl -X POST http://127.0.0.1:4000/v1/chat/completions \ - -H "Authorization: Bearer *** \ - -H "Content-Type: application/json" \ - -d '{"model":"syslog-auto","messages":[{"role":"user","content":"test"}]}' - ``` +### Phase 3: Production Hardening (Week 3+) -3. **Run 24-hour shadow**: Both :4000 and :9000 active, agents use :9000 - - Monitor LiteLLM spend logs vs router metrics confirm parity - - Verify GPU health metrics unaffected - - Check guardrails not generating false positives - -**Zero-Downtime Guarantee:** Router :9000 remains the primary agent endpoint. LiteLLM :4000 is tested in parallel with no impact on production traffic. - -### Phase 2: Cutover (Week 2) Gradual Agent Migration, Zero Cumulative Downtime - -**Goal:** Move agents one-by-one to LiteLLM endpoint. Each agent is migrated individually other agents unaffected. +**Goal:** Lock down, optimize, monitor. **Tasks:** -1. **Migrate API keys to LiteLLM virtual keys:** - - Create virtual key per agent in LiteLLM UI - - Set model access: `syslog-auto` (default), plus individual GPU models - - Set per-agent budget limits - - Create teams: - - "Core Agents" (Abiba, Mumuni, Tanko) enterprise tier - - "Dev Agents" (Kagenz0, Koby, Koonimo) professional tier - -2. **Update agent configs one at a time:** - - Change `OPENAI_API_BASE` from `http://192.168.68.116:9000/v1` `http://192.168.68.116:4000/v1` - - Replace agent API keys with LiteLLM virtual keys - - Test each agent individually verify routing works - - **Per-agent downtime: <2 minutes** - -3. **Migrate admin functions:** - - Key creation/revocation LiteLLM UI - - Rate limit management LiteLLM per-key RPM + router hardware RPM (dual enforcement) - - Deprecated key tracking LiteLLM UI key list - -4. **Enable SSO via Authentik + custom_sso.py:** - - Authentik OAuth2 provider configured - - NGINX auth_request for /ui/* paths - - Custom UI SSO sign-in handler reads x-authentik-* headers - - Users authenticate with existing Authentik credentials - -5. **Keep router :9000 accessible** as emergency fallback for 48 hours - - NGINX configured with router_fallback for 502 errors - - Agents can revert by changing `OPENAI_API_BASE` back to :9000 - -**Zero-Downtime Guarantee:** Each agent has <2 minutes downtime during config update. Router :9000 stays online throughout. Fallback path available for instant rollback. - -### Phase 3: Production Hardening (Week 3+) Optimize & Scale - -**Goal:** Lock down, optimize, monitor. Prepare for client-facing services. - -**Tasks:** - -1. **Remove deprecated router endpoints** (only after all agents migrated): - - Drop `/admin/keys/*` fully migrated to LiteLLM UI - - Drop Phase 0 dual-key logic (LiteLLM handles key rotation) - - Simplify `API_KEYS` to single `ROUTER_API_KEY` - -2. **Add LiteLLM observability:** - - Prometheus metrics export - - Slack/email budget alerts via webhook Hermes - - Daily spend report webhook - -3. **Enable LiteLLM caching** (Redis, shared with router): - ```yaml -router_settings: - redis_host: os.environ/REDIS_HOST - redis_port: 6379 - cache: true - cache_ttl: 3600 -``` - -4. **Add external model fallbacks** for client-facing services: - - Add Anthropic Claude as fallback for code-heavy requests - - Add OpenAI GPT-4o as fallback for reasoning overflow - - LiteLLM's native fallback chains handle this cleanly - -5. **Router slim-down:** Extract GPU health metrics to dedicated /health only - - Keep: routing, slots, health checks, performance recording - - Remove: key management, dual-key logic, admin endpoints - -6. **Multi-tenancy setup** for client-facing inference services: - - Organization Team User hierarchy per client - - Per-client budget limits and guardrail policies - - Self-service onboarding via Authentik SSO LiteLLM UI - -**Zero-Downtime Guarantee:** All changes are additive new features added while existing routing continues uninterrupted. Router remains the GPU intelligence layer throughout. - ---- +1. Remove deprecated router endpoints (after all agents migrated) +2. Add LiteLLM observability (Prometheus, Slack/email alerts) +3. Enable LiteLLM caching (shared Redis) +4. Add external model fallbacks for client-facing services +5. Router slim-down: keep routing/slots/health/perf, remove key management +6. Multi-tenancy setup for client-facing inference services --- ## 6. Nginx Configuration (with Authentik OIDC Forward Auth) -The existing nginx config routes `/admin/` router :9000. This MUST change: - ```nginx # OLD (remove) # location /admin/ { @@ -526,8 +458,6 @@ The existing nginx config routes `/admin/` router :9000. This MUST change: # === Authentik auth subrequest endpoint === location /authentik/auth { internal; - # Proxy to Authentik's outpost on acerpve (192.168.68.11) - # Uses internal /etc/hosts resolution: auth.sysloggh.net 192.168.68.11 proxy_pass https://auth.sysloggh.net/outpost.goauthentik.io/auth/nginx; proxy_pass_request_body off; proxy_set_header Content-Length ""; @@ -550,19 +480,17 @@ location /ui/ { proxy_set_header Connection "upgrade"; } -# === LiteLLM SSO callback (for OIDC redirect flow) === +# === LiteLLM SSO callback === location /sso/callback { proxy_pass http://127.0.0.1:4000/sso/callback; proxy_set_header Host $host; } -# === API endpoint Bearer token auth (no Authentik) === +# === API endpoint Bearer token auth === location /v1/ { - # Primary: LiteLLM gateway proxy_pass http://127.0.0.1:4000/v1/; proxy_set_header Host $host; proxy_read_timeout 600s; - # Fallback: direct router (if LiteLLM down) error_page 502 = @router_fallback; } @@ -571,7 +499,7 @@ location @router_fallback { proxy_set_header Host $host; } -# === Key management API (needs master_key, not Authentik) === +# === Key management API === location /key/ { proxy_pass http://127.0.0.1:4000/key/; proxy_set_header Host $host; @@ -583,7 +511,6 @@ location /router/ { proxy_pass http://127.0.0.1:9000/; } -# Health check combines both layers location /health { proxy_pass http://127.0.0.1:4000/health; } @@ -591,12 +518,8 @@ location /health { --- ---- - ## 7. Docker Compose (`docker-compose.yml` on CT 116) -**Deployment host:** CT 116 `syslog-api` (192.168.68.116) on minipve. All services co-located. - ```yaml services: # Layer 1: LiteLLM Gateway (Policy & Admin) @@ -613,8 +536,8 @@ services: - DATABASE_URL=postgresql://litellm:***@localhost:5432/litellm - STORE_MODEL_IN_DB=True - ROUTER_API_KEY=${ROUTER_API_KEY} - - OPENAI_API_KEY=*** # Optional: external fallback - - ANTHROPIC_API_KEY=${ANTH...KEY} # Optional: external fallback + - OPENAI_API_KEY=*** + - ANTHROPIC_API_KEY=${ANTH...KEY} - PROXY_BASE_URL=https://litellm.sysloggh.net command: - --config @@ -626,7 +549,6 @@ services: condition: service_healthy restart: unless-stopped - # Database for LiteLLM postgres: image: postgres:16-alpine network_mode: "host" @@ -643,7 +565,6 @@ services: retries: 5 restart: unless-stopped - # Nginx also needs auth resolution nginx: image: nginx:alpine extra_hosts: @@ -654,31 +575,25 @@ services: - "80:80" restart: unless-stopped - # Layer 2: Custom Router (Intelligence & Hardware) - # Already deployed separately not in this compose file - # The router is managed by the existing harness deployment on CT 116 - volumes: pgdata: ``` --- ---- - ## 8. Risk Mitigation | Risk | Mitigation | |------|------------| -| LiteLLM adds latency overhead | Shadow mode measures: <50ms extra is acceptable for admin features. LiteLLM is a thin proxy. | -| LiteLLM down = all agents down | NGINX fallback to router :9000 direct (see 6). Agents can also be configured with dual endpoints. | -| Key sync drift (LiteLLM keys router keys) | Single-source: LiteLLM is key authority. Router uses one `ROUTER_API_KEY` from LiteLLM's perspective. Agent keys live in LiteLLM only. | -| Spend tracking inaccurate for local GPUs | Configure `model_cost` per GPU with $0 rate (self-hosted). Optionally track "internal cost" via custom pricing. | -| Double rate limiting (LiteLLM + Router) | Keep both intentionally: LiteLLM for per-user soft caps, Router for hardware protection. Non-overlapping concerns. | -| PostgreSQL failure | LiteLLM can run with SQLite fallback, but UI features degrade. Postgres is the recommended path. | -| Router custom logic becomes a black box to LiteLLM | Acceptable trade-off. LiteLLM sees router as opaque OpenAI endpoint. GPU-level routing decisions are router's domain. | - ---- +| LiteLLM adds latency overhead | Shadow mode measures: <50ms extra is acceptable | +| LiteLLM down = all agents down | NGINX fallback to router :9000 direct (see 6) | +| Explicit GPU saturated no fallback available | LiteLLM fallback chains try all 3 GPUs in order before failing | +| Fallback chain masking real GPU failures | Router returns `saturated: true` only for capacity, `down` returns different error | +| Key sync drift | Single-source: LiteLLM is key authority. Router uses one `ROUTER_API_KEY` | +| Spend tracking inaccurate for local GPUs | `model_cost` per GPU with $0 rate; optional symbolic pricing for internal billing | +| Double rate limiting | Intentional: LiteLLM for per-user caps, Router for hardware protection | +| PostgreSQL failure | LiteLLM can run with SQLite fallback; UI features degrade | +| Per-model metrics accuracy with syslog-auto | `syslog-auto` is opaque by design (content-based routing). Explicit models are accurate. Use explicit models for per-GPU billing. | --- @@ -692,85 +607,54 @@ volumes: | New agent onboarding | Generate key, update env var, redeploy router | Create in UI, share key | | Admin UX | curl + JSON responses | Visual dashboard, graphs, search | | Audit trail | Router logs (stdout only) | Database-backed with UI search | -| SSO | None | Google/GitHub/Microsoft OIDC | +| SSO | None | Authentik OIDC | | Budget enforcement | None | Automatic: key suspended at $limit | +| GPU failover | Silent (inaccurate metrics) | Explicit (LiteLLM fallback chains, per-model logs) | | GPU routing intelligence | Full (unchanged) | Full (unchanged) | -| GPU health monitoring | Full (unchanged) | Full (unchanged) | --- ---- - -## 10. Migration Commands (Quick Reference) CT 116 - -**Target:** CT 116 `syslog-api` (192.168.68.116), minipve. All services co-located. +## 10. Migration Commands (Quick Reference) ```bash # On CT 116 (SSH via minipve: pct exec 116 bash): -# === Phase 0: Infrastructure Prep === - -# 0. DNS split-horizon fix +# Phase 0: Infrastructure Prep echo "192.168.68.11 auth.sysloggh.net" >> /etc/hosts -getent hosts auth.sysloggh.net # Verify 192.168.68.11 -# 1. Navigate to harness deployment directory cd /opt/litellm - -# 2. Deploy Postgres (LiteLLM DB) docker compose up -d postgres - -# 3. Replace LiteLLM config with production config.yaml (see 4.3) -# - All models http://router:9000/v1 -# - Add guardrails (pre-call, post-call, content filter) -# - Set num_retries: 0 (router handles retry) -# - Set request_timeout: 600 - -# 4. Deploy custom_sso.py -# - Mount to LiteLLM container volume -# - Reference in config.yaml: custom_ui_sso_sign_in_handler: custom_sso.custom_ui_sso_sign_in_handler - -# 5. Restart LiteLLM container +# Replace config.yaml with production version (see 4.3) docker compose restart litellm -# 6. Verify internal routing +# Verify +curl http://127.0.0.1:4000/health curl -X POST http://127.0.0.1:4000/v1/chat/completions \ - -H "Authorization: Bearer *** \ - -H "Content-Type: application/json" \ - -d '{"model":"syslog-auto","messages":[{"role":"user","content":"test"}]}' + -H "Authorization: Bearer ***" \ + -H "Content-Type: application/json" \ + -d '{"model":"syslog-auto","messages":[{"role":"user","content":"test"}]}' ``` --- ## 11. GitOps Workflow -**Branching Strategy:** - `main` production-ready code - `feature/litellm-migration` current development branch -- `deploy/phase-0`, `deploy/phase-1`, etc. deployment-specific branches **Conventional Commits:** ``` -(): - -- feat(plan): update LiteLLM migration plan for CT 116 deployment with Authentik OIDC + zero-downtime strategy -- fix(router): handle None temp_c/vram_pct in gpu_health_score -- chore(docker): add postgres volume for LiteLLM database persistence -- docs: LiteLLM migration plan two-layer architecture with model identity gap analysis +feat(plan): add fallback chains and strict passthrough for model identity gap +fix(router): strict passthrough for explicit models no silent rerouting +feat(plan): update LiteLLM migration plan for CT 116 deployment with Authentik OIDC +docs: LiteLLM migration plan two-layer architecture with model identity gap analysis ``` -**Deployment Flow:** -1. Commit changes to `feature/litellm-migration` -2. Open PR to `main` -3. Review with `git diff --stat` -4. Merge to `main` after approval -5. Run deployment scripts in order: Phase 0 Phase 1 Phase 2 Phase 3 - --- ## 12. Zero-Downtime Migration Strategy -**Per-Agent Cutover (2 minutes):** +**Per-Agent Cutover (<2 minutes):** 1. Create LiteLLM virtual key in UI 2. Update agent's `OPENAI_API_BASE` to `:4000` 3. Verify routing works @@ -781,33 +665,64 @@ curl -X POST http://127.0.0.1:4000/v1/chat/completions \ 2. NGINX `router_fallback` handles automatic failover 3. Monitor GPU metrics for health checks -**Verification:** -- Check LiteLLM `/health` endpoint -- Verify GPU metrics via router `/metrics` -- Test each agent individually -- Confirm no 502/503 errors in logs - --- -## Appendix A: Model Identity Gap Analysis +## Appendix A: Model Identity Gap Analysis (RESOLVED) -| Model | GPU | VRAM | Context | Status | -|-------|-----|------|---------|--------| -| qwen3.6-35B-A3B | MoE/Strix | 65GB | 262K | Healthy | -| qwen3.6-27B-code | Dense/RTX3090 | 24GB | 262K | Healthy | -| gemma-4-12b | VLM/RTX 5070 | 12GB | 262K | Healthy | +### Problem Identified (Abiba, June 14) -**Notes:** -- All three GPUs are operational and available for LiteLLM routing -- GPU health scoring (40/30/30) prevents routing to unhealthy GPUs -- Router handles all GPU-level routing decisions LiteLLM sees them as opaque endpoints +The original architecture had a metrics accuracy gap: when an agent requested `qwen3.6-35B-A3B` and MoE was busy, the router silently rerouted to Dense. LiteLLM logged it as MoE usage, corrupting per-model spend/usage tracking. + +### Root Cause + +```python +# OLD router code (router-fixed.py): +if is_gpu_busy(target) and req in allowed: + alts = [m for m in avail if m != target and m in allowed] + if alts: + alt = select_best_gpu(alts, "explicit", agent) + if alt: return alt # silently changed GPU +``` + +### Resolution: Strict Passthrough + LiteLLM Fallback Chains + +Two changes deployed: + +**1. Router strict passthrough:** +```python +# NEW: strict mode no silent fallback +if req != "auto": + target = req if req in avail else avail[0] + if req not in avail: + return {"model": req, "reason": "explicit_unavailable", "saturated": True} + if is_gpu_busy(target): + return {"model": target, "reason": "explicit_saturated", "saturated": True} + return {"model": target, "reason": "explicit"} +``` + +**2. LiteLLM fallback chains (in config.yaml):** +```yaml +router_settings: + allowed_fails: 100 + fallbacks: + - qwen3.6-35B-A3B: ["qwen3.6-27B-code", "gemma-4-12b"] + - qwen3.6-27B-code: ["qwen3.6-35B-A3B", "gemma-4-12b"] + - gemma-4-12b: ["qwen3.6-27B-code", "qwen3.6-35B-A3B"] +``` + +### Result + +| Scenario | Before | After | +|----------|--------|-------| +| Agent asks for MoE, MoE available | MoE used, metrics OK | MoE used, metrics OK | +| Agent asks for MoE, MoE busy | Router Dense silently, metrics WRONG | Router 503, LiteLLM Dense, metrics show BOTH attempts | +| Agent uses syslog-auto | Router picks GPU, LiteLLM sees opaque | Same (syslog-auto is opaque by design) | +| All 3 GPUs saturated | Router queues (30s), then 503 | Same, LiteLLM sees 503 after fallback chain exhausted | --- ## Appendix B: LiteLLM Virtual Key Migration -**Current API Keys LiteLLM Virtual Keys:** - | Agent | Old Key | New LiteLLM Key | Tier | Budget | |-------|---------|-----------------|------|--------| | Abiba | sk-***-*** | sk-litellm-*** | enterprise | $1000 | @@ -836,10 +751,9 @@ curl -X POST http://127.0.0.1:4000/v1/chat/completions \ - LiteLLM metrics `http://127.0.0.1:4000/metrics` - Router metrics `http://127.0.0.1:9000/metrics` - GPU health metrics `http://127.0.0.1:9000/metrics/gpu` -- Circuit breaker metrics `http://127.0.0.1:9000/metrics/circuit-breaker` **Alerts:** -- GPU health score > 70 alert to Hermes -- Circuit breaker trip alert to Hermes -- LiteLLM spend > $100/day alert to Hermes -- LiteLLM latency > 1000ms alert to Hermes \ No newline at end of file +- GPU health score > 70 alert +- Circuit breaker trip alert +- LiteLLM spend > $100/day alert +- LiteLLM latency > 1000ms alert