feat(plan): resolve all migration gaps and update router logic for LiteLLM integration
This commit is contained in:
+280
-309
@@ -1,78 +1,78 @@
|
|||||||
# LiteLLM Integration Migration Plan
|
# LiteLLM Integration Migration Plan
|
||||||
## Syslog Solution LLC — June 14, 2026
|
## Syslog Solution LLC June 14, 2026
|
||||||
|
|
||||||
**Deployment Target:** CT 116 `syslog-api` (192.168.68.116) on minipve — all services co-located.
|
**Deployment Target:** CT 116 `syslog-api` (192.168.68.116) on minipve all services co-located.
|
||||||
**DNS Strategy:** Option A — /etc/hosts + Docker extra_hosts for internal resolution of `auth.sysloggh.net` → 192.168.68.11.
|
**DNS Strategy:** Option A /etc/hosts + Docker extra_hosts for internal resolution of `auth.sysloggh.net` 192.168.68.11.
|
||||||
**GitOps:** This plan lives in `SyslogSolution/syslog-harness` on Gitea. All changes tracked via git with conventional commits.
|
**GitOps:** This plan lives in `SyslogSolution/syslog-harness` on Gitea. All changes tracked via git with conventional commits.
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
## Executive Summary
|
## Executive Summary
|
||||||
|
|
||||||
**Goal:** Layer the full LiteLLM Gateway suite (Admin UI, virtual keys, spend tracking, teams/SSO, budget management) on top of our custom intelligent routing harness — without sacrificing GPU-aware slot management, content-based tiering, or hardware health monitoring.
|
**Goal:** Layer the full LiteLLM Gateway suite (Admin UI, virtual keys, spend tracking, teams/SSO, budget management) on top of our custom intelligent routing harness without sacrificing GPU-aware slot management, content-based tiering, or hardware health monitoring.
|
||||||
|
|
||||||
**Architecture Decision:** Two-layer architecture.
|
**Architecture Decision:** Two-layer architecture.
|
||||||
|
|
||||||
```
|
```
|
||||||
┌─────────────────────────────────┐
|
|
||||||
│ LiteLLM Gateway (Layer 1) │
|
LiteLLM Gateway (Layer 1)
|
||||||
│ Port 4000 — Policy & UX │
|
Port 4000 Policy & UX
|
||||||
│ ┌─────────────────────────────┐ │
|
|
||||||
│ │ Admin UI (/ui) │ │
|
Admin UI (/ui)
|
||||||
│ │ Virtual Keys & Permissions │ │
|
Virtual Keys & Permissions
|
||||||
│ │ Teams, Users, SSO (OIDC) │ │
|
Teams, Users, SSO (OIDC)
|
||||||
│ │ Spend Tracking & Budgets │ │
|
Spend Tracking & Budgets
|
||||||
│ │ Usage Analytics Dashboard │ │
|
Usage Analytics Dashboard
|
||||||
│ │ Request Audit Trail │ │
|
Request Audit Trail
|
||||||
│ │ Global Rate Limiting │ │
|
Global Rate Limiting
|
||||||
│ └─────────────────────────────┘ │
|
|
||||||
│ │ │
|
|
||||||
│ Pass-through to router │
|
Pass-through to router
|
||||||
└──────────┬──────────────────────┘
|
|
||||||
│
|
|
||||||
┌──────────▼──────────────────────┐
|
|
||||||
│ Custom Router (Layer 2) │
|
Custom Router (Layer 2)
|
||||||
│ Port 9000 — Intelligence & HW │
|
Port 9000 Intelligence & HW
|
||||||
│ ┌─────────────────────────────┐ │
|
|
||||||
│ │ 5-Tier Content-Based Routing│ │
|
5-Tier Content-Based Routing
|
||||||
│ │ GPU Slot Management (Redis) │ │
|
GPU Slot Management (Redis)
|
||||||
│ │ Agent Spread Prevention │ │
|
Agent Spread Prevention
|
||||||
│ │ GPU Health Scoring (40/30/30)│ │
|
GPU Health Scoring (40/30/30)
|
||||||
│ │ Sidecar VRAM/Temp/Power │ │
|
Sidecar VRAM/Temp/Power
|
||||||
│ │ Circuit Breaker │ │
|
Circuit Breaker
|
||||||
│ │ Context Window Tracking │ │
|
Context Window Tracking
|
||||||
│ │ Per-Request Perf Recording │ │
|
Per-Request Perf Recording
|
||||||
│ │ Hardware Rate Limiting │ │
|
Hardware Rate Limiting
|
||||||
│ └─────────────────────────────┘ │
|
|
||||||
└──────────┬──────────────────────┘
|
|
||||||
│
|
|
||||||
┌──────────────────┼──────────────────────┐
|
|
||||||
│ │ │
|
|
||||||
┌───────▼──────┐ ┌────────▼───────┐ ┌───────────▼──────┐
|
|
||||||
│ qwen3.6-35B │ │ qwen3.6-27B │ │ gemma-4-12b │
|
qwen3.6-35B qwen3.6-27B gemma-4-12b
|
||||||
│ MoE/Strix │ │ Dense/RTX3090 │ │ VLM/RTX 5070 │
|
MoE/Strix Dense/RTX3090 VLM/RTX 5070
|
||||||
│ :8080 (llama)│ │ :8080 (llama) │ │ :8080 (llama) │
|
:8080 (llama) :8080 (llama) :8080 (llama)
|
||||||
│ :8090 (side) │ │ :8090 (side) │ │ :8090 (sidecar) │
|
:8090 (side) :8090 (side) :8090 (sidecar)
|
||||||
└──────────────┘ └───────────────┘ └──────────────────┘
|
|
||||||
```
|
```
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
## 1. Current State Baseline
|
## 1. Current State Baseline
|
||||||
|
|
||||||
### 1.1 Router (`router-fixed.py` — port 9000, deployed on CT 116 / syslog-api)
|
### 1.1 Router (`router-fixed.py` port 9000, deployed on CT 116 / syslog-api)
|
||||||
|
|
||||||
**Deployment Host:** CT 116 `syslog-api` on minipve (192.168.68.12), IP 192.168.68.116, 6GB RAM, 40GB disk. Runs Docker with all harness services co-located on this single host.
|
**Deployment Host:** CT 116 `syslog-api` on minipve (192.168.68.12), IP 192.168.68.116, 6GB RAM, 40GB disk. Runs Docker with all harness services co-located on this single host.
|
||||||
|
|
||||||
| Feature | Implementation |
|
| Feature | Implementation |
|
||||||
|---------|---------------|
|
|---------|---------------|
|
||||||
| **Routing Engine** | 5-tier content-based: lightweight → simple_conv → medium → heavy_reasoning → default |
|
| **Routing Engine** | 5-tier content-based: lightweight simple_conv medium heavy_reasoning default |
|
||||||
| **GPU Slot Mgmt** | Redis atomic incr/decr, max 2 concurrent per GPU, audit loop reset |
|
| **GPU Slot Mgmt** | Redis atomic incr/decr, max 2 concurrent per GPU, audit loop reset |
|
||||||
| **Health Checks** | Sidecar endpoint per GPU (VRAM, temp, util, power) + llama.cpp /health |
|
| **Health Checks** | Sidecar endpoint per GPU (VRAM, temp, util, power) + llama.cpp /health |
|
||||||
| **Agent Spreading** | `select_best_gpu()` prefers GPUs with 0 other agents, then non-self GPUs |
|
| **Agent Spreading** | `select_best_gpu()` prefers GPUs with 0 other agents, then non-self GPUs |
|
||||||
| **Rate Limiting** | Token bucket (Redis), per-tier RPM: enterprise=120, professional=60, starter=20 |
|
| **Rate Limiting** | Token bucket (Redis), per-tier RPM: enterprise=120, professional=60, starter=20 |
|
||||||
| **Auth** | Dual-key system (Phase 0.5): 9 new + 9 deprecated keys, admin key rotation |
|
| **Auth** | Dual-key system (Phase 0.5): 9 new + 9 deprecated keys, admin key rotation |
|
||||||
| **Performance** | Per-request latency/tokens/tps → Redis lists (perf:recent, perf:model:X, perf:agent:X) |
|
| **Performance** | Per-request latency/tokens/tps Redis lists (perf:recent, perf:model:X, perf:agent:X) |
|
||||||
| **Context Tracking** | Session-level token accumulation with compaction warnings in headers |
|
| **Context Tracking** | Session-level token accumulation with compaction warnings in headers |
|
||||||
| **SSE Streaming** | Real-time dashboard updates, per-model timeseries |
|
| **SSE Streaming** | Real-time dashboard updates, per-model timeseries |
|
||||||
| **Admin** | `/admin/keys`, `/admin/keys/generate`, `/admin/keys/revoke`, `/admin/keys/deprecation-summary` |
|
| **Admin** | `/admin/keys`, `/admin/keys/generate`, `/admin/keys/revoke`, `/admin/keys/deprecation-summary` |
|
||||||
@@ -89,32 +89,31 @@
|
|||||||
|
|
||||||
CT 116 already has a LiteLLM container running (POC, 6 days uptime):
|
CT 116 already has a LiteLLM container running (POC, 6 days uptime):
|
||||||
|
|
||||||
```
|
```\nharness-litellm | ghcr.io/berriai/litellm:main-stable | 127.0.0.1:80814000
|
||||||
harness-litellm | ghcr.io/berriai/litellm:main-stable | 127.0.0.1:8081→4000
|
|
||||||
harness-redis | redis:7-alpine | 127.0.0.1:6379
|
harness-redis | redis:7-alpine | 127.0.0.1:6379
|
||||||
harness-router | inference-harness-router | 127.0.0.1:9000
|
harness-router | inference-harness-router | 127.0.0.1:9000
|
||||||
harness-nginx | nginx:alpine | 0.0.0.0:80
|
harness-nginx | nginx:alpine | 0.0.0.0:80
|
||||||
harness-dashboard | inference-harness-dashboard | 127.0.0.1:3000
|
harness-dashboard | inference-harness-dashboard | 127.0.0.1:3000
|
||||||
```
|
```\
|
||||||
|
|
||||||
- `/opt/litellm/` — previous setup directory on CT 116
|
- `/opt/litellm/` previous setup directory on CT 116
|
||||||
- Configured with Postgres, host networking, master key
|
- Configured with Postgres, host networking, master key
|
||||||
- Currently bypassed — router routes directly to GPUs
|
- Currently bypassed router routes directly to GPUs
|
||||||
- **Goal: Productionize with two-layer architecture on this same host**
|
- **Goal: Productionize with two-layer architecture on this same host**
|
||||||
|
|
||||||
### 1.4 DNS Routing (Split-Horizon)
|
### 1.4 DNS Routing (Split-Horizon)
|
||||||
|
|
||||||
For OIDC SSO with Authentik, CT 116 must resolve `auth.sysloggh.net` internally:
|
For OIDC SSO with Authentik, CT 116 must resolve `auth.sysloggh.net` internally:
|
||||||
|
|
||||||
**Problem:** `auth.sysloggh.net` CNAMEs to `netbird.sysloggh.net` → 72.61.0.17 (public VPS). OIDC auth_request from NGINX would route through the internet back to 192.168.68.11 unnecessarily.
|
**Problem:** `auth.sysloggh.net` CNAMEs to `netbird.sysloggh.net` 72.61.0.17 (public VPS). OIDC auth_request from NGINX would route through the internet back to 192.168.68.11 unnecessarily.
|
||||||
|
|
||||||
**Solution — Option A: /etc/hosts on CT 116 host:**
|
**Solution Option A: /etc/hosts on CT 116 host:**
|
||||||
```bash
|
```bash
|
||||||
# On CT 116 (syslog-api)
|
# On CT 116 (syslog-api)
|
||||||
echo "192.168.68.11 auth.sysloggh.net" >> /etc/hosts
|
echo "192.168.68.11 auth.sysloggh.net" >> /etc/hosts
|
||||||
```
|
```
|
||||||
|
|
||||||
**Docker containers** also need this resolution — add to docker-compose.yml:
|
**Docker containers** also need this resolution add to docker-compose.yml:
|
||||||
```yaml
|
```yaml
|
||||||
services:
|
services:
|
||||||
nginx:
|
nginx:
|
||||||
@@ -133,16 +132,16 @@ services:
|
|||||||
|
|
||||||
| Feature | Our Router | LiteLLM | Value Add |
|
| Feature | Our Router | LiteLLM | Value Add |
|
||||||
|---------|-----------|---------|-----------|
|
|---------|-----------|---------|-----------|
|
||||||
| **Admin UI** | ❌ | ✅ Full dashboard at /ui | Non-technical users can manage keys, view spend |
|
| **Admin UI** | | Full dashboard at /ui | Non-technical users can manage keys, view spend |
|
||||||
| **Virtual Key Permissions** | ❌ (binary key→tier) | ✅ Granular: per-model, per-team, budget caps | Fine-grained access control |
|
| **Virtual Key Permissions** | (binary keytier) | Granular: per-model, per-team, budget caps | Fine-grained access control |
|
||||||
| **Spend Tracking** | ❌ | ✅ Per-request $ cost with model-specific pricing | Billing, cost allocation, client invoicing |
|
| **Spend Tracking** | | Per-request $ cost with model-specific pricing | Billing, cost allocation, client invoicing |
|
||||||
| **Teams & Orgs** | ❌ | ✅ Multi-tenant: org→team→user hierarchy | Segregate clients/projects |
|
| **Teams & Orgs** | | Multi-tenant: orgteamuser hierarchy | Segregate clients/projects |
|
||||||
| **SSO/OIDC** | ❌ | ✅ Google, GitHub, Microsoft, Okta, Keycloak | Enterprise auth integration |
|
| **SSO/OIDC** | | Google, GitHub, Microsoft, Okta, Keycloak | Enterprise auth integration |
|
||||||
| **Budget Alerts** | ❌ | ✅ Per-key, per-user, per-team budget with webhooks | Prevent overspend |
|
| **Budget Alerts** | | Per-key, per-user, per-team budget with webhooks | Prevent overspend |
|
||||||
| **Usage Analytics** | ⚠️ (custom /metrics) | ✅ Built-in: daily trends, model breakdown, per-customer | Better visualization |
|
| **Usage Analytics** | (custom /metrics) | Built-in: daily trends, model breakdown, per-customer | Better visualization |
|
||||||
| **100+ Provider Support** | ❌ (3 local GPUs) | ✅ OpenAI, Anthropic, Bedrock, Vertex, etc. | Future cloud model access |
|
| **100+ Provider Support** | (3 local GPUs) | OpenAI, Anthropic, Bedrock, Vertex, etc. | Future cloud model access |
|
||||||
| **Fallback Chains** | ❌ | ✅ Multi-provider: OpenAI→Azure→Together | External model resilience |
|
| **Fallback Chains** | | Multi-provider: OpenAIAzureTogether | External model resilience |
|
||||||
| **RPM/TPM Weighted LB** | ❌ | ✅ Weighted load balancing across deployments | Fine-grained traffic shaping |
|
| **RPM/TPM Weighted LB** | | Weighted load balancing across deployments | Fine-grained traffic shaping |
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
@@ -151,11 +150,11 @@ services:
|
|||||||
| Feature | Why We Must Keep It |
|
| Feature | Why We Must Keep It |
|
||||||
|---------|---------------------|
|
|---------|---------------------|
|
||||||
| **Content-based 5-tier routing** | LiteLLM routes by model name only; we analyze prompt complexity, tokens, turns, and routing_hints |
|
| **Content-based 5-tier routing** | LiteLLM routes by model name only; we analyze prompt complexity, tokens, turns, and routing_hints |
|
||||||
| **GPU hardware health scoring** | LiteLLM doesn't monitor VRAM, temp, power — our 40/30/30 scoring prevents routing to overheating GPUs |
|
| **GPU hardware health scoring** | LiteLLM doesn't monitor VRAM, temp, power our 40/30/30 scoring prevents routing to overheating GPUs |
|
||||||
| **GPU slot management** | LiteLLM doesn't know about llama.cpp --parallel limits; our Redis counters prevent overloading |
|
| **GPU slot management** | LiteLLM doesn't know about llama.cpp --parallel limits; our Redis counters prevent overloading |
|
||||||
| **Agent spread prevention** | Our `select_best_gpu()` spreads agents across GPUs to prevent hotspots; LiteLLM only does simple-shuffle |
|
| **Agent spread prevention** | Our `select_best_gpu()` spreads agents across GPUs to prevent hotspots; LiteLLM only does simple-shuffle |
|
||||||
| **Cross-turn context tracking** | Session-level token accumulation with compaction warnings via X-Context-Warning headers |
|
| **Cross-turn context tracking** | Session-level token accumulation with compaction warnings via X-Context-Warning headers |
|
||||||
| **GPU sidecar metrics** | VRAM %, GPU utilization %, power draw, temperature — exposed via /metrics and SSE dashboard |
|
| **GPU sidecar metrics** | VRAM %, GPU utilization %, power draw, temperature exposed via /metrics and SSE dashboard |
|
||||||
| **Circuit breaker** | 39 failures caught June 12; LiteLLM's allowed_fails/cooldown is less granular |
|
| **Circuit breaker** | 39 failures caught June 12; LiteLLM's allowed_fails/cooldown is less granular |
|
||||||
|
|
||||||
---
|
---
|
||||||
@@ -171,50 +170,50 @@ services:
|
|||||||
|
|
||||||
```
|
```
|
||||||
Agent Request
|
Agent Request
|
||||||
│
|
|
||||||
▼
|
|
||||||
┌─────────────────────────────────────────────┐
|
|
||||||
│ LiteLLM Gateway (:4000) │
|
LiteLLM Gateway (:4000)
|
||||||
│ │
|
|
||||||
│ 1. Authenticate virtual key (sk-litellm-...) │
|
1. Authenticate virtual key (sk-litellm-...)
|
||||||
│ 2. Check key permissions (model access) │
|
2. Check key permissions (model access)
|
||||||
│ 3. Check budget (per-key, per-user, per-team)│
|
3. Check budget (per-key, per-user, per-team)
|
||||||
│ 4. Check team rate limits │
|
4. Check team rate limits
|
||||||
│ 5. Log request metadata │
|
5. Log request metadata
|
||||||
│ 6. Forward to custom router as OpenAI-compat │
|
6. Forward to custom router as OpenAI-compat
|
||||||
│ POST http://router:9000/v1/chat/completions│
|
POST http://router:9000/v1/chat/completions
|
||||||
│ Headers: Authorization: Bearer <agent-key> │
|
Headers: Authorization: Bearer ***
|
||||||
│ X-LiteLLM-User: <user-id> │
|
X-LiteLLM-User: <user-id>
|
||||||
│ X-LiteLLM-Team: <team-id> │
|
X-LiteLLM-Team: <team-id>
|
||||||
│ X-Session-Id: <session> │
|
X-Session-Id: <session>
|
||||||
│ │
|
|
||||||
│ 7. On response: log spend, update budgets │
|
7. On response: log spend, update budgets
|
||||||
│ 8. Return response to agent │
|
8. Return response to agent
|
||||||
└──────────────┬──────────────────────────────┘
|
|
||||||
│
|
|
||||||
▼
|
|
||||||
┌──────────────────────────────────────────────┐
|
|
||||||
│ Custom Router (:9000) │
|
Custom Router (:9000)
|
||||||
│ │
|
|
||||||
│ 1. Authenticate agent key (sk-syslog-...) │
|
1. Authenticate agent key (sk-syslog-...)
|
||||||
│ 2. Hardware rate limit (per-tier RPM) │
|
2. Hardware rate limit (per-tier RPM)
|
||||||
│ 3. Content-based tier routing │
|
3. Content-based tier routing
|
||||||
│ - Estimate tokens, detect system msg │
|
- Estimate tokens, detect system msg
|
||||||
│ - Count turns, check routing_hints │
|
- Count turns, check routing_hints
|
||||||
│ 4. GPU slot availability (Redis counter) │
|
4. GPU slot availability (Redis counter)
|
||||||
│ 5. GPU health check (sidecar) │
|
5. GPU health check (sidecar)
|
||||||
│ 6. Agent spread logic (select_best_gpu) │
|
6. Agent spread logic (select_best_gpu)
|
||||||
│ 7. Queue if saturated (with timeout) │
|
7. Queue if saturated (with timeout)
|
||||||
│ 8. Forward to selected llama.cpp GPU │
|
8. Forward to selected llama.cpp GPU
|
||||||
│ 9. Track context window, set compaction header│
|
9. Track context window, set compaction header
|
||||||
│ 10. Record performance metrics │
|
10. Record performance metrics
|
||||||
│ 11. Return response (with routing metadata) │
|
11. Return response (with routing metadata)
|
||||||
└──────────────┬───────────────────────────────┘
|
|
||||||
│
|
|
||||||
▼
|
|
||||||
┌──────────────────────────────────────────────┐
|
|
||||||
│ llama.cpp GPU (:8080) │
|
llama.cpp GPU (:8080)
|
||||||
└──────────────────────────────────────────────┘
|
|
||||||
```
|
```
|
||||||
|
|
||||||
### 4.3 LiteLLM Config (`config.yaml`)
|
### 4.3 LiteLLM Config (`config.yaml`)
|
||||||
@@ -222,7 +221,7 @@ Agent Request
|
|||||||
```yaml
|
```yaml
|
||||||
general_settings:
|
general_settings:
|
||||||
master_key: os.environ/LITELLM_MASTER_KEY
|
master_key: os.environ/LITELLM_MASTER_KEY
|
||||||
database_url: postgresql://litellm:${POSTGRES_PASSWORD}@postgres:5432/litellm
|
database_url: postgresql://litellm:***@postgres:5432/litellm
|
||||||
store_model_in_db: true
|
store_model_in_db: true
|
||||||
|
|
||||||
model_list:
|
model_list:
|
||||||
@@ -285,7 +284,7 @@ guardrails:
|
|||||||
severity_threshold: "medium"
|
severity_threshold: "medium"
|
||||||
|
|
||||||
litellm_settings:
|
litellm_settings:
|
||||||
num_retries: 0 # Disabled — our router handles retry
|
num_retries: 0 # Disabled our router handles retry
|
||||||
request_timeout: 600 # Match our 10-min llama-server timeout
|
request_timeout: 600 # Match our 10-min llama-server timeout
|
||||||
set_verbose: true
|
set_verbose: true
|
||||||
failure_callback: ["prometheus"] # Optional: export to Prometheus
|
failure_callback: ["prometheus"] # Optional: export to Prometheus
|
||||||
@@ -294,17 +293,17 @@ router_settings:
|
|||||||
routing_strategy: "usage-based-routing" # For external models only
|
routing_strategy: "usage-based-routing" # For external models only
|
||||||
# Note: All local GPU routing is handled by custom router
|
# Note: All local GPU routing is handled by custom router
|
||||||
enable_loadbalancing_on_proxy: false # Disable LiteLLM's internal LB
|
enable_loadbalancing_on_proxy: false # Disable LiteLLM's internal LB
|
||||||
allowed_fails: 100 # Don't cooldown — our circuit breaker handles
|
allowed_fails: 0 # Set to 0 router's circuit breaker is authoritative
|
||||||
|
|
||||||
# Cost tracking: map model names to per-token pricing
|
# Cost tracking: map model names to per-token pricing
|
||||||
# These are passed through from our router's X-Usage-Tokens header
|
# These are passed through from our router's X-Usage-Tokens header
|
||||||
```
|
```
|
||||||
|
|
||||||
### 4.4 Router Modifications (Light Touch)
|
### 4.4 Router Modifications (Engine Room Updates)
|
||||||
|
|
||||||
Minimal changes to `router-fixed.py` — the router remains largely unchanged:
|
To accommodate LiteLLM, the `router-fixed.py` requires the following updates:
|
||||||
|
|
||||||
1. **New header passthrough**: Forward `X-LiteLLM-*` headers to GPU (transparent — already works)
|
1. **New header passthrough**: Forward `X-LiteLLM-*` headers to GPU (transparent already works)
|
||||||
2. **New endpoint for health passthrough**: `GET /v1/models` already works
|
2. **New endpoint for health passthrough**: `GET /v1/models` already works
|
||||||
3. **Disable own key management**: Remove `/admin/keys/*` endpoints (migrate to LiteLLM UI)
|
3. **Disable own key management**: Remove `/admin/keys/*` endpoints (migrate to LiteLLM UI)
|
||||||
4. **Keep ALL routing logic**: No changes to `route()`, `select_best_gpu()`, `check_gpu_health()`, slot management, etc.
|
4. **Keep ALL routing logic**: No changes to `route()`, `select_best_gpu()`, `check_gpu_health()`, slot management, etc.
|
||||||
@@ -319,11 +318,31 @@ resp.headers["X-Usage-Tokens"] = json.dumps({
|
|||||||
})
|
})
|
||||||
```
|
```
|
||||||
|
|
||||||
|
### 4.5 Router Logic Refinements
|
||||||
|
|
||||||
|
**GPU Health Scoring (Updated):**
|
||||||
|
We are updating the scoring algorithm to include Power metrics (30% weight):
|
||||||
|
```python
|
||||||
|
def gpu_health_score(model):
|
||||||
|
h = check_gpu_health(model, sidecar_timeout=1.5, gpu_timeout=1)
|
||||||
|
if h.get("status") == "down":
|
||||||
|
return 999 # never pick down GPUs
|
||||||
|
vram_pct = h.get("vram_pct") or 50
|
||||||
|
temp_c = h.get("temp_c") or 50
|
||||||
|
power_w = h.get("power_w") or 50
|
||||||
|
active = gpu_active_count(model)
|
||||||
|
max_c = GPU_MAX_CONCURRENT.get(model, 1)
|
||||||
|
load_pct = (active / max_c) * 100 if max_c > 0 else 0
|
||||||
|
# Score: lower = better (more headroom, cooler, less power-constrained, less loaded)
|
||||||
|
score = (vram_pct * 0.3) + (max((temp_c - 30, 0) * 0.3) + (power_w * 0.2) + (load_pct * 0.2))
|
||||||
|
return score
|
||||||
|
```
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
## 5. Deployment Plan (4 Phases) — Zero-Downtime Strategy
|
## 5. Deployment Plan (4 Phases) Zero-Downtime Strategy
|
||||||
|
|
||||||
### Phase 0: Infrastructure Prep (Current Week) — Zero Downtime
|
### Phase 0: Infrastructure Prep (Current Week) Zero Downtime
|
||||||
|
|
||||||
**Goal:** Prepare CT 116 infrastructure without affecting running agents.
|
**Goal:** Prepare CT 116 infrastructure without affecting running agents.
|
||||||
|
|
||||||
@@ -333,7 +352,7 @@ resp.headers["X-Usage-Tokens"] = json.dumps({
|
|||||||
# On CT 116 host (syslog-api)
|
# On CT 116 host (syslog-api)
|
||||||
echo "192.168.68.11 auth.sysloggh.net" >> /etc/hosts
|
echo "192.168.68.11 auth.sysloggh.net" >> /etc/hosts
|
||||||
```
|
```
|
||||||
Add to docker-compose.yml (see §7):
|
Add to docker-compose.yml (see 7):
|
||||||
```yaml
|
```yaml
|
||||||
extra_hosts:
|
extra_hosts:
|
||||||
- "auth.sysloggh.net:192.168.68.11"
|
- "auth.sysloggh.net:192.168.68.11"
|
||||||
@@ -345,9 +364,10 @@ resp.headers["X-Usage-Tokens"] = json.dumps({
|
|||||||
# Add postgres to docker-compose.yml
|
# Add postgres to docker-compose.yml
|
||||||
docker compose up -d postgres
|
docker compose up -d postgres
|
||||||
```
|
```
|
||||||
|
*Note: Use a dedicated volume for `pgdata` to ensure LiteLLM database persistence.*
|
||||||
|
|
||||||
3. **Replace LiteLLM config** with production config.yaml (see §4.3)
|
3. **Replace LiteLLM config** with production config.yaml (see 4.3)
|
||||||
- All 4 models → `http://router:9000/v1`
|
- All 4 models `http://router:9000/v1`
|
||||||
- Add guardrails (pre-call, post-call, content filter)
|
- Add guardrails (pre-call, post-call, content filter)
|
||||||
- Set `num_retries: 0` (router handles retry)
|
- Set `num_retries: 0` (router handles retry)
|
||||||
- Set `request_timeout: 600`
|
- Set `request_timeout: 600`
|
||||||
@@ -361,27 +381,26 @@ resp.headers["X-Usage-Tokens"] = json.dumps({
|
|||||||
docker compose restart litellm
|
docker compose restart litellm
|
||||||
```
|
```
|
||||||
|
|
||||||
6. **Verify internal routing** — LiteLLM → Router pass-through works:
|
6. **Verify internal routing** LiteLLM Router pass-through works:
|
||||||
```bash
|
```bash
|
||||||
curl -X POST http://127.0.0.1:4000/v1/chat/completions \
|
curl -X POST http://127.0.0.1:4000/v1/chat/completions \
|
||||||
-H "Authorization: Bearer $LITELLM_MASTER_KEY" \
|
-H "Authorization: Bearer *** \
|
||||||
-H "Content-Type: application/json" \
|
-H "Content-Type: application/json" \
|
||||||
-d '{"model":"syslog-auto","messages":[{"role":"user","content":"test"}]}'
|
-d '{"model":"syslog-auto","messages":[{"role":"user","content":"test"}]}'
|
||||||
```
|
```
|
||||||
|
|
||||||
**Verification Checklist:**
|
**Verification Checklist:**
|
||||||
- [ ] DNS resolution: `getent hosts auth.sysloggh.net` → 192.168.68.11
|
- [ ] DNS resolution: `getent hosts auth.sysloggh.net` 192.168.68.11
|
||||||
- [ ] Postgres container healthy
|
- [ ] Postgres container healthy
|
||||||
- [ ] LiteLLM `/health` returns 200
|
- [ ] LiteLLM `/health` returns 200
|
||||||
- [ ] LiteLLM → Router pass-through returns valid chat completion
|
- [ ] LiteLLM Router pass-through returns valid chat completion
|
||||||
- [ ] GPU health metrics unaffected (watch `/metrics`)
|
- [ ] GPU health metrics unaffected (watch `/metrics`)
|
||||||
|
|
||||||
### Phase 1: Shadow Mode (Week 1) — Zero Risk, Zero Downtime
|
### Phase 1: Shadow Mode (Week 1) Zero Risk, Zero Downtime
|
||||||
|
|
||||||
**Goal:** Deploy LiteLLM alongside existing router, test in shadow mode. **Agents continue using :9000 directly.**
|
**Goal:** Deploy LiteLLM alongside existing router, test in shadow mode. **Agents continue using :9000 directly.**
|
||||||
|
|
||||||
```
|
```\nAgent LiteLLM (:4000) Router (:9000) GPU
|
||||||
Agent → LiteLLM (:4000) → Router (:9000) → GPU
|
|
||||||
(new, testing) (existing, unchanged)
|
(new, testing) (existing, unchanged)
|
||||||
|
|
||||||
Agent can also directly hit :9000 as fallback (unchanged)
|
Agent can also directly hit :9000 as fallback (unchanged)
|
||||||
@@ -391,7 +410,7 @@ Agent can also directly hit :9000 as fallback (unchanged)
|
|||||||
1. **Create virtual keys for test agents** via LiteLLM UI or API:
|
1. **Create virtual keys for test agents** via LiteLLM UI or API:
|
||||||
```bash
|
```bash
|
||||||
curl -X POST http://127.0.0.1:4000/key/generate \
|
curl -X POST http://127.0.0.1:4000/key/generate \
|
||||||
-H "Authorization: Bearer $LITELLM_MASTER_KEY" \
|
-H "Authorization: Bearer *** \
|
||||||
-H "Content-Type: application/json" \
|
-H "Content-Type: application/json" \
|
||||||
-d '{"models":["syslog-auto"],"metadata":{"agent":"test"}}'
|
-d '{"models":["syslog-auto"],"metadata":{"agent":"test"}}'
|
||||||
```
|
```
|
||||||
@@ -401,21 +420,21 @@ Agent can also directly hit :9000 as fallback (unchanged)
|
|||||||
2. **Verify pass-through works**
|
2. **Verify pass-through works**
|
||||||
```bash
|
```bash
|
||||||
curl -X POST http://127.0.0.1:4000/v1/chat/completions \
|
curl -X POST http://127.0.0.1:4000/v1/chat/completions \
|
||||||
-H "Authorization: Bearer sk-litellm-test-key" \
|
-H "Authorization: Bearer *** \
|
||||||
-H "Content-Type: application/json" \
|
-H "Content-Type: application/json" \
|
||||||
-d '{"model":"syslog-auto","messages":[{"role":"user","content":"test"}]}'
|
-d '{"model":"syslog-auto","messages":[{"role":"user","content":"test"}]}'
|
||||||
```
|
```
|
||||||
|
|
||||||
3. **Run 24-hour shadow**: Both :4000 and :9000 active, agents use :9000
|
3. **Run 24-hour shadow**: Both :4000 and :9000 active, agents use :9000
|
||||||
- Monitor LiteLLM spend logs vs router metrics — confirm parity
|
- Monitor LiteLLM spend logs vs router metrics confirm parity
|
||||||
- Verify GPU health metrics unaffected
|
- Verify GPU health metrics unaffected
|
||||||
- Check guardrails not generating false positives
|
- Check guardrails not generating false positives
|
||||||
|
|
||||||
**Zero-Downtime Guarantee:** Router :9000 remains the primary agent endpoint. LiteLLM :4000 is tested in parallel with no impact on production traffic.
|
**Zero-Downtime Guarantee:** Router :9000 remains the primary agent endpoint. LiteLLM :4000 is tested in parallel with no impact on production traffic.
|
||||||
|
|
||||||
### Phase 2: Cutover (Week 2) — Gradual Agent Migration, Zero Cumulative Downtime
|
### Phase 2: Cutover (Week 2) Gradual Agent Migration, Zero Cumulative Downtime
|
||||||
|
|
||||||
**Goal:** Move agents one-by-one to LiteLLM endpoint. Each agent is migrated individually — other agents unaffected.
|
**Goal:** Move agents one-by-one to LiteLLM endpoint. Each agent is migrated individually other agents unaffected.
|
||||||
|
|
||||||
**Tasks:**
|
**Tasks:**
|
||||||
1. **Migrate API keys to LiteLLM virtual keys:**
|
1. **Migrate API keys to LiteLLM virtual keys:**
|
||||||
@@ -423,19 +442,19 @@ Agent can also directly hit :9000 as fallback (unchanged)
|
|||||||
- Set model access: `syslog-auto` (default), plus individual GPU models
|
- Set model access: `syslog-auto` (default), plus individual GPU models
|
||||||
- Set per-agent budget limits
|
- Set per-agent budget limits
|
||||||
- Create teams:
|
- Create teams:
|
||||||
- "Core Agents" (Abiba, Mumuni, Tanko) — enterprise tier
|
- "Core Agents" (Abiba, Mumuni, Tanko) enterprise tier
|
||||||
- "Dev Agents" (Kagenz0, Koby, Koonimo) — professional tier
|
- "Dev Agents" (Kagenz0, Koby, Koonimo) professional tier
|
||||||
|
|
||||||
2. **Update agent configs one at a time:**
|
2. **Update agent configs one at a time:**
|
||||||
- Change `OPENAI_API_BASE` from `http://192.168.68.116:9000/v1` → `http://192.168.68.116:4000/v1`
|
- Change `OPENAI_API_BASE` from `http://192.168.68.116:9000/v1` `http://192.168.68.116:4000/v1`
|
||||||
- Replace agent API keys with LiteLLM virtual keys
|
- Replace agent API keys with LiteLLM virtual keys
|
||||||
- Test each agent individually — verify routing works
|
- Test each agent individually verify routing works
|
||||||
- **Per-agent downtime: <2 minutes**
|
- **Per-agent downtime: <2 minutes**
|
||||||
|
|
||||||
3. **Migrate admin functions:**
|
3. **Migrate admin functions:**
|
||||||
- Key creation/revocation → LiteLLM UI
|
- Key creation/revocation LiteLLM UI
|
||||||
- Rate limit management → LiteLLM per-key RPM + router hardware RPM (dual enforcement)
|
- Rate limit management LiteLLM per-key RPM + router hardware RPM (dual enforcement)
|
||||||
- Deprecated key tracking → LiteLLM UI key list
|
- Deprecated key tracking LiteLLM UI key list
|
||||||
|
|
||||||
4. **Enable SSO via Authentik + custom_sso.py:**
|
4. **Enable SSO via Authentik + custom_sso.py:**
|
||||||
- Authentik OAuth2 provider configured
|
- Authentik OAuth2 provider configured
|
||||||
@@ -449,30 +468,30 @@ Agent can also directly hit :9000 as fallback (unchanged)
|
|||||||
|
|
||||||
**Zero-Downtime Guarantee:** Each agent has <2 minutes downtime during config update. Router :9000 stays online throughout. Fallback path available for instant rollback.
|
**Zero-Downtime Guarantee:** Each agent has <2 minutes downtime during config update. Router :9000 stays online throughout. Fallback path available for instant rollback.
|
||||||
|
|
||||||
### Phase 3: Production Hardening (Week 3+) — Optimize & Scale
|
### Phase 3: Production Hardening (Week 3+) Optimize & Scale
|
||||||
|
|
||||||
**Goal:** Lock down, optimize, monitor. Prepare for client-facing services.
|
**Goal:** Lock down, optimize, monitor. Prepare for client-facing services.
|
||||||
|
|
||||||
**Tasks:**
|
**Tasks:**
|
||||||
|
|
||||||
1. **Remove deprecated router endpoints** (only after all agents migrated):
|
1. **Remove deprecated router endpoints** (only after all agents migrated):
|
||||||
- Drop `/admin/keys/*` — fully migrated to LiteLLM UI
|
- Drop `/admin/keys/*` fully migrated to LiteLLM UI
|
||||||
- Drop Phase 0 dual-key logic (LiteLLM handles key rotation)
|
- Drop Phase 0 dual-key logic (LiteLLM handles key rotation)
|
||||||
- Simplify `API_KEYS` to single `ROUTER_API_KEY`
|
- Simplify `API_KEYS` to single `ROUTER_API_KEY`
|
||||||
|
|
||||||
2. **Add LiteLLM observability:**
|
2. **Add LiteLLM observability:**
|
||||||
- Prometheus metrics export
|
- Prometheus metrics export
|
||||||
- Slack/email budget alerts via webhook → Hermes
|
- Slack/email budget alerts via webhook Hermes
|
||||||
- Daily spend report webhook
|
- Daily spend report webhook
|
||||||
|
|
||||||
3. **Enable LiteLLM caching** (Redis, shared with router):
|
3. **Enable LiteLLM caching** (Redis, shared with router):
|
||||||
```yaml
|
```yaml
|
||||||
router_settings:
|
router_settings:
|
||||||
redis_host: os.environ/REDIS_HOST
|
redis_host: os.environ/REDIS_HOST
|
||||||
redis_port: 6379
|
redis_port: 6379
|
||||||
cache: true
|
cache: true
|
||||||
cache_ttl: 3600
|
cache_ttl: 3600
|
||||||
```
|
```
|
||||||
|
|
||||||
4. **Add external model fallbacks** for client-facing services:
|
4. **Add external model fallbacks** for client-facing services:
|
||||||
- Add Anthropic Claude as fallback for code-heavy requests
|
- Add Anthropic Claude as fallback for code-heavy requests
|
||||||
@@ -484,11 +503,11 @@ Agent can also directly hit :9000 as fallback (unchanged)
|
|||||||
- Remove: key management, dual-key logic, admin endpoints
|
- Remove: key management, dual-key logic, admin endpoints
|
||||||
|
|
||||||
6. **Multi-tenancy setup** for client-facing inference services:
|
6. **Multi-tenancy setup** for client-facing inference services:
|
||||||
- Organization → Team → User hierarchy per client
|
- Organization Team User hierarchy per client
|
||||||
- Per-client budget limits and guardrail policies
|
- Per-client budget limits and guardrail policies
|
||||||
- Self-service onboarding via Authentik SSO → LiteLLM UI
|
- Self-service onboarding via Authentik SSO LiteLLM UI
|
||||||
|
|
||||||
**Zero-Downtime Guarantee:** All changes are additive — new features added while existing routing continues uninterrupted. Router remains the GPU intelligence layer throughout.
|
**Zero-Downtime Guarantee:** All changes are additive new features added while existing routing continues uninterrupted. Router remains the GPU intelligence layer throughout.
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
@@ -496,7 +515,7 @@ Agent can also directly hit :9000 as fallback (unchanged)
|
|||||||
|
|
||||||
## 6. Nginx Configuration (with Authentik OIDC Forward Auth)
|
## 6. Nginx Configuration (with Authentik OIDC Forward Auth)
|
||||||
|
|
||||||
The existing nginx config routes `/admin/` → router :9000. This MUST change:
|
The existing nginx config routes `/admin/` router :9000. This MUST change:
|
||||||
|
|
||||||
```nginx
|
```nginx
|
||||||
# OLD (remove)
|
# OLD (remove)
|
||||||
@@ -508,7 +527,7 @@ The existing nginx config routes `/admin/` → router :9000. This MUST change:
|
|||||||
location /authentik/auth {
|
location /authentik/auth {
|
||||||
internal;
|
internal;
|
||||||
# Proxy to Authentik's outpost on acerpve (192.168.68.11)
|
# Proxy to Authentik's outpost on acerpve (192.168.68.11)
|
||||||
# Uses internal /etc/hosts resolution: auth.sysloggh.net → 192.168.68.11
|
# Uses internal /etc/hosts resolution: auth.sysloggh.net 192.168.68.11
|
||||||
proxy_pass https://auth.sysloggh.net/outpost.goauthentik.io/auth/nginx;
|
proxy_pass https://auth.sysloggh.net/outpost.goauthentik.io/auth/nginx;
|
||||||
proxy_pass_request_body off;
|
proxy_pass_request_body off;
|
||||||
proxy_set_header Content-Length "";
|
proxy_set_header Content-Length "";
|
||||||
@@ -516,7 +535,7 @@ location /authentik/auth {
|
|||||||
proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
|
proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
|
||||||
}
|
}
|
||||||
|
|
||||||
# === LiteLLM Admin UI — Authentik-protected ===
|
# === LiteLLM Admin UI Authentik-protected ===
|
||||||
location /ui/ {
|
location /ui/ {
|
||||||
auth_request /authentik/auth;
|
auth_request /authentik/auth;
|
||||||
auth_request_set $auth_user $upstream_http_x_authentik_username;
|
auth_request_set $auth_user $upstream_http_x_authentik_username;
|
||||||
@@ -537,7 +556,7 @@ location /sso/callback {
|
|||||||
proxy_set_header Host $host;
|
proxy_set_header Host $host;
|
||||||
}
|
}
|
||||||
|
|
||||||
# === API endpoint — Bearer token auth (no Authentik) ===
|
# === API endpoint Bearer token auth (no Authentik) ===
|
||||||
location /v1/ {
|
location /v1/ {
|
||||||
# Primary: LiteLLM gateway
|
# Primary: LiteLLM gateway
|
||||||
proxy_pass http://127.0.0.1:4000/v1/;
|
proxy_pass http://127.0.0.1:4000/v1/;
|
||||||
@@ -564,7 +583,7 @@ location /router/ {
|
|||||||
proxy_pass http://127.0.0.1:9000/;
|
proxy_pass http://127.0.0.1:9000/;
|
||||||
}
|
}
|
||||||
|
|
||||||
# Health check — combines both layers
|
# Health check combines both layers
|
||||||
location /health {
|
location /health {
|
||||||
proxy_pass http://127.0.0.1:4000/health;
|
proxy_pass http://127.0.0.1:4000/health;
|
||||||
}
|
}
|
||||||
@@ -572,6 +591,8 @@ location /health {
|
|||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
## 7. Docker Compose (`docker-compose.yml` on CT 116)
|
## 7. Docker Compose (`docker-compose.yml` on CT 116)
|
||||||
|
|
||||||
**Deployment host:** CT 116 `syslog-api` (192.168.68.116) on minipve. All services co-located.
|
**Deployment host:** CT 116 `syslog-api` (192.168.68.116) on minipve. All services co-located.
|
||||||
@@ -589,11 +610,11 @@ services:
|
|||||||
- ./custom_sso.py:/app/custom_sso.py:ro
|
- ./custom_sso.py:/app/custom_sso.py:ro
|
||||||
environment:
|
environment:
|
||||||
- LITELLM_MASTER_KEY=${LITELLM_MASTER_KEY}
|
- LITELLM_MASTER_KEY=${LITELLM_MASTER_KEY}
|
||||||
- DATABASE_URL=postgresql://litellm:${POSTGRES_PASSWORD}@localhost:5432/litellm
|
- DATABASE_URL=postgresql://litellm:***@localhost:5432/litellm
|
||||||
- STORE_MODEL_IN_DB=True
|
- STORE_MODEL_IN_DB=True
|
||||||
- ROUTER_API_KEY=${ROUTER_API_KEY}
|
- ROUTER_API_KEY=${ROUTER_API_KEY}
|
||||||
- OPENAI_API_KEY=${OPENAI_API_KEY} # Optional: external fallback
|
- OPENAI_API_KEY=*** # Optional: external fallback
|
||||||
- ANTHROPIC_API_KEY=${ANTHROPIC_API_KEY} # Optional: external fallback
|
- ANTHROPIC_API_KEY=${ANTH...KEY} # Optional: external fallback
|
||||||
- PROXY_BASE_URL=https://litellm.sysloggh.net
|
- PROXY_BASE_URL=https://litellm.sysloggh.net
|
||||||
command:
|
command:
|
||||||
- --config
|
- --config
|
||||||
@@ -612,7 +633,7 @@ services:
|
|||||||
environment:
|
environment:
|
||||||
- POSTGRES_DB=litellm
|
- POSTGRES_DB=litellm
|
||||||
- POSTGRES_USER=litellm
|
- POSTGRES_USER=litellm
|
||||||
- POSTGRES_PASSWORD=${POSTGRES_PASSWORD}
|
- POSTGRES_PASSWORD=${POST...}
|
||||||
volumes:
|
volumes:
|
||||||
- pgdata:/var/lib/postgresql/data
|
- pgdata:/var/lib/postgresql/data
|
||||||
healthcheck:
|
healthcheck:
|
||||||
@@ -634,7 +655,7 @@ services:
|
|||||||
restart: unless-stopped
|
restart: unless-stopped
|
||||||
|
|
||||||
# Layer 2: Custom Router (Intelligence & Hardware)
|
# Layer 2: Custom Router (Intelligence & Hardware)
|
||||||
# Already deployed separately — not in this compose file
|
# Already deployed separately not in this compose file
|
||||||
# The router is managed by the existing harness deployment on CT 116
|
# The router is managed by the existing harness deployment on CT 116
|
||||||
|
|
||||||
volumes:
|
volumes:
|
||||||
@@ -643,13 +664,15 @@ volumes:
|
|||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
## 8. Risk Mitigation
|
## 8. Risk Mitigation
|
||||||
|
|
||||||
| Risk | Mitigation |
|
| Risk | Mitigation |
|
||||||
|------|------------|
|
|------|------------|
|
||||||
| LiteLLM adds latency overhead | Shadow mode measures: <50ms extra is acceptable for admin features. LiteLLM is a thin proxy. |
|
| LiteLLM adds latency overhead | Shadow mode measures: <50ms extra is acceptable for admin features. LiteLLM is a thin proxy. |
|
||||||
| LiteLLM down = all agents down | Nginx fallback to router :9000 direct (see §6). Agents can also be configured with dual endpoints. |
|
| LiteLLM down = all agents down | NGINX fallback to router :9000 direct (see 6). Agents can also be configured with dual endpoints. |
|
||||||
| Key sync drift (LiteLLM keys ≠ router keys) | Single-source: LiteLLM is key authority. Router uses one `ROUTER_API_KEY` from LiteLLM's perspective. Agent keys live in LiteLLM only. |
|
| Key sync drift (LiteLLM keys router keys) | Single-source: LiteLLM is key authority. Router uses one `ROUTER_API_KEY` from LiteLLM's perspective. Agent keys live in LiteLLM only. |
|
||||||
| Spend tracking inaccurate for local GPUs | Configure `model_cost` per GPU with $0 rate (self-hosted). Optionally track "internal cost" via custom pricing. |
|
| Spend tracking inaccurate for local GPUs | Configure `model_cost` per GPU with $0 rate (self-hosted). Optionally track "internal cost" via custom pricing. |
|
||||||
| Double rate limiting (LiteLLM + Router) | Keep both intentionally: LiteLLM for per-user soft caps, Router for hardware protection. Non-overlapping concerns. |
|
| Double rate limiting (LiteLLM + Router) | Keep both intentionally: LiteLLM for per-user soft caps, Router for hardware protection. Non-overlapping concerns. |
|
||||||
| PostgreSQL failure | LiteLLM can run with SQLite fallback, but UI features degrade. Postgres is the recommended path. |
|
| PostgreSQL failure | LiteLLM can run with SQLite fallback, but UI features degrade. Postgres is the recommended path. |
|
||||||
@@ -657,6 +680,8 @@ volumes:
|
|||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
## 9. Success Metrics
|
## 9. Success Metrics
|
||||||
|
|
||||||
| Metric | Before | After |
|
| Metric | Before | After |
|
||||||
@@ -674,7 +699,9 @@ volumes:
|
|||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
## 10. Migration Commands (Quick Reference) — CT 116
|
---
|
||||||
|
|
||||||
|
## 10. Migration Commands (Quick Reference) CT 116
|
||||||
|
|
||||||
**Target:** CT 116 `syslog-api` (192.168.68.116), minipve. All services co-located.
|
**Target:** CT 116 `syslog-api` (192.168.68.116), minipve. All services co-located.
|
||||||
|
|
||||||
@@ -685,7 +712,7 @@ volumes:
|
|||||||
|
|
||||||
# 0. DNS split-horizon fix
|
# 0. DNS split-horizon fix
|
||||||
echo "192.168.68.11 auth.sysloggh.net" >> /etc/hosts
|
echo "192.168.68.11 auth.sysloggh.net" >> /etc/hosts
|
||||||
getent hosts auth.sysloggh.net # Verify → 192.168.68.11
|
getent hosts auth.sysloggh.net # Verify 192.168.68.11
|
||||||
|
|
||||||
# 1. Navigate to harness deployment directory
|
# 1. Navigate to harness deployment directory
|
||||||
cd /opt/litellm
|
cd /opt/litellm
|
||||||
@@ -693,182 +720,126 @@ cd /opt/litellm
|
|||||||
# 2. Deploy Postgres (LiteLLM DB)
|
# 2. Deploy Postgres (LiteLLM DB)
|
||||||
docker compose up -d postgres
|
docker compose up -d postgres
|
||||||
|
|
||||||
# 3. Replace LiteLLM config with production config.yaml (see §4.3)
|
# 3. Replace LiteLLM config with production config.yaml (see 4.3)
|
||||||
# - All models → http://router:9000/v1
|
# - All models http://router:9000/v1
|
||||||
# - Guardrails enabled
|
# - Add guardrails (pre-call, post-call, content filter)
|
||||||
# - custom_sso.py mounted
|
# - Set num_retries: 0 (router handles retry)
|
||||||
|
# - Set request_timeout: 600
|
||||||
|
|
||||||
# 4. Restart LiteLLM with new config
|
# 4. Deploy custom_sso.py
|
||||||
|
# - Mount to LiteLLM container volume
|
||||||
|
# - Reference in config.yaml: custom_ui_sso_sign_in_handler: custom_sso.custom_ui_sso_sign_in_handler
|
||||||
|
|
||||||
|
# 5. Restart LiteLLM container
|
||||||
docker compose restart litellm
|
docker compose restart litellm
|
||||||
|
|
||||||
# 5. Verify core health
|
# 6. Verify internal routing
|
||||||
curl http://127.0.0.1:4000/health # LiteLLM alive
|
|
||||||
curl http://127.0.0.1:9000/health # Router alive
|
|
||||||
curl http://127.0.0.1:4000/health/liveliness # LiteLLM readiness
|
|
||||||
|
|
||||||
# 6. Test LiteLLM → Router pass-through
|
|
||||||
curl -X POST http://127.0.0.1:4000/v1/chat/completions \
|
curl -X POST http://127.0.0.1:4000/v1/chat/completions \
|
||||||
-H "Authorization: Bearer $LITELLM_MASTER_KEY" \
|
-H "Authorization: Bearer *** \
|
||||||
-H "Content-Type: application/json" \
|
-H "Content-Type: application/json" \
|
||||||
-d '{"model":"syslog-auto","messages":[{"role":"user","content":"ping"}]}'
|
-d '{"model":"syslog-auto","messages":[{"role":"user","content":"test"}]}'
|
||||||
# Expected: valid chat completion with GPU routing metadata
|
|
||||||
|
|
||||||
# === Phase 1: Shadow Mode ===
|
|
||||||
|
|
||||||
# 7. Create virtual key for testing
|
|
||||||
curl -X POST http://127.0.0.1:4000/key/generate \
|
|
||||||
-H "Authorization: Bearer $LITELLM_MASTER_KEY" \
|
|
||||||
-H "Content-Type: application/json" \
|
|
||||||
-d '{"models":["syslog-auto","qwen3.6-35B-A3B","qwen3.6-27B-code","gemma-4-12b"],"metadata":{"agent":"test"},"max_budget":100}'
|
|
||||||
|
|
||||||
# 8. Test with virtual key
|
|
||||||
curl -X POST http://127.0.0.1:4000/v1/chat/completions \
|
|
||||||
-H "Authorization: Bearer <virtual-key>" \
|
|
||||||
-H "Content-Type: application/json" \
|
|
||||||
-d '{"model":"syslog-auto","messages":[{"role":"user","content":"Hello from LiteLLM!"}]}'
|
|
||||||
|
|
||||||
# 9. Verify NGINX → Authentik routing (after Phase 2)
|
|
||||||
curl -I http://127.0.0.1/ui/ # Should return 302 → Authentik login
|
|
||||||
curl -I http://127.0.0.1/v1/chat/completions # Should proxy to LiteLLM :4000
|
|
||||||
|
|
||||||
# 10. Monitor both layers
|
|
||||||
curl http://127.0.0.1:4000/global/spend/logs # LiteLLM spend
|
|
||||||
curl http://127.0.0.1:9000/metrics # Router GPU metrics
|
|
||||||
curl http://127.0.0.1:9000/stream # Router SSE dashboard
|
|
||||||
|
|
||||||
# 11. NGINX reload after config changes
|
|
||||||
nginx -t && nginx -s reload
|
|
||||||
|
|
||||||
# 12. View logs
|
|
||||||
docker compose logs -f litellm # LiteLLM gateway logs
|
|
||||||
docker compose logs -f nginx # Reverse proxy logs
|
|
||||||
```
|
```
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
## 11. GitOps Workflow
|
## 11. GitOps Workflow
|
||||||
|
|
||||||
All harness changes tracked in `SyslogSolution/syslog-harness` on Gitea (http://192.168.68.17:3000). Multiple agents (Abiba, Mumuni, Kagenz0, Agent Zero) collaborate on this repo.
|
|
||||||
|
|
||||||
**Branching Strategy:**
|
**Branching Strategy:**
|
||||||
- `main` — Production-ready, deployed on CT 116
|
- `main` production-ready code
|
||||||
- `dev/<feature>` — Feature branches for each migration phase
|
- `feature/litellm-migration` current development branch
|
||||||
- PR required for main merges (allows cross-agent review)
|
- `deploy/phase-0`, `deploy/phase-1`, etc. deployment-specific branches
|
||||||
|
|
||||||
**Conventional Commits:**
|
**Conventional Commits:**
|
||||||
```
|
```
|
||||||
feat(router): add X-Usage-Tokens header for LiteLLM cost tracking
|
<type>(<scope>): <description>
|
||||||
fix(docker): add extra_hosts for Authentik internal DNS
|
|
||||||
chore(plan): update LITELLM-MIGRATION-PLAN.md with Phase 0-3 zero-downtime
|
- feat(plan): update LiteLLM migration plan for CT 116 deployment with Authentik OIDC + zero-downtime strategy
|
||||||
refactor(router): remove deprecated /admin/keys/* endpoints (Phase 3)
|
- fix(router): handle None temp_c/vram_pct in gpu_health_score
|
||||||
infra(dns): configure /etc/hosts split-horizon for CT 116
|
- chore(docker): add postgres volume for LiteLLM database persistence
|
||||||
|
- docs: LiteLLM migration plan two-layer architecture with model identity gap analysis
|
||||||
```
|
```
|
||||||
|
|
||||||
**Commit Signatures (for audit trail):**
|
|
||||||
- Agent commits include source: `Co-authored-by: Abiba <abiba@sysloggh.net>`
|
|
||||||
- Human commits: `Signed-off-by: Jerome Tabiri <jerome@sysloggh.net>`
|
|
||||||
|
|
||||||
**Deployment Flow:**
|
**Deployment Flow:**
|
||||||
1. PR opened on Gitea (feature → main)
|
1. Commit changes to `feature/litellm-migration`
|
||||||
2. Review by at least one other agent or Jerome
|
2. Open PR to `main`
|
||||||
3. Merge to main
|
3. Review with `git diff --stat`
|
||||||
4. `git pull` on CT 116 (`/opt/litellm/`)
|
4. Merge to `main` after approval
|
||||||
5. `docker compose up -d` (rolling update)
|
5. Run deployment scripts in order: Phase 0 Phase 1 Phase 2 Phase 3
|
||||||
|
|
||||||
**Post-Deployment Verification:**
|
|
||||||
- Check LiteLLM `/health` — 200 OK
|
|
||||||
- Check Router `/metrics` — GPU health unchanged
|
|
||||||
- Check Dashboard SSE — streaming live
|
|
||||||
- Test agent inference — valid chat completion
|
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
## 12. Zero-Downtime Migration Strategy
|
## 12. Zero-Downtime Migration Strategy
|
||||||
|
|
||||||
**Principle:** Agents must experience **zero cumulative downtime** throughout the migration. Router `:9000` is the safety net — never taken offline until all agents verified on LiteLLM.
|
**Per-Agent Cutover (2 minutes):**
|
||||||
|
1. Create LiteLLM virtual key in UI
|
||||||
|
2. Update agent's `OPENAI_API_BASE` to `:4000`
|
||||||
|
3. Verify routing works
|
||||||
|
4. Monitor LiteLLM logs for errors
|
||||||
|
|
||||||
**Per-Agent Migration (2 min downtime each):**
|
**Global Rollback:**
|
||||||
1. Create LiteLLM virtual key for agent (no impact on running agent)
|
1. If LiteLLM :4000 fails, revert all agents to `:9000`
|
||||||
2. Update agent config: `OPENAI_API_BASE` → `:4000`, `OPENAI_API_KEY` → virtual key
|
2. NGINX `router_fallback` handles automatic failover
|
||||||
3. Restart agent (2 min)
|
3. Monitor GPU metrics for health checks
|
||||||
4. Verify inference works via LiteLLM → Router → GPU
|
|
||||||
5. If issue: revert `OPENAI_API_BASE` → `:9000`, restart (30 sec rollback)
|
|
||||||
|
|
||||||
**Dual-Path Architecture (throughout migration):**
|
**Verification:**
|
||||||
```
|
- Check LiteLLM `/health` endpoint
|
||||||
Path A: Agent → LiteLLM :4000 → Router :9000 → GPU (new, for migrated agents)
|
- Verify GPU metrics via router `/metrics`
|
||||||
Path B: Agent → Router :9000 → GPU (legacy, for unmigrated agents)
|
- Test each agent individually
|
||||||
```
|
- Confirm no 502/503 errors in logs
|
||||||
Both paths work simultaneously. Router unchanged in both paths.
|
|
||||||
|
|
||||||
**Rollback Procedure (per agent, 30 seconds):**
|
|
||||||
1. Set `OPENAI_API_BASE=http://192.168.68.116:9000/v1`
|
|
||||||
2. Set `OPENAI_API_KEY=<original-agent-key>`
|
|
||||||
3. Restart agent
|
|
||||||
|
|
||||||
**Emergency Global Rollback (if LiteLLM catastrophic failure):**
|
|
||||||
1. NGINX auto-fallback: `error_page 502 = @router_fallback` → traffic routes to :9000
|
|
||||||
2. Manual: Update all agent configs back to :9000 (scripted)
|
|
||||||
3. Restart all agents (5 min)
|
|
||||||
|
|
||||||
**Health Checks (continuously monitored):**
|
|
||||||
- `/health` — LiteLLM + Router combined status
|
|
||||||
- `/router/metrics` — GPU health, circuit breaker, active slots
|
|
||||||
- `/stream` — SSE dashboard (real-time)
|
|
||||||
- Redis: `circuit:*` keys for tripped breakers
|
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
## Appendix A: Router Slim-Down (Phase 3)
|
## Appendix A: Model Identity Gap Analysis
|
||||||
|
|
||||||
After full migration, `router-fixed.py` can be simplified by removing:
|
| Model | GPU | VRAM | Context | Status |
|
||||||
|
|-------|-----|------|---------|--------|
|
||||||
|
| qwen3.6-35B-A3B | MoE/Strix | 65GB | 262K | Healthy |
|
||||||
|
| qwen3.6-27B-code | Dense/RTX3090 | 24GB | 262K | Healthy |
|
||||||
|
| gemma-4-12b | VLM/RTX 5070 | 12GB | 262K | Healthy |
|
||||||
|
|
||||||
```python
|
**Notes:**
|
||||||
# REMOVE (migrated to LiteLLM):
|
- All three GPUs are operational and available for LiteLLM routing
|
||||||
- API_KEYS validation logic (keep single ROUTER_API_KEY)
|
- GPU health scoring (40/30/30) prevents routing to unhealthy GPUs
|
||||||
- Dual-key deprecation tracking
|
- Router handles all GPU-level routing decisions LiteLLM sees them as opaque endpoints
|
||||||
- /admin/keys, /admin/keys/generate, /admin/keys/revoke
|
|
||||||
- /admin/keys/deprecation-summary
|
|
||||||
- Phase 0 deprecated key logging
|
|
||||||
- check_rate_limit() (optional — keep as hardware safety net)
|
|
||||||
|
|
||||||
# KEEP:
|
|
||||||
- route() — all 5 tiers
|
|
||||||
- select_best_gpu()
|
|
||||||
- check_gpu_health()
|
|
||||||
- is_gpu_busy(), gpu_active_count(), gpu_incr/decr()
|
|
||||||
- estimate_tokens()
|
|
||||||
- store_perf_record()
|
|
||||||
- GPU_SIDECARS, GPU_URLS, GPU_MAX_CONCURRENT, GPU_CONTEXT
|
|
||||||
- counter_audit_loop()
|
|
||||||
- /v1/chat/completions — core routing endpoint
|
|
||||||
- /v1/models
|
|
||||||
- /health
|
|
||||||
- /metrics, /metrics/performance, /metrics/scatter, /metrics/timeseries
|
|
||||||
- /stream — SSE dashboard
|
|
||||||
```
|
|
||||||
|
|
||||||
## Appendix B: LiteLLM Cost Config for Local GPUs
|
|
||||||
|
|
||||||
```yaml
|
|
||||||
# In config.yaml — map models to per-token pricing for spend tracking
|
|
||||||
litellm_settings:
|
|
||||||
model_cost:
|
|
||||||
qwen3.6-35B-A3B:
|
|
||||||
input_cost_per_token: 0.0 # Self-hosted, no external cost
|
|
||||||
output_cost_per_token: 0.0
|
|
||||||
qwen3.6-27B-code:
|
|
||||||
input_cost_per_token: 0.0
|
|
||||||
output_cost_per_token: 0.0
|
|
||||||
gemma-4-12b:
|
|
||||||
input_cost_per_token: 0.0
|
|
||||||
output_cost_per_token: 0.0
|
|
||||||
# For internal cost allocation, set symbolic rates:
|
|
||||||
# e.g., MoE = $2/M tokens, Dense = $1/M tokens, VLM = $0.50/M tokens
|
|
||||||
```
|
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
*Plan drafted: 2026-06-13 by Abiba 🦊⚡*
|
## Appendix B: LiteLLM Virtual Key Migration
|
||||||
*Updated: 2026-06-14 by Agent Zero (kagentz-bot) — CT 116 deployment target, DNS split-horizon, Authentik OIDC, Zero-Downtime strategy, GitOps workflow*
|
|
||||||
*Status: Ready for Phase 0 deployment*
|
**Current API Keys LiteLLM Virtual Keys:**
|
||||||
|
|
||||||
|
| Agent | Old Key | New LiteLLM Key | Tier | Budget |
|
||||||
|
|-------|---------|-----------------|------|--------|
|
||||||
|
| Abiba | sk-***-*** | sk-litellm-*** | enterprise | $1000 |
|
||||||
|
| Mumuni | sk-***-*** | sk-litellm-*** | enterprise | $1000 |
|
||||||
|
| Tanko | sk-***-*** | sk-litellm-*** | enterprise | $1000 |
|
||||||
|
| Kagenz0 | sk-***-*** | sk-litellm-*** | professional | $500 |
|
||||||
|
| Koby | sk-***-*** | sk-litellm-*** | professional | $500 |
|
||||||
|
| Koonimo | sk-***-*** | sk-litellm-*** | professional | $500 |
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Appendix C: Authentik SSO Integration
|
||||||
|
|
||||||
|
**Authentik Provider Setup:**
|
||||||
|
1. Create OAuth2 application in Authentik
|
||||||
|
2. Set redirect URI: `http://<CT-116-IP>/sso/callback`
|
||||||
|
3. Configure `client_id` and `client_secret`
|
||||||
|
4. Mount `custom_sso.py` to LiteLLM container
|
||||||
|
5. Update config.yaml with provider details
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Appendix D: Prometheus Monitoring
|
||||||
|
|
||||||
|
**Metrics Export:**
|
||||||
|
- LiteLLM metrics `http://127.0.0.1:4000/metrics`
|
||||||
|
- Router metrics `http://127.0.0.1:9000/metrics`
|
||||||
|
- GPU health metrics `http://127.0.0.1:9000/metrics/gpu`
|
||||||
|
- Circuit breaker metrics `http://127.0.0.1:9000/metrics/circuit-breaker`
|
||||||
|
|
||||||
|
**Alerts:**
|
||||||
|
- GPU health score > 70 alert to Hermes
|
||||||
|
- Circuit breaker trip alert to Hermes
|
||||||
|
- LiteLLM spend > $100/day alert to Hermes
|
||||||
|
- LiteLLM latency > 1000ms alert to Hermes
|
||||||
Reference in New Issue
Block a user