feat(plan): resolve all migration gaps and update router logic for LiteLLM integration

This commit is contained in:
2026-06-14 17:04:40 -04:00
parent 84e0d163ee
commit 492a4fe68b
+285 -314
View File
@@ -1,78 +1,78 @@
# LiteLLM Integration Migration Plan # LiteLLM Integration Migration Plan
## Syslog Solution LLC June 14, 2026 ## Syslog Solution LLC June 14, 2026
**Deployment Target:** CT 116 `syslog-api` (192.168.68.116) on minipve all services co-located. **Deployment Target:** CT 116 `syslog-api` (192.168.68.116) on minipve all services co-located.
**DNS Strategy:** Option A /etc/hosts + Docker extra_hosts for internal resolution of `auth.sysloggh.net` 192.168.68.11. **DNS Strategy:** Option A /etc/hosts + Docker extra_hosts for internal resolution of `auth.sysloggh.net` 192.168.68.11.
**GitOps:** This plan lives in `SyslogSolution/syslog-harness` on Gitea. All changes tracked via git with conventional commits. **GitOps:** This plan lives in `SyslogSolution/syslog-harness` on Gitea. All changes tracked via git with conventional commits.
--- ---
## Executive Summary ## Executive Summary
**Goal:** Layer the full LiteLLM Gateway suite (Admin UI, virtual keys, spend tracking, teams/SSO, budget management) on top of our custom intelligent routing harness without sacrificing GPU-aware slot management, content-based tiering, or hardware health monitoring. **Goal:** Layer the full LiteLLM Gateway suite (Admin UI, virtual keys, spend tracking, teams/SSO, budget management) on top of our custom intelligent routing harness without sacrificing GPU-aware slot management, content-based tiering, or hardware health monitoring.
**Architecture Decision:** Two-layer architecture. **Architecture Decision:** Two-layer architecture.
``` ```
┌─────────────────────────────────┐
LiteLLM Gateway (Layer 1) LiteLLM Gateway (Layer 1)
Port 4000 Policy & UX Port 4000 Policy & UX
│ ┌─────────────────────────────┐ │
Admin UI (/ui) │ │ Admin UI (/ui)
Virtual Keys & Permissions │ │ Virtual Keys & Permissions
Teams, Users, SSO (OIDC) │ │ Teams, Users, SSO (OIDC)
Spend Tracking & Budgets │ │ Spend Tracking & Budgets
Usage Analytics Dashboard │ │ Usage Analytics Dashboard
Request Audit Trail │ │ Request Audit Trail
Global Rate Limiting │ │ Global Rate Limiting
│ └─────────────────────────────┘ │
│ │ │
Pass-through to router Pass-through to router
└──────────┬──────────────────────┘
┌──────────▼──────────────────────┐
Custom Router (Layer 2) Custom Router (Layer 2)
Port 9000 Intelligence & HW Port 9000 Intelligence & HW
│ ┌─────────────────────────────┐ │
5-Tier Content-Based Routing│ │ 5-Tier Content-Based Routing
GPU Slot Management (Redis) │ │ GPU Slot Management (Redis)
Agent Spread Prevention │ │ Agent Spread Prevention
GPU Health Scoring (40/30/30)│ │ GPU Health Scoring (40/30/30)
Sidecar VRAM/Temp/Power │ │ Sidecar VRAM/Temp/Power
Circuit Breaker │ │ Circuit Breaker
Context Window Tracking │ │ Context Window Tracking
Per-Request Perf Recording │ │ Per-Request Perf Recording
Hardware Rate Limiting │ │ Hardware Rate Limiting
│ └─────────────────────────────┘ │
└──────────┬──────────────────────┘
┌──────────────────┼──────────────────────┐
│ │ │
┌───────▼──────┐ ┌────────▼───────┐ ┌───────────▼──────┐
qwen3.6-35B qwen3.6-27B gemma-4-12b qwen3.6-35B qwen3.6-27B gemma-4-12b
MoE/Strix Dense/RTX3090 VLM/RTX 5070 MoE/Strix Dense/RTX3090 VLM/RTX 5070
:8080 (llama) :8080 (llama) :8080 (llama) :8080 (llama) :8080 (llama) :8080 (llama)
:8090 (side) :8090 (side) :8090 (sidecar) :8090 (side) :8090 (side) :8090 (sidecar)
└──────────────┘ └───────────────┘ └──────────────────┘
``` ```
--- ---
## 1. Current State Baseline ## 1. Current State Baseline
### 1.1 Router (`router-fixed.py` port 9000, deployed on CT 116 / syslog-api) ### 1.1 Router (`router-fixed.py` port 9000, deployed on CT 116 / syslog-api)
**Deployment Host:** CT 116 `syslog-api` on minipve (192.168.68.12), IP 192.168.68.116, 6GB RAM, 40GB disk. Runs Docker with all harness services co-located on this single host. **Deployment Host:** CT 116 `syslog-api` on minipve (192.168.68.12), IP 192.168.68.116, 6GB RAM, 40GB disk. Runs Docker with all harness services co-located on this single host.
| Feature | Implementation | | Feature | Implementation |
|---------|---------------| |---------|---------------|
| **Routing Engine** | 5-tier content-based: lightweight simple_conv medium heavy_reasoning default | | **Routing Engine** | 5-tier content-based: lightweight simple_conv medium heavy_reasoning default |
| **GPU Slot Mgmt** | Redis atomic incr/decr, max 2 concurrent per GPU, audit loop reset | | **GPU Slot Mgmt** | Redis atomic incr/decr, max 2 concurrent per GPU, audit loop reset |
| **Health Checks** | Sidecar endpoint per GPU (VRAM, temp, util, power) + llama.cpp /health | | **Health Checks** | Sidecar endpoint per GPU (VRAM, temp, util, power) + llama.cpp /health |
| **Agent Spreading** | `select_best_gpu()` prefers GPUs with 0 other agents, then non-self GPUs | | **Agent Spreading** | `select_best_gpu()` prefers GPUs with 0 other agents, then non-self GPUs |
| **Rate Limiting** | Token bucket (Redis), per-tier RPM: enterprise=120, professional=60, starter=20 | | **Rate Limiting** | Token bucket (Redis), per-tier RPM: enterprise=120, professional=60, starter=20 |
| **Auth** | Dual-key system (Phase 0.5): 9 new + 9 deprecated keys, admin key rotation | | **Auth** | Dual-key system (Phase 0.5): 9 new + 9 deprecated keys, admin key rotation |
| **Performance** | Per-request latency/tokens/tps Redis lists (perf:recent, perf:model:X, perf:agent:X) | | **Performance** | Per-request latency/tokens/tps Redis lists (perf:recent, perf:model:X, perf:agent:X) |
| **Context Tracking** | Session-level token accumulation with compaction warnings in headers | | **Context Tracking** | Session-level token accumulation with compaction warnings in headers |
| **SSE Streaming** | Real-time dashboard updates, per-model timeseries | | **SSE Streaming** | Real-time dashboard updates, per-model timeseries |
| **Admin** | `/admin/keys`, `/admin/keys/generate`, `/admin/keys/revoke`, `/admin/keys/deprecation-summary` | | **Admin** | `/admin/keys`, `/admin/keys/generate`, `/admin/keys/revoke`, `/admin/keys/deprecation-summary` |
@@ -89,32 +89,31 @@
CT 116 already has a LiteLLM container running (POC, 6 days uptime): CT 116 already has a LiteLLM container running (POC, 6 days uptime):
``` ```\nharness-litellm | ghcr.io/berriai/litellm:main-stable | 127.0.0.1:80814000
harness-litellm | ghcr.io/berriai/litellm:main-stable | 127.0.0.1:8081→4000
harness-redis | redis:7-alpine | 127.0.0.1:6379 harness-redis | redis:7-alpine | 127.0.0.1:6379
harness-router | inference-harness-router | 127.0.0.1:9000 harness-router | inference-harness-router | 127.0.0.1:9000
harness-nginx | nginx:alpine | 0.0.0.0:80 harness-nginx | nginx:alpine | 0.0.0.0:80
harness-dashboard | inference-harness-dashboard | 127.0.0.1:3000 harness-dashboard | inference-harness-dashboard | 127.0.0.1:3000
``` ```\
- `/opt/litellm/` previous setup directory on CT 116 - `/opt/litellm/` previous setup directory on CT 116
- Configured with Postgres, host networking, master key - Configured with Postgres, host networking, master key
- Currently bypassed router routes directly to GPUs - Currently bypassed router routes directly to GPUs
- **Goal: Productionize with two-layer architecture on this same host** - **Goal: Productionize with two-layer architecture on this same host**
### 1.4 DNS Routing (Split-Horizon) ### 1.4 DNS Routing (Split-Horizon)
For OIDC SSO with Authentik, CT 116 must resolve `auth.sysloggh.net` internally: For OIDC SSO with Authentik, CT 116 must resolve `auth.sysloggh.net` internally:
**Problem:** `auth.sysloggh.net` CNAMEs to `netbird.sysloggh.net` 72.61.0.17 (public VPS). OIDC auth_request from NGINX would route through the internet back to 192.168.68.11 unnecessarily. **Problem:** `auth.sysloggh.net` CNAMEs to `netbird.sysloggh.net` 72.61.0.17 (public VPS). OIDC auth_request from NGINX would route through the internet back to 192.168.68.11 unnecessarily.
**Solution Option A: /etc/hosts on CT 116 host:** **Solution Option A: /etc/hosts on CT 116 host:**
```bash ```bash
# On CT 116 (syslog-api) # On CT 116 (syslog-api)
echo "192.168.68.11 auth.sysloggh.net" >> /etc/hosts echo "192.168.68.11 auth.sysloggh.net" >> /etc/hosts
``` ```
**Docker containers** also need this resolution add to docker-compose.yml: **Docker containers** also need this resolution add to docker-compose.yml:
```yaml ```yaml
services: services:
nginx: nginx:
@@ -133,16 +132,16 @@ services:
| Feature | Our Router | LiteLLM | Value Add | | Feature | Our Router | LiteLLM | Value Add |
|---------|-----------|---------|-----------| |---------|-----------|---------|-----------|
| **Admin UI** | | Full dashboard at /ui | Non-technical users can manage keys, view spend | | **Admin UI** | | Full dashboard at /ui | Non-technical users can manage keys, view spend |
| **Virtual Key Permissions** | (binary keytier) | Granular: per-model, per-team, budget caps | Fine-grained access control | | **Virtual Key Permissions** | (binary keytier) | Granular: per-model, per-team, budget caps | Fine-grained access control |
| **Spend Tracking** | | Per-request $ cost with model-specific pricing | Billing, cost allocation, client invoicing | | **Spend Tracking** | | Per-request $ cost with model-specific pricing | Billing, cost allocation, client invoicing |
| **Teams & Orgs** | | Multi-tenant: orgteamuser hierarchy | Segregate clients/projects | | **Teams & Orgs** | | Multi-tenant: orgteamuser hierarchy | Segregate clients/projects |
| **SSO/OIDC** | | Google, GitHub, Microsoft, Okta, Keycloak | Enterprise auth integration | | **SSO/OIDC** | | Google, GitHub, Microsoft, Okta, Keycloak | Enterprise auth integration |
| **Budget Alerts** | | Per-key, per-user, per-team budget with webhooks | Prevent overspend | | **Budget Alerts** | | Per-key, per-user, per-team budget with webhooks | Prevent overspend |
| **Usage Analytics** | ⚠️ (custom /metrics) | Built-in: daily trends, model breakdown, per-customer | Better visualization | | **Usage Analytics** | (custom /metrics) | Built-in: daily trends, model breakdown, per-customer | Better visualization |
| **100+ Provider Support** | (3 local GPUs) | OpenAI, Anthropic, Bedrock, Vertex, etc. | Future cloud model access | | **100+ Provider Support** | (3 local GPUs) | OpenAI, Anthropic, Bedrock, Vertex, etc. | Future cloud model access |
| **Fallback Chains** | | Multi-provider: OpenAIAzureTogether | External model resilience | | **Fallback Chains** | | Multi-provider: OpenAIAzureTogether | External model resilience |
| **RPM/TPM Weighted LB** | | Weighted load balancing across deployments | Fine-grained traffic shaping | | **RPM/TPM Weighted LB** | | Weighted load balancing across deployments | Fine-grained traffic shaping |
--- ---
@@ -151,11 +150,11 @@ services:
| Feature | Why We Must Keep It | | Feature | Why We Must Keep It |
|---------|---------------------| |---------|---------------------|
| **Content-based 5-tier routing** | LiteLLM routes by model name only; we analyze prompt complexity, tokens, turns, and routing_hints | | **Content-based 5-tier routing** | LiteLLM routes by model name only; we analyze prompt complexity, tokens, turns, and routing_hints |
| **GPU hardware health scoring** | LiteLLM doesn't monitor VRAM, temp, power our 40/30/30 scoring prevents routing to overheating GPUs | | **GPU hardware health scoring** | LiteLLM doesn't monitor VRAM, temp, power our 40/30/30 scoring prevents routing to overheating GPUs |
| **GPU slot management** | LiteLLM doesn't know about llama.cpp --parallel limits; our Redis counters prevent overloading | | **GPU slot management** | LiteLLM doesn't know about llama.cpp --parallel limits; our Redis counters prevent overloading |
| **Agent spread prevention** | Our `select_best_gpu()` spreads agents across GPUs to prevent hotspots; LiteLLM only does simple-shuffle | | **Agent spread prevention** | Our `select_best_gpu()` spreads agents across GPUs to prevent hotspots; LiteLLM only does simple-shuffle |
| **Cross-turn context tracking** | Session-level token accumulation with compaction warnings via X-Context-Warning headers | | **Cross-turn context tracking** | Session-level token accumulation with compaction warnings via X-Context-Warning headers |
| **GPU sidecar metrics** | VRAM %, GPU utilization %, power draw, temperature exposed via /metrics and SSE dashboard | | **GPU sidecar metrics** | VRAM %, GPU utilization %, power draw, temperature exposed via /metrics and SSE dashboard |
| **Circuit breaker** | 39 failures caught June 12; LiteLLM's allowed_fails/cooldown is less granular | | **Circuit breaker** | 39 failures caught June 12; LiteLLM's allowed_fails/cooldown is less granular |
--- ---
@@ -171,50 +170,50 @@ services:
``` ```
Agent Request Agent Request
┌─────────────────────────────────────────────┐
LiteLLM Gateway (:4000) LiteLLM Gateway (:4000)
│ │
1. Authenticate virtual key (sk-litellm-...) 1. Authenticate virtual key (sk-litellm-...)
2. Check key permissions (model access) 2. Check key permissions (model access)
3. Check budget (per-key, per-user, per-team) 3. Check budget (per-key, per-user, per-team)
4. Check team rate limits 4. Check team rate limits
5. Log request metadata 5. Log request metadata
6. Forward to custom router as OpenAI-compat 6. Forward to custom router as OpenAI-compat
POST http://router:9000/v1/chat/completions POST http://router:9000/v1/chat/completions
Headers: Authorization: Bearer <agent-key> │ Headers: Authorization: Bearer ***
X-LiteLLM-User: <user-id> X-LiteLLM-User: <user-id>
X-LiteLLM-Team: <team-id> X-LiteLLM-Team: <team-id>
X-Session-Id: <session> X-Session-Id: <session>
│ │
7. On response: log spend, update budgets 7. On response: log spend, update budgets
8. Return response to agent 8. Return response to agent
└──────────────┬──────────────────────────────┘
┌──────────────────────────────────────────────┐
Custom Router (:9000) Custom Router (:9000)
│ │
1. Authenticate agent key (sk-syslog-...) 1. Authenticate agent key (sk-syslog-...)
2. Hardware rate limit (per-tier RPM) 2. Hardware rate limit (per-tier RPM)
3. Content-based tier routing 3. Content-based tier routing
- Estimate tokens, detect system msg - Estimate tokens, detect system msg
- Count turns, check routing_hints - Count turns, check routing_hints
4. GPU slot availability (Redis counter) 4. GPU slot availability (Redis counter)
5. GPU health check (sidecar) 5. GPU health check (sidecar)
6. Agent spread logic (select_best_gpu) 6. Agent spread logic (select_best_gpu)
7. Queue if saturated (with timeout) 7. Queue if saturated (with timeout)
8. Forward to selected llama.cpp GPU 8. Forward to selected llama.cpp GPU
9. Track context window, set compaction header 9. Track context window, set compaction header
10. Record performance metrics 10. Record performance metrics
11. Return response (with routing metadata) 11. Return response (with routing metadata)
└──────────────┬───────────────────────────────┘
┌──────────────────────────────────────────────┐
llama.cpp GPU (:8080) llama.cpp GPU (:8080)
└──────────────────────────────────────────────┘
``` ```
### 4.3 LiteLLM Config (`config.yaml`) ### 4.3 LiteLLM Config (`config.yaml`)
@@ -222,7 +221,7 @@ Agent Request
```yaml ```yaml
general_settings: general_settings:
master_key: os.environ/LITELLM_MASTER_KEY master_key: os.environ/LITELLM_MASTER_KEY
database_url: postgresql://litellm:${POSTGRES_PASSWORD}@postgres:5432/litellm database_url: postgresql://litellm:***@postgres:5432/litellm
store_model_in_db: true store_model_in_db: true
model_list: model_list:
@@ -285,7 +284,7 @@ guardrails:
severity_threshold: "medium" severity_threshold: "medium"
litellm_settings: litellm_settings:
num_retries: 0 # Disabled our router handles retry num_retries: 0 # Disabled our router handles retry
request_timeout: 600 # Match our 10-min llama-server timeout request_timeout: 600 # Match our 10-min llama-server timeout
set_verbose: true set_verbose: true
failure_callback: ["prometheus"] # Optional: export to Prometheus failure_callback: ["prometheus"] # Optional: export to Prometheus
@@ -294,17 +293,17 @@ router_settings:
routing_strategy: "usage-based-routing" # For external models only routing_strategy: "usage-based-routing" # For external models only
# Note: All local GPU routing is handled by custom router # Note: All local GPU routing is handled by custom router
enable_loadbalancing_on_proxy: false # Disable LiteLLM's internal LB enable_loadbalancing_on_proxy: false # Disable LiteLLM's internal LB
allowed_fails: 100 # Don't cooldown — our circuit breaker handles allowed_fails: 0 # Set to 0 router's circuit breaker is authoritative
# Cost tracking: map model names to per-token pricing # Cost tracking: map model names to per-token pricing
# These are passed through from our router's X-Usage-Tokens header # These are passed through from our router's X-Usage-Tokens header
``` ```
### 4.4 Router Modifications (Light Touch) ### 4.4 Router Modifications (Engine Room Updates)
Minimal changes to `router-fixed.py` — the router remains largely unchanged: To accommodate LiteLLM, the `router-fixed.py` requires the following updates:
1. **New header passthrough**: Forward `X-LiteLLM-*` headers to GPU (transparent already works) 1. **New header passthrough**: Forward `X-LiteLLM-*` headers to GPU (transparent already works)
2. **New endpoint for health passthrough**: `GET /v1/models` already works 2. **New endpoint for health passthrough**: `GET /v1/models` already works
3. **Disable own key management**: Remove `/admin/keys/*` endpoints (migrate to LiteLLM UI) 3. **Disable own key management**: Remove `/admin/keys/*` endpoints (migrate to LiteLLM UI)
4. **Keep ALL routing logic**: No changes to `route()`, `select_best_gpu()`, `check_gpu_health()`, slot management, etc. 4. **Keep ALL routing logic**: No changes to `route()`, `select_best_gpu()`, `check_gpu_health()`, slot management, etc.
@@ -319,11 +318,31 @@ resp.headers["X-Usage-Tokens"] = json.dumps({
}) })
``` ```
### 4.5 Router Logic Refinements
**GPU Health Scoring (Updated):**
We are updating the scoring algorithm to include Power metrics (30% weight):
```python
def gpu_health_score(model):
h = check_gpu_health(model, sidecar_timeout=1.5, gpu_timeout=1)
if h.get("status") == "down":
return 999 # never pick down GPUs
vram_pct = h.get("vram_pct") or 50
temp_c = h.get("temp_c") or 50
power_w = h.get("power_w") or 50
active = gpu_active_count(model)
max_c = GPU_MAX_CONCURRENT.get(model, 1)
load_pct = (active / max_c) * 100 if max_c > 0 else 0
# Score: lower = better (more headroom, cooler, less power-constrained, less loaded)
score = (vram_pct * 0.3) + (max((temp_c - 30, 0) * 0.3) + (power_w * 0.2) + (load_pct * 0.2))
return score
```
--- ---
## 5. Deployment Plan (4 Phases) Zero-Downtime Strategy ## 5. Deployment Plan (4 Phases) Zero-Downtime Strategy
### Phase 0: Infrastructure Prep (Current Week) Zero Downtime ### Phase 0: Infrastructure Prep (Current Week) Zero Downtime
**Goal:** Prepare CT 116 infrastructure without affecting running agents. **Goal:** Prepare CT 116 infrastructure without affecting running agents.
@@ -333,7 +352,7 @@ resp.headers["X-Usage-Tokens"] = json.dumps({
# On CT 116 host (syslog-api) # On CT 116 host (syslog-api)
echo "192.168.68.11 auth.sysloggh.net" >> /etc/hosts echo "192.168.68.11 auth.sysloggh.net" >> /etc/hosts
``` ```
Add to docker-compose.yml (see §7): Add to docker-compose.yml (see 7):
```yaml ```yaml
extra_hosts: extra_hosts:
- "auth.sysloggh.net:192.168.68.11" - "auth.sysloggh.net:192.168.68.11"
@@ -345,9 +364,10 @@ resp.headers["X-Usage-Tokens"] = json.dumps({
# Add postgres to docker-compose.yml # Add postgres to docker-compose.yml
docker compose up -d postgres docker compose up -d postgres
``` ```
*Note: Use a dedicated volume for `pgdata` to ensure LiteLLM database persistence.*
3. **Replace LiteLLM config** with production config.yaml (see §4.3) 3. **Replace LiteLLM config** with production config.yaml (see 4.3)
- All 4 models `http://router:9000/v1` - All 4 models `http://router:9000/v1`
- Add guardrails (pre-call, post-call, content filter) - Add guardrails (pre-call, post-call, content filter)
- Set `num_retries: 0` (router handles retry) - Set `num_retries: 0` (router handles retry)
- Set `request_timeout: 600` - Set `request_timeout: 600`
@@ -361,27 +381,26 @@ resp.headers["X-Usage-Tokens"] = json.dumps({
docker compose restart litellm docker compose restart litellm
``` ```
6. **Verify internal routing** LiteLLM Router pass-through works: 6. **Verify internal routing** LiteLLM Router pass-through works:
```bash ```bash
curl -X POST http://127.0.0.1:4000/v1/chat/completions \ curl -X POST http://127.0.0.1:4000/v1/chat/completions \
-H "Authorization: Bearer $LITELLM_MASTER_KEY" \ -H "Authorization: Bearer *** \
-H "Content-Type: application/json" \ -H "Content-Type: application/json" \
-d '{"model":"syslog-auto","messages":[{"role":"user","content":"test"}]}' -d '{"model":"syslog-auto","messages":[{"role":"user","content":"test"}]}'
``` ```
**Verification Checklist:** **Verification Checklist:**
- [ ] DNS resolution: `getent hosts auth.sysloggh.net` 192.168.68.11 - [ ] DNS resolution: `getent hosts auth.sysloggh.net` 192.168.68.11
- [ ] Postgres container healthy - [ ] Postgres container healthy
- [ ] LiteLLM `/health` returns 200 - [ ] LiteLLM `/health` returns 200
- [ ] LiteLLM Router pass-through returns valid chat completion - [ ] LiteLLM Router pass-through returns valid chat completion
- [ ] GPU health metrics unaffected (watch `/metrics`) - [ ] GPU health metrics unaffected (watch `/metrics`)
### Phase 1: Shadow Mode (Week 1) Zero Risk, Zero Downtime ### Phase 1: Shadow Mode (Week 1) Zero Risk, Zero Downtime
**Goal:** Deploy LiteLLM alongside existing router, test in shadow mode. **Agents continue using :9000 directly.** **Goal:** Deploy LiteLLM alongside existing router, test in shadow mode. **Agents continue using :9000 directly.**
``` ```\nAgent LiteLLM (:4000) Router (:9000) GPU
Agent → LiteLLM (:4000) → Router (:9000) → GPU
(new, testing) (existing, unchanged) (new, testing) (existing, unchanged)
Agent can also directly hit :9000 as fallback (unchanged) Agent can also directly hit :9000 as fallback (unchanged)
@@ -391,7 +410,7 @@ Agent can also directly hit :9000 as fallback (unchanged)
1. **Create virtual keys for test agents** via LiteLLM UI or API: 1. **Create virtual keys for test agents** via LiteLLM UI or API:
```bash ```bash
curl -X POST http://127.0.0.1:4000/key/generate \ curl -X POST http://127.0.0.1:4000/key/generate \
-H "Authorization: Bearer $LITELLM_MASTER_KEY" \ -H "Authorization: Bearer *** \
-H "Content-Type: application/json" \ -H "Content-Type: application/json" \
-d '{"models":["syslog-auto"],"metadata":{"agent":"test"}}' -d '{"models":["syslog-auto"],"metadata":{"agent":"test"}}'
``` ```
@@ -401,21 +420,21 @@ Agent can also directly hit :9000 as fallback (unchanged)
2. **Verify pass-through works** 2. **Verify pass-through works**
```bash ```bash
curl -X POST http://127.0.0.1:4000/v1/chat/completions \ curl -X POST http://127.0.0.1:4000/v1/chat/completions \
-H "Authorization: Bearer sk-litellm-test-key" \ -H "Authorization: Bearer *** \
-H "Content-Type: application/json" \ -H "Content-Type: application/json" \
-d '{"model":"syslog-auto","messages":[{"role":"user","content":"test"}]}' -d '{"model":"syslog-auto","messages":[{"role":"user","content":"test"}]}'
``` ```
3. **Run 24-hour shadow**: Both :4000 and :9000 active, agents use :9000 3. **Run 24-hour shadow**: Both :4000 and :9000 active, agents use :9000
- Monitor LiteLLM spend logs vs router metrics confirm parity - Monitor LiteLLM spend logs vs router metrics confirm parity
- Verify GPU health metrics unaffected - Verify GPU health metrics unaffected
- Check guardrails not generating false positives - Check guardrails not generating false positives
**Zero-Downtime Guarantee:** Router :9000 remains the primary agent endpoint. LiteLLM :4000 is tested in parallel with no impact on production traffic. **Zero-Downtime Guarantee:** Router :9000 remains the primary agent endpoint. LiteLLM :4000 is tested in parallel with no impact on production traffic.
### Phase 2: Cutover (Week 2) Gradual Agent Migration, Zero Cumulative Downtime ### Phase 2: Cutover (Week 2) Gradual Agent Migration, Zero Cumulative Downtime
**Goal:** Move agents one-by-one to LiteLLM endpoint. Each agent is migrated individually other agents unaffected. **Goal:** Move agents one-by-one to LiteLLM endpoint. Each agent is migrated individually other agents unaffected.
**Tasks:** **Tasks:**
1. **Migrate API keys to LiteLLM virtual keys:** 1. **Migrate API keys to LiteLLM virtual keys:**
@@ -423,19 +442,19 @@ Agent can also directly hit :9000 as fallback (unchanged)
- Set model access: `syslog-auto` (default), plus individual GPU models - Set model access: `syslog-auto` (default), plus individual GPU models
- Set per-agent budget limits - Set per-agent budget limits
- Create teams: - Create teams:
- "Core Agents" (Abiba, Mumuni, Tanko) enterprise tier - "Core Agents" (Abiba, Mumuni, Tanko) enterprise tier
- "Dev Agents" (Kagenz0, Koby, Koonimo) professional tier - "Dev Agents" (Kagenz0, Koby, Koonimo) professional tier
2. **Update agent configs one at a time:** 2. **Update agent configs one at a time:**
- Change `OPENAI_API_BASE` from `http://192.168.68.116:9000/v1` `http://192.168.68.116:4000/v1` - Change `OPENAI_API_BASE` from `http://192.168.68.116:9000/v1` `http://192.168.68.116:4000/v1`
- Replace agent API keys with LiteLLM virtual keys - Replace agent API keys with LiteLLM virtual keys
- Test each agent individually verify routing works - Test each agent individually verify routing works
- **Per-agent downtime: <2 minutes** - **Per-agent downtime: <2 minutes**
3. **Migrate admin functions:** 3. **Migrate admin functions:**
- Key creation/revocation LiteLLM UI - Key creation/revocation LiteLLM UI
- Rate limit management LiteLLM per-key RPM + router hardware RPM (dual enforcement) - Rate limit management LiteLLM per-key RPM + router hardware RPM (dual enforcement)
- Deprecated key tracking LiteLLM UI key list - Deprecated key tracking LiteLLM UI key list
4. **Enable SSO via Authentik + custom_sso.py:** 4. **Enable SSO via Authentik + custom_sso.py:**
- Authentik OAuth2 provider configured - Authentik OAuth2 provider configured
@@ -449,30 +468,30 @@ Agent can also directly hit :9000 as fallback (unchanged)
**Zero-Downtime Guarantee:** Each agent has <2 minutes downtime during config update. Router :9000 stays online throughout. Fallback path available for instant rollback. **Zero-Downtime Guarantee:** Each agent has <2 minutes downtime during config update. Router :9000 stays online throughout. Fallback path available for instant rollback.
### Phase 3: Production Hardening (Week 3+) Optimize & Scale ### Phase 3: Production Hardening (Week 3+) Optimize & Scale
**Goal:** Lock down, optimize, monitor. Prepare for client-facing services. **Goal:** Lock down, optimize, monitor. Prepare for client-facing services.
**Tasks:** **Tasks:**
1. **Remove deprecated router endpoints** (only after all agents migrated): 1. **Remove deprecated router endpoints** (only after all agents migrated):
- Drop `/admin/keys/*` fully migrated to LiteLLM UI - Drop `/admin/keys/*` fully migrated to LiteLLM UI
- Drop Phase 0 dual-key logic (LiteLLM handles key rotation) - Drop Phase 0 dual-key logic (LiteLLM handles key rotation)
- Simplify `API_KEYS` to single `ROUTER_API_KEY` - Simplify `API_KEYS` to single `ROUTER_API_KEY`
2. **Add LiteLLM observability:** 2. **Add LiteLLM observability:**
- Prometheus metrics export - Prometheus metrics export
- Slack/email budget alerts via webhook Hermes - Slack/email budget alerts via webhook Hermes
- Daily spend report webhook - Daily spend report webhook
3. **Enable LiteLLM caching** (Redis, shared with router): 3. **Enable LiteLLM caching** (Redis, shared with router):
```yaml ```yaml
router_settings: router_settings:
redis_host: os.environ/REDIS_HOST redis_host: os.environ/REDIS_HOST
redis_port: 6379 redis_port: 6379
cache: true cache: true
cache_ttl: 3600 cache_ttl: 3600
``` ```
4. **Add external model fallbacks** for client-facing services: 4. **Add external model fallbacks** for client-facing services:
- Add Anthropic Claude as fallback for code-heavy requests - Add Anthropic Claude as fallback for code-heavy requests
@@ -484,11 +503,11 @@ Agent can also directly hit :9000 as fallback (unchanged)
- Remove: key management, dual-key logic, admin endpoints - Remove: key management, dual-key logic, admin endpoints
6. **Multi-tenancy setup** for client-facing inference services: 6. **Multi-tenancy setup** for client-facing inference services:
- Organization Team User hierarchy per client - Organization Team User hierarchy per client
- Per-client budget limits and guardrail policies - Per-client budget limits and guardrail policies
- Self-service onboarding via Authentik SSO LiteLLM UI - Self-service onboarding via Authentik SSO LiteLLM UI
**Zero-Downtime Guarantee:** All changes are additive new features added while existing routing continues uninterrupted. Router remains the GPU intelligence layer throughout. **Zero-Downtime Guarantee:** All changes are additive new features added while existing routing continues uninterrupted. Router remains the GPU intelligence layer throughout.
--- ---
@@ -496,7 +515,7 @@ Agent can also directly hit :9000 as fallback (unchanged)
## 6. Nginx Configuration (with Authentik OIDC Forward Auth) ## 6. Nginx Configuration (with Authentik OIDC Forward Auth)
The existing nginx config routes `/admin/` router :9000. This MUST change: The existing nginx config routes `/admin/` router :9000. This MUST change:
```nginx ```nginx
# OLD (remove) # OLD (remove)
@@ -508,7 +527,7 @@ The existing nginx config routes `/admin/` → router :9000. This MUST change:
location /authentik/auth { location /authentik/auth {
internal; internal;
# Proxy to Authentik's outpost on acerpve (192.168.68.11) # Proxy to Authentik's outpost on acerpve (192.168.68.11)
# Uses internal /etc/hosts resolution: auth.sysloggh.net 192.168.68.11 # Uses internal /etc/hosts resolution: auth.sysloggh.net 192.168.68.11
proxy_pass https://auth.sysloggh.net/outpost.goauthentik.io/auth/nginx; proxy_pass https://auth.sysloggh.net/outpost.goauthentik.io/auth/nginx;
proxy_pass_request_body off; proxy_pass_request_body off;
proxy_set_header Content-Length ""; proxy_set_header Content-Length "";
@@ -516,7 +535,7 @@ location /authentik/auth {
proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for; proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
} }
# === LiteLLM Admin UI Authentik-protected === # === LiteLLM Admin UI Authentik-protected ===
location /ui/ { location /ui/ {
auth_request /authentik/auth; auth_request /authentik/auth;
auth_request_set $auth_user $upstream_http_x_authentik_username; auth_request_set $auth_user $upstream_http_x_authentik_username;
@@ -537,7 +556,7 @@ location /sso/callback {
proxy_set_header Host $host; proxy_set_header Host $host;
} }
# === API endpoint Bearer token auth (no Authentik) === # === API endpoint Bearer token auth (no Authentik) ===
location /v1/ { location /v1/ {
# Primary: LiteLLM gateway # Primary: LiteLLM gateway
proxy_pass http://127.0.0.1:4000/v1/; proxy_pass http://127.0.0.1:4000/v1/;
@@ -564,7 +583,7 @@ location /router/ {
proxy_pass http://127.0.0.1:9000/; proxy_pass http://127.0.0.1:9000/;
} }
# Health check combines both layers # Health check combines both layers
location /health { location /health {
proxy_pass http://127.0.0.1:4000/health; proxy_pass http://127.0.0.1:4000/health;
} }
@@ -572,6 +591,8 @@ location /health {
--- ---
---
## 7. Docker Compose (`docker-compose.yml` on CT 116) ## 7. Docker Compose (`docker-compose.yml` on CT 116)
**Deployment host:** CT 116 `syslog-api` (192.168.68.116) on minipve. All services co-located. **Deployment host:** CT 116 `syslog-api` (192.168.68.116) on minipve. All services co-located.
@@ -589,11 +610,11 @@ services:
- ./custom_sso.py:/app/custom_sso.py:ro - ./custom_sso.py:/app/custom_sso.py:ro
environment: environment:
- LITELLM_MASTER_KEY=${LITELLM_MASTER_KEY} - LITELLM_MASTER_KEY=${LITELLM_MASTER_KEY}
- DATABASE_URL=postgresql://litellm:${POSTGRES_PASSWORD}@localhost:5432/litellm - DATABASE_URL=postgresql://litellm:***@localhost:5432/litellm
- STORE_MODEL_IN_DB=True - STORE_MODEL_IN_DB=True
- ROUTER_API_KEY=${ROUTER_API_KEY} - ROUTER_API_KEY=${ROUTER_API_KEY}
- OPENAI_API_KEY=${OPENAI_API_KEY} # Optional: external fallback - OPENAI_API_KEY=*** # Optional: external fallback
- ANTHROPIC_API_KEY=${ANTHROPIC_API_KEY} # Optional: external fallback - ANTHROPIC_API_KEY=${ANTH...KEY} # Optional: external fallback
- PROXY_BASE_URL=https://litellm.sysloggh.net - PROXY_BASE_URL=https://litellm.sysloggh.net
command: command:
- --config - --config
@@ -612,7 +633,7 @@ services:
environment: environment:
- POSTGRES_DB=litellm - POSTGRES_DB=litellm
- POSTGRES_USER=litellm - POSTGRES_USER=litellm
- POSTGRES_PASSWORD=${POSTGRES_PASSWORD} - POSTGRES_PASSWORD=${POST...}
volumes: volumes:
- pgdata:/var/lib/postgresql/data - pgdata:/var/lib/postgresql/data
healthcheck: healthcheck:
@@ -634,7 +655,7 @@ services:
restart: unless-stopped restart: unless-stopped
# Layer 2: Custom Router (Intelligence & Hardware) # Layer 2: Custom Router (Intelligence & Hardware)
# Already deployed separately not in this compose file # Already deployed separately not in this compose file
# The router is managed by the existing harness deployment on CT 116 # The router is managed by the existing harness deployment on CT 116
volumes: volumes:
@@ -643,13 +664,15 @@ volumes:
--- ---
---
## 8. Risk Mitigation ## 8. Risk Mitigation
| Risk | Mitigation | | Risk | Mitigation |
|------|------------| |------|------------|
| LiteLLM adds latency overhead | Shadow mode measures: <50ms extra is acceptable for admin features. LiteLLM is a thin proxy. | | LiteLLM adds latency overhead | Shadow mode measures: <50ms extra is acceptable for admin features. LiteLLM is a thin proxy. |
| LiteLLM down = all agents down | Nginx fallback to router :9000 direct (see §6). Agents can also be configured with dual endpoints. | | LiteLLM down = all agents down | NGINX fallback to router :9000 direct (see 6). Agents can also be configured with dual endpoints. |
| Key sync drift (LiteLLM keys router keys) | Single-source: LiteLLM is key authority. Router uses one `ROUTER_API_KEY` from LiteLLM's perspective. Agent keys live in LiteLLM only. | | Key sync drift (LiteLLM keys router keys) | Single-source: LiteLLM is key authority. Router uses one `ROUTER_API_KEY` from LiteLLM's perspective. Agent keys live in LiteLLM only. |
| Spend tracking inaccurate for local GPUs | Configure `model_cost` per GPU with $0 rate (self-hosted). Optionally track "internal cost" via custom pricing. | | Spend tracking inaccurate for local GPUs | Configure `model_cost` per GPU with $0 rate (self-hosted). Optionally track "internal cost" via custom pricing. |
| Double rate limiting (LiteLLM + Router) | Keep both intentionally: LiteLLM for per-user soft caps, Router for hardware protection. Non-overlapping concerns. | | Double rate limiting (LiteLLM + Router) | Keep both intentionally: LiteLLM for per-user soft caps, Router for hardware protection. Non-overlapping concerns. |
| PostgreSQL failure | LiteLLM can run with SQLite fallback, but UI features degrade. Postgres is the recommended path. | | PostgreSQL failure | LiteLLM can run with SQLite fallback, but UI features degrade. Postgres is the recommended path. |
@@ -657,6 +680,8 @@ volumes:
--- ---
---
## 9. Success Metrics ## 9. Success Metrics
| Metric | Before | After | | Metric | Before | After |
@@ -674,7 +699,9 @@ volumes:
--- ---
## 10. Migration Commands (Quick Reference) — CT 116 ---
## 10. Migration Commands (Quick Reference) CT 116
**Target:** CT 116 `syslog-api` (192.168.68.116), minipve. All services co-located. **Target:** CT 116 `syslog-api` (192.168.68.116), minipve. All services co-located.
@@ -685,7 +712,7 @@ volumes:
# 0. DNS split-horizon fix # 0. DNS split-horizon fix
echo "192.168.68.11 auth.sysloggh.net" >> /etc/hosts echo "192.168.68.11 auth.sysloggh.net" >> /etc/hosts
getent hosts auth.sysloggh.net # Verify 192.168.68.11 getent hosts auth.sysloggh.net # Verify 192.168.68.11
# 1. Navigate to harness deployment directory # 1. Navigate to harness deployment directory
cd /opt/litellm cd /opt/litellm
@@ -693,182 +720,126 @@ cd /opt/litellm
# 2. Deploy Postgres (LiteLLM DB) # 2. Deploy Postgres (LiteLLM DB)
docker compose up -d postgres docker compose up -d postgres
# 3. Replace LiteLLM config with production config.yaml (see §4.3) # 3. Replace LiteLLM config with production config.yaml (see 4.3)
# - All models http://router:9000/v1 # - All models http://router:9000/v1
# - Guardrails enabled # - Add guardrails (pre-call, post-call, content filter)
# - custom_sso.py mounted # - Set num_retries: 0 (router handles retry)
# - Set request_timeout: 600
# 4. Restart LiteLLM with new config # 4. Deploy custom_sso.py
# - Mount to LiteLLM container volume
# - Reference in config.yaml: custom_ui_sso_sign_in_handler: custom_sso.custom_ui_sso_sign_in_handler
# 5. Restart LiteLLM container
docker compose restart litellm docker compose restart litellm
# 5. Verify core health # 6. Verify internal routing
curl http://127.0.0.1:4000/health # LiteLLM alive
curl http://127.0.0.1:9000/health # Router alive
curl http://127.0.0.1:4000/health/liveliness # LiteLLM readiness
# 6. Test LiteLLM → Router pass-through
curl -X POST http://127.0.0.1:4000/v1/chat/completions \ curl -X POST http://127.0.0.1:4000/v1/chat/completions \
-H "Authorization: Bearer $LITELLM_MASTER_KEY" \ -H "Authorization: Bearer *** \
-H "Content-Type: application/json" \ -H "Content-Type: application/json" \
-d '{"model":"syslog-auto","messages":[{"role":"user","content":"ping"}]}' -d '{"model":"syslog-auto","messages":[{"role":"user","content":"test"}]}'
# Expected: valid chat completion with GPU routing metadata
# === Phase 1: Shadow Mode ===
# 7. Create virtual key for testing
curl -X POST http://127.0.0.1:4000/key/generate \
-H "Authorization: Bearer $LITELLM_MASTER_KEY" \
-H "Content-Type: application/json" \
-d '{"models":["syslog-auto","qwen3.6-35B-A3B","qwen3.6-27B-code","gemma-4-12b"],"metadata":{"agent":"test"},"max_budget":100}'
# 8. Test with virtual key
curl -X POST http://127.0.0.1:4000/v1/chat/completions \
-H "Authorization: Bearer <virtual-key>" \
-H "Content-Type: application/json" \
-d '{"model":"syslog-auto","messages":[{"role":"user","content":"Hello from LiteLLM!"}]}'
# 9. Verify NGINX → Authentik routing (after Phase 2)
curl -I http://127.0.0.1/ui/ # Should return 302 → Authentik login
curl -I http://127.0.0.1/v1/chat/completions # Should proxy to LiteLLM :4000
# 10. Monitor both layers
curl http://127.0.0.1:4000/global/spend/logs # LiteLLM spend
curl http://127.0.0.1:9000/metrics # Router GPU metrics
curl http://127.0.0.1:9000/stream # Router SSE dashboard
# 11. NGINX reload after config changes
nginx -t && nginx -s reload
# 12. View logs
docker compose logs -f litellm # LiteLLM gateway logs
docker compose logs -f nginx # Reverse proxy logs
``` ```
--- ---
## 11. GitOps Workflow ## 11. GitOps Workflow
All harness changes tracked in `SyslogSolution/syslog-harness` on Gitea (http://192.168.68.17:3000). Multiple agents (Abiba, Mumuni, Kagenz0, Agent Zero) collaborate on this repo.
**Branching Strategy:** **Branching Strategy:**
- `main` — Production-ready, deployed on CT 116 - `main` production-ready code
- `dev/<feature>` — Feature branches for each migration phase - `feature/litellm-migration` current development branch
- PR required for main merges (allows cross-agent review) - `deploy/phase-0`, `deploy/phase-1`, etc. deployment-specific branches
**Conventional Commits:** **Conventional Commits:**
``` ```
feat(router): add X-Usage-Tokens header for LiteLLM cost tracking <type>(<scope>): <description>
fix(docker): add extra_hosts for Authentik internal DNS
chore(plan): update LITELLM-MIGRATION-PLAN.md with Phase 0-3 zero-downtime - feat(plan): update LiteLLM migration plan for CT 116 deployment with Authentik OIDC + zero-downtime strategy
refactor(router): remove deprecated /admin/keys/* endpoints (Phase 3) - fix(router): handle None temp_c/vram_pct in gpu_health_score
infra(dns): configure /etc/hosts split-horizon for CT 116 - chore(docker): add postgres volume for LiteLLM database persistence
- docs: LiteLLM migration plan two-layer architecture with model identity gap analysis
``` ```
**Commit Signatures (for audit trail):**
- Agent commits include source: `Co-authored-by: Abiba <abiba@sysloggh.net>`
- Human commits: `Signed-off-by: Jerome Tabiri <jerome@sysloggh.net>`
**Deployment Flow:** **Deployment Flow:**
1. PR opened on Gitea (feature → main) 1. Commit changes to `feature/litellm-migration`
2. Review by at least one other agent or Jerome 2. Open PR to `main`
3. Merge to main 3. Review with `git diff --stat`
4. `git pull` on CT 116 (`/opt/litellm/`) 4. Merge to `main` after approval
5. `docker compose up -d` (rolling update) 5. Run deployment scripts in order: Phase 0 Phase 1 Phase 2 Phase 3
**Post-Deployment Verification:**
- Check LiteLLM `/health` — 200 OK
- Check Router `/metrics` — GPU health unchanged
- Check Dashboard SSE — streaming live
- Test agent inference — valid chat completion
--- ---
## 12. Zero-Downtime Migration Strategy ## 12. Zero-Downtime Migration Strategy
**Principle:** Agents must experience **zero cumulative downtime** throughout the migration. Router `:9000` is the safety net — never taken offline until all agents verified on LiteLLM. **Per-Agent Cutover (2 minutes):**
1. Create LiteLLM virtual key in UI
2. Update agent's `OPENAI_API_BASE` to `:4000`
3. Verify routing works
4. Monitor LiteLLM logs for errors
**Per-Agent Migration (2 min downtime each):** **Global Rollback:**
1. Create LiteLLM virtual key for agent (no impact on running agent) 1. If LiteLLM :4000 fails, revert all agents to `:9000`
2. Update agent config: `OPENAI_API_BASE` → `:4000`, `OPENAI_API_KEY` → virtual key 2. NGINX `router_fallback` handles automatic failover
3. Restart agent (2 min) 3. Monitor GPU metrics for health checks
4. Verify inference works via LiteLLM → Router → GPU
5. If issue: revert `OPENAI_API_BASE` → `:9000`, restart (30 sec rollback)
**Dual-Path Architecture (throughout migration):** **Verification:**
``` - Check LiteLLM `/health` endpoint
Path A: Agent → LiteLLM :4000 → Router :9000 → GPU (new, for migrated agents) - Verify GPU metrics via router `/metrics`
Path B: Agent → Router :9000 → GPU (legacy, for unmigrated agents) - Test each agent individually
``` - Confirm no 502/503 errors in logs
Both paths work simultaneously. Router unchanged in both paths.
**Rollback Procedure (per agent, 30 seconds):**
1. Set `OPENAI_API_BASE=http://192.168.68.116:9000/v1`
2. Set `OPENAI_API_KEY=<original-agent-key>`
3. Restart agent
**Emergency Global Rollback (if LiteLLM catastrophic failure):**
1. NGINX auto-fallback: `error_page 502 = @router_fallback` → traffic routes to :9000
2. Manual: Update all agent configs back to :9000 (scripted)
3. Restart all agents (5 min)
**Health Checks (continuously monitored):**
- `/health` — LiteLLM + Router combined status
- `/router/metrics` — GPU health, circuit breaker, active slots
- `/stream` — SSE dashboard (real-time)
- Redis: `circuit:*` keys for tripped breakers
--- ---
## Appendix A: Router Slim-Down (Phase 3) ## Appendix A: Model Identity Gap Analysis
After full migration, `router-fixed.py` can be simplified by removing: | Model | GPU | VRAM | Context | Status |
|-------|-----|------|---------|--------|
| qwen3.6-35B-A3B | MoE/Strix | 65GB | 262K | Healthy |
| qwen3.6-27B-code | Dense/RTX3090 | 24GB | 262K | Healthy |
| gemma-4-12b | VLM/RTX 5070 | 12GB | 262K | Healthy |
```python **Notes:**
# REMOVE (migrated to LiteLLM): - All three GPUs are operational and available for LiteLLM routing
- API_KEYS validation logic (keep single ROUTER_API_KEY) - GPU health scoring (40/30/30) prevents routing to unhealthy GPUs
- Dual-key deprecation tracking - Router handles all GPU-level routing decisions LiteLLM sees them as opaque endpoints
- /admin/keys, /admin/keys/generate, /admin/keys/revoke
- /admin/keys/deprecation-summary
- Phase 0 deprecated key logging
- check_rate_limit() (optional — keep as hardware safety net)
# KEEP:
- route() — all 5 tiers
- select_best_gpu()
- check_gpu_health()
- is_gpu_busy(), gpu_active_count(), gpu_incr/decr()
- estimate_tokens()
- store_perf_record()
- GPU_SIDECARS, GPU_URLS, GPU_MAX_CONCURRENT, GPU_CONTEXT
- counter_audit_loop()
- /v1/chat/completions — core routing endpoint
- /v1/models
- /health
- /metrics, /metrics/performance, /metrics/scatter, /metrics/timeseries
- /stream — SSE dashboard
```
## Appendix B: LiteLLM Cost Config for Local GPUs
```yaml
# In config.yaml — map models to per-token pricing for spend tracking
litellm_settings:
model_cost:
qwen3.6-35B-A3B:
input_cost_per_token: 0.0 # Self-hosted, no external cost
output_cost_per_token: 0.0
qwen3.6-27B-code:
input_cost_per_token: 0.0
output_cost_per_token: 0.0
gemma-4-12b:
input_cost_per_token: 0.0
output_cost_per_token: 0.0
# For internal cost allocation, set symbolic rates:
# e.g., MoE = $2/M tokens, Dense = $1/M tokens, VLM = $0.50/M tokens
```
--- ---
*Plan drafted: 2026-06-13 by Abiba 🦊⚡* ## Appendix B: LiteLLM Virtual Key Migration
*Updated: 2026-06-14 by Agent Zero (kagentz-bot) — CT 116 deployment target, DNS split-horizon, Authentik OIDC, Zero-Downtime strategy, GitOps workflow*
*Status: Ready for Phase 0 deployment* **Current API Keys LiteLLM Virtual Keys:**
| Agent | Old Key | New LiteLLM Key | Tier | Budget |
|-------|---------|-----------------|------|--------|
| Abiba | sk-***-*** | sk-litellm-*** | enterprise | $1000 |
| Mumuni | sk-***-*** | sk-litellm-*** | enterprise | $1000 |
| Tanko | sk-***-*** | sk-litellm-*** | enterprise | $1000 |
| Kagenz0 | sk-***-*** | sk-litellm-*** | professional | $500 |
| Koby | sk-***-*** | sk-litellm-*** | professional | $500 |
| Koonimo | sk-***-*** | sk-litellm-*** | professional | $500 |
---
## Appendix C: Authentik SSO Integration
**Authentik Provider Setup:**
1. Create OAuth2 application in Authentik
2. Set redirect URI: `http://<CT-116-IP>/sso/callback`
3. Configure `client_id` and `client_secret`
4. Mount `custom_sso.py` to LiteLLM container
5. Update config.yaml with provider details
---
## Appendix D: Prometheus Monitoring
**Metrics Export:**
- LiteLLM metrics `http://127.0.0.1:4000/metrics`
- Router metrics `http://127.0.0.1:9000/metrics`
- GPU health metrics `http://127.0.0.1:9000/metrics/gpu`
- Circuit breaker metrics `http://127.0.0.1:9000/metrics/circuit-breaker`
**Alerts:**
- GPU health score > 70 alert to Hermes
- Circuit breaker trip alert to Hermes
- LiteLLM spend > $100/day alert to Hermes
- LiteLLM latency > 1000ms alert to Hermes