Author SHA1 Message Date
SyslogBot b09a93f45c feat: Smart Queue Consumer implementation draft + architecture review
- SMART_QUEUE_IMPLEMENTATION.md: Complete implementation draft (1572 lines)
  with 10 quick-win fixes and full smart queue consumer rewrite
- ARCHITECTURE_REVIEW.md: 26-issue audit with prioritized findings
- Verified all 3 GPUs live: amdpve (73% util), llmgpu (idle), ocu_llm (idle)
- Redis 7.4.9 confirmed streams support
- GPU sidecar metrics verified on all hosts

Key fixes:
- QW-1: Dockerfile path mismatch (Dockerfile.queue -> queue-service/Dockerfile)
- QW-2: Nginx fallback only on ALL-GPU failure (not single GPU)
- QW-3: Container names fixed to Docker service names
- QW-4: Redis host default fixed (192.168.68.7 -> redis)
- QW-5: Dependency version pinning
- QW-7-10: Health checks, restart policy, Gunicorn, single-process collector

Smart queue features:
- Redis Streams + consumer groups
- GPU-aware load balancing via sidecar metrics
- Per-GPU circuit breakers with half-open recovery
- Adaptive backpressure (0-30 normal, 30-40 warn, 40-50 503, >50 open)
- Dead letter queue with retry endpoint
- Job ID tracking and /status/<job_id> API
2026-05-17 03:55:20 +00:00
SyslogBot e95475f431 Add GPU dashboard container + Nginx routing 2026-05-15 22:25:56 +00:00
57 changed files with 4273 additions and 12798 deletions
-8
View File
@@ -1,8 +0,0 @@
# Syslog Harness Environment
REDIS_HOST=192.168.68.8
REDIS_PORT=6379
AMDPVE_ENDPOINT=http://192.168.68.15:8080
LLMGPU_ENDPOINT=http://192.168.68.8:8080
OCU_LLM_ENDPOINT=http://192.168.68.110:8080
CIRCUIT_BREAKER_THRESHOLD=5
CIRCUIT_BREAKER_TIMEOUT=30
-6
View File
@@ -1,6 +0,0 @@
.git
__pycache__/
*.pyc
*.bak
*.backup*
.env
+390
View File
@@ -0,0 +1,390 @@
# Syslog Harness Architecture Review & Improvement Recommendations
**Date:** 2026-05-17
**Commit:** `e95475f` "Add GPU dashboard container + Nginx routing"
**Repo:** http://192.168.68.17:3000/SyslogSolution/syslog-harness.git
---
## 1. Current Architecture Overview
```
Host (192.168.68.123)
Agent :8080> Nginx Router > Queue Service > Dashboard
:8080 :8091 :3001
GPU Pool Redis > GPU Dashboard
:8080 :6379 :8092
amdpve llmgpu ocu_llm
.15:8080 .8:8080 .110:8080
MoE 35B Dense 27B Light 4B
```
### Services
| Service | Port | Container | Image | Purpose |
|---|---|---|---|---|
| **Nginx Router** | 8080 | Host-level | OS nginx | Routes by `X-Syslog-Model` header |
| **Queue Service** | 8091 | `syslog-queue` | `python:3.13-slim` | Request queue + circuit breaker |
| **Dashboard** | 3001 | `syslog-dashboard` | `python:3.11-slim` | Observability UI + GPU health |
| **GPU Dashboard** | 8092 | `syslog-gpu-dashboard` | `python:3.11-slim` | Hardware metrics (temp, VRAM, power) |
| **Redis** | 6379 | `syslog-redis` | `redis:7-alpine` | Queue storage |
### GPU Backends
| Host | GPU | Model | Capacity |
|---|---|---|---|
| 192.168.68.15 | AMD Strix Halo | qwen3.6-35B-A3B (MoE) | 65GB VRAM |
| 192.168.68.8 | RTX 3090 | qwen3.5-27B (Dense) | 24GB VRAM |
| 192.168.68.110 | RTX 5070 | gemma-4-E4B (Light) | 12GB VRAM |
### Data Flow
1. **Agent** sends request with `X-Syslog-Model` header Nginx :8080
2. **Nginx** routes to appropriate GPU based on header mapping
3. **GPU backend** (llama.cpp) processes request
4. **Fallback:** If GPU returns 502/503/timeout Nginx redirects to queue-service :8091
5. **Queue** stores request in Redis `inference:requests` LPUSH
6. **Dashboard** :3001 polls queue-service + GPU health for display
7. **GPU Dashboard** :8092 collects hardware metrics every 10s
---
## 2. File Inventory
```
docker-compose.yml # Main compose (Docker networking)
gpu-router-docker.conf # Nginx config for Docker deployment
Dockerfile.gpu # GPU dashboard container
Dockerfile.dashboard # Dashboard container (root-level)
queue-service/Dockerfile # Queue service container
queue-service/queue-service.py # Queue logic (121 lines)
dashboard/harness-dashboard.py # Dashboard app (133 lines)
dashboard/Dockerfile # Dashboard container (subdir)
dashboard/Dockerfile.dashboard # Dashboard container (duplicate)
gpu-dashboard/gpu_collector.py # GPU hardware collector (115 lines)
gpu-dashboard/gpu.html # GPU dashboard UI (183 lines)
gpu-dashboard/collector.py # Duplicate collector (hermes-workspace path)
gpu-dashboard/start.sh # Legacy startup script
MIGRATION_PLAN.md # Production migration plan
README.md # Documentation
syslog-harness-check/ # Checkpoint subdirectory (mirror)
```
---
## 3. Detailed Findings
### 3.1 Queue Service (`queue-service/queue-service.py`)
**Architecture:** Simple Flask app using Redis LPUSH/RPUSH for a FIFO queue. A basic circuit breaker prevents queue overflow at 50 messages.
**Issues Found:**
| # | Severity | Location | Issue |
|---|---|---|---|
| Q1 | **CRITICAL** | Lines 82-88 | **Queue is fire-and-forget with no consumer.** Requests are pushed to Redis but nothing dequeues or processes them. The queue is a dead storage pit. |
| Q2 | **CRITICAL** | Lines 28-32 | **Hardcoded GPU IPs** in the queue service duplicate the Nginx config. No configuration source of truth. |
| Q3 | **HIGH** | Lines 21-22 | **Redis host fallback to `192.168.68.7`** (line 21) conflicts with docker-compose which sets `REDIS_HOST=redis` (line 24). The default is unreachable inside Docker. |
| Q4 | **HIGH** | Lines 66-95 | **No job result retrieval mechanism.** Once enqueued, there's no API to poll for completion, get a job ID, or retrieve results. |
| Q5 | **HIGH** | Lines 73-79 | **Circuit breaker is a simple depth threshold.** No backoff, no recovery window, no sliding window. Once closed, it stays closed until manually drained. |
| Q6 | **MEDIUM** | Lines 50-57 | **GPU health check is synchronous and blocks** the `/status` endpoint. Checking 3 GPUs sequentially with 3s timeout means `/status` can take up to 9s. |
| Q7 | **MEDIUM** | Lines 35-40 | **`get_redis()` swallows all exceptions** and returns `None`. This makes Redis failures silent queue depth returns 0 on failure (line 47), potentially allowing overflow. |
| Q8 | **MEDIUM** | Lines 83-84 | **Headers filtered to only X-* prefixed** the `Content-Type` header is dropped entirely, meaning the receiver can't determine payload format. |
| Q9 | **LOW** | Line 121 | **No graceful shutdown.** Flask development server doesn't handle SIGTERM gracefully. |
### 3.2 Nginx Gateway (`gpu-router-docker.conf`)
**Architecture:** Nginx routes requests to GPU backends based on `X-Syslog-Model` header value. Has rate limiting, streaming support, and queue fallback.
**Issues Found:**
| # | Severity | Location | Issue |
|---|---|---|---|
| N1 | **HIGH** | Lines 79-80 | **`burst=20 nodelay`** means 20 requests are served immediately beyond the rate limit, then throttled. This defeats the purpose of rate limiting under burst traffic all 20 could still overwhelm a GPU. |
| N2 | **HIGH** | Lines 99-100 | **`proxy_next_upstream` with `tries 2`** means on error/timeout/502/503, Nginx retries once. But it retries against the *same GPU pool*, not a different one. The same GPU that failed gets hit again. |
| N3 | **HIGH** | Lines 106, 112-121 | **Queue fallback (`@queue_fallback`) is triggered for ANY 502/503/504**, including when a single GPU is overloaded. This means individual GPU slowness causes queue fallback instead of just queuing when ALL GPUs are down. |
| N4 | **MEDIUM** | Line 90 | **`proxy_pass_header X-Syslog-Model`** is non-standard. Nginx automatically passes request headers; this directive is for response headers. The model header is already passed implicitly via `proxy_set_header` inheritance. |
| N5 | **MEDIUM** | Lines 27, 32 | **Hardcoded container names** (`syslog-harness-dashboard-1`, `syslog-harness-gpu-dashboard-1`). These change based on docker-compose project prefix. Should use service names. |
| N6 | **LOW** | Lines 67-73 | **GPU dashboard at `/gpu` path** has `X-Forwarded-Proto` but the dashboard service (simple HTTP server) doesn't use it. Inconsistent header handling across locations. |
### 3.3 Dashboard (`dashboard/harness-dashboard.py`)
**Architecture:** Simple HTTP server using Python's `http.server`. Fetches queue status and GPU health, renders HTML.
**Issues Found:**
| # | Severity | Location | Issue |
|---|---|---|---|
| D1 | **HIGH** | Lines 34-40 | **`get_queue_status()` calls queue-service synchronously.** Combined with per-GPU health checks (lines 18-31), the `/api/status` endpoint makes 4 sequential HTTP calls. Worst case: 2 + 33s = 11s response time. |
| D2 | **MEDIUM** | Lines 101-127 | **Uses `SimpleHTTPRequestHandler`** which is single-threaded. Under concurrent dashboard access, requests queue up. Should use `ThreadingHTTPServer`. |
| D3 | **MEDIUM** | Lines 16-18 | **GPU endpoints hardcoded** in dashboard, separate from queue-service and Nginx. Three separate sources of truth for GPU addresses. |
| D4 | **LOW** | Line 127 | **Silent log suppression.** While intentional, this makes debugging impossible without modifying the source. |
### 3.4 GPU Dashboard (`gpu-dashboard/`)
**Architecture:** `gpu_collector.py` polls sidecar (port 8090) and llama.cpp (port 8080) endpoints every 10s, writes JSON to `gpu_metrics.json`. Static HTTP server serves the dashboard.
**Issues Found:**
| # | Severity | Location | Issue |
|---|---|---|---|
| G1 | **HIGH** | Lines 97-98 | **Sequential collection.** All 3 GPUs are polled sequentially (line 98: list comprehension). If one host is unreachable, it blocks collection for all three. |
| G2 | **HIGH** | Line 105-107 | **`/app/public/gpu_metrics.json` path is hardcoded** and differs from `collector.py` (line 11: `/root/hermes-workspace/public/gpu_metrics.json`). Inconsistent between the two collector files. |
| G3 | **MEDIUM** | Lines 19-25 | **`fetch_json` swallows all exceptions.** A timeout on one GPU's sidecar is silently ignored, making it impossible to distinguish "no data" from "collector error". |
| G4 | **MEDIUM** | Line 14 | **`DEAD_THRESHOLD = 60` seconds is aggressive.** A GPU that restarts takes 60s before reappearing as online, even if it's back in 5s. |
| G5 | **LOW** | Lines 10-14 | **`start.sh` references `/root/hermes-workspace/public`** but `Dockerfile.gpu` creates `/app/public`. Inconsistent between legacy and current deployment. |
### 3.5 Docker Compose (`docker-compose.yml`)
**Issues Found:**
| # | Severity | Location | Issue |
|---|---|---|---|
| C1 | **HIGH** | Lines 19-20 | **Queue service exposes port 8091 externally.** In a multi-tenant or public-facing deployment, the queue API should be internal-only. |
| C2 | **MEDIUM** | Lines 13-15 | **`Dockerfile.queue` referenced but doesn't exist at root level.** The file is at `queue-service/Dockerfile`. The compose build context is `.` (root) but the dockerfile path doesn't match. |
| C3 | **MEDIUM** | Lines 6, 16, 26, 31, 43 | **`restart: always`** instead of `restart: unless-stopped`. On crash, `always` restarts even after manual stop, making maintenance harder. |
| C4 | **LOW** | Lines 23-25 | **No health checks defined** for any service. Docker can't detect if a service is actually healthy, only if the container is running. |
| C5 | **LOW** | Line 10 | **Redis has no password.** Unauthenticated Redis exposed on the Docker network. |
| C6 | **LOW** | Lines 49-51 | **No network driver specified** for the bridge network (minor defaults to bridge). No IPAM configuration for large deployments. |
### 3.6 Container Images
**Issues Found:**
| # | Severity | Location | Issue |
|---|---|---|---|
| I1 | **HIGH** | All Dockerfiles | **No `requirements.txt` or dependency pinning.** All dependencies (`flask`, `redis`, `requests`) are installed without version pins. Builds are non-reproducible. |
| I2 | **MEDIUM** | `Dockerfile.gpu` line 3 | **`pip install requests`** unnecessary dependency for the GPU dashboard (only uses `urllib`). Adds ~300KB to the image. |
| I3 | **MEDIUM** | `Dockerfile.gpu` line 14 | **Multi-process CMD with `&`** no process supervisor. If the collector crashes, it won't restart. The `http.server` also won't receive SIGTERM properly. |
| I4 | **LOW** | All Dockerfiles | **No `.dockerignore` file.** The entire context is sent to the Docker daemon, including `.git` directories and any local artifacts. |
| I5 | **LOW** | `Dockerfile.dashboard` (root) vs `dashboard/Dockerfile.dashboard` | **Duplicate Dockerfiles** with slight differences (Python 3.11 vs 3.13, WORKDIR differences). |
---
## 4. Smart Queuing Analysis & Recommendations
### Current State: No Smart Queuing
The queue service is a **passive storage mechanism** it stores requests but has no intelligence:
- **No load balancing** no awareness of GPU load (slots_busy, VRAM usage, queue depth per GPU)
- **No job prioritization** FIFO only, no priority levels
- **No backpressure** simple threshold, no exponential backoff or adaptive limits
- **No retry logic** failed GPU requests go to queue but are never reprocessed
- **No dead letter handling** stuck or failed jobs have no lifecycle management
- **No consumer** nothing dequeues and forwards to GPUs
- **No job tracking** no job IDs, no status updates, no result retrieval
### Recommended Architecture: Smart Queue with Consumer
```
Agent > Nginx > Smart Queue API > Redis Streams (with consumers)
Consumer
Pool
GPU 1 (load) GPU 2 (load) GPU 3 (load)
Health Health Health
Update GPU scores
Priority Queue (sorted by urgency)
Dead Letter Queue (failed jobs)
Backpressure (adaptive rate limit)
```
### Specific Recommendations
#### R1: Implement Redis Streams as Queue Backend
- Replace `LPUSH/RPUSH` (FIFO list) with **Redis Streams** (`XADD/XREADGROUP`)
- Streams support consumer groups, message acknowledgment, and pending messages
- Enables proper dead letter queue handling and retry logic
- **File:** `queue-service/queue-service.py`
```python
# Before: Simple list
r.rpush(QUEUE_KEY, json.dumps(job))
# After: Redis Stream with consumer group
stream_key = "inference:stream"
consumer_group = "gpu-workers"
r.xadd(stream_key, {"job": json.dumps(job)}, maxlen=10000, approx=True)
```
#### R2: Build a Queue Consumer Pool
- Deploy 1+ consumer containers that poll the stream and forward to GPUs
- Consumer selects GPU based on: health status, current load (slots_busy), and VRAM availability
- **File:** New `queue-service/consumer.py`
```python
class LoadBalancedConsumer:
def select_gpu(self, job):
"""Select GPU based on load, health, and model compatibility."""
candidates = [g for g in self.gpus if g.health == "up" and not g.full]
if not candidates:
return None
# Sort by: slots_idle (descending), VRAM_available (descending)
candidates.sort(key=lambda g: (g.slots_idle, g.vram_free_mb), reverse=True)
return candidates[0]
```
#### R3: Implement Priority Queuing
- Add priority field to job payload: `high`, `normal`, `low`
- Use Redis Streams with multiple stream keys per priority level
- Consumer checks `high` `normal` `low` in order
- **File:** `queue-service/queue-service.py` enqueue endpoint
#### R4: Add Backpressure Mechanism
- Instead of hard threshold at 50, implement **adaptive backpressure**:
- Queue depth 0-30: normal operation
- Queue depth 30-40: return `retry-after` header with increasing delay
- Queue depth 40-50: return 503 with exponential retry-after
- Queue depth >50: circuit breaker open
- **File:** `queue-service/queue-service.py`
#### R5: Dead Letter Queue (DLQ)
- Move failed/unprocessable jobs to a `inference:dead-letter` stream
- Include failure reason, attempt count, and original payload
- Provide admin API to inspect, retry, or discard DLQ entries
- **File:** `queue-service/queue-service.py`
```python
# New endpoint
@app.route("/dlq", methods=["GET"])
def list_dlq():
return r.xrange("inference:dead-letter")
@app.route("/dlq/retry/<message_id>", methods=["POST"])
def retry_dlq(message_id):
job = r.xget("inference:dead-letter", message_id)
r.xadd("inference:stream", {"job": job})
```
#### R6: GPU-Aware Routing
- Queue consumer should check GPU `slots_busy` before routing
- If a GPU is busy, try the next available GPU
- Track per-GPU queue depth and avoid overloading a single GPU
- **File:** New consumer logic
#### R7: Job Status API
- Add job ID generation on enqueue
- Provide `/status/<job_id>` endpoint to check progress
- Store job state in Redis: `queued` `processing` `completed`/`failed`
- **File:** `queue-service/queue-service.py`
```python
@app.route("/enqueue", methods=["POST"])
def enqueue():
job_id = str(uuid.uuid4())
job = {"id": job_id, "payload": ..., "status": "queued", "created_at": time.time()}
r.xadd(stream_key, {"job": json.dumps(job)})
r.hset("job:status", job_id, json.dumps({"status": "queued"}))
return jsonify({"job_id": job_id, "status": "queued"}), 202
@app.route("/status/<job_id>")
def job_status(job_id):
status = r.hget("job:status", job_id)
return jsonify(json.loads(status)) if status else {"error": "not found"}, 404
```
#### R8: Health-Based Circuit Breaker
- Replace simple depth threshold with **per-GPU circuit breakers**
- Track consecutive failures per GPU
- Implement half-open state: after cooldown, probe one GPU to test recovery
- **File:** `queue-service/queue-service.py`
#### R9: Centralized Configuration
- Move GPU endpoints from 3 locations (queue-service, dashboard, Nginx) to:
- Redis config key: `config:gpus`
- Or environment file mounted to all containers
- Nginx can use Lua/variable from config instead of static upstreams
- **File:** New `config/` directory or Redis-based config
---
## 5. Priority Issue Summary
### Critical (Fix Immediately)
1. **Q1** Queue has no consumer; enqueued requests are never processed
2. **Q4** No job ID or result retrieval mechanism
3. **N3** Queue fallback triggers on individual GPU failure, not all-down
### High (Fix Before Production)
4. **Q5** Circuit breaker has no recovery mechanism
5. **Q6** `/status` endpoint blocks on GPU health checks
6. **D1** Dashboard `/api/status` makes 4 sequential calls, up to 11s
7. **C2** `Dockerfile.queue` path mismatch in docker-compose
8. **I1** No dependency pinning in any Dockerfile
9. **I3** Multi-process CMD without supervisor in GPU dashboard
### Medium (Improve in Next Iteration)
10. **Q3** Redis host default conflicts with Docker networking
11. **Q7** Silent exception swallowing in Redis access
12. **Q8** Content-Type header dropped in queue
13. **D2** Single-threaded dashboard server
14. **D3** Three separate sources of truth for GPU addresses
15. **G1** Sequential GPU collection blocks on single failure
16. **N1** Rate limit burst of 20 nodelay defeats protection
17. **N5** Hardcoded container names in Nginx
18. **C1** Queue API exposed externally
19. **C4** No Docker health checks
### Low (Nice to Have)
20. **Q9** No graceful shutdown
21. **C3** `restart: always` vs `unless-stopped`
22. **C5** No Redis authentication
23. **G4** 60s dead threshold is too aggressive
24. **I2** Unnecessary `requests` dependency
25. **I4** No `.dockerignore`
26. **I5** Duplicate Dockerfiles
---
## 6. Deployment Architecture Summary
### What Works Well
- Clean separation of concerns: routing (Nginx), queuing (Redis + queue-service), observability (two dashboards)
- Good GPU hardware monitoring with temperature, VRAM, power, fan metrics
- SSE streaming support in Nginx for LLM response streaming
- Rate limiting at the gateway layer
- Circuit breaker pattern implemented (even if basic)
### What Needs Work
- **Queue is incomplete** storage without processing is the most critical gap
- **No job lifecycle** requests go in and never come out
- **Duplicated configuration** GPU addresses in 3+ places
- **No monitoring/alerting** no Prometheus metrics, no alerting rules
- **Single point of failure** no Redis replication, no container redundancy
- **No logging** Flask dev server logs are minimal; no structured logging
### Recommended Next Steps
1. **Priority 1:** Implement queue consumer with GPU load-based routing
2. **Priority 2:** Add job status tracking and result retrieval
3. **Priority 3:** Fix Nginx fallback to only trigger when ALL GPUs are down
4. **Priority 4:** Add Docker health checks and proper dependency management
5. **Priority 5:** Centralize GPU configuration in Redis or environment
6. **Priority 6:** Add Prometheus metrics endpoint for observability
+5
View File
@@ -0,0 +1,5 @@
FROM python:3.11-slim
WORKDIR /app
COPY dashboard/harness-dashboard.py .
EXPOSE 3001
CMD ["python3", "harness-dashboard.py"]
+14
View File
@@ -0,0 +1,14 @@
FROM python:3.11-slim
RUN pip install requests
COPY gpu-dashboard/ /app/
WORKDIR /app
RUN mkdir -p /app/public && \
cp gpu.html /app/public/ && \
touch /app/public/gpu_metrics.json
EXPOSE 8092
CMD ["sh", "-c", "python3 gpu_collector.py & python3 -m http.server 8092 --directory /app/public & wait"]
-56
View File
@@ -1,56 +0,0 @@
# LiteLLM Virtual Key Distribution — Phase 2 Cutover
## Syslog Solution LLC — July 1, 2026 (keys regenerated)
### Active Agent Keys (Update agent configs with these)
| Agent | LiteLLM Virtual Key | Tier | Budget | Models | Verified |
|-------|-------------------|------|--------|--------|----------|
| **Abiba** | `sk-i7F7rCfgpS0oouOTA2LgMw` | enterprise | $1,000 | All 4 | ✅ Jul 1 |
| **Mumuni** | `sk-XY2aUfvy2BIs6kp1ZPh6VA` | enterprise | $1,000 | All 4 | ✅ Jul 1 |
| **Tanko** | `sk-3ZsdWJbbNSo9zSnJN2OsJw` | enterprise | $1,000 | All 4 | ✅ Jul 1 |
| **Kagenz0** | `sk-VYKRAVYEG1rKVEMDKEUggg` | enterprise | $1,000 | All 4 | ✅ Jul 1 |
| **Koby** | `sk-gWDaPAp-FavgKdqKJpzqhQ` | professional | $500 | All 4 | ✅ Jul 1 |
| **Koonimo** | `sk-1kLC8ZxW-tEZueV3NU4rCQ` | professional | $500 | All 4 | ✅ Jul 1 |
### Configuration Changes Required Per Agent
Each agent must update their `OPENAI_API_BASE` and `OPENAI_API_KEY`:
```
OPENAI_API_BASE=http://192.168.68.116/v1
OPENAI_API_KEY=<agent's LiteLLM key from table above>
```
### LiteLLM Admin Access
| Resource | Key |
|----------|-----|
| LiteLLM Admin UI | http://192.168.68.116/ui/ (Authentik SSO — pending) |
| LiteLLM Master Key | `sk-litellm-7f96080dd99b15c36bd4b333b58a6796` |
| Router Admin Key | `sk-admin-ee09fffd04978b61a1569ac670c68814` |
### Fallback (if LiteLLM fails)
If LiteLLM is down, NGINX automatically falls back to the router on :9000.
In that case, agents can also directly use:
```
OPENAI_API_BASE=http://192.168.68.116/v1
OPENAI_API_KEY=<agent's original sk-syslog-* key>
```
(Only works when NGINX @router_fallback is active)
### Deprecated Router Keys
These keys still work through the @router_fallback but NOT through LiteLLM:
- sk-syslog-abiba, sk-syslog-mumuni, sk-syslog-tanko
- sk-syslog-kagenz0, sk-syslog-koby, sk-syslog-koonimo
- sk-starter-abc123, sk-professional-xyz789
- sk-syslog-local-master-key
These have been fully revoked as of July 1, 2026.
### Key Rotation History
- **2026-07-01**: Keys regenerated for Abiba, Kagenz0, Koby, Koonimo (old values unrecoverable).
Mumuni and Tanko keys confirmed working, left unchanged. All blocked/stale keys purged.
- **2026-06-17**: Initial key creation (values in this doc were placeholder/incorrect).
-759
View File
@@ -1,759 +0,0 @@
# LiteLLM Integration Migration Plan
## Syslog Solution LLC June 14, 2026
**Deployment Target:** CT 116 `syslog-api` (192.168.68.116) on minipve all services co-located.
**DNS Strategy:** Option A /etc/hosts + Docker extra_hosts for internal resolution of `auth.sysloggh.net` 192.168.68.11.
**GitOps:** This plan lives in `SyslogSolution/syslog-harness` on Gitea. All changes tracked via git with conventional commits.
---
## Executive Summary
**Goal:** Layer the full LiteLLM Gateway suite (Admin UI, virtual keys, spend tracking, teams/SSO, budget management) on top of our custom intelligent routing harness without sacrificing GPU-aware slot management, content-based tiering, or hardware health monitoring.
**Architecture Decision:** Two-layer architecture.
```
LiteLLM Gateway (Layer 1)
Port 4000 Policy & UX
Admin UI (/ui)
Virtual Keys & Permissions
Teams, Users, SSO (OIDC)
Spend Tracking & Budgets
Usage Analytics Dashboard
Request Audit Trail
Global Rate Limiting
Pass-through to router
Custom Router (Layer 2)
Port 9000 Intelligence & HW
5-Tier Content-Based Routing
GPU Slot Management (Redis)
Agent Spread Prevention
GPU Health Scoring
Sidecar VRAM/Temp/Power
Circuit Breaker
Context Window Tracking
Per-Request Perf Recording
Hardware Rate Limiting
qwen3.6-35B qwen3.6-27B gpu-vision
MoE/Strix Dense/RTX3090 VLM/RTX 5070
:8080 (llama) :8080 (llama) :8080 (llama)
:8090 (side) :8090 (side) :8090 (sidecar)
```
---
## 1. Current State Baseline
### 1.1 Router (`router-fixed.py` port 9000, deployed on CT 116 / syslog-api)
**Deployment Host:** CT 116 `syslog-api` on minipve (192.168.68.12), IP 192.168.68.116, 6GB RAM, 40GB disk. Runs Docker with all harness services co-located on this single host.
| Feature | Implementation |
|---------|---------------|
| **Routing Engine** | 5-tier content-based: lightweight simple_conv medium heavy_reasoning default |
| **GPU Slot Mgmt** | Redis atomic incr/decr, max 2 concurrent per GPU, audit loop reset |
| **Health Checks** | Sidecar endpoint per GPU (VRAM, temp, util, power) + llama.cpp /health |
| **Agent Spreading** | `select_best_gpu()` prefers GPUs with 0 other agents, then non-self GPUs |
| **Rate Limiting** | Token bucket (Redis), per-tier RPM: enterprise=120, professional=60, starter=20 |
| **Auth** | Dual-key system (Phase 0.5): 9 new + 9 deprecated keys, admin key rotation |
| **Performance** | Per-request latency/tokens/tps Redis lists (perf:recent, perf:model:X, perf:agent:X) |
| **Context Tracking** | Session-level token accumulation with compaction warnings in headers |
| **SSE Streaming** | Real-time dashboard updates, per-model timeseries |
| **Admin** | `/admin/keys`, `/admin/keys/generate`, `/admin/keys/revoke`, `/admin/keys/deprecation-summary` |
| **Strict Passthrough** | Explicit model requests go to that GPU exactly (no silent fallback). LiteLLM owns failover. |
### 1.2 GPU Backends
| GPU | Host | llama.cpp | Sidecar | VRAM | Context |
|-----|------|-----------|---------|------|---------|
| qwen3.6-35B-A3B (MoE) | 192.168.68.15 | :8080 | :8090 | Strix Halo | 262K |
| qwen3.6-27B-code (Dense) | 192.168.68.8 | :8080 | :8090 | RTX 3090 | 262K |
| gpu-vision (VLM) | 192.168.68.110 | :8080 | :8090 | RTX 5070 | 262K |
### 1.3 Existing LiteLLM POC on CT 116
CT 116 already has a LiteLLM container running (POC, 6 days uptime):
```
harness-litellm | ghcr.io/berriai/litellm:main-stable | 127.0.0.1:8081->4000
harness-redis | redis:7-alpine | 127.0.0.1:6379
harness-router | inference-harness-router | 127.0.0.1:9000
harness-nginx | nginx:alpine | 0.0.0.0:80
harness-dashboard | inference-harness-dashboard | 127.0.0.1:3000
```
- `/opt/litellm/` previous setup directory on CT 116
- Configured with Postgres, host networking, master key
- Currently bypassed router routes directly to GPUs
- **Goal: Productionize with two-layer architecture on this same host**
### 1.4 DNS Routing (Split-Horizon)
For OIDC SSO with Authentik, CT 116 must resolve `auth.sysloggh.net` internally:
**Problem:** `auth.sysloggh.net` CNAMEs to `netbird.sysloggh.net` 72.61.0.17 (public VPS). OIDC auth_request from NGINX would route through the internet back to 192.168.68.11 unnecessarily.
**Solution Option A: /etc/hosts on CT 116 host:**
```bash
# On CT 116 (syslog-api)
echo "192.168.68.11 auth.sysloggh.net" >> /etc/hosts
```
**Docker containers** also need this resolution add to docker-compose.yml:
```yaml
services:
nginx:
extra_hosts:
- "auth.sysloggh.net:192.168.68.11"
litellm:
extra_hosts:
- "auth.sysloggh.net:192.168.68.11"
```
**DNS Servers:** CT 116 uses 192.168.68.10 for DNS. AdGuard (192.168.68.11) is the long-term solution for LAN-wide split-horizon DNS.
---
## 2. What LiteLLM Brings (That We Don't Have)
| Feature | Our Router | LiteLLM | Value Add |
|---------|-----------|---------|-----------|
| **Admin UI** | | Full dashboard at /ui | Non-technical users can manage keys, view spend |
| **Virtual Key Permissions** | (binary key->tier) | Granular: per-model, per-team, budget caps | Fine-grained access control |
| **Spend Tracking** | | Per-request $ cost with model-specific pricing | Billing, cost allocation, client invoicing |
| **Teams & Orgs** | | Multi-tenant: org->team->user hierarchy | Segregate clients/projects |
| **SSO/OIDC** | | Google, GitHub, Microsoft, Okta, Keycloak | Enterprise auth integration |
| **Budget Alerts** | | Per-key, per-user, per-team budget with webhooks | Prevent overspend |
| **Usage Analytics** | (custom /metrics) | Built-in: daily trends, model breakdown, per-customer | Better visualization |
| **100+ Provider Support** | (3 local GPUs) | OpenAI, Anthropic, Bedrock, Vertex, etc. | Future cloud model access |
| **Fallback Chains** | (silent rerouting) | Explicit multi-provider failover with per-model logging | Accurate per-model tracking, visible failover |
| **RPM/TPM Weighted LB** | | Weighted load balancing across deployments | Fine-grained traffic shaping |
---
## 3. What We Keep (That LiteLLM Doesn't Have)
| Feature | Why We Must Keep It |
|---------|---------------------|
| **Content-based 5-tier routing** | LiteLLM routes by model name only; we analyze prompt complexity, tokens, turns, and routing_hints |
| **GPU hardware health scoring** | LiteLLM doesn't monitor VRAM, temp, power our scoring prevents routing to overheating GPUs |
| **GPU slot management** | LiteLLM doesn't know about llama.cpp --parallel limits; our Redis counters prevent overloading |
| **Agent spread prevention** | Our `select_best_gpu()` spreads agents across GPUs to prevent hotspots; LiteLLM only does simple-shuffle |
| **Cross-turn context tracking** | Session-level token accumulation with compaction warnings via X-Context-Warning headers |
| **GPU sidecar metrics** | VRAM %, GPU utilization %, power draw, temperature exposed via /metrics and SSE dashboard |
| **Circuit breaker** | 39 failures caught June 12; LiteLLM's allowed_fails/cooldown is less granular |
---
## 4. Migration Architecture
### 4.1 Principle: "LiteLLM is the lobby, our router is the engine room"
- **LiteLLM** handles everything a **user/admin** touches: keys, teams, budgets, spend logs, SSO, the UI
- **Custom Router** handles everything the **GPUs** need: health checks, slot booking, content-based routing, hardware monitoring, circuit breaking
### 4.2 Flow
```
Agent Request
LiteLLM Gateway (:4000)
1. Authenticate virtual key (sk-litellm-...)
2. Check key permissions (model access)
3. Check budget (per-key, per-user, per-team)
4. Check team rate limits
5. Log request metadata
6. Forward to custom router as OpenAI-compat
POST http://router:9000/v1/chat/completions
Headers: Authorization: Bearer ***
X-LiteLLM-User: <user-id>
X-LiteLLM-Team: <team-id>
X-Session-Id: <session>
7. On response: log spend, update budgets
8. If router returns 503 (GPU saturated):
consult fallback chain, retry next model
9. Return response to agent
Custom Router (:9000)
1. Authenticate agent key (sk-syslog-...)
2. Hardware rate limit (per-tier RPM)
3. Content-based tier routing (for syslog-auto)
OR strict passthrough (for explicit models)
4. GPU slot availability (Redis counter)
5. GPU health check (sidecar)
6. Agent spread logic (select_best_gpu)
7. Queue if saturated (with timeout)
8. Forward to selected llama.cpp GPU
9. Track context window, set compaction header
10. Record performance metrics
11. Return response (with routing metadata)
llama.cpp GPU (:8080)
```
### 4.3 LiteLLM Config (`config.yaml`)
```yaml
general_settings:
master_key: os.environ/LITELLM_MASTER_KEY
database_url: postgresql://litellm:***@postgres:5432/litellm
store_model_in_db: true
model_list:
# Content-based auto-routing (router picks GPU via 5-tier analysis)
- model_name: syslog-auto
litellm_params:
model: openai/syslog-auto
api_base: http://router:9000/v1
api_key: os.environ/ROUTER_API_KEY
rpm: 600
# Individual GPU strict passthrough (exact GPU, no silent fallback)
- model_name: qwen3.6-35B-A3B
litellm_params:
model: openai/qwen3.6-35B-A3B
api_base: http://router:9000/v1
api_key: os.environ/ROUTER_API_KEY
- model_name: qwen3.6-27B-code
litellm_params:
model: openai/qwen3.6-27B-code
api_base: http://router:9000/v1
api_key: os.environ/ROUTER_API_KEY
- model_name: gpu-vision
litellm_params:
model: openai/gpu-vision
api_base: http://router:9000/v1
api_key: os.environ/ROUTER_API_KEY
# Guardrails: Pre-call and post-call content moderation
guardrails:
- guardrail_name: "input-moderation"
litellm_params:
guardrail: openai_moderation
mode: "pre_call"
- guardrail_name: "output-moderation"
litellm_params:
guardrail: openai_moderation
mode: "post_call"
- guardrail_name: "harmful-content-filter"
litellm_params:
guardrail: litellm_content_filter
mode: "pre_call"
categories:
- category: "harmful_self_harm"
enabled: true
action: "BLOCK"
severity_threshold: "medium"
- category: "harmful_violence"
enabled: true
action: "BLOCK"
severity_threshold: "medium"
- category: "harmful_illegal_weapons"
enabled: true
action: "BLOCK"
severity_threshold: "medium"
litellm_settings:
num_retries: 0 # Disabled our router handles retry
request_timeout: 600 # Match our 10-min llama-server timeout
set_verbose: true
failure_callback: ["prometheus"] # Optional: export to Prometheus
router_settings:
routing_strategy: "usage-based-routing" # For external models only
enable_loadbalancing_on_proxy: false # Disable LiteLLM's internal LB
allowed_fails: 100 # Router returns 503 on saturated GPUs cooldown disabled
# Fallback chains: LiteLLM retries down the chain when router returns saturated
# This gives accurate per-model metrics because router no longer silently reroutes
fallbacks:
- qwen3.6-35B-A3B: ["qwen3.6-27B-code", "gpu-vision"]
- qwen3.6-27B-code: ["qwen3.6-35B-A3B", "gpu-vision"]
- gpu-vision: ["qwen3.6-27B-code", "qwen3.6-35B-A3B"]
# Cost tracking: map model names to per-token pricing for spend tracking
litellm_settings:
model_cost:
qwen3.6-35B-A3B:
input_cost_per_token: 0.0
output_cost_per_token: 0.0
qwen3.6-27B-code:
input_cost_per_token: 0.0
output_cost_per_token: 0.0
gpu-vision:
input_cost_per_token: 0.0
output_cost_per_token: 0.0
# For internal cost allocation, set symbolic rates:
# e.g., MoE = $2/M tokens, Dense = $1/M tokens, VLM = $0.50/M tokens
```
### 4.4 Router Modifications
To accommodate LiteLLM, `router-fixed.py` requires the following updates:
1. **Strict passthrough for explicit models** (DEPLOYED):
```python
# In route(), the explicit model section changed from silent fallback to strict:
req = rd.get("model","auto")
if req != "auto":
# STRICT MODE: no silent fallback LiteLLM handles failover chains.
# This keeps per-model metrics accurate. Returns saturated if busy.
target = req if req in avail else avail[0]
if req not in avail:
return {"model": req, "reason": "explicit_unavailable", "saturated": True}
if is_gpu_busy(target):
return {"model": target, "reason": "explicit_saturated", "saturated": True}
return {"model": target, "reason": "explicit"}
```
2. **New header passthrough**: Forward `X-LiteLLM-*` headers to GPU (transparent already works)
3. **New endpoint for health passthrough**: `GET /v1/models` already works
4. **Keep ALL routing logic**: No changes to `select_best_gpu()`, `check_gpu_health()`, slot management, etc. Content-based routing for `syslog-auto` is fully intact.
5. **Add LiteLLM-compatible response**: Return `X-Usage-Tokens` header so LiteLLM can track token costs
```python
resp.headers["X-Usage-Tokens"] = json.dumps({
"prompt_tokens": prompt_tokens,
"completion_tokens": completion_tokens,
"model": model
})
```
### 4.5 Router Logic Refinements
**GPU Health Scoring (Updated):**
We are updating the scoring algorithm to include Power metrics:
```python
def gpu_health_score(model):
h = check_gpu_health(model, sidecar_timeout=1.5, gpu_timeout=1)
if h.get("status") == "down":
return 999 # never pick down GPUs
vram_pct = h.get("vram_pct") or 50
temp_c = h.get("temp_c") or 50
power_w = h.get("power_w") or 50
active = gpu_active_count(model)
max_c = GPU_MAX_CONCURRENT.get(model, 1)
load_pct = (active / max_c) * 100 if max_c > 0 else 0
# Score: lower = better
score = (vram_pct * 0.3) + (max((temp_c - 30, 0) * 0.3) + (power_w * 0.2) + (load_pct * 0.2))
return score
```
---
## 5. Deployment Plan (4 Phases) Zero-Downtime Strategy
### Phase 0: Infrastructure Prep (Current Week) Zero Downtime
**Goal:** Prepare CT 116 infrastructure without affecting running agents.
**Tasks:**
1. **Set up DNS split-horizon on CT 116**
```bash
echo "192.168.68.11 auth.sysloggh.net" >> /etc/hosts
```
2. **Deploy Postgres container** alongside existing services
3. **Replace LiteLLM config** with production config.yaml (see 4.3)
- All 4 models `http://router:9000/v1`
- Fallback chains for explicit models
- Guardrails (pre-call, post-call, content filter)
- `allowed_fails: 100` (router returns 503 on saturated)
- `num_retries: 0` (LiteLLM retries handled by fallback chains)
4. **Deploy custom_sso.py** for Authentik OIDC integration
5. **Restart LiteLLM container** with new config
6. **Verify internal routing**
**Verification Checklist:**
- [ ] DNS resolution: `getent hosts auth.sysloggh.net` 192.168.68.11
- [ ] Postgres container healthy
- [ ] LiteLLM `/health` returns 200
- [ ] LiteLLM Router pass-through returns valid chat completion
- [ ] GPU health metrics unaffected
- [ ] Explicit model request returns saturated (not silently rerouted) when GPU busy
### Phase 1: Shadow Mode (Week 1) Zero Risk, Zero Downtime
**Goal:** Deploy LiteLLM alongside existing router, test in shadow mode. **Agents continue using :9000 directly.**
**Tasks:**
1. Create virtual keys for test agents via LiteLLM UI
2. Verify pass-through works for all 4 models
3. Validate fallback chains: saturate MoE confirm LiteLLM retries Dense confirm VLM
4. Run 24-hour shadow: monitor LiteLLM spend logs vs router metrics
5. Verify GPU health metrics unaffected
6. Check guardrails not generating false positives
### Phase 2: Cutover (Week 2) Gradual Agent Migration
**Goal:** Move agents one-by-one to LiteLLM endpoint.
**Tasks:**
1. Migrate API keys to LiteLLM virtual keys
2. Create teams: "Core Agents" (enterprise), "Dev Agents" (professional)
3. Update agent configs one at a time: `OPENAI_API_BASE` `:4000`
4. Test each agent individually
5. Enable SSO via Authentik + custom_sso.py
6. Keep router :9000 as emergency fallback for 48 hours
### Phase 3: Production Hardening (Week 3+)
**Goal:** Lock down, optimize, monitor.
**Tasks:**
1. Remove deprecated router endpoints (after all agents migrated)
2. Add LiteLLM observability (Prometheus, Slack/email alerts)
3. Enable LiteLLM caching (shared Redis)
4. Add external model fallbacks for client-facing services
5. Router slim-down: keep routing/slots/health/perf, remove key management
6. Multi-tenancy setup for client-facing inference services
---
## 6. Nginx Configuration (with Authentik OIDC Forward Auth)
```nginx
# OLD (remove)
# location /admin/ {
# proxy_pass http://127.0.0.1:9000/admin/;
# }
# === Authentik auth subrequest endpoint ===
location /authentik/auth {
internal;
proxy_pass https://auth.sysloggh.net/outpost.goauthentik.io/auth/nginx;
proxy_pass_request_body off;
proxy_set_header Content-Length "";
proxy_set_header X-Original-URL $scheme://$http_host$request_uri;
proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
}
# === LiteLLM Admin UI Authentik-protected ===
location /ui/ {
auth_request /authentik/auth;
auth_request_set $auth_user $upstream_http_x_authentik_username;
auth_request_set $auth_email $upstream_http_x_authentik_email;
proxy_set_header X-Authentik-Username $auth_user;
proxy_set_header X-Authentik-Email $auth_email;
proxy_pass http://127.0.0.1:4000/ui/;
proxy_http_version 1.1;
proxy_set_header Upgrade $http_upgrade;
proxy_set_header Connection "upgrade";
}
# === LiteLLM SSO callback ===
location /sso/callback {
proxy_pass http://127.0.0.1:4000/sso/callback;
proxy_set_header Host $host;
}
# === API endpoint Bearer token auth ===
location /v1/ {
proxy_pass http://127.0.0.1:4000/v1/;
proxy_set_header Host $host;
proxy_read_timeout 600s;
error_page 502 = @router_fallback;
}
location @router_fallback {
proxy_pass http://127.0.0.1:9000/v1/;
proxy_set_header Host $host;
}
# === Key management API ===
location /key/ {
proxy_pass http://127.0.0.1:4000/key/;
proxy_set_header Host $host;
proxy_set_header Authorization $http_authorization;
}
# Keep router metrics accessible (not behind LiteLLM)
location /router/ {
proxy_pass http://127.0.0.1:9000/;
}
location /health {
proxy_pass http://127.0.0.1:4000/health;
}
```
---
## 7. Docker Compose (`docker-compose.yml` on CT 116)
```yaml
services:
# Layer 1: LiteLLM Gateway (Policy & Admin)
litellm:
image: ghcr.io/berriai/litellm:main-stable
network_mode: "host"
extra_hosts:
- "auth.sysloggh.net:192.168.68.11"
volumes:
- ./config.yaml:/app/config.yaml:ro
- ./custom_sso.py:/app/custom_sso.py:ro
environment:
- LITELLM_MASTER_KEY=${LITELLM_MASTER_KEY}
- DATABASE_URL=postgresql://litellm:***@localhost:5432/litellm
- STORE_MODEL_IN_DB=True
- ROUTER_API_KEY=${ROUTER_API_KEY}
- OPENAI_API_KEY=***
- ANTHROPIC_API_KEY=${ANTH...KEY}
- PROXY_BASE_URL=https://litellm.sysloggh.net
command:
- --config
- /app/config.yaml
- --port
- "4000"
depends_on:
postgres:
condition: service_healthy
restart: unless-stopped
postgres:
image: postgres:16-alpine
network_mode: "host"
environment:
- POSTGRES_DB=litellm
- POSTGRES_USER=litellm
- POSTGRES_PASSWORD=${POST...}
volumes:
- pgdata:/var/lib/postgresql/data
healthcheck:
test: ["CMD-SHELL", "pg_isready -U litellm"]
interval: 5s
timeout: 3s
retries: 5
restart: unless-stopped
nginx:
image: nginx:alpine
extra_hosts:
- "auth.sysloggh.net:192.168.68.11"
volumes:
- ./nginx/nginx.conf:/etc/nginx/nginx.conf:ro
ports:
- "80:80"
restart: unless-stopped
volumes:
pgdata:
```
---
## 8. Risk Mitigation
| Risk | Mitigation |
|------|------------|
| LiteLLM adds latency overhead | Shadow mode measures: <50ms extra is acceptable |
| LiteLLM down = all agents down | NGINX fallback to router :9000 direct (see 6) |
| Explicit GPU saturated no fallback available | LiteLLM fallback chains try all 3 GPUs in order before failing |
| Fallback chain masking real GPU failures | Router returns `saturated: true` only for capacity, `down` returns different error |
| Key sync drift | Single-source: LiteLLM is key authority. Router uses one `ROUTER_API_KEY` |
| Spend tracking inaccurate for local GPUs | `model_cost` per GPU with $0 rate; optional symbolic pricing for internal billing |
| Double rate limiting | Intentional: LiteLLM for per-user caps, Router for hardware protection |
| PostgreSQL failure | LiteLLM can run with SQLite fallback; UI features degrade |
| Per-model metrics accuracy with syslog-auto | `syslog-auto` is opaque by design (content-based routing). Explicit models are accurate. Use explicit models for per-GPU billing. |
---
## 9. Success Metrics
| Metric | Before | After |
|--------|--------|-------|
| Key management | Manual CLI + env vars + redeploy | UI-based, instant, no redeploy |
| Spend visibility | None | Per-agent, per-team, per-model $ tracking |
| Access control | Tier-based (3 levels) | Per-key, per-model, budget-capped |
| New agent onboarding | Generate key, update env var, redeploy router | Create in UI, share key |
| Admin UX | curl + JSON responses | Visual dashboard, graphs, search |
| Audit trail | Router logs (stdout only) | Database-backed with UI search |
| SSO | None | Authentik OIDC |
| Budget enforcement | None | Automatic: key suspended at $limit |
| GPU failover | Silent (inaccurate metrics) | Explicit (LiteLLM fallback chains, per-model logs) |
| GPU routing intelligence | Full (unchanged) | Full (unchanged) |
---
## 10. Migration Commands (Quick Reference)
```bash
# On CT 116 (SSH via minipve: pct exec 116 bash):
# Phase 0: Infrastructure Prep
echo "192.168.68.11 auth.sysloggh.net" >> /etc/hosts
cd /opt/litellm
docker compose up -d postgres
# Replace config.yaml with production version (see 4.3)
docker compose restart litellm
# Verify
curl http://127.0.0.1:4000/health
curl -X POST http://127.0.0.1:4000/v1/chat/completions \
-H "Authorization: Bearer ***" \
-H "Content-Type: application/json" \
-d '{"model":"syslog-auto","messages":[{"role":"user","content":"test"}]}'
```
---
## 11. GitOps Workflow
- `main` production-ready code
- `feature/litellm-migration` current development branch
**Conventional Commits:**
```
feat(plan): add fallback chains and strict passthrough for model identity gap
fix(router): strict passthrough for explicit models no silent rerouting
feat(plan): update LiteLLM migration plan for CT 116 deployment with Authentik OIDC
docs: LiteLLM migration plan two-layer architecture with model identity gap analysis
```
---
## 12. Zero-Downtime Migration Strategy
**Per-Agent Cutover (<2 minutes):**
1. Create LiteLLM virtual key in UI
2. Update agent's `OPENAI_API_BASE` to `:4000`
3. Verify routing works
4. Monitor LiteLLM logs for errors
**Global Rollback:**
1. If LiteLLM :4000 fails, revert all agents to `:9000`
2. NGINX `router_fallback` handles automatic failover
3. Monitor GPU metrics for health checks
---
## Appendix A: Model Identity Gap Analysis (RESOLVED)
### Problem Identified (Abiba, June 14)
The original architecture had a metrics accuracy gap: when an agent requested `qwen3.6-35B-A3B` and MoE was busy, the router silently rerouted to Dense. LiteLLM logged it as MoE usage, corrupting per-model spend/usage tracking.
### Root Cause
```python
# OLD router code (router-fixed.py):
if is_gpu_busy(target) and req in allowed:
alts = [m for m in avail if m != target and m in allowed]
if alts:
alt = select_best_gpu(alts, "explicit", agent)
if alt: return alt # silently changed GPU
```
### Resolution: Strict Passthrough + LiteLLM Fallback Chains
Two changes deployed:
**1. Router strict passthrough:**
```python
# NEW: strict mode no silent fallback
if req != "auto":
target = req if req in avail else avail[0]
if req not in avail:
return {"model": req, "reason": "explicit_unavailable", "saturated": True}
if is_gpu_busy(target):
return {"model": target, "reason": "explicit_saturated", "saturated": True}
return {"model": target, "reason": "explicit"}
```
**2. LiteLLM fallback chains (in config.yaml):**
```yaml
router_settings:
allowed_fails: 100
fallbacks:
- qwen3.6-35B-A3B: ["qwen3.6-27B-code", "gpu-vision"]
- qwen3.6-27B-code: ["qwen3.6-35B-A3B", "gpu-vision"]
- gpu-vision: ["qwen3.6-27B-code", "qwen3.6-35B-A3B"]
```
### Result
| Scenario | Before | After |
|----------|--------|-------|
| Agent asks for MoE, MoE available | MoE used, metrics OK | MoE used, metrics OK |
| Agent asks for MoE, MoE busy | Router Dense silently, metrics WRONG | Router 503, LiteLLM Dense, metrics show BOTH attempts |
| Agent uses syslog-auto | Router picks GPU, LiteLLM sees opaque | Same (syslog-auto is opaque by design) |
| All 3 GPUs saturated | Router queues (30s), then 503 | Same, LiteLLM sees 503 after fallback chain exhausted |
---
## Appendix B: LiteLLM Virtual Key Migration
| Agent | Old Key | New LiteLLM Key | Tier | Budget |
|-------|---------|-----------------|------|--------|
| Abiba | sk-***-*** | sk-litellm-*** | enterprise | $1000 |
| Mumuni | sk-***-*** | sk-litellm-*** | enterprise | $1000 |
| Tanko | sk-***-*** | sk-litellm-*** | enterprise | $1000 |
| Kagenz0 | sk-***-*** | sk-litellm-*** | professional | $500 |
| Koby | sk-***-*** | sk-litellm-*** | professional | $500 |
| Koonimo | sk-***-*** | sk-litellm-*** | professional | $500 |
---
## Appendix C: Authentik SSO Integration
**Authentik Provider Setup:**
1. Create OAuth2 application in Authentik
2. Set redirect URI: `http://<CT-116-IP>/sso/callback`
3. Configure `client_id` and `client_secret`
4. Mount `custom_sso.py` to LiteLLM container
5. Update config.yaml with provider details
---
## Appendix D: Prometheus Monitoring
**Metrics Export:**
- LiteLLM metrics `http://127.0.0.1:4000/metrics`
- Router metrics `http://127.0.0.1:9000/metrics`
- GPU health metrics `http://127.0.0.1:9000/metrics/gpu`
**Alerts:**
- GPU health score > 70 alert
- Circuit breaker trip alert
- LiteLLM spend > $100/day alert
- LiteLLM latency > 1000ms alert
+45 -57
View File
@@ -1,75 +1,63 @@
# syslog-harness — Inference API Harness
# Syslog Harness
CT 116 Docker stack for routing local GPU models through a unified OpenAI-compatible API.
Operational orchestration layer for Syslog's internal AI agents.
## Architecture
```
nginx :80 → router :9000 → GPU backends
├─ strix-moe (Qwen3.6-MoE-35B-A3B) @ 192.168.68.15:8080 [2 slots, 262K ctx]
├─ gpu-dense (Qwen3.8-27B-U) @ 192.168.68.8:8080 [1 slot, 131K ctx]
└─ gpu-vision (Qwen3.5-9B VLM) @ 192.168.68.110:8080 [1 slot, 131K ctx]
Total: 4 concurrent slots
LiteLLM :8081 (fallback) | Dashboard :3000 | Redis :6379 (local)
┌─────────────┐ ┌──────────────┐ ┌─────────────┐
│ Agent │────>│ Nginx │────>│ GPU Pool │
│ (Hermes) │ │ Router │ │ (MoE/Dense)│
└─────────────┘ └──────────────┘ └─────────────┘
│
├──> :8091 Queue Service (Docker)
│
└──> :3001 Dashboard (Docker)
```
## Deploy
## Components
| Service | Port | Container | Purpose |
|---|---|---|---|
| Nginx Router | 8080 | Host | Routes requests to GPU backends |
| Queue Service | 8091 | `syslog-queue` | Enqueues requests when GPUs are down |
| Dashboard | 3001 | `syslog-dashboard` | Observability UI + API |
## GPU Routing
| Header `X-Syslog-Model` | Backend | Model |
|---|---|---|
| (none) / `standard` | amdpve (.15) | qwen3.6-35B-A3B (MoE) |
| `heavy` / `qwen3.5-27B` | llmgpu (.8) | qwen3.5-27B (Dense) |
| `light` / `gemma-4` | ocu_llm (.110) | gemma-4-E4B (Light) |
## Quick Start
```bash
cd /opt/inference-harness
# Build & start
docker compose build
docker compose up -d
# Verify
curl http://localhost:8091/health
curl http://localhost:3001/api/status
```
## Endpoints
## Dashboard
| URL | Purpose |
|-----|---------|
| `/v1/chat/completions` | Inference API (OpenAI-compatible) — **API key required** |
| `/v1/models` | Available models |
| `/` | Dashboard (GPU health, routing, agents, timeseries) |
- **UI:** `http://<host>:8080/dashboard/harness.html`
- **API:** `http://<host>:8080/dashboard/api/status`
## Authentication
## Circuit Breaker
**All `/v1/chat/completions` requests require a valid API key** via `Authorization: Bearer <key>`. Missing or invalid keys return **401 Unauthorized**.
- Rate limit: 10 req/s per IP
- Burst: 20 requests
- Excess returns 503
- Queue fallback on GPU 502/503
## Agent API Keys
## Production Migration
| Agent | Key |
|-------|-----|
| Abiba | `sk-syslog-abiba` |
| Mumuni | `sk-syslog-mumuni` |
| Tanko | `sk-syslog-tanko` |
| Koby | `sk-syslog-koby` |
| Kagenz0 | `sk-syslog-kagenz0` |
| Koonimo | `sk-syslog-koonimo` |
See [MIGRATION_PLAN.md](./MIGRATION_PLAN.md)
## Routing Tiers
| Tier | Trigger | Priority |
|------|---------|----------|
| Lightweight | No system prompt, ≤1 turn, ≤100 words | VLM → MoE → Dense |
| Simple Conv | ≤1000 tokens, ≤4 turns | VLM → MoE → Dense |
| Heavy | >4000 tokens OR >8 turns | Dense → MoE → VLM |
| Default | Everything else | MoE → VLM → Dense |
## Queue
When all GPUs are saturated, requests enter a polling queue (500ms intervals) instead of returning 503 immediately. Timeout: 30s (configurable via `QUEUE_TIMEOUT` env or `X-Queue-Timeout` header).
## Models
| GPU | Model | VRAM | Slots | Context | Best For |
|-----|-------|------|-------|
| Strix Halo | qwen3.6-35B-A3B (MoE) | 65GB | 2 | 262K | General quality |
| RTX 3090 | gpu-dense (Qwen3.8-27B-Uncensored) | 24GB | 1 | 131K | Dense, general |
| RTX 5070 | gpu-vision (VLM) | 12GB | 2 | 262K | Speed, vision |
## Maintenance
Automated cron job runs daily at 3:00 AM UTC (`/opt/inference-harness/maintenance.sh`):
- Cleans Redis timeseries keys >60 days
- Prunes Docker build cache >7 days
- Logs container health and Redis memory
Logs: `/var/log/harness-maintenance.log`
---
*Built for Syslog Solution LLC — Quality over speed.*
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
-356
View File
@@ -1,356 +0,0 @@
"""SyslogAI Harness Dashboard — Modern Design."""
import os, json, time, queue, threading
import requests
from flask import Flask, request, render_template_string, Response, stream_with_context
ROUTER_METRICS = os.environ.get("ROUTER_METRICS_URL", "http://router:9000/metrics")
app = Flask(__name__)
sse_subscribers = []; sse_lock = threading.Lock()
def fetch_state():
try:
r = requests.get(ROUTER_METRICS, timeout=5)
if r.status_code == 200: return r.json()
except Exception: pass
return {"gpus":[],"route_counts":{},"agent_counts":{},"recent":[],"timestamp":time.time()}
def broadcast_loop():
while True:
time.sleep(3)
data = fetch_state(); payload = json.dumps(data)
with sse_lock:
dead = [q for q in sse_subscribers if not q.put(payload)]
for q in dead: sse_subscribers.remove(q)
threading.Thread(target=broadcast_loop, daemon=True).start()
DASHBOARD_HTML = r"""<!DOCTYPE html>
<html lang="en" data-bs-theme="dark">
<head>
<meta charset="UTF-8"><meta name="viewport" content="width=device-width, initial-scale=1.0">
<title>SyslogAI Harness</title>
<link href="https://cdn.jsdelivr.net/npm/bootstrap@5.3.3/dist/css/bootstrap.min.css" rel="stylesheet">
<style>
body { background: #0b0f17; color: #bcc3cd; font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', system-ui, sans-serif; padding: 20px 24px; }
.card { background: #111827; border: 1px solid #1e293b; border-radius: 10px; height: 100%; }
.stat-card { background: #111827; border: 1px solid #1e293b; border-radius: 10px; padding: 18px 20px; text-align: center; }
.stat-value { font-size: 28px; font-weight: 700; line-height: 1.1; }
.stat-label { font-size: 11px; text-transform: uppercase; letter-spacing: 0.6px; color: #64748b; margin-top: 4px; }
.gpu-card { background: #111827; border: 1px solid #1e293b; border-radius: 10px; padding: 16px 18px; height: 100%; }
.gpu-card .title { font-size: 13px; font-weight: 600; color: #e2e8f0; margin-bottom: 12px; display: flex; align-items: center; gap: 8px; }
.gpu-card .status-dot { width: 8px; height: 8px; border-radius: 50%; flex-shrink: 0; }
.gpu-card .row-metric { display: flex; justify-content: space-between; font-size: 12px; padding: 2px 0; }
.gpu-card .row-metric .lbl { color: #64748b; }
.gpu-card .row-metric .val { color: #e2e8f0; font-variant-numeric: tabular-nums; }
.gpu-card .slot-bar { display: flex; gap: 3px; margin-top: 8px; }
.gpu-card .slot-bar .s { flex: 1; height: 5px; border-radius: 2px; background: #1e293b; }
.gpu-card .slot-bar .s.active { background: #38bdf8; }
.chart-card { background: #111827; border: 1px solid #1e293b; border-radius: 10px; padding: 16px 18px; height: 100%; display: flex; flex-direction: column; }
.chart-card .title { font-size: 13px; font-weight: 600; color: #e2e8f0; margin-bottom: 12px; }
.bar-row { margin-bottom: 8px; }
.bar-label { display: flex; justify-content: space-between; font-size: 11px; margin-bottom: 3px; color: #64748b; }
.bar-label .name { color: #cbd5e1; }
.bar-track { height: 5px; background: #1e293b; border-radius: 3px; overflow: hidden; }
.bar-fill { height: 100%; border-radius: 3px; transition: width 0.6s ease; }
.table-custom { font-size: 11px; margin: 0; }
.table-custom th { color: #64748b; font-weight: 500; font-size: 10px; text-transform: uppercase; border-color: #1e293b; padding: 8px 10px; }
.table-custom td { color: #94a3b8; border-color: rgba(30,41,59,0.5); padding: 6px 10px; }
.agent-badge { font-size: 10px; padding: 2px 7px; border-radius: 8px; font-weight: 600; }
.btn-sm-period { font-size: 10px; padding: 3px 10px; border-radius: 6px; border: 1px solid #1e293b; color: #64748b; background: transparent; cursor: pointer; }
.btn-sm-period.active { background: #1d4ed8; color: #fff; border-color: #1d4ed8; }
.ring-label { font-size: 22px; font-weight: 700; }
.ring-sublabel { font-size: 10px; color: #64748b; }
</style>
</head>
<body>
<!-- HEADER -->
<div class="d-flex justify-content-between align-items-center mb-4">
<div>
<h5 class="mb-0 text-white fw-bold">&#x26A1; SyslogAI Harness</h5>
<div class="small text-secondary" id="live-indicator">
<span class="status-dot" id="live-dot" style="width:6px;height:6px;border-radius:50%;display:inline-block;background:#22c55e;animation:pulse 2s infinite"></span>
<span id="connection-status">live</span> &middot; <span id="update-time"></span>
</div>
</div>
<div class="d-flex gap-2">
<div class="stat-card" style="min-width:100px"><div class="stat-value text-info" id="kpi-total">0</div><div class="stat-label">Requests</div></div>
<div class="stat-card" style="min-width:100px"><div class="stat-value text-warning" id="kpi-active">0</div><div class="stat-label">Active</div></div>
<div class="stat-card" style="min-width:100px"><div class="stat-value" style="color:#a78bfa" id="kpi-agents">0</div><div class="stat-label">Agents</div></div>
</div>
</div>
<div class="row g-3 align-items-stretch">
<!-- ROW 1: Usage Chart (8) + GPU Metrics (4) -->
<div class="col-md-8"><div class="chart-card"><div class="title d-flex justify-content-between align-items-center">
<span>Usage Over Time</span>
<div class="d-flex gap-1">
<button class="btn-sm-period active" onclick="switchPeriod('day')">24h</button>
<button class="btn-sm-period" onclick="switchPeriod('week')">7d</button>
<button class="btn-sm-period" onclick="switchPeriod('month')">30d</button>
</div>
</div><div id="timeseries-chart" style="height:150px"></div><div id="timeseries-legend" class="d-flex justify-content-center gap-3 mt-2 flex-wrap small"></div></div></div>
<div class="col-md-4"><div class="chart-card"><div class="title">GPU Metrics</div><div id="gpu-metrics-card"></div></div></div>
<!-- ROW 2: 3 GPU Cards -->
<div class="col-md-4"><div class="gpu-card" id="gpu-moe"><div class="text-secondary small">Loading...</div></div></div>
<div class="col-md-4"><div class="gpu-card" id="gpu-dense"><div class="text-secondary small">Loading...</div></div></div>
<div class="col-md-4"><div class="gpu-card" id="gpu-light"><div class="text-secondary small">Loading...</div></div></div>
<!-- ROW 3: Queue + Model + Agent -->
<div class="col-md-4"><div class="chart-card"><div class="title">Queue Status</div><div class="text-center" id="queue-viz"></div></div></div>
<div class="col-md-4"><div class="chart-card"><div class="title">Model Distribution</div><div id="route-bars"></div></div></div>
<div class="col-md-4"><div class="chart-card"><div class="title">Agent Activity</div><div id="agent-bars"></div></div></div>
<!-- ROW 4: Performance Analytics -->
<div class="col-12 mb-2"><div class="d-flex align-items-center gap-2"><span class="fw-bold text-white" style="font-size:14px">&#x1F4CA; Performance Analytics</span>
<div class="d-flex gap-1 ms-auto">
<button class="btn-sm-period active" onclick="switchPerfWindow('1')">1h</button>
<button class="btn-sm-period" onclick="switchPerfWindow('24')">24h</button>
</div>
</div></div>
<div class="col-md-6"><div class="chart-card"><div class="title">Latency — P50 / P95 / P99 (ms)</div><div id="perf-latency"></div></div></div>
<div class="col-md-6"><div class="chart-card"><div class="title">Throughput — Tokens / sec</div><div id="perf-throughput"></div></div></div>
<div class="col-md-6"><div class="chart-card"><div class="title">Routing Effectiveness — by Reason</div><div id="perf-reasons"></div></div></div>
<div class="col-md-6"><div class="chart-card"><div class="title">Agent Performance</div><div id="perf-agents"></div></div></div>
<!-- ROW 5: Latency vs Context Scatter -->
<div class="col-12"><div class="chart-card"><div class="title d-flex justify-content-between align-items-center">
<span>Latency vs Prompt Size — by Model</span>
<div class="d-flex gap-2">
<select id="scatter-model" onchange="loadScatter()" style="font-size:10px;background:#1e293b;color:#94a3b8;border:1px solid #334155;border-radius:4px;padding:2px 6px">
<option value="all">All Models</option>
<option value="qwen3.5-9b-vlm">9B VLM</option>
<option value="qwen3.6-27B-code">27B Dense</option>
<option value="qwen3.6-35B-A3B">35B MoE</option>
</select>
</div>
</div><div id="scatter-plot" style="height:200px;position:relative"></div><div id="scatter-legend" class="d-flex justify-content-center gap-3 mt-2 flex-wrap small"></div></div></div>
<!-- ROW 6: Live Stream -->
<div class="col-12"><div class="chart-card"><div class="title">Live Stream</div>
<div class="table-responsive"><table class="table table-custom mb-0">
<thead><tr><th>Time</th><th>Agent</th><th>Model</th><th>Reason</th><th>Tier</th></tr></thead>
<tbody id="route-tbody"></tbody>
</table></div>
</div></div>
</div>
<script>
var MC={'qwen3.5-9b-vlm':'#22c55e','qwen3.6-27B-code':'#f59e0b','qwen3.6-35B-A3B':'#a78bfa'};
var ML={'qwen3.5-9b-vlm':'Qwen3.5 9B VLM','qwen3.6-27B-code':'Qwen Code','qwen3.6-35B-A3B':'Qwen MoE'};
var GL={'qwen3.6-35B-A3B':'MoE - Strix Halo','qwen3.6-27B-code':'Dense - RTX 3090','qwen3.5-9b-vlm':'VLM - RTX 5070'};
function $(id){return document.getElementById(id);}
function render(data){
if(!data||!data.gpus)return;
var t=Object.values(data.route_counts||{}).reduce((a,b)=>a+b,0);
var ta=0,tm=0;data.gpus.forEach(function(g){ta+=(g.active_requests||0);tm+=(g.max_concurrent||1)});
$('kpi-total').textContent=t;$('kpi-active').textContent=ta+'/'+tm;$('kpi-agents').textContent=Object.keys(data.agent_counts||{}).length;
$('update-time').textContent=new Date().toLocaleTimeString();
var ids={'qwen3.6-35B-A3B':'gpu-moe','qwen3.6-27B-code':'gpu-dense','qwen3.5-9b-vlm':'gpu-light'};
data.gpus.forEach(function(g){
var el=$(ids[g.id]);if(!el)return;
var a=g.active_requests||0,mx=g.max_concurrent||1;
var sc=g.status==='healthy'?'#22c55e':g.status==='saturated'?'#f59e0b':'#ef4444';
var ss=g.status==='healthy'?'Online':g.status==='saturated'?'Busy':'Offline';
var slots='';for(var i=0;i<mx;i++)slots+='<span class=\"s'+(i<a?' active':'')+'\"></span>';
var h='<div class=\"title\"><span class=\"status-dot\" style=\"background:'+sc+'\"></span>'+GL[g.id]+'<span class=\"ms-auto small\" style=\"color:'+sc+'\">'+ss+'</span></div>';
h+='<div class=\"row-metric\"><span class=\"lbl\">VRAM</span><span class=\"val\">'+g.vram_used_mb+' / '+g.vram_total_mb+' MB</span></div>';
h+='<div class=\"row-metric\"><span class=\"lbl\">Utilization</span><span class=\"val\">'+g.gpu_util_pct+'%</span></div>';
h+='<div class=\"row-metric\"><span class=\"lbl\">Temperature</span><span class=\"val\" style=\"color:'+(g.temp_c>85?'#ef4444':g.temp_c>70?'#f59e0b':'#22c55e')+'\">'+g.temp_c+'C</span></div>';
if(g.power_w)h+='<div class=\"row-metric\"><span class=\"lbl\">Power</span><span class=\"val\">'+g.power_w+'W'+(g.power_limit_w?'/'+g.power_limit_w+'W':'')+'</span></div>';
h+='<div class=\"row-metric\"><span class=\"lbl\">Slots</span><span class=\"val\" style=\"color:'+(a>=mx?'#ef4444':'#e2e8f0')+'\">'+a+' / '+mx+'</span></div>';
h+='<div class=\"slot-bar\">'+slots+'</div>';el.innerHTML=h;
});
renderQueue(data);renderGPUMetrics(data);
var rc=data.route_counts||{},mr=Math.max(1,...Object.values(rc));
$('route-bars').innerHTML=Object.entries(rc).length?Object.entries(rc).sort((a,b)=>b[1]-a[1]).map(function(e){var m=e[0],c=e[1];return'<div class=\"bar-row\"><div class=\"bar-label\"><span class=\"name\">'+(ML[m]||m)+'</span><span>'+c+' ('+(t?Math.round(c/t*100):0)+'%)</span></div><div class=\"bar-track\"><div class=\"bar-fill\" style=\"width:'+(c/mr*100)+'%;background:'+(MC[m]||'#38bdf8')+'\"></div></div></div>';}).join(''):'<div class=\"text-secondary small\">-</div>';
var ac=data.agent_counts||{},ma=Math.max(1,...Object.values(ac));
$('agent-bars').innerHTML=Object.entries(ac).length?Object.entries(ac).sort((a,b)=>b[1]-a[1]).map(function(e){return'<div class=\"bar-row\"><div class=\"bar-label\"><span class=\"name\">'+e[0]+'</span><span>'+e[1]+'</span></div><div class=\"bar-track\"><div class=\"bar-fill\" style=\"width:'+(e[1]/ma*100)+'%;background:#38bdf8\"></div></div></div>';}).join(''):'<div class=\"text-secondary small\">-</div>';
var recent=data.recent||[];
$('route-tbody').innerHTML=recent.length?recent.slice(0,20).map(function(r){var d=new Date(r.ts*1000),ag=r.agent||'?';return'<tr><td class=\"text-secondary\">'+d.toLocaleTimeString()+'</td><td><span class=\"agent-badge\" style=\"background:rgba(56,189,248,0.12);color:#38bdf8\">'+ag+'</span></td><td>'+(ML[r.model]||r.model)+'</td><td class=\"text-secondary\">'+(r.reason||'')+'</td><td class=\"text-uppercase\" style=\"font-size:10px;color:'+(r.tier==='enterprise'?'#a78bfa':'#64748b')+'\">'+(r.tier||'')+'</td></tr>';}).join(''):'<tr><td colspan=\"5\" class=\"text-secondary\">Waiting...</td></tr>';
}
function renderQueue(data){
var el=$('queue-viz');if(!el)return;
var ta=0,tm=0;data.gpus.forEach(function(g){ta+=(g.active_requests||0);tm+=(g.max_concurrent||1)});
var pct=tm>0?Math.round(ta/tm*100):0,st=pct>=100?'SATURATED':pct>=50?'BUSY':'IDLE';
var sc=pct>=100?'#ef4444':pct>=50?'#f59e0b':'#22c55e';
var circ=188.5,dash=(pct/100)*circ;
var h='<div class=\"d-inline-block position-relative mb-2\"><svg width=\"72\" height=\"72\"><circle cx=\"36\" cy=\"36\" r=\"30\" fill=\"none\" stroke=\"#1e293b\" stroke-width=\"6\"/><circle cx=\"36\" cy=\"36\" r=\"30\" fill=\"none\" stroke=\"'+sc+'\" stroke-width=\"6\" stroke-dasharray=\"'+dash+' '+(circ-dash)+'\" stroke-linecap=\"round\" transform=\"rotate(-90 36 36)\"/></svg><div style=\"position:absolute;top:50%;left:50%;transform:translate(-50%,-50%);text-align:center\"><div class=\"ring-label\" style=\"color:'+sc+'\">'+ta+'</div><div class=\"ring-sublabel\">/ '+tm+' slots</div></div></div>';
h+='<div class=\"fw-bold mb-2 small\" style=\"color:'+sc+'\">'+st+'</div>';
var lb={'qwen3.6-35B-A3B':'MoE','qwen3.6-27B-code':'Dense','qwen3.5-9b-vlm':'VLM'};
data.gpus.forEach(function(g){var a=g.active_requests||0,mx=g.max_concurrent||1,gp=mx>0?Math.round(a/mx*100):0;h+='<div class=\"d-flex align-items-center gap-2 mb-1 justify-content-center\"><span class=\"small\" style=\"min-width:32px;text-align:right;font-size:10px\">'+(lb[g.id]||g.id)+'</span><div style=\"flex:1;max-width:70px;height:3px;background:#1e293b;border-radius:2px;overflow:hidden\"><div style=\"height:100%;width:'+gp+'%;background:'+sc+';border-radius:2px\"></div></div><span class=\"small\" style=\"min-width:22px;font-size:10px\">'+a+'/'+mx+'</span></div>'});
el.innerHTML=h;
}
function renderGPUMetrics(data){
var el=$('gpu-metrics-card');if(!el)return;
var lb={'qwen3.6-35B-A3B':'MoE','qwen3.6-27B-code':'Dense','qwen3.5-9b-vlm':'VLM'};
var h='';data.gpus.forEach(function(g){
var nm=lb[g.id]||g.id,tp=g.temp_c||0,ut=g.gpu_util_pct||0,pw=g.power_w||0,pl=g.power_limit_w||0;
var tc=tp>85?'#ef4444':tp>70?'#f59e0b':'#22c55e',uc=ut>90?'#ef4444':ut>70?'#f59e0b':'#22c55e';
h+='<div class=\"mb-3\"><div class=\"fw-bold small text-white-50 mb-1\">'+nm+'</div>';
h+='<div class=\"d-flex align-items-center gap-2 mb-1\"><span class=\"small text-secondary\" style=\"min-width:30px\">T</span><div class=\"flex-grow-1\" style=\"height:3px;background:#1e293b;border-radius:2px;overflow:hidden\"><div style=\"height:100%;width:'+Math.min(tp,100)+'%;background:'+tc+';border-radius:2px\"></div></div><span class=\"small\" style=\"color:'+tc+';min-width:30px;text-align:right\">'+tp+'C</span></div>';
h+='<div class=\"d-flex align-items-center gap-2 mb-1\"><span class=\"small text-secondary\" style=\"min-width:30px\">U</span><div class=\"flex-grow-1\" style=\"height:3px;background:#1e293b;border-radius:2px;overflow:hidden\"><div style=\"height:100%;width:'+ut+'%;background:'+uc+';border-radius:2px\"></div></div><span class=\"small\" style=\"color:'+uc+';min-width:30px;text-align:right\">'+ut+'%</span></div>';
if(pw>0){var pp=pl>0?Math.round(pw/pl*100):0,pc=pp>90?'#ef4444':pp>70?'#f59e0b':'#22c55e';h+='<div class=\"d-flex align-items-center gap-2\"><span class=\"small text-secondary\" style=\"min-width:30px\">P</span><div class=\"flex-grow-1\" style=\"height:3px;background:#1e293b;border-radius:2px;overflow:hidden\"><div style=\"height:100%;width:'+pp+'%;background:'+pc+';border-radius:2px\"></div></div><span class=\"small\" style=\"color:'+pc+';min-width:30px;text-align:right\">'+pw+'W</span></div>';}
h+='</div>';});
el.innerHTML=h;
}
var cp='day';
function switchPeriod(p){cp=p;document.querySelectorAll('.btn-sm-period').forEach(function(b){b.classList.remove('active')});event.target.classList.add('active');loadTS();}
function loadTS(){fetch('/api/timeseries?period='+cp).then(function(r){return r.json()}).then(renderTS).catch(function(){})}
function renderTS(d){
var models=d.models||{},labels=d.labels||[];
if(!labels.length)return;
var cn=$('timeseries-chart'),lg=$('timeseries-legend'),mn=Object.keys(models);
if(!mn.length){cn.innerHTML='<div class=\"text-secondary small text-center py-4\">-</div>';return;}
var mv=1;for(var m in models)for(var i=0;i<models[m].length;i++)if(models[m][i]>mv)mv=models[m][i];mv=Math.ceil(mv*1.15)||1;
var W=labels.length>1?100/(labels.length-1):100,H=130;
var paths='';for(var mi=0;mi<mn.length;mi++){var m=mn[mi],vals=models[m]||[],d='';for(var i=0;i<vals.length;i++){var x=i*W,y=H-(vals[i]/mv)*H;d+=(i===0?'M':'L')+x.toFixed(1)+','+y.toFixed(1)+' ';}paths+='<path d=\"'+d+'\" fill=\"none\" stroke=\"'+(MC[m]||'#38bdf8')+'\" stroke-width=\"2\" stroke-linecap=\"round\" opacity=\"0.8\"/>';}
var grid='';for(var g=0;g<=4;g++){var y=(g/4)*H;grid+='<line x1=\"0\" y1=\"'+y.toFixed(1)+'\" x2=\"100\" y2=\"'+y.toFixed(1)+'\" stroke=\"#1e293b\" stroke-width=\"1\"/>';}
cn.innerHTML='<svg viewBox=\"0 0 100 '+(H+16)+'\" style=\"width:100%;height:'+(H+20)+'px;display:block\" preserveAspectRatio=\"none\">'+grid+paths+'</svg>';
lg.innerHTML=mn.map(function(m){return'<span class=\"d-flex align-items-center gap-1\"><svg width=\"14\" height=\"8\"><line x1=\"0\" y1=\"4\" x2=\"14\" y2=\"4\" stroke=\"'+(MC[m]||'#38bdf8')+'\" stroke-width=\"2\"/></svg>'+(ML[m]||m)+'</span>';}).join('');
}
var perfWindow='24';
function switchPerfWindow(w){perfWindow=w;document.querySelectorAll('.btn-sm-period').forEach(function(b,i){if(i>=4)b.classList.toggle('active',b.textContent.trim().replace('h','')===w)});loadPerf();}
function loadPerf(){fetch('/api/performance?window='+perfWindow).then(function(r){return r.json()}).then(renderPerf).catch(function(){})}
function renderPerf(d){
var models=d.models||[],reasons=d.reasons||[],agents=d.agents||[],sum=d.summary||{};
// Latency bars: p50/p95/p99 per model
var mlab={'qwen3.6-35B-A3B':'35B MoE','qwen3.6-27B-code':'27B Dense','qwen3.5-9b-vlm':'9B VLM'};
var mcol={'qwen3.6-35B-A3B':'#a78bfa','qwen3.6-27B-code':'#f59e0b','qwen3.5-9b-vlm':'#22c55e'};
if(!models.length){$('perf-latency').innerHTML='<div class="text-secondary small text-center py-4">Accumulating data...</div>';return;}
var maxLat=Math.max(...models.map(function(m){return m.latency.p99||0}),1);
var latHTML=models.map(function(m){
var l=m.latency||{},p50=l.p50||0,p95=l.p95||0,p99=l.p99||0,c=mcol[m.model]||'#38bdf8';
return'<div class="mb-2" style="font-size:11px"><div class="d-flex justify-content-between mb-1"><span style="color:#e2e8f0">'+mlab[m.model]+'</span><span class="text-secondary">'+m.count+' reqs</span></div>'+
'<div class="d-flex align-items-center gap-2 mb-1"><span class="text-secondary" style="min-width:28px">p50</span><div class="flex-grow-1" style="height:14px;background:#1e293b;border-radius:4px;overflow:hidden;position:relative"><div style="position:absolute;left:0;top:0;height:100%;width:'+(p50/maxLat*100)+'%;background:'+c+';opacity:0.3;border-radius:4px"></div><div style="position:absolute;left:0;top:0;height:100%;width:'+(p95/maxLat*100)+'%;background:'+c+';opacity:0.5;border-radius:4px"></div><div style="position:absolute;left:0;top:0;height:100%;width:'+(p99/maxLat*100)+'%;background:'+c+';border-radius:4px"></div></div><span style="color:'+c+';min-width:48px;text-align:right;font-variant-numeric:tabular-nums">'+p99+'ms</span></div>'+
'<div class="d-flex gap-3" style="font-size:10px;color:#64748b;padding-left:32px"><span>p50: '+p50+'ms</span><span>p95: '+p95+'ms</span><span>p99: '+p99+'ms</span></div></div>';
}).join('');
$('perf-latency').innerHTML=latHTML;
// Throughput comparison
var maxTps=Math.max(...models.map(function(m){return m.throughput.avg_tokens_per_sec||0}),1);
var tpsHTML=models.map(function(m){
var t=m.throughput||{},avg=t.avg_tokens_per_sec||0,p50=t.p50||0,c=mcol[m.model]||'#38bdf8';
var isAllStreaming = avg===0 && p50===0;
if(isAllStreaming){
return'<div class="mb-2" style="font-size:11px"><div class="d-flex justify-content-between mb-1"><span style="color:#e2e8f0">'+mlab[m.model]+'</span><span style="color:#64748b;font-style:italic">streaming only</span></div><div class="text-secondary" style="font-size:10px">t/s available for non-streaming requests only</div></div>';
}
return'<div class="mb-2" style="font-size:11px"><div class="d-flex justify-content-between mb-1"><span style="color:#e2e8f0">'+mlab[m.model]+'</span><span style="color:'+c+'" class="fw-bold">'+avg+' tok/s</span></div>'+
'<div class="d-flex align-items-center gap-2"><span class="text-secondary" style="min-width:28px">avg</span><div class="flex-grow-1" style="height:6px;background:#1e293b;border-radius:3px;overflow:hidden"><div style="height:100%;width:'+(Math.max(avg/maxTps*100,6))+'%;background:'+c+';border-radius:3px"></div></div><span class="small" style="color:'+c+';min-width:54px;text-align:right">'+avg+' tok/s</span></div>'+
'<div class="d-flex align-items-center gap-2 mt-1"><span class="text-secondary" style="min-width:28px;font-size:10px">p50</span><div class="flex-grow-1" style="height:4px;background:#1e293b;border-radius:2px;overflow:hidden"><div style="height:100%;width:'+(Math.max(p50/maxTps*100,4))+'%;background:'+c+';opacity:0.5;border-radius:2px"></div></div><span style="font-size:10px;color:#64748b">'+p50+' tok/s</span></div></div>';
}).join('');
$('perf-throughput').innerHTML=tpsHTML;
// Routing reasons table
if(reasons.length){
var rHTML='<table class="table table-custom mb-0"><thead><tr><th>Reason</th><th>Count</th><th>Avg Lat</th><th>P95 Lat</th></tr></thead><tbody>';
reasons.forEach(function(r){rHTML+='<tr><td>'+r.reason+'</td><td>'+r.count+'</td><td>'+r.avg_total_ms+'ms</td><td>'+r.p95_total_ms+'ms</td></tr>';});
rHTML+='</tbody></table>';$('perf-reasons').innerHTML=rHTML;
}else{$('perf-reasons').innerHTML='<div class="text-secondary small text-center py-3">-</div>';}
// Agent performance
if(agents.length){
var maxAc=Math.max(...agents.map(function(a){return a.count||0}),1);
var aHTML=agents.map(function(a){return'<div class="mb-2" style="font-size:11px"><div class="d-flex justify-content-between mb-1"><span style="color:#e2e8f0">'+a.agent+'</span><span class="text-secondary">'+a.count+' reqs</span></div><div class="d-flex align-items-center gap-2"><div class="flex-grow-1" style="height:4px;background:#1e293b;border-radius:2px;overflow:hidden"><div style="height:100%;width:'+(a.count/maxAc*100)+'%;background:#38bdf8;border-radius:2px"></div></div><span class="small" style="color:#38bdf8;min-width:60px;text-align:right">'+a.avg_total_ms+'ms avg</span></div></div>';}).join('');
$('perf-agents').innerHTML=aHTML;
}else{$('perf-agents').innerHTML='<div class="text-secondary small text-center py-3">-</div>';}
}
function poll(){fetch('/api/state').then(function(r){return r.json()}).then(function(data){render(data);$('connection-status').textContent='live';}).catch(function(){$('connection-status').textContent='reconnecting';});}
function loadScatter(){
var m=$('scatter-model').value;
fetch('/api/scatter?window=24&model='+m).then(function(r){return r.json()}).then(renderScatter).catch(function(){});
}
function renderScatter(d){
var pts=d.points||[],el=$('scatter-plot'),lg=$('scatter-legend');
if(!pts.length){el.innerHTML='<div class="text-secondary small text-center py-5">No data yet</div>';return;}
var mcol={'qwen3.6-35B-A3B':'#a78bfa','qwen3.6-27B-code':'#f59e0b','qwen3.5-9b-vlm':'#22c55e','unknown':'#38bdf8'};
var mlab={'qwen3.6-35B-A3B':'35B MoE','qwen3.6-27B-code':'27B Dense','qwen3.5-9b-vlm':'9B VLM'};
var maxX=Math.max.apply(null,pts.map(function(p){return p.prompt_tokens||0}))||1000;
var maxY=Math.max.apply(null,pts.map(function(p){return p.inference_ms||0}))||5000;
// Log scale for X axis (prompt tokens vary widely)
var toX=function(t){return Math.log10(Math.max(t,1))/Math.log10(Math.max(maxX,10))*100;};
var toY=function(t){return (t/maxY)*100;};
var dots='';
pts.forEach(function(p){
var x=toX(p.prompt_tokens),y=toY(p.inference_ms),c=mcol[p.model]||'#38bdf8';
var r=p.stream?1.5:2.5,o=p.stream?0.4:0.8;
dots+='<circle cx="'+x+'" cy="'+(100-y)+'" r="'+r+'" fill="'+c+'" opacity="'+o+'"><title>'+mlab[p.model]+' | '+p.prompt_tokens+' tok | '+p.inference_ms+'ms | '+p.agent+'</title></circle>';
});
// Grid lines
var grid='';
for(var i=1;i<=4;i++){grid+='<line x1="0" y1="'+(i*20)+'" x2="100" y2="'+(i*20)+'" stroke="#1e293b" stroke-width="0.5"/>';}
for(var i=1;i<=4;i++){grid+='<line x1="'+(i*20)+'" y1="0" x2="'+(i*20)+'" y2="100" stroke="#1e293b" stroke-width="0.5"/>';}
// Axis labels
var xTicks='';
var xVals=[10,100,1000,10000,100000];
xVals.forEach(function(v){if(v<=maxX)xTicks+='<text x="'+toX(v)+'" y="103" text-anchor="middle" font-size="8" fill="#64748b">'+(v>=1000?(v/1000)+'k':v)+'</text>';});
var yTicks='';
var yVals=[500,1000,5000,10000,50000,100000];
yVals.forEach(function(v){if(v<=maxY)yTicks+='<text x="-2" y="'+(97-toY(v))+'" text-anchor="end" font-size="8" fill="#64748b">'+(v>=1000?(v/1000)+'s':v+'ms')+'</text>';});
el.innerHTML='<svg viewBox="-35 0 140 115" style="width:100%;height:200px">'+grid+dots+xTicks+yTicks+'<text x="50" y="112" text-anchor="middle" font-size="9" fill="#475569">Prompt Tokens (log scale)</text><text x="-38" y="50" text-anchor="middle" font-size="9" fill="#475569" transform="rotate(-90,-38,50)">Inference Time</text></svg>';
// Legend
var models=[];pts.forEach(function(p){if(models.indexOf(p.model)===-1)models.push(p.model);});
lg.innerHTML=models.map(function(m){return'<span class="d-flex align-items-center gap-1 small"><svg width="10" height="10"><circle cx="5" cy="5" r="3.5" fill="'+(mcol[m]||'#38bdf8')+'"/></svg>'+mlab[m]+'</span>';}).join('');
}
poll();setInterval(poll,3000);loadTS();loadPerf();setInterval(loadPerf,15000);loadScatter();setInterval(loadScatter,30000);
</script>
</body>
</html>"""
@app.route("/")
def dashboard(): return render_template_string(DASHBOARD_HTML)
@app.route("/api/state")
def api_state(): return fetch_state()
@app.route("/api/scatter")
def api_scatter():
window = request.args.get("window", "24")
model = request.args.get("model", "all")
try:
r = requests.get(f"http://router:9000/metrics/scatter?window={window}&model={model}", timeout=10)
if r.status_code == 200: return r.json()
except Exception: pass
return {"points": [], "count": 0}
@app.route("/api/performance")
def api_performance():
window = request.args.get("window", "24")
model = request.args.get("model", "all")
try:
r = requests.get(f"http://router:9000/metrics/performance?window={window}&model={model}", timeout=10)
if r.status_code == 200: return r.json()
except Exception: pass
return {"models": [], "reasons": [], "agents": [], "summary": {"total_requests": 0}}
@app.route("/api/timeseries")
def api_timeseries():
period = request.args.get("period", "day")
try:
r = requests.get("http://router:9000/metrics/timeseries?period=" + period, timeout=5)
if r.status_code == 200: return r.json()
except Exception: pass
return {"models": {}, "labels": []}
@app.route("/api/stream")
def api_stream():
def ev():
q = queue.Queue()
with sse_lock: sse_subscribers.append(q)
try:
yield "data: "+json.dumps(fetch_state())+"\n\n"
while True:
try: msg = q.get(timeout=3); yield "data: "+msg+"\n\n"
except queue.Empty: yield "data: "+json.dumps(fetch_state())+"\n\n"
except GeneratorExit: pass
finally:
with sse_lock:
if q in sse_subscribers: sse_subscribers.remove(q)
return Response(stream_with_context(ev()), mimetype="text/event-stream", headers={"Cache-Control":"no-cache","X-Accel-Buffering":"no","Access-Control-Allow-Origin":"*"})
@app.route("/health")
def health(): return {"status":"healthy","service":"harness-dashboard"}
if __name__ == "__main__":
app.run(host="0.0.0.0", port=3000, debug=False)
-659
View File
@@ -1,659 +0,0 @@
import os, json, time, logging, traceback, threading, queue, statistics, math
import requests, redis
from flask import Flask, request, jsonify, Response, stream_with_context
REDIS_URL = os.environ.get("REDIS_URL", "redis://redis:6379")
GPU_MOE_URL = os.environ.get("GPU_MOE_URL", "http://192.168.68.15:8080/v1")
GPU_DENSE_URL = os.environ.get("GPU_DENSE_URL", "http://192.168.68.8:8080/v1")
GPU_LIGHT_URL = os.environ.get("GPU_LIGHT_URL", "http://192.168.68.110:8080/v1")
GPU_SIDECARS = {
"qwen3.6-35B-A3B": "http://192.168.68.15:8090",
"qwen3.6-27B-code": "http://192.168.68.8:8090",
"qwen3.5-9b-vlm": "http://192.168.68.110:8090",
}
GPU_URLS = {
"qwen3.6-35B-A3B": GPU_MOE_URL,
"qwen3.6-27B-code": GPU_DENSE_URL,
"qwen3.5-9b-vlm": GPU_LIGHT_URL,
}
# Max concurrent requests per GPU (based on llama.cpp --parallel)
GPU_MAX_CONCURRENT = {
"qwen3.6-35B-A3B": 2, # 2 slots (cross-agent spread prevents overheating)
"qwen3.6-27B-code": 2, # 2 slots (128K context frees VRAM)
"qwen3.5-9b-vlm": 2, # 2 slots (12GB VRAM, 4GB headroom)
}
# Context window sizes (tokens) — used for compaction signals
GPU_CONTEXT = {
"qwen3.6-35B-A3B": 262144,
"qwen3.6-27B-code": 131072,
"qwen3.5-9b-vlm": 262144,
}
TIER_MODELS = {
"starter": ["qwen3.5-9b-vlm"],
"professional": ["qwen3.6-35B-A3B", "qwen3.6-27B-code", "qwen3.5-9b-vlm"],
"enterprise": ["qwen3.6-35B-A3B", "qwen3.6-27B-code", "qwen3.5-9b-vlm"],
}
API_KEYS = {
"sk-syslog-local-master-key": {"tier": "enterprise", "agent": "admin"},
"sk-syslog-abiba": {"tier": "enterprise", "agent": "Abiba"},
"sk-syslog-mumuni": {"tier": "enterprise", "agent": "Mumuni"},
"sk-syslog-tanko": {"tier": "enterprise", "agent": "Tanko"},
"sk-syslog-koby": {"tier": "enterprise", "agent": "Koby"},
"sk-syslog-kagenz0": {"tier": "enterprise", "agent": "Kagenz0"},
"sk-syslog-koonimo": {"tier": "enterprise", "agent": "Koonimo"},
"sk-starter-abc123": {"tier": "starter", "agent": "test-starter"},
"sk-professional-xyz789": {"tier": "professional", "agent": "test-pro"},
}
logging.basicConfig(level=logging.INFO, format="%(asctime)s [ROUTER] %(levelname)s %(message)s")
log = logging.getLogger("router")
try: r = redis.from_url(REDIS_URL, decode_responses=True); r.ping()
except Exception: r = None
def counter_audit_loop():
"""Every 30s, check GPU slots and reset counters if all slots idle."""
while True:
time.sleep(30)
if not r: continue
for model, url in GPU_URLS.items():
try:
resp = requests.get(url.replace("/v1","") + "/slots",
headers={"Authorization": "Bearer not-needed"}, timeout=5)
if resp.status_code == 200:
slots = resp.json()
all_idle = all(not s.get("is_processing", False) for s in slots)
if all_idle:
current = int(r.get("active:" + model) or 0)
if current > 0:
r.set("active:" + model, 0)
log.info("AUDIT: Reset stuck counter for %s (was %d)", model, current)
except Exception:
pass
threading.Thread(target=counter_audit_loop, daemon=True).start()
app = Flask(__name__)
sse_subscribers = []; sse_lock = threading.Lock()
def gpu_active_count(model):
"""Get number of in-flight requests for a GPU."""
if r:
return int(r.get("active:" + model) or 0)
return 0
def gpu_incr(model):
if r: r.incr("active:" + model)
def gpu_decr(model):
if r:
v = r.decr("active:" + model)
if v and int(v) < 0:
r.set("active:" + model, 0) # never go negative
def check_gpu_health(model, sidecar_timeout=5, gpu_timeout=3):
url = GPU_SIDECARS.get(model)
if not url: return {"status": "unknown"}
try:
resp = requests.get(url, timeout=sidecar_timeout)
if resp.status_code == 200:
d = resp.json()
pct = (d.get("vram_used_mb",0) / max(d.get("vram_total_mb",1), 1)) * 100
status = "healthy" # VRAM usage != saturation; busy slots handled by is_gpu_busy()
vram_warning = pct >= 95
# Also check if llama.cpp endpoint is actually responding
gpu_url = GPU_URLS.get(model, "")
try:
hr = requests.get(gpu_url.replace("/v1","") + "/health", headers={"Authorization": "Bearer not-needed"}, timeout=gpu_timeout)
if hr.status_code != 200:
status = "down"
except Exception:
status = "down"
return {"status": status, "vram_warning": vram_warning, "vram_used_mb": d.get("vram_used_mb"), "vram_total_mb": d.get("vram_total_mb"), "vram_pct": round(pct,1), "temp_c": d.get("temp_c"), "gpu_util_pct": d.get("gpu_util_pct"), "gpu_name": d.get("gpu_name"), "power_w": d.get("power_w"), "power_limit_w": d.get("power_limit_w")}
except Exception: pass
return {"status": "down"}
def available_models(): return [m for m in GPU_URLS if check_gpu_health(m)["status"] in ("healthy","saturated")]
def estimate_tokens(msgs):
"""Estimate token count from messages. Uses JSON length / 3.5 (closer to real tokenizer ratios for dense text)."""
return len(json.dumps(msgs, default=str)) // 3.5
def store_perf_record(model, agent, tier, reason, queue_ms, inference_ms, prompt_tokens, completion_tokens, stream):
"""Store detailed performance record in Redis for analytics."""
if not r: return
try:
total_ms = queue_ms + inference_ms
tps = completion_tokens / (inference_ms / 1000) if inference_ms > 0 and completion_tokens > 0 else 0
rec = json.dumps({
"ts": time.time(),
"model": model, "agent": agent, "tier": tier, "reason": reason,
"queue_ms": round(queue_ms, 1),
"inference_ms": round(inference_ms, 1),
"total_ms": round(total_ms, 1),
"prompt_tokens": prompt_tokens,
"completion_tokens": completion_tokens,
"tokens_per_sec": round(tps, 1),
"stream": stream
})
# Global recent list (last 500)
r.lpush("perf:recent", rec)
r.ltrim("perf:recent", 0, 499)
# Per-model list (last 200)
r.lpush("perf:model:" + model, rec)
r.ltrim("perf:model:" + model, 0, 199)
# Per-reason list (last 200)
r.lpush("perf:reason:" + reason, rec)
r.ltrim("perf:reason:" + reason, 0, 199)
# Per-agent list (last 200)
r.lpush("perf:agent:" + agent, rec)
r.ltrim("perf:agent:" + agent, 0, 199)
except Exception:
pass
def is_gpu_busy(model):
"""Check if GPU is at or near max concurrent capacity."""
active = gpu_active_count(model)
max_c = GPU_MAX_CONCURRENT.get(model, 1)
return active >= max_c
def select_best_gpu(candidates, reason, agent=""):
"""Pick best GPU, spreading agents across GPUs to prevent hotspots."""
# Count how many distinct agents are on each GPU
gpu_agent_counts = {}
if r:
for m in GPU_URLS:
count = 0
for ak in API_KEYS.values():
if r.get("agent_gpu:" + ak["agent"] + ":" + m):
count += 1
gpu_agent_counts[m] = count
# First pass: prefer GPUs with 0 other agents (fresh GPU for this agent)
for m in candidates:
if not is_gpu_busy(m) and gpu_agent_counts.get(m, 0) == 0:
return {"model": m, "reason": reason}
# Second pass: prefer GPU this agent is NOT already on (skip own GPU)
if agent:
for m in candidates:
if not is_gpu_busy(m) and not r.get("agent_gpu:" + agent + ":" + m):
return {"model": m, "reason": reason}
# Third pass: any non-busy GPU
for m in candidates:
if not is_gpu_busy(m):
return {"model": m, "reason": reason}
# All busy — pick least loaded
best = None
best_load = 999
for m in candidates:
load = gpu_active_count(m)
if load < best_load:
best_load = load
best = m
if best:
return {"model": best, "reason": "load_balanced_" + reason}
return None
def route(rd, tier, agent=""):
msgs = rd.get("messages",[]); t = estimate_tokens(msgs)
sys = any(m.get("role")=="system" for m in msgs)
turns = len([m for m in msgs if m.get("role") in ("user","assistant")])
hints = rd.get("routing_hints",{})
allowed = TIER_MODELS.get(tier, ["qwen3.5-9b-vlm"])
avail = [m for m in available_models() if m in allowed]
if not avail: return {"model": allowed[0], "reason": "all_saturated", "saturated": True}
# Check if all available GPUs are at max capacity
if all(is_gpu_busy(m) for m in avail):
return {"model": avail[0], "reason": "all_saturated", "saturated": True}
req = rd.get("model","auto")
if req != "auto":
target = req if req in avail else avail[0]
# If explicit model is busy, check if another can take it
if is_gpu_busy(target) and req in allowed:
alts = [m for m in avail if m != target and m in allowed]
if alts:
alt = select_best_gpu(alts, "explicit", agent)
if alt: return alt
return {"model": target, "reason": "explicit"}
if hints:
if hints.get("priority")=="speed" and "qwen3.5-9b-vlm" in avail:
return select_best_gpu(["qwen3.5-9b-vlm"], "hint_speed", agent) or {"model":"qwen3.5-9b-vlm","reason":"hint_speed"}
if hints.get("priority")=="quality" and "qwen3.6-35B-A3B" in avail:
return select_best_gpu(["qwen3.6-35B-A3B"], "hint_quality", agent) or {"model":"qwen3.6-35B-A3B","reason":"hint_quality"}
first_msg = msgs[0].get("content","") if msgs else ""
words = len(first_msg.split()) if isinstance(first_msg, str) else 99
# TIER 1: Lightweight — single-turn short queries → VLM (fastest)
if not sys and turns <= 1 and t <= 500 and words <= 100 and "qwen3.5-9b-vlm" in avail:
if not is_gpu_busy("qwen3.5-9b-vlm"):
return {"model":"qwen3.5-9b-vlm","reason":"lightweight"}
# VLM busy — Dense is faster for short queries than MoE
fallback = [m for m in ["qwen3.6-27B-code","qwen3.6-35B-A3B"] if m in avail]
result = select_best_gpu(fallback, "lightweight_fallback", agent)
if result: return result
# TIER 2: Simple conversations — VLM primary (up to 15K tok), fastest for moderate chat
if t <= 15000 and turns <= 12 and "qwen3.5-9b-vlm" in avail:
if not is_gpu_busy("qwen3.5-9b-vlm"):
return {"model":"qwen3.5-9b-vlm","reason":"simple_conv"}
# VLM busy — fall back to Dense, then MoE
fallback = [m for m in ["qwen3.6-27B-code","qwen3.6-35B-A3B"] if m in avail]
result = select_best_gpu(fallback, "simple_conv_fallback", agent)
if result: return result
# TIER 3: Medium complexity — Dense primary, VLM fallback (quality + speed balance)
if t <= 25000:
candidates = [m for m in ["qwen3.6-27B-code","qwen3.5-9b-vlm","qwen3.6-35B-A3B"] if m in avail]
result = select_best_gpu(candidates, "medium", agent)
if result: return result
# TIER 4: Heavy reasoning — MoE primary (workhorse), Dense fallback
if t > 25000:
candidates = [m for m in ["qwen3.6-35B-A3B","qwen3.6-27B-code","qwen3.5-9b-vlm"] if m in avail]
result = select_best_gpu(candidates, "heavy_reasoning", agent)
if result: return result
# TIER 5: Default — Dense primary, MoE fallback
candidates = [m for m in ["qwen3.6-27B-code","qwen3.5-9b-vlm","qwen3.6-35B-A3B"] if m in avail]
result = select_best_gpu(candidates, "default", agent)
if result: return result
return {"model":avail[0],"reason":"last_resort"}
def clean_unicode(text):
if not isinstance(text, str): return text
text = text.replace(chr(0x2014), "-"); text = text.replace(chr(0x2013), "-")
text = text.replace(chr(0x2018), "'"); text = text.replace(chr(0x2019), "'")
text = text.replace(chr(0x201C), '"'); text = text.replace(chr(0x201D), '"')
text = text.replace(chr(0x2026), "..."); text = text.replace(chr(0x00A0), " ")
return text.encode("ascii", "ignore").decode("ascii")
def clean_response(d):
if isinstance(d, dict): return {k: clean_response(v) for k,v in d.items()}
if isinstance(d, list): return [clean_response(v) for v in d]
if isinstance(d, str): return clean_unicode(d)
return d
def get_metrics():
d = {"gpus":[],"route_counts":{},"agent_counts":{},"tier_counts":{},"recent":[],"timestamp":time.time(),"active_requests":{}}
for m in GPU_URLS:
h = check_gpu_health(m)
d["gpus"].append({"id":m,"gpu_name":h.get("gpu_name",m),"status":h.get("status"),"vram_used_mb":h.get("vram_used_mb"),"vram_total_mb":h.get("vram_total_mb"),"vram_pct":h.get("vram_pct"),"temp_c":h.get("temp_c"),"gpu_util_pct":h.get("gpu_util_pct"),"power_w":h.get("power_w"),"power_limit_w":h.get("power_limit_w"),"active_requests":gpu_active_count(m), "max_concurrent": GPU_MAX_CONCURRENT.get(m, 1)})
d["active_requests"][m] = gpu_active_count(m)
if r:
try:
for m in GPU_URLS: d["route_counts"][m] = int(r.get("routes:"+m) or 0)
for k,v in API_KEYS.items():
c = int(r.get("routes:agent:"+v["agent"]) or 0)
if c>0: d["agent_counts"][v["agent"]] = c
for t in TIER_MODELS: d["tier_counts"][t] = int(r.get("routes:tier:"+t) or 0)
raw = r.lrange("routes:recent",0,49)
d["recent"] = [json.loads(x) for x in raw] if raw else []
except Exception: pass
return d
def bcast():
data = get_metrics(); payload = json.dumps(data)
with sse_lock:
dead = []
for q in sse_subscribers:
try: q.put(payload)
except Exception: dead.append(q)
for q in dead: sse_subscribers.remove(q)
QUEUE_TIMEOUT = int(os.environ.get("QUEUE_TIMEOUT", "30")) # max seconds to queue before 503
@app.route("/v1/chat/completions", methods=["POST"])
def chat():
try:
rd = request.get_json(force=True)
ak = request.headers.get("Authorization","").replace("Bearer ","")
if not ak or ak not in API_KEYS:
log.warning("AUTH_REJECTED: no/invalid API key from %s", request.remote_addr)
return jsonify({"error": "Unauthorized — valid API key required"}), 401
ki = API_KEYS[ak]
tier, agent = ki["tier"], ki["agent"]
# Allow agent to override queue timeout via header
q_timeout = int(request.headers.get("X-Queue-Timeout", str(QUEUE_TIMEOUT)))
# Cross-turn context tracking: accumulate tokens per session
session_id = request.headers.get("X-Session-Id", "")
session_tokens = 0
if session_id and r:
try:
prev = int(r.get("session:" + session_id) or 0)
current = estimate_tokens(rd.get("messages",[]))
session_tokens = max(prev, current) # context only grows
r.set("session:" + session_id, session_tokens, ex=86400) # TTL 24h
except Exception: pass
d = route(rd, tier, agent)
queue_start = time.time()
# Queue loop: wait for a GPU slot instead of immediate 503
while d.get("saturated"):
elapsed = time.time() - queue_start
if elapsed > q_timeout:
resp = jsonify({"error": "All GPUs saturated", "queued_s": round(elapsed,1), "retry_after_s": 5})
resp.headers["Retry-After"] = "5"
log.warning("QUEUE_TIMEOUT: %s waited %.1fs, all GPUs saturated", agent, elapsed)
return resp, 503
time.sleep(0.5) # poll every 500ms
d = route(rd, tier, agent)
queue_ms = (time.time() - queue_start) * 1000
if queue_ms > 500:
log.info("QUEUED: %s waited %.0fms before slot opened", agent, queue_ms)
model, reason, url = d["model"], d["reason"], GPU_URLS[d["model"]]
is_stream = rd.get("stream", False)
gpu_incr(model)
log.info("ROUTE: %s -> %s (%s) stream=%s active=%d/%d", agent, model, reason, is_stream, gpu_active_count(model), GPU_MAX_CONCURRENT.get(model,1))
# Track which GPU this agent is using (TTL 120s covers typical request)
if r and agent:
try: r.setex("agent_gpu:" + agent + ":" + model, 120, "1")
except: pass
if r:
try:
r.incr("routes:"+model); r.incr("routes:tier:"+tier); r.incr("routes:agent:"+agent)
r.incr("ts:"+model+":"+time.strftime("%Y%m%d%H"))
r.lpush("routes:recent", json.dumps({"ts":time.time(),"model":model,"reason":reason,"tier":tier,"agent":agent,"queue_ms": round(queue_ms,1)}))
r.ltrim("routes:recent",0,999)
except Exception: pass
start = time.time()
resp = requests.post(url+"/chat/completions", json=rd,
headers={"Content-Type":"application/json","Authorization":"Bearer not-needed"}, timeout=300, stream=is_stream)
lat = int((time.time()-start)*1000)
gpu_decr(model)
if resp.status_code != 200: return jsonify({"error":"GPU error "+str(resp.status_code)}), 502
if is_stream:
# Buffer SSE chunks, handle split lines for large responses
chunks = []
stream_timings = {}
buf = "" # accumulate partial lines
for raw in resp.iter_content(chunk_size=None, decode_unicode=True):
if raw:
cleaned = clean_unicode(raw)
chunks.append(cleaned)
buf += cleaned
# Process complete lines from buffer
while "\n" in buf:
line, buf = buf.split("\n", 1)
line = line.strip()
if line.startswith("data: ") and not stream_timings:
js = line[6:].strip()
if js.startswith("{") and "timings" in js and "predicted_n" in js:
try:
tj = json.loads(js).get("timings", {})
if tj:
stream_timings = tj
except: pass
# Store perf record with real token counts from stream
if stream_timings:
pt = stream_timings.get("prompt_n", 0)
ct = stream_timings.get("predicted_n", 0)
tps = stream_timings.get("predicted_per_second", 0)
gen_ms = stream_timings.get("predicted_ms", lat)
store_perf_record(model, agent, tier, reason, queue_ms, gen_ms, pt, ct, True)
else:
store_perf_record(model, agent, tier, reason, queue_ms, lat, estimate_tokens(rd.get("messages",[])), 0, True)
# Yield all chunks to client
def gen():
for c in chunks: yield c
bcast()
ctx_remaining = GPU_CONTEXT.get(model, 65536) - max(session_tokens, estimate_tokens(rd.get("messages",[])))
ctx_pct = ctx_remaining / GPU_CONTEXT.get(model, 65536) * 100
ctx_warning = "compact_urgent" if ctx_pct < 5 else ("compact_recommended" if ctx_pct < 15 else ("compact_soon" if ctx_pct < 30 else "ok"))
sse_resp = Response(stream_with_context(gen()), mimetype="text/event-stream")
sse_resp.headers["X-Context-Remaining"] = str(max(0, ctx_remaining))
sse_resp.headers["X-Context-Warning"] = ctx_warning
sse_resp.headers["X-Context-Model"] = model
return sse_resp
data = clean_response(resp.json())
for c in data.get("choices",[]):
msg = c.get("message",{})
if not msg.get("content") and msg.get("reasoning_content"):
msg["content"] = msg["reasoning_content"]
# Extract performance data from llama.cpp response
usage = data.get("usage", {})
timings = data.get("timings", {})
prompt_tokens = usage.get("prompt_tokens", 0)
completion_tokens = usage.get("completion_tokens", 0)
inference_ms = lat # total GPU round-trip
store_perf_record(model, agent, tier, reason, queue_ms, inference_ms, prompt_tokens, completion_tokens, False)
ctx_remaining = GPU_CONTEXT.get(model, 65536) - max(session_tokens, estimate_tokens(rd.get("messages",[])))
ctx_pct = ctx_remaining / GPU_CONTEXT.get(model, 65536) * 100
ctx_warning = "compact_urgent" if ctx_pct < 5 else ("compact_recommended" if ctx_pct < 15 else ("compact_soon" if ctx_pct < 30 else "ok"))
data["routing"] = {"model":model,"reason":reason,"gpu":url,"tier":tier,"agent":agent,"latency_ms":lat,"queue_ms": round(queue_ms,1),"active_gpu":gpu_active_count(model),"context_remaining": max(0, ctx_remaining),"context_pct": round(ctx_pct,1),"context_warning": ctx_warning}
resp = jsonify(data)
resp.headers["X-Context-Remaining"] = str(max(0, ctx_remaining))
resp.headers["X-Context-Warning"] = ctx_warning
resp.headers["X-Context-Model"] = model
bcast()
return resp
except requests.Timeout:
gpu_decr(model)
log.error("TIMEOUT: %s -> %s", agent, model)
return jsonify({"error":"timeout"}), 504
except Exception as e:
gpu_decr(model)
log.error("Error: %s\n%s", e, traceback.format_exc())
return jsonify({"error":str(e)}), 500
@app.route("/metrics/performance")
def performance():
"""Per-request performance analytics with percentiles per model/reason/agent."""
if not r: return jsonify({"error": "Redis unavailable"}), 503
try:
window_hours = int(request.args.get("window", "24"))
model_filter = request.args.get("model", "all")
# Load recent records
cutoff = time.time() - (window_hours * 3600)
raw = r.lrange("perf:recent", 0, -1)
records = []
for x in raw:
try:
rec = json.loads(x)
if rec["ts"] >= cutoff:
records.append(rec)
except: pass
# Filter by model if specified
if model_filter != "all":
records = [r for r in records if r["model"] == model_filter]
if not records:
return jsonify({"models": [], "reasons": [], "agents": [], "summary": {"total_requests": 0}})
def pct(values, p):
if len(values) < 2: return round(values[0], 1) if values else 0
return round(statistics.quantiles(sorted(values), n=100, method='inclusive')[min(p-1, 98)], 1)
# Per-model stats
model_groups = {}
for rec in records:
m = rec["model"]
if m not in model_groups: model_groups[m] = []
model_groups[m].append(rec)
models = []
for m, recs in sorted(model_groups.items()):
latencies = [r["total_ms"] for r in recs]
tps_vals = [r["tokens_per_sec"] for r in recs if r["tokens_per_sec"] > 0]
non_stream = [r for r in recs if not r["stream"]]
queue_times = [r["queue_ms"] for r in non_stream]
models.append({
"model": m,
"count": len(recs),
"stream_pct": round(len([r for r in recs if r["stream"]]) / len(recs) * 100, 1),
"latency": {
"avg": round(statistics.mean(latencies), 1),
"p50": pct(latencies, 50),
"p95": pct(latencies, 95),
"p99": pct(latencies, 99)
},
"throughput": {
"avg_tokens_per_sec": round(statistics.mean(tps_vals), 1) if tps_vals else 0,
"p50": pct(tps_vals, 50) if tps_vals else 0,
"p95": pct(tps_vals, 95) if tps_vals else 0,
},
"queue": {
"avg_ms": round(statistics.mean(queue_times), 1) if queue_times else 0,
"p95_ms": pct(queue_times, 95) if queue_times else 0,
} if queue_times else None
})
# Per-reason stats
reason_groups = {}
for rec in records:
rsn = rec["reason"]
if rsn not in reason_groups: reason_groups[rsn] = []
reason_groups[rsn].append(rec)
reasons = []
for rsn, recs in sorted(reason_groups.items(), key=lambda x: -len(x[1])):
latencies = [r["total_ms"] for r in recs]
reasons.append({
"reason": rsn,
"count": len(recs),
"avg_total_ms": round(statistics.mean(latencies), 1),
"p95_total_ms": pct(latencies, 95)
})
# Per-agent stats
agent_groups = {}
for rec in records:
ag = rec["agent"]
if ag not in agent_groups: agent_groups[ag] = []
agent_groups[ag].append(rec)
agents = []
for ag, recs in sorted(agent_groups.items(), key=lambda x: -len(x[1])):
latencies = [r["total_ms"] for r in recs]
tps_vals = [r["tokens_per_sec"] for r in recs if r["tokens_per_sec"] > 0]
agents.append({
"agent": ag,
"count": len(recs),
"avg_total_ms": round(statistics.mean(latencies), 1),
"avg_tokens_per_sec": round(statistics.mean(tps_vals), 1) if tps_vals else 0
})
all_lat = [r["total_ms"] for r in records]
all_tps = [r["tokens_per_sec"] for r in records if r["tokens_per_sec"] > 0]
summary = {
"total_requests": len(records),
"window_hours": window_hours,
"latency": {
"avg_ms": round(statistics.mean(all_lat), 1),
"p50_ms": pct(all_lat, 50),
"p95_ms": pct(all_lat, 95),
"p99_ms": pct(all_lat, 99)
},
"throughput_avg_tps": round(statistics.mean(all_tps), 1) if all_tps else 0
}
return jsonify({"models": models, "reasons": reasons, "agents": agents, "summary": summary})
except Exception as e:
return jsonify({"error": str(e)}), 500
@app.route("/metrics/scatter")
def scatter():
"""Return individual data points for scatter plots (prompt_tokens vs latency)."""
if not r: return jsonify({"error": "Redis unavailable"}), 503
try:
window_hours = int(request.args.get("window", "24"))
model_filter = request.args.get("model", "all")
cutoff = time.time() - (window_hours * 3600)
raw = r.lrange("perf:recent", 0, -1)
points = []
for x in raw:
try:
rec = json.loads(x)
if rec["ts"] >= cutoff:
if model_filter == "all" or rec["model"] == model_filter:
points.append({
"model": rec["model"],
"agent": rec["agent"],
"reason": rec["reason"],
"prompt_tokens": int(rec.get("prompt_tokens", 0)),
"completion_tokens": rec.get("completion_tokens", 0),
"inference_ms": round(rec["inference_ms"], 1),
"tokens_per_sec": rec.get("tokens_per_sec", 0),
"stream": rec.get("stream", False)
})
except: pass
return jsonify({"points": points, "count": len(points)})
except Exception as e:
return jsonify({"error": str(e)}), 500
@app.route("/v1/models")
def models():
def _h(m): return check_gpu_health(m, sidecar_timeout=1.5, gpu_timeout=1)
return jsonify({"object":"list","data":[{"id":m,"object":"model","owned_by":"syslog","status":_h(m).get("status"),"gpu":_h(m).get("gpu_name")} for m in GPU_URLS]})
@app.route("/health")
def health():
gpus = {}
for m in GPU_URLS:
h = check_gpu_health(m, sidecar_timeout=1.5, gpu_timeout=1)
h["active_requests"] = gpu_active_count(m)
h["max_concurrent"] = GPU_MAX_CONCURRENT.get(m, 1)
gpus[m] = h
return jsonify({"status":"healthy","redis":"connected" if r else "down","gpus":gpus,"available_models":available_models()})
@app.route("/metrics")
def metrics(): return jsonify(get_metrics())
@app.route("/metrics/timeseries")
def metrics_timeseries():
period = request.args.get("period", "day"); models_list = list(GPU_URLS.keys())
data = {"models": {}, "labels": []}
if period == "day":
buckets = [time.strftime("%Y%m%d%H", time.gmtime(time.time()-h*3600)) for h in range(23,-1,-1)]
data["labels"] = [time.strftime("%H:00", time.gmtime(time.time()-h*3600)) for h in range(23,-1,-1)]
elif period == "week":
buckets = [time.strftime("%Y%m%d", time.gmtime(time.time()-d*86400)) for d in range(6,-1,-1)]
data["labels"] = [time.strftime("%a", time.gmtime(time.time()-d*86400)) for d in range(6,-1,-1)]
else:
buckets = [time.strftime("%Y%m%d", time.gmtime(time.time()-d*86400)) for d in range(29,-1,-1)]
data["labels"] = [time.strftime("%m/%d", time.gmtime(time.time()-d*86400)) for d in range(29,-1,-1)]
if r:
for model in models_list:
counts = []
for bucket in buckets:
total = 0
if period in ("week","month"):
for hh in range(24): total += int(r.get("ts:"+model+":"+bucket+"{:02d}".format(hh)) or 0)
else: total = int(r.get("ts:"+model+":"+bucket) or 0)
counts.append(total)
data["models"][model] = counts
return jsonify(data)
@app.route("/stream")
def stream():
def ev():
q = queue.Queue()
with sse_lock: sse_subscribers.append(q)
try:
yield "data: "+json.dumps(get_metrics())+"\n\n"
while True:
try: yield "data: "+q.get(timeout=3)+"\n\n"
except queue.Empty: yield "data: "+json.dumps(get_metrics())+"\n\n"
except GeneratorExit: pass
finally:
with sse_lock:
if q in sse_subscribers: sse_subscribers.remove(q)
return Response(stream_with_context(ev()), mimetype="text/event-stream",
headers={"Cache-Control":"no-cache","X-Accel-Buffering":"no","Access-Control-Allow-Origin":"*"})
if __name__ == "__main__":
log.info("Router on :9000 (load-aware)")
app.run(host="0.0.0.0", port=9000, debug=False)
-80
View File
@@ -1,80 +0,0 @@
# Add time-series tracking and endpoint to router
with open('/opt/inference-harness/router/router.py') as f:
code = f.read()
# Add time-series tracking in the chat handler (after Redis incr)
old_track = '''r.incr('routes:'+model); r.incr('routes:tier:'+tier); r.incr('routes:agent:'+agent)
r.lpush('routes:recent', json.dumps'''
new_track = '''r.incr('routes:'+model); r.incr('routes:tier:'+tier); r.incr('routes:agent:'+agent)
# Time-series: hourly bucket
hour_key = 'ts:'+model+':'+time.strftime('%Y%m%d%H')
r.incr(hour_key)
r.expire(hour_key, 86400*31) # keep 31 days
r.lpush('routes:recent', json.dumps'''
code = code.replace(old_track, new_track)
# Add /metrics/timeseries endpoint before if __name__
ts_endpoint = '''
@app.route('/metrics/timeseries')
def metrics_timeseries():
period = request.args.get('period', 'day')
models = list(GPU_URLS.keys())
data = {'models': {}, 'labels': []}
if period == 'day':
# Last 24 hours, hourly buckets
buckets = []
for h in range(23, -1, -1):
t = time.time() - h * 3600
buckets.append(time.strftime('%Y%m%d%H', time.gmtime(t)))
data['labels'] = [time.strftime('%H:00', time.gmtime(time.time() - h*3600)) for h in range(23, -1, -1)]
elif period == 'week':
# Last 7 days, daily buckets
buckets = []
for d in range(6, -1, -1):
t = time.time() - d * 86400
buckets.append(time.strftime('%Y%m%d', time.gmtime(t)))
data['labels'] = [time.strftime('%a', time.gmtime(time.time() - d*86400)) for d in range(6, -1, -1)]
else:
# Month — last 30 days, 3-day buckets
buckets = []
for d in range(29, -1, -3):
t = time.time() - d * 86400
buckets.append(time.strftime('%Y%m%d', time.gmtime(t)))
data['labels'] = [time.strftime('%m/%d', time.gmtime(time.time() - d*86400)) for d in range(29, -1, -3)]
if r:
for model in models:
counts = []
for bucket in buckets:
if period == 'month':
# Sum 3 consecutive days per bucket
total = 0
base = time.strptime(bucket, '%Y%m%d')
for offset in range(3):
d = time.strftime('%Y%m%d', time.gmtime(time.mktime(base) + offset*86400))
total += int(r.get('ts:'+model+':'+d) or 0)
# Also check hourly keys for today
for hh in range(24):
total += int(r.get('ts:'+model+':'+d+'{:02d}'.format(hh)) or 0)
counts.append(total)
else:
key = 'ts:'+model+':'+bucket
if period == 'week':
# Sum all hours in the day
total = sum(int(r.get(key+'{:02d}'.format(h)) or 0) for h in range(24))
else:
total = int(r.get(key) or 0)
counts.append(total)
data['models'][model] = counts
return jsonify(data)
'''
# Insert before if __name__
code = code.replace(if __name__ == __main__:, ts_endpoint + nif __name__ == __main__:)
with open('/opt/inference-harness/router/router.py', 'w') as f:
f.write(code)
print('Time-series tracking and endpoint added')
+7 -6
View File
@@ -1,7 +1,8 @@
FROM python:3.12-slim
FROM python:3.13-slim
COPY harness-dashboard.py /app/harness-dashboard.py
WORKDIR /app
COPY requirements.txt .
RUN pip install --no-cache-dir -r requirements.txt
COPY dashboard.py .
EXPOSE 3000
CMD ["python", "dashboard.py"]
EXPOSE 3001
CMD ["python3", "harness-dashboard.py"]
+5
View File
@@ -0,0 +1,5 @@
FROM python:3.11-slim
WORKDIR /app
COPY harness-dashboard.py .
EXPOSE 3001
CMD ["python3", "harness-dashboard.py"]
-145
View File
@@ -1,145 +0,0 @@
<!DOCTYPE html>
<html lang="en">
<head>
<meta charset="UTF-8"><meta name="viewport" content="width=device-width, initial-scale=1.0">
<title>Inference Harness - Dashboard</title>
<style>
:root{--bg:#0b0f17;--text:#bcc3cd;--panel:rgba(31,41,55,0.7);--border:rgba(75,85,99,0.4)}
body{background:var(--bg);color:var(--text);font-family:-apple-system,BlinkMacSystemFont,'Segoe UI',sans-serif;margin:0;padding:1.5rem}
.glass-panel{background:var(--panel);backdrop-filter:blur(12px);border:1px solid var(--border);border-radius:0.75rem;overflow:hidden}
.stat-value{font-size:28px;font-weight:700;line-height:1.1}
.stat-label{font-size:11px;text-transform:uppercase;letter-spacing:0.6px;color:#64748b}
.health-bar{height:0.5rem;background:#374151;border-radius:0.375rem;overflow:hidden}
.health-fill{height:100%;transition:width .5s ease}
.dot-green{background:#10b981;animation:pulse 2s infinite}
.dot-yellow{background:#f59e0b;animation:pulse 2s infinite}
.dot-red{background:#ef4444;animation:pulse 2s infinite}
@keyframes pulse{0%,100%{opacity:1}50%{opacity:.8}}
.container{max-width:1400px;margin:0 auto}
.grid-5{display:grid;grid-template-columns:repeat(5,1fr);gap:1rem}
.grid-3{display:grid;grid-template-columns:repeat(3,1fr);gap:1rem}
.grid-2{display:grid;grid-template-columns:1fr 1fr;gap:1.5rem}
.mb-6{margin-bottom:1.5rem}
.p-4{padding:1rem}
.text-center{text-align:center}
.font-bold{font-weight:700}
.text-white{color:#fff}
.text-sm{font-size:.875rem}
.text-xs{font-size:.75rem}
.text-gray-400{color:#9ca3af}
.text-gray-500{color:#6b7280}
.flex{display:flex}
.items-center{align-items:center}
.justify-between{justify-content:space-between}
.gap-4{gap:1rem}
.w-3{width:.75rem}.h-3{height:.75rem}.rounded-full{border-radius:9999px}
.bg-blue-600{background:#2563eb}.bg-blue-600:hover{background:#1d4ed8}
.px-4{padding-left:1rem;padding-right:1rem}.py-2{padding-top:.5rem;padding-bottom:.5rem}
.rounded-lg{border-radius:.5rem}.text-white{color:#fff}
.mt-2{margin-top:.5rem}.mb-2{margin-bottom:.5rem}.mb-3{margin-bottom:.75rem}
.border-t{border-top:1px solid #374151}.pt-4{padding-top:1rem}
.text-xl{font-size:1.25rem}.text-2xl{font-size:1.5rem}.text-lg{font-size:1.125rem}
.font-semibold{font-weight:600}.font-normal{font-weight:400}
.text-emerald-400{color:#34d399}.text-amber-400{color:#fbbf24}.text-red-400{color:#f87171}
</style>
</head>
<body>
<div class="container">
<div class="flex justify-between items-center mb-6">
<div><h1 class="text-2xl font-bold text-white">Inference Harness</h1><p class="text-sm text-gray-400">Syslog Solution LLC · Real-time Monitoring</p></div>
<div class="flex items-center gap-4">
<span id="statusBadge" class="flex items-center gap-2"><span id="statusDot" class="w-3 h-3 rounded-full dot-green"></span><span id="statusText" class="text-emerald-400 text-lg font-semibold">healthy</span></span>
<span id="lastUpdate" class="text-sm text-gray-500"></span>
</div>
</div>
<div class="grid-5 mb-6">
<div class="glass-panel p-4 text-center"><p class="stat-value text-white" id="kpiGpus">-</p><p class="stat-label">GPUs Online</p></div>
<div class="glass-panel p-4 text-center"><p class="stat-value text-white" id="kpiTrips">-</p><p class="stat-label">Circuit Trips</p></div>
<div class="glass-panel p-4 text-center"><p class="stat-value text-white" id="kpiLatency">-</p><p class="stat-label">Avg Latency</p></div>
<div class="glass-panel p-4 text-center"><p class="stat-value text-white" id="kpiReqs">-</p><p class="stat-label">Requests/min</p></div>
<div class="glass-panel p-4 text-center"><p class="stat-value text-white" id="kpiActive">-</p><p class="stat-label">Active Requests</p></div>
</div>
<h2 class="text-xl font-semibold text-white mb-3">GPU Health Scoring <span class="text-sm text-gray-400 font-normal">Live: VRAM 40% · Temp 30% · Load 30%</span></h2>
<div id="gpuCards" class="grid-3 mb-6"></div>
<div class="glass-panel p-4 mb-6"><h3 class="text-sm font-semibold text-gray-400 mb-3">Health Score History (60s rolling)</h3><div id="healthChart" style="height:220px"></div></div>
<div class="text-center text-sm text-gray-500 pt-4 border-t"><p>Inference Harness Dashboard · Syslog Solution LLC · Auto-refresh 15s</p></div>
</div>
<script>
const COLORS={'strix-moe':'#10b981','gpu-dense':'#8b5cf6','gpu-vision':'#3b82f6'};
const HISTORY=[]; // rolling 60 sample history for chart
function Q(id){return document.getElementById(id)}
function updateStatus(trips,degraded){
const d=Q('statusDot'),t=Q('statusText');
if(degraded){d.className='w-3 h-3 rounded-full dot-red';t.textContent='degraded';t.className='text-red-400 text-lg font-semibold'}
else if(trips>0){d.className='w-3 h-3 rounded-full dot-yellow';t.textContent='trips:'+trips;t.className='text-amber-400 text-lg font-semibold'}
else{d.className='w-3 h-3 rounded-full dot-green';t.textContent='healthy';t.className='text-emerald-400 text-lg font-semibold'}
}
function fetchAll(){
Promise.all([
fetch('/metrics/gpu-health').then(r=>r.json()),
fetch('/metrics/latency').then(r=>r.json())
]).then(function(_a){var health=_a[0],latency=_a[1];
Q('lastUpdate').textContent=new Date().toLocaleTimeString();
var gpus=health.gpus||[],kpi=health.kpi||{};
// KPIs
Q('kpiGpus').textContent=kpi.gpus_online+'/'+kpi.total_gpus;
Q('kpiTrips').textContent=kpi.total_trips;
Q('kpiLatency').textContent=(latency.avg_ms||0)+'ms';
Q('kpiReqs').textContent=latency.requests_per_min||0;
var active=0;gpus.forEach(function(g){active+=g.active_requests||0});
Q('kpiActive').textContent=active;
// Status
var degraded=gpus.some(function(g){return g.status==='down'||g.circuit_tripped});
updateStatus(kpi.total_trips||0,degraded);
// GPU cards
var html='';
gpus.forEach(function(g,i){
var s=g.health_score||0,clr=s>=70?'#10b981':(s>=40?'#f59e0b':'#ef4444');
var badge='';
if(g.circuit_tripped)badge='<span class="text-xs px-2 py-1 bg-red-900 rounded text-red-300">TRIPPED</span>';
else if(i===0)badge='<span class="text-xs px-2 py-1 bg-emerald-900 rounded text-emerald-300">Best</span>';
html+='<div class="glass-panel p-4"><div class="flex justify-between items-center mb-2"><p class="text-lg font-bold text-white">'+g.label+'</p>'+badge+'</div>'+
'<p class="text-4xl font-bold mb-2" style="color:'+clr+'">'+Math.round(s)+'</p>'+
'<div class="health-bar"><div class="health-fill" style="width:'+s+'%;background:'+clr+'"></div></div>'+
'<div class="flex justify-between mt-2 text-xs text-gray-400"><span>VRAM '+g.vram_pct+'%</span><span>'+g.temp_c+'°C</span><span>Active '+g.active_requests+'/'+g.max_concurrent+'</span><span>Trips '+g.circuit_trip_count+'</span></div></div>';
});
Q('gpuCards').innerHTML=html;
// Rolling history
var now=Date.now();HISTORY.push({ts:now,gpus:gpus.map(function(g){return{id:g.id,score:g.health_score}})});
if(HISTORY.length>60)HISTORY.shift();
renderChart();
}).catch(function(){});
}
function renderChart(){
var W=800,H=220,svg='<svg viewBox="0 0 '+W+' '+H+'" style="width:100%;height:220px" preserveAspectRatio="none">';
// Grid lines
for(var i=0;i<=4;i++){var y=(i/4)*H;svg+='<line x1="0" y1="'+y+'" x2="'+W+'" y2="'+y+'" stroke="#1e293b" stroke-width="1"/>';svg+='<text x="5" y="'+(y+10)+'" font-size="9" fill="#64748b">'+(100-i*25)+'</text>'}
// Time labels
for(var i=0;i<=4;i++){var lx=(i/4)*W,lt=HISTORY.length>0?new Date(HISTORY[Math.floor(i/4*(HISTORY.length-1))].ts).toLocaleTimeString():'';svg+='<text x="'+lx+'" y="'+(H-2)+'" font-size="8" fill="#475569" text-anchor="middle">'+lt+'</text>'}
// Plot lines per GPU
var ids=HISTORY.length>0?HISTORY[0].gpus.map(function(g){return g.id}):[];
ids.forEach(function(id){
var c=COLORS[id]||'#94a3b8',pts='',first=true;
HISTORY.forEach(function(h,i){
var gpu=h.gpus.find(function(g){return g.id===id});
if(!gpu)return;
var x=(i/(HISTORY.length-1||1))*W,y=H-(gpu.score/100*H);
pts+=(first?'M':'L')+x.toFixed(1)+','+y.toFixed(1);first=false;
});
if(pts)svg+='<path d="'+pts+'" fill="none" stroke="'+c+'" stroke-width="2" opacity="0.9"/>';
});
svg+='</svg>';Q('healthChart').innerHTML=svg;
}
fetchAll();setInterval(fetchAll,15000);
</script>
</body>
</html>
-362
View File
@@ -1,362 +0,0 @@
"""SyslogAI Harness Dashboard — Modern Design."""
import os, json, time, queue, threading
import requests
from flask import Flask, request, render_template_string, Response, stream_with_context
ROUTER_METRICS = os.environ.get("ROUTER_METRICS_URL", "http://router:9000/metrics")
app = Flask(__name__)
sse_subscribers = []; sse_lock = threading.Lock()
def fetch_state():
try:
r = requests.get(ROUTER_METRICS, timeout=5)
if r.status_code == 200: return r.json()
except Exception: pass
return {"gpus":[],"route_counts":{},"agent_counts":{},"recent":[],"timestamp":time.time()}
def broadcast_loop():
while True:
time.sleep(5)
data = fetch_state(); payload = json.dumps(data)
with sse_lock:
dead = [q for q in sse_subscribers if not q.put(payload)]
for q in dead: sse_subscribers.remove(q)
threading.Thread(target=broadcast_loop, daemon=True).start()
DASHBOARD_HTML = r"""<!DOCTYPE html>
<html lang="en" data-bs-theme="dark">
<head>
<meta charset="UTF-8"><meta name="viewport" content="width=device-width, initial-scale=1.0">
<title>SyslogAI Harness</title>
<link href="https://cdn.jsdelivr.net/npm/bootstrap@5.3.3/dist/css/bootstrap.min.css" rel="stylesheet">
<style>
body { background: #0b0f17; color: #bcc3cd; font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', system-ui, sans-serif; padding: 20px 24px; }
.card { background: #111827; border: 1px solid #1e293b; border-radius: 10px; height: 100%; }
.stat-card { background: #111827; border: 1px solid #1e293b; border-radius: 10px; padding: 18px 20px; text-align: center; }
.stat-value { font-size: 28px; font-weight: 700; line-height: 1.1; }
.stat-label { font-size: 11px; text-transform: uppercase; letter-spacing: 0.6px; color: #64748b; margin-top: 4px; }
.gpu-card { background: #111827; border: 1px solid #1e293b; border-radius: 10px; padding: 16px 18px; height: 100%; }
.gpu-card .title { font-size: 13px; font-weight: 600; color: #e2e8f0; margin-bottom: 12px; display: flex; align-items: center; gap: 8px; }
.gpu-card .status-dot { width: 8px; height: 8px; border-radius: 50%; flex-shrink: 0; }
.gpu-card .row-metric { display: flex; justify-content: space-between; font-size: 12px; padding: 2px 0; }
.gpu-card .row-metric .lbl { color: #64748b; }
.gpu-card .row-metric .val { color: #e2e8f0; font-variant-numeric: tabular-nums; }
.gpu-card .slot-bar { display: flex; gap: 3px; margin-top: 8px; }
.gpu-card .slot-bar .s { flex: 1; height: 5px; border-radius: 2px; background: #1e293b; }
.gpu-card .slot-bar .s.active { background: #38bdf8; }
.chart-card { background: #111827; border: 1px solid #1e293b; border-radius: 10px; padding: 16px 18px; height: 100%; display: flex; flex-direction: column; }
.chart-card .title { font-size: 13px; font-weight: 600; color: #e2e8f0; margin-bottom: 12px; }
.bar-row { margin-bottom: 8px; }
.bar-label { display: flex; justify-content: space-between; font-size: 11px; margin-bottom: 3px; color: #64748b; }
.bar-label .name { color: #cbd5e1; }
.bar-track { height: 5px; background: #1e293b; border-radius: 3px; overflow: hidden; }
.bar-fill { height: 100%; border-radius: 3px; transition: width 0.6s ease; }
.table-custom { font-size: 11px; margin: 0; }
.table-custom th { color: #64748b; font-weight: 500; font-size: 10px; text-transform: uppercase; border-color: #1e293b; padding: 8px 10px; }
.table-custom td { color: #94a3b8; border-color: rgba(30,41,59,0.5); padding: 6px 10px; }
.agent-badge { font-size: 10px; padding: 2px 7px; border-radius: 8px; font-weight: 600; }
.btn-sm-period { font-size: 10px; padding: 3px 10px; border-radius: 6px; border: 1px solid #1e293b; color: #64748b; background: transparent; cursor: pointer; }
.btn-sm-period.active { background: #1d4ed8; color: #fff; border-color: #1d4ed8; }
.ring-label { font-size: 22px; font-weight: 700; }
.ring-sublabel { font-size: 10px; color: #64748b; }
</style>
</head>
<body>
<!-- HEADER -->
<div class="d-flex justify-content-between align-items-center mb-4">
<div>
<h5 class="mb-0 text-white fw-bold">&#x26A1; SyslogAI Harness</h5>
<div class="small text-secondary" id="live-indicator">
<span class="status-dot" id="live-dot" style="width:6px;height:6px;border-radius:50%;display:inline-block;background:#22c55e;animation:pulse 2s infinite"></span>
<span id="connection-status">live</span> &middot; <span id="update-time"></span>
</div>
</div>
<div class="d-flex gap-2">
<div class="stat-card" style="min-width:100px"><div class="stat-value text-info" id="kpi-total">0</div><div class="stat-label">Requests</div></div>
<div class="stat-card" style="min-width:100px"><div class="stat-value text-warning" id="kpi-active">0</div><div class="stat-label">Active</div></div>
<div class="stat-card" style="min-width:100px"><div class="stat-value" style="color:#a78bfa" id="kpi-agents">0</div><div class="stat-label">Agents</div></div>
</div>
</div>
<div class="row g-3 align-items-stretch">
<!-- ROW 1: Usage Chart (8) + GPU Metrics (4) -->
<div class="col-md-8"><div class="chart-card"><div class="title d-flex justify-content-between align-items-center">
<span>Usage Over Time</span>
<div class="d-flex gap-1">
<button class="btn-sm-period active" onclick="switchPeriod('day')">24h</button>
<button class="btn-sm-period" onclick="switchPeriod('week')">7d</button>
<button class="btn-sm-period" onclick="switchPeriod('month')">30d</button>
</div>
</div><div id="timeseries-chart" style="height:150px"></div><div id="timeseries-legend" class="d-flex justify-content-center gap-3 mt-2 flex-wrap small"></div></div></div>
<div class="col-md-4"><div class="chart-card"><div class="title">GPU Metrics</div><div id="gpu-metrics-card"></div></div></div>
<!-- ROW 2: 3 GPU Cards -->
<div class="col-md-4"><div class="gpu-card" id="gpu-moe"><div class="text-secondary small">Loading...</div></div></div>
<div class="col-md-4"><div class="gpu-card" id="gpu-dense"><div class="text-secondary small">Loading...</div></div></div>
<div class="col-md-4"><div class="gpu-card" id="gpu-light"><div class="text-secondary small">Loading...</div></div></div>
<!-- ROW 3: Queue + Model + Agent -->
<div class="col-md-4"><div class="chart-card"><div class="title">Queue Status</div><div class="text-center" id="queue-viz"></div></div></div>
<div class="col-md-4"><div class="chart-card"><div class="title">Model Distribution</div><div id="route-bars"></div></div></div>
<div class="col-md-4"><div class="chart-card"><div class="title">Agent Activity</div><div id="agent-bars"></div></div></div>
<!-- ROW 4: Performance Analytics -->
<div class="col-12 mb-2"><div class="d-flex align-items-center gap-2"><span class="fw-bold text-white" style="font-size:14px">&#x1F4CA; Performance Analytics</span>
<div class="d-flex gap-1 ms-auto">
<button class="btn-sm-period active" onclick="switchPerfWindow('1')">1h</button>
<button class="btn-sm-period" onclick="switchPerfWindow('24')">24h</button>
</div>
</div></div>
<div class="col-md-6"><div class="chart-card"><div class="title">Latency — P50 / P95 / P99 (ms)</div><div id="perf-latency"></div></div></div>
<div class="col-md-6"><div class="chart-card"><div class="title">Throughput — Tokens / sec</div><div id="perf-throughput"></div></div></div>
<div class="col-md-6"><div class="chart-card"><div class="title">Routing Effectiveness — by Reason</div><div id="perf-reasons"></div></div></div>
<div class="col-md-6"><div class="chart-card"><div class="title">Agent Performance</div><div id="perf-agents"></div></div></div>
<!-- ROW 5: Latency vs Context Scatter -->
<div class="col-12"><div class="chart-card"><div class="title d-flex justify-content-between align-items-center">
<span>Latency vs Prompt Size — by Model</span>
<div class="d-flex gap-2">
<select id="scatter-model" onchange="loadScatter()" style="font-size:10px;background:#1e293b;color:#94a3b8;border:1px solid #334155;border-radius:4px;padding:2px 6px">
<option value="all">All Models</option>
<option value="gpu-vision">9B VLM</option>
<option value="gpu-dense">27B Dense</option>
<option value="strix-moe">35B MoE</option>
</select>
</div>
</div><div id="scatter-plot" style="height:200px;position:relative"></div><div id="scatter-legend" class="d-flex justify-content-center gap-3 mt-2 flex-wrap small"></div></div></div>
<!-- ROW 6: Live Stream -->
<div class="col-12"><div class="chart-card"><div class="title">Live Stream</div>
<div class="table-responsive"><table class="table table-custom mb-0">
<thead><tr><th>Time</th><th>Agent</th><th>Model</th><th>Reason</th><th>Tier</th></tr></thead>
<tbody id="route-tbody"></tbody>
</table></div>
</div></div>
</div>
<script>
var MC={'gpu-vision':'#22c55e','gpu-dense':'#f59e0b','strix-moe':'#a78bfa'};
var ML={'gpu-vision':'Qwen3.5-9B VLM','gpu-dense':'Qwen3.8-27B','strix-moe':'Qwen MoE'};
var GL={'strix-moe':'MoE - Strix Halo','gpu-dense':'Dense - RTX 3090','gpu-vision':'VLM - RTX 5070'};
function $(id){return document.getElementById(id);}
function render(data){
if(!data||!data.gpus)return;
var t=Object.values(data.route_counts||{}).reduce((a,b)=>a+b,0);
var ta=0,tm=0;data.gpus.forEach(function(g){ta+=(g.active_requests||0);tm+=(g.max_concurrent||1)});
$('kpi-total').textContent=t;$('kpi-active').textContent=ta+'/'+tm;$('kpi-agents').textContent=Object.keys(data.agent_counts||{}).length;
$('update-time').textContent=new Date().toLocaleTimeString();
var ids={'strix-moe':'gpu-moe','gpu-dense':'gpu-dense','gpu-vision':'gpu-vision'};
data.gpus.forEach(function(g){
var el=$(ids[g.id]);if(!el)return;
var a=g.active_requests||0,mx=g.max_concurrent||1;
var sc=g.status==='healthy'?'#22c55e':g.status==='saturated'?'#f59e0b':'#ef4444';
var ss=g.status==='healthy'?'Online':g.status==='saturated'?'Busy':'Offline';
var slots='';for(var i=0;i<mx;i++)slots+='<span class=\"s'+(i<a?' active':'')+'\"></span>';
var h='<div class=\"title\"><span class=\"status-dot\" style=\"background:'+sc+'\"></span>'+GL[g.id]+'<span class=\"ms-auto small\" style=\"color:'+sc+'\">'+ss+'</span></div>';
h+='<div class=\"row-metric\"><span class=\"lbl\">VRAM</span><span class=\"val\">'+g.vram_used_mb+' / '+g.vram_total_mb+' MB</span></div>';
h+='<div class=\"row-metric\"><span class=\"lbl\">Utilization</span><span class=\"val\">'+g.gpu_util_pct+'%</span></div>';
h+='<div class=\"row-metric\"><span class=\"lbl\">Temperature</span><span class=\"val\" style=\"color:'+(g.temp_c>85?'#ef4444':g.temp_c>70?'#f59e0b':'#22c55e')+'\">'+g.temp_c+'C</span></div>';
if(g.power_w)h+='<div class=\"row-metric\"><span class=\"lbl\">Power</span><span class=\"val\">'+g.power_w+'W'+(g.power_limit_w?'/'+g.power_limit_w+'W':'')+'</span></div>';
h+='<div class=\"row-metric\"><span class=\"lbl\">Slots</span><span class=\"val\" style=\"color:'+(a>=mx?'#ef4444':'#e2e8f0')+'\">'+a+' / '+mx+'</span></div>';
h+='<div class=\"slot-bar\">'+slots+'</div>';el.innerHTML=h;
});
renderQueue(data);renderGPUMetrics(data);
var rc=data.route_counts||{},mr=Math.max(1,...Object.values(rc));
$('route-bars').innerHTML=Object.entries(rc).length?Object.entries(rc).sort((a,b)=>b[1]-a[1]).map(function(e){var m=e[0],c=e[1];return'<div class=\"bar-row\"><div class=\"bar-label\"><span class=\"name\">'+(ML[m]||m)+'</span><span>'+c+' ('+(t?Math.round(c/t*100):0)+'%)</span></div><div class=\"bar-track\"><div class=\"bar-fill\" style=\"width:'+(c/mr*100)+'%;background:'+(MC[m]||'#38bdf8')+'\"></div></div></div>';}).join(''):'<div class=\"text-secondary small\">-</div>';
var ac=data.agent_counts||{},ma=Math.max(1,...Object.values(ac));
$('agent-bars').innerHTML=Object.entries(ac).length?Object.entries(ac).sort((a,b)=>b[1]-a[1]).map(function(e){return'<div class=\"bar-row\"><div class=\"bar-label\"><span class=\"name\">'+e[0]+'</span><span>'+e[1]+'</span></div><div class=\"bar-track\"><div class=\"bar-fill\" style=\"width:'+(e[1]/ma*100)+'%;background:#38bdf8\"></div></div></div>';}).join(''):'<div class=\"text-secondary small\">-</div>';
var recent=data.recent||[];
$('route-tbody').innerHTML=recent.length?recent.slice(0,20).map(function(r){var d=new Date(r.ts*1000),ag=r.agent||'?';return'<tr><td class=\"text-secondary\">'+d.toLocaleTimeString()+'</td><td><span class=\"agent-badge\" style=\"background:rgba(56,189,248,0.12);color:#38bdf8\">'+ag+'</span></td><td>'+(ML[r.model]||r.model)+'</td><td class=\"text-secondary\">'+(r.reason||'')+'</td><td class=\"text-uppercase\" style=\"font-size:10px;color:'+(r.tier==='enterprise'?'#a78bfa':'#64748b')+'\">'+(r.tier||'')+'</td></tr>';}).join(''):'<tr><td colspan=\"5\" class=\"text-secondary\">Waiting...</td></tr>';
}
function renderQueue(data){
var el=$('queue-viz');if(!el)return;
var ta=0,tm=0;data.gpus.forEach(function(g){ta+=(g.active_requests||0);tm+=(g.max_concurrent||1)});
var pct=tm>0?Math.round(ta/tm*100):0,st=pct>=100?'SATURATED':pct>=50?'BUSY':'IDLE';
var sc=pct>=100?'#ef4444':pct>=50?'#f59e0b':'#22c55e';
var circ=188.5,dash=(pct/100)*circ;
var h='<div class=\"d-inline-block position-relative mb-2\"><svg width=\"72\" height=\"72\"><circle cx=\"36\" cy=\"36\" r=\"30\" fill=\"none\" stroke=\"#1e293b\" stroke-width=\"6\"/><circle cx=\"36\" cy=\"36\" r=\"30\" fill=\"none\" stroke=\"'+sc+'\" stroke-width=\"6\" stroke-dasharray=\"'+dash+' '+(circ-dash)+'\" stroke-linecap=\"round\" transform=\"rotate(-90 36 36)\"/></svg><div style=\"position:absolute;top:50%;left:50%;transform:translate(-50%,-50%);text-align:center\"><div class=\"ring-label\" style=\"color:'+sc+'\">'+ta+'</div><div class=\"ring-sublabel\">/ '+tm+' slots</div></div></div>';
h+='<div class=\"fw-bold mb-2 small\" style=\"color:'+sc+'\">'+st+'</div>';
var lb={'strix-moe':'MoE','gpu-dense':'Dense','gpu-vision':'VLM'};
data.gpus.forEach(function(g){var a=g.active_requests||0,mx=g.max_concurrent||1,gp=mx>0?Math.round(a/mx*100):0;h+='<div class=\"d-flex align-items-center gap-2 mb-1 justify-content-center\"><span class=\"small\" style=\"min-width:32px;text-align:right;font-size:10px\">'+(lb[g.id]||g.id)+'</span><div style=\"flex:1;max-width:70px;height:3px;background:#1e293b;border-radius:2px;overflow:hidden\"><div style=\"height:100%;width:'+gp+'%;background:'+sc+';border-radius:2px\"></div></div><span class=\"small\" style=\"min-width:22px;font-size:10px\">'+a+'/'+mx+'</span></div>'});
el.innerHTML=h;
}
function renderGPUMetrics(data){
var el=$('gpu-metrics-card');if(!el)return;
var lb={'strix-moe':'MoE','gpu-dense':'Dense','gpu-vision':'VLM'};
var h='';data.gpus.forEach(function(g){
var nm=lb[g.id]||g.id,tp=g.temp_c||0,ut=g.gpu_util_pct||0,pw=g.power_w||0,pl=g.power_limit_w||0;
var tc=tp>85?'#ef4444':tp>70?'#f59e0b':'#22c55e',uc=ut>90?'#ef4444':ut>70?'#f59e0b':'#22c55e';
h+='<div class=\"mb-3\"><div class=\"fw-bold small text-white-50 mb-1\">'+nm+'</div>';
h+='<div class=\"d-flex align-items-center gap-2 mb-1\"><span class=\"small text-secondary\" style=\"min-width:30px\">T</span><div class=\"flex-grow-1\" style=\"height:3px;background:#1e293b;border-radius:2px;overflow:hidden\"><div style=\"height:100%;width:'+Math.min(tp,100)+'%;background:'+tc+';border-radius:2px\"></div></div><span class=\"small\" style=\"color:'+tc+';min-width:30px;text-align:right\">'+tp+'C</span></div>';
h+='<div class=\"d-flex align-items-center gap-2 mb-1\"><span class=\"small text-secondary\" style=\"min-width:30px\">U</span><div class=\"flex-grow-1\" style=\"height:3px;background:#1e293b;border-radius:2px;overflow:hidden\"><div style=\"height:100%;width:'+ut+'%;background:'+uc+';border-radius:2px\"></div></div><span class=\"small\" style=\"color:'+uc+';min-width:30px;text-align:right\">'+ut+'%</span></div>';
if(pw>0){var pp=pl>0?Math.round(pw/pl*100):0,pc=pp>90?'#ef4444':pp>70?'#f59e0b':'#22c55e';h+='<div class=\"d-flex align-items-center gap-2\"><span class=\"small text-secondary\" style=\"min-width:30px\">P</span><div class=\"flex-grow-1\" style=\"height:3px;background:#1e293b;border-radius:2px;overflow:hidden\"><div style=\"height:100%;width:'+pp+'%;background:'+pc+';border-radius:2px\"></div></div><span class=\"small\" style=\"color:'+pc+';min-width:30px;text-align:right\">'+pw+'W</span></div>';}
h+='</div>';});
el.innerHTML=h;
}
var cp='day';
function switchPeriod(p){cp=p;document.querySelectorAll('.btn-sm-period').forEach(function(b){b.classList.remove('active')});event.target.classList.add('active');loadTS();}
function loadTS(){fetch('/api/timeseries?period='+cp).then(function(r){return r.json()}).then(renderTS).catch(function(){})}
function renderTS(d){
var models=d.models||{},labels=d.labels||[];
if(!labels.length)return;
var cn=$('timeseries-chart'),lg=$('timeseries-legend'),mn=Object.keys(models);
if(!mn.length){cn.innerHTML='<div class=\"text-secondary small text-center py-4\">-</div>';return;}
var mv=1;for(var m in models)for(var i=0;i<models[m].length;i++)if(models[m][i]>mv)mv=models[m][i];mv=Math.ceil(mv*1.15)||1;
var W=labels.length>1?100/(labels.length-1):100,H=130;
var paths='';for(var mi=0;mi<mn.length;mi++){var m=mn[mi],vals=models[m]||[],d='';for(var i=0;i<vals.length;i++){var x=i*W,y=H-(vals[i]/mv)*H;d+=(i===0?'M':'L')+x.toFixed(1)+','+y.toFixed(1)+' ';}paths+='<path d=\"'+d+'\" fill=\"none\" stroke=\"'+(MC[m]||'#38bdf8')+'\" stroke-width=\"2\" stroke-linecap=\"round\" opacity=\"0.8\"/>';}
var grid='';for(var g=0;g<=4;g++){var y=(g/4)*H;grid+='<line x1=\"0\" y1=\"'+y.toFixed(1)+'\" x2=\"100\" y2=\"'+y.toFixed(1)+'\" stroke=\"#1e293b\" stroke-width=\"1\"/>';}
cn.innerHTML='<svg viewBox=\"0 0 100 '+(H+16)+'\" style=\"width:100%;height:'+(H+20)+'px;display:block\" preserveAspectRatio=\"none\">'+grid+paths+'</svg>';
lg.innerHTML=mn.map(function(m){return'<span class=\"d-flex align-items-center gap-1\"><svg width=\"14\" height=\"8\"><line x1=\"0\" y1=\"4\" x2=\"14\" y2=\"4\" stroke=\"'+(MC[m]||'#38bdf8')+'\" stroke-width=\"2\"/></svg>'+(ML[m]||m)+'</span>';}).join('');
}
var perfWindow='24';
function switchPerfWindow(w){perfWindow=w;document.querySelectorAll('.btn-sm-period').forEach(function(b,i){if(i>=4)b.classList.toggle('active',b.textContent.trim().replace('h','')===w)});loadPerf();}
function loadPerf(){fetch('/api/performance?window='+perfWindow).then(function(r){return r.json()}).then(renderPerf).catch(function(){})}
function renderPerf(d){
var models=d.models||[],reasons=d.reasons||[],agents=d.agents||[],sum=d.summary||{};
// Latency bars: p50/p95/p99 per model
var mlab={'strix-moe':'35B MoE','gpu-dense':'27B Dense','gpu-vision':'9B VLM'};
var mcol={'strix-moe':'#a78bfa','gpu-dense':'#f59e0b','gpu-vision':'#22c55e'};
if(!models.length){$('perf-latency').innerHTML='<div class="text-secondary small text-center py-4">Accumulating data...</div>';return;}
var maxLat=Math.max(...models.map(function(m){return m.latency.p99||0}),1);
var latHTML=models.map(function(m){
var l=m.latency||{},p50=l.p50||0,p95=l.p95||0,p99=l.p99||0,c=mcol[m.model]||'#38bdf8';
return'<div class="mb-2" style="font-size:11px"><div class="d-flex justify-content-between mb-1"><span style="color:#e2e8f0">'+mlab[m.model]+'</span><span class="text-secondary">'+m.count+' reqs</span></div>'+
'<div class="d-flex align-items-center gap-2 mb-1"><span class="text-secondary" style="min-width:28px">p50</span><div class="flex-grow-1" style="height:14px;background:#1e293b;border-radius:4px;overflow:hidden;position:relative"><div style="position:absolute;left:0;top:0;height:100%;width:'+(p50/maxLat*100)+'%;background:'+c+';opacity:0.3;border-radius:4px"></div><div style="position:absolute;left:0;top:0;height:100%;width:'+(p95/maxLat*100)+'%;background:'+c+';opacity:0.5;border-radius:4px"></div><div style="position:absolute;left:0;top:0;height:100%;width:'+(p99/maxLat*100)+'%;background:'+c+';border-radius:4px"></div></div><span style="color:'+c+';min-width:48px;text-align:right;font-variant-numeric:tabular-nums">'+p99+'ms</span></div>'+
'<div class="d-flex gap-3" style="font-size:10px;color:#64748b;padding-left:32px"><span>p50: '+p50+'ms</span><span>p95: '+p95+'ms</span><span>p99: '+p99+'ms</span></div></div>';
}).join('');
$('perf-latency').innerHTML=latHTML;
// Throughput comparison
var maxTps=Math.max(...models.map(function(m){return m.throughput.avg_tokens_per_sec||0}),1);
var tpsHTML=models.map(function(m){
var t=m.throughput||{},avg=t.avg_tokens_per_sec||0,p50=t.p50||0,c=mcol[m.model]||'#38bdf8';
var isAllStreaming = avg===0 && p50===0;
if(isAllStreaming){
return'<div class="mb-2" style="font-size:11px"><div class="d-flex justify-content-between mb-1"><span style="color:#e2e8f0">'+mlab[m.model]+'</span><span style="color:#64748b;font-style:italic">streaming only</span></div><div class="text-secondary" style="font-size:10px">t/s available for non-streaming requests only</div></div>';
}
return'<div class="mb-2" style="font-size:11px"><div class="d-flex justify-content-between mb-1"><span style="color:#e2e8f0">'+mlab[m.model]+'</span><span style="color:'+c+'" class="fw-bold">'+avg+' tok/s</span></div>'+
'<div class="d-flex align-items-center gap-2"><span class="text-secondary" style="min-width:28px">avg</span><div class="flex-grow-1" style="height:6px;background:#1e293b;border-radius:3px;overflow:hidden"><div style="height:100%;width:'+(Math.max(avg/maxTps*100,6))+'%;background:'+c+';border-radius:3px"></div></div><span class="small" style="color:'+c+';min-width:54px;text-align:right">'+avg+' tok/s</span></div>'+
'<div class="d-flex align-items-center gap-2 mt-1"><span class="text-secondary" style="min-width:28px;font-size:10px">p50</span><div class="flex-grow-1" style="height:4px;background:#1e293b;border-radius:2px;overflow:hidden"><div style="height:100%;width:'+(Math.max(p50/maxTps*100,4))+'%;background:'+c+';opacity:0.5;border-radius:2px"></div></div><span style="font-size:10px;color:#64748b">'+p50+' tok/s</span></div></div>';
}).join('');
$('perf-throughput').innerHTML=tpsHTML;
// Routing reasons table
if(reasons.length){
var rHTML='<table class="table table-custom mb-0"><thead><tr><th>Reason</th><th>Count</th><th>Avg Lat</th><th>P95 Lat</th></tr></thead><tbody>';
reasons.forEach(function(r){rHTML+='<tr><td>'+r.reason+'</td><td>'+r.count+'</td><td>'+r.avg_total_ms+'ms</td><td>'+r.p95_total_ms+'ms</td></tr>';});
rHTML+='</tbody></table>';$('perf-reasons').innerHTML=rHTML;
}else{$('perf-reasons').innerHTML='<div class="text-secondary small text-center py-3">-</div>';}
// Agent performance
if(agents.length){
var maxAc=Math.max(...agents.map(function(a){return a.count||0}),1);
var aHTML=agents.map(function(a){return'<div class="mb-2" style="font-size:11px"><div class="d-flex justify-content-between mb-1"><span style="color:#e2e8f0">'+a.agent+'</span><span class="text-secondary">'+a.count+' reqs</span></div><div class="d-flex align-items-center gap-2"><div class="flex-grow-1" style="height:4px;background:#1e293b;border-radius:2px;overflow:hidden"><div style="height:100%;width:'+(a.count/maxAc*100)+'%;background:#38bdf8;border-radius:2px"></div></div><span class="small" style="color:#38bdf8;min-width:60px;text-align:right">'+a.avg_total_ms+'ms avg</span></div></div>';}).join('');
$('perf-agents').innerHTML=aHTML;
}else{$('perf-agents').innerHTML='<div class="text-secondary small text-center py-3">-</div>';}
}
var sseConnected = false;
function poll(){if(sseConnected)return;fetch('/api/state').then(function(r){return r.json()}).then(function(data){render(data);$('connection-status').textContent='live';}).catch(function(){$('connection-status').textContent='reconnecting';});}
// SSE connection tracking
var es = new EventSource('/api/stream');
es.onopen = function(){sseConnected = true; $('connection-status').textContent = 'live (SSE)';};
es.onmessage = function(e){try{render(JSON.parse(e.data));}catch(ex){}};
es.onerror = function(){sseConnected = false; $('connection-status').textContent = 'polling';};
function loadScatter(){
var m=$('scatter-model').value;
fetch('/api/scatter?window=24&model='+m).then(function(r){return r.json()}).then(renderScatter).catch(function(){});
}
function renderScatter(d){
var pts=d.points||[],el=$('scatter-plot'),lg=$('scatter-legend');
if(!pts.length){el.innerHTML='<div class="text-secondary small text-center py-5">No data yet</div>';return;}
var mcol={'strix-moe':'#a78bfa','gpu-dense':'#f59e0b','gpu-vision':'#22c55e','unknown':'#38bdf8'};
var mlab={'strix-moe':'35B MoE','gpu-dense':'27B Dense','gpu-vision':'9B VLM'};
var maxX=Math.max.apply(null,pts.map(function(p){return p.prompt_tokens||0}))||1000;
var maxY=Math.max.apply(null,pts.map(function(p){return p.inference_ms||0}))||5000;
// Log scale for X axis (prompt tokens vary widely)
var toX=function(t){return Math.log10(Math.max(t,1))/Math.log10(Math.max(maxX,10))*100;};
var toY=function(t){return (t/maxY)*100;};
var dots='';
pts.forEach(function(p){
var x=toX(p.prompt_tokens),y=toY(p.inference_ms),c=mcol[p.model]||'#38bdf8';
var r=p.stream?1.5:2.5,o=p.stream?0.4:0.8;
dots+='<circle cx="'+x+'" cy="'+(100-y)+'" r="'+r+'" fill="'+c+'" opacity="'+o+'"><title>'+mlab[p.model]+' | '+p.prompt_tokens+' tok | '+p.inference_ms+'ms | '+p.agent+'</title></circle>';
});
// Grid lines
var grid='';
for(var i=1;i<=4;i++){grid+='<line x1="0" y1="'+(i*20)+'" x2="100" y2="'+(i*20)+'" stroke="#1e293b" stroke-width="0.5"/>';}
for(var i=1;i<=4;i++){grid+='<line x1="'+(i*20)+'" y1="0" x2="'+(i*20)+'" y2="100" stroke="#1e293b" stroke-width="0.5"/>';}
// Axis labels
var xTicks='';
var xVals=[10,100,1000,10000,100000];
xVals.forEach(function(v){if(v<=maxX)xTicks+='<text x="'+toX(v)+'" y="103" text-anchor="middle" font-size="8" fill="#64748b">'+(v>=1000?(v/1000)+'k':v)+'</text>';});
var yTicks='';
var yVals=[500,1000,5000,10000,50000,100000];
yVals.forEach(function(v){if(v<=maxY)yTicks+='<text x="-2" y="'+(97-toY(v))+'" text-anchor="end" font-size="8" fill="#64748b">'+(v>=1000?(v/1000)+'s':v+'ms')+'</text>';});
el.innerHTML='<svg viewBox="-35 0 140 115" style="width:100%;height:200px">'+grid+dots+xTicks+yTicks+'<text x="50" y="112" text-anchor="middle" font-size="9" fill="#475569">Prompt Tokens (log scale)</text><text x="-38" y="50" text-anchor="middle" font-size="9" fill="#475569" transform="rotate(-90,-38,50)">Inference Time</text></svg>';
// Legend
var models=[];pts.forEach(function(p){if(models.indexOf(p.model)===-1)models.push(p.model);});
lg.innerHTML=models.map(function(m){return'<span class="d-flex align-items-center gap-1 small"><svg width="10" height="10"><circle cx="5" cy="5" r="3.5" fill="'+(mcol[m]||'#38bdf8')+'"/></svg>'+mlab[m]+'</span>';}).join('');
}
poll();setInterval(poll,10000);loadTS();loadPerf();setInterval(loadPerf,15000);loadScatter();setInterval(loadScatter,30000);
</script>
</body>
</html>"""
@app.route("/")
def dashboard(): return render_template_string(DASHBOARD_HTML)
@app.route("/api/state")
def api_state(): return fetch_state()
@app.route("/api/scatter")
def api_scatter():
window = request.args.get("window", "24")
model = request.args.get("model", "all")
try:
r = requests.get(f"http://router:9000/metrics/scatter?window={window}&model={model}", timeout=10)
if r.status_code == 200: return r.json()
except Exception: pass
return {"points": [], "count": 0}
@app.route("/api/performance")
def api_performance():
window = request.args.get("window", "24")
model = request.args.get("model", "all")
try:
r = requests.get(f"http://router:9000/metrics/performance?window={window}&model={model}", timeout=10)
if r.status_code == 200: return r.json()
except Exception: pass
return {"models": [], "reasons": [], "agents": [], "summary": {"total_requests": 0}}
@app.route("/api/timeseries")
def api_timeseries():
period = request.args.get("period", "day")
try:
r = requests.get("http://router:9000/metrics/timeseries?period=" + period, timeout=5)
if r.status_code == 200: return r.json()
except Exception: pass
return {"models": {}, "labels": []}
@app.route("/api/stream")
def api_stream():
def ev():
q = queue.Queue()
with sse_lock: sse_subscribers.append(q)
try:
yield "data: "+json.dumps(fetch_state())+"\n\n"
while True:
try: msg = q.get(timeout=3); yield "data: "+msg+"\n\n"
except queue.Empty: yield "data: "+json.dumps(fetch_state())+"\n\n"
except GeneratorExit: pass
finally:
with sse_lock:
if q in sse_subscribers: sse_subscribers.remove(q)
return Response(stream_with_context(ev()), mimetype="text/event-stream", headers={"Cache-Control":"no-cache","X-Accel-Buffering":"no","Access-Control-Allow-Origin":"*"})
@app.route("/health")
def health(): return {"status":"healthy","service":"harness-dashboard"}
if __name__ == "__main__":
app.run(host="0.0.0.0", port=3000, debug=False)
+133
View File
@@ -0,0 +1,133 @@
#!/usr/bin/env python3
"""Syslog Harness Dashboard — Simple HTTP server exposing GPU health + metrics."""
import json
import os
import time
import urllib.request
from http.server import HTTPServer, SimpleHTTPRequestHandler
from datetime import datetime
GPUS = {
"amdpve": {"endpoint": os.getenv("AMDVE_EP", "192.168.68.15:8080"), "model": "qwen3.6-35B-A3B (MoE)", "vram": "65GB"},
"llmgpu": {"endpoint": os.getenv("LLMGPU_EP", "192.168.68.8:8080"), "model": "qwen3.5-27B (Dense)", "vram": "24GB"},
"ocu_llm": {"endpoint": os.getenv("OCU_LLM_EP", "192.168.68.110:8080"), "model": "gemma-4-E4B (Light)", "vram": "12GB"},
}
def check_gpu(name, info):
try:
start = time.time()
# Use simple HTTP GET to check if the GPU endpoint is alive
resp = urllib.request.urlopen(f"http://{info['endpoint']}/", timeout=3)
latency = (time.time() - start) * 1000
return {
"status": "up",
"latency_ms": round(latency, 1),
"model": info["model"],
"vram": info["vram"],
}
except Exception as e:
return {"status": "down", "error": str(e)[:50], "model": info["model"], "vram": info["vram"]}
def get_queue_status():
try:
req = urllib.request.Request("http://queue-service:8091/status")
resp = urllib.request.urlopen(req, timeout=2)
return json.loads(resp.read())
except Exception:
return {"queue_depth": -1, "circuit_breaker": "unknown", "gpu_health": {}}
DASHBOARD_HTML = """
<!DOCTYPE html>
<html><head><meta charset="utf-8"><title>🦅 Syslog Harness</title>
<style>
body { background: #1a1a2e; color: #e0e0e0; font-family: monospace; margin: 0; padding: 20px; }
.card { background: #16213e; border-radius: 8px; padding: 16px; margin: 10px 0; border-left: 4px solid #0f3460; }
.up { border-left-color: #00d26a; } .down { border-left-color: #ff4757; }
.warn { border-left-color: #ffa502; }
h1 { color: #00d26a; font-size: 24px; } h2 { color: #0f3460; font-size: 16px; }
.metric { display: inline-block; margin: 4px 12px; }
.value { font-weight: bold; color: #00d26a; }
#refresh { position: fixed; top: 10px; right: 10px; background: #0f3460; color: white;
border: none; padding: 8px 16px; border-radius: 4px; cursor: pointer; }
table { width: 100%; border-collapse: collapse; margin: 10px 0; }
th, td { text-align: left; padding: 8px; border-bottom: 1px solid #0f3460; }
th { color: #00d26a; }
</style></head><body>
<button id="refresh" onclick="location.reload()">↻ Refresh</button>
<h1>🦅 Syslog Harness Dashboard</h1>
<h2>Updated: <span id="ts"></span></h2>
<div class="card" id="queue-card">
<h2>Queue & Circuit Breaker</h2>
<div class="metric">Depth: <span class="value" id="depth">--</span></div>
<div class="metric">Circuit: <span class="value" id="circuit">--</span></div>
<div class="metric">Threshold: <span class="value" id="threshold">--</span></div>
</div>
<div class="card">
<h2>GPU Endpoints</h2>
<table><tr><th>GPU</th><th>Model</th><th>VRAM</th><th>Status</th><th>Latency</th></tr>
<tbody id="gpu-table"></tbody></table>
</div>
<script>
document.getElementById('ts').textContent = new Date().toISOString();
fetch('/api/status').then(r => r.json()).then(data => {
document.getElementById('depth').textContent = data.queue_depth;
document.getElementById('circuit').textContent = data.circuit_breaker;
document.getElementById('threshold').textContent = 'warn:' + data.thresholds.warn + ' / open:' + data.thresholds.open;
const card = document.getElementById('queue-card');
if (data.circuit_breaker === 'open') card.className = 'card warn';
else if (data.circuit_breaker === 'warn') card.className = 'card warn';
else card.className = 'card up';
let html = '';
for (const [name, gpu] of Object.entries(data.gpu_health)) {
const status = gpu.status === 'up' ? '✅' : '❌';
const latency = gpu.status === 'up' ? gpu.latency_ms + 'ms' : gpu.error;
const rowClass = gpu.status === 'up' ? '' : 'down';
html += `<tr class="${rowClass}"><td>${name}</td><td>${gpu.model}</td><td>${gpu.vram}</td><td>${status}</td><td>${latency}</td></tr>`;
}
document.getElementById('gpu-table').innerHTML = html;
});
setInterval(() => location.reload(), 10000);
</script></body></html>
"""
class Handler(SimpleHTTPRequestHandler):
def do_GET(self):
if self.path == "/" or self.path == "/harness.html":
self.send_response(200)
self.send_header("Content-Type", "text/html; charset=utf-8")
self.end_headers()
self.wfile.write(DASHBOARD_HTML.encode())
elif self.path == "/api/status":
status = get_queue_status()
enriched = {
"queue_depth": status.get("queue_depth", -1),
"circuit_breaker": status.get("circuit_breaker", "unknown"),
"thresholds": status.get("thresholds", {"warn": 30, "open": 50}),
"gpu_health": {},
}
for name, info in GPUS.items():
enriched["gpu_health"][name] = check_gpu(name, info)
self.send_response(200)
self.send_header("Content-Type", "application/json")
self.end_headers()
self.wfile.write(json.dumps(enriched).encode())
else:
self.send_response(404)
self.end_headers()
def log_message(self, format, *args):
pass # Suppress request logs
if __name__ == "__main__":
server = HTTPServer(("0.0.0.0", 3001), Handler)
print("Dashboard running on :3001/harness.html")
server.serve_forever()
-445
View File
@@ -1,445 +0,0 @@
<!DOCTYPE html>
<html lang="en" x-data="dashboard()" x-init="init()">
<head>
<meta charset="UTF-8">
<meta name="viewport" content="width=device-width, initial-scale=1.0">
<title>Inference Harness - Dashboard</title>
<!-- Tailwind CSS -->
<script src="https://cdn.tailwindcss.com"></script>
<!-- Alpine.js -->
<script defer src="https://cdn.jsdelivr.net/npm/alpinejs@3.14.8/dist/cdn.min.js"></script>
<!-- Chart.js -->
<script src="https://cdn.jsdelivr.net/npm/chart.js@4.4.0/dist/chart.umd.min.js"></script>
<!-- Custom Styles -->
<style>
/* Custom scrollbar */
::-webkit-scrollbar { width: 8px; height: 8px; }
::-webkit-scrollbar-track { background: #1f2937; }
::-webkit-scrollbar-thumb { background: #374151; border-radius: 4px; }
::-webkit-scrollbar-thumb:hover { background: #4b5563; }
/* Status dots with pulse animation */
.dot-green { background: #10b981; animation: pulse-green 2s infinite; }
.dot-yellow { background: #f59e0b; animation: pulse-yellow 2s infinite; }
.dot-red { background: #ef4444; animation: pulse-red 2s infinite; }
@keyframes pulse-green {
0%, 100% { opacity: 1; box-shadow: 0 0 0 0 rgba(16, 185, 129, 0.7); }
50% { opacity: 0.8; box-shadow: 0 0 0 6px rgba(16, 185, 129, 0); }
}
@keyframes pulse-yellow {
0%, 100% { opacity: 1; box-shadow: 0 0 0 0 rgba(245, 158, 11, 0.7); }
50% { opacity: 0.8; box-shadow: 0 0 0 6px rgba(245, 158, 11, 0); }
}
@keyframes pulse-red {
0%, 100% { opacity: 1; box-shadow: 0 0 0 0 rgba(239, 68, 68, 0.7); }
50% { opacity: 0.8; box-shadow: 0 0 0 6px rgba(239, 68, 68, 0); }
}
/* Glassmorphism panels */
.glass-panel {
background: rgba(31, 41, 55, 0.7);
backdrop-filter: blur(12px);
border: 1px solid rgba(75, 85, 99, 0.4);
}
/* Smooth transitions */
.transition-all-300 { transition: all 0.3s ease; }
/* Status badges */
.status-badge {
display: inline-flex;
align-items: center;
padding: 0.125rem 0.5rem;
border-radius: 0.375rem;
font-size: 0.75rem;
font-weight: 600;
}
/* Health bar gradient */
.health-bar {
height: 0.5rem;
background-color: #374151;
border-radius: 0.375rem;
overflow: hidden;
}
.health-fill {
height: 100%;
transition: width 0.3s ease;
}
</style>
</head>
<body class="bg-gradient-to-br from-gray-900 via-gray-800 to-gray-900 min-h-screen text-white">
<!-- Loading Overlay -->
<div x-show="isLoading" class="fixed inset-0 bg-gray-900 bg-opacity-90 z-50 flex items-center justify-center">
<div class="text-center">
<div class="w-16 h-16 border-4 border-blue-600 border-t-transparent rounded-full animate-spin mx-auto mb-4"></div>
<p class="text-blue-400 text-lg font-semibold">Loading Dashboard...</p>
</div>
</div>
<!-- Main Container -->
<div class="container mx-auto px-4 py-6 max-w-[1920px]">
<!-- Header Section -->
<div class="flex flex-col lg:flex-row justify-between items-start lg:items-center gap-4 mb-6">
<div class="flex items-center gap-3">
<img src="/favicon.svg" class="w-10 h-10" alt="Logo">
<div>
<h1 class="text-2xl font-bold text-white">Inference Harness</h1>
<p class="text-sm text-gray-400">Syslog Solution LLC Real-time Monitoring</p>
</div>
</div>
<div class="flex items-center gap-6">
<div class="flex items-center gap-2">
<div x-text="globalStatus" x-class="{
'dot-green': globalStatus === 'healthy',
'dot-yellow': globalStatus === 'degraded',
'dot-red': globalStatus === 'critical'
}" class="w-4 h-4 rounded-full"></div>
<span x-text="globalStatus" x-bind:class="{
'text-emerald-400': globalStatus === 'healthy',
'text-amber-400': globalStatus === 'degraded',
'text-red-400': globalStatus === 'critical'
}" class="font-semibold text-lg"></span>
</div>
<button @click="refreshAll()" class="px-4 py-2 bg-blue-600 hover:bg-blue-700 text-white rounded-lg text-sm flex items-center gap-2 transition-all-300">
<svg class="w-5 h-5" fill="none" stroke="currentColor" viewBox="0 0 24 24">
<path stroke-linecap="round" stroke-linejoin="round" stroke-width="2" d="M4 4v5h.582m15.356 2A8.001 8.001 0 004.582 9m0 0H9m11 11v-5h-.581m0 0a8.003 8.003 0 01-15.357-2m15.357 2H15"></path>
</svg>
Refresh
</button>
<span x-text="lastUpdate" class="text-sm text-gray-500"></span>
</div>
</div>
<!-- KPI Cards Row -->
<div class="grid grid-cols-1 md:grid-cols-2 lg:grid-cols-5 gap-4 mb-6">
<div class="glass-panel rounded-xl p-4 transition-all-300 hover:shadow-lg hover:shadow-blue-500/20">
<div class="flex items-center justify-between mb-2">
<span class="text-2xl"></span>
<span x-text="kpi.gpu_count_trend || ''" class="text-sm text-gray-400"></span>
</div>
<p class="text-3xl font-bold text-white" x-text="kpi.gpu_count || 0"></p>
<p class="text-sm text-gray-400 mt-1">GPUs Online</p>
</div>
<div class="glass-panel rounded-xl p-4 transition-all-300 hover:shadow-lg hover:shadow-purple-500/20">
<div class="flex items-center justify-between mb-2">
<span class="text-2xl"></span>
<span x-text="kpi.sessions_trend || ''" class="text-sm text-gray-400"></span>
</div>
<p class="text-3xl font-bold text-white" x-text="kpi.active_sessions || 0"></p>
<p class="text-sm text-gray-400 mt-1">Active Sessions</p>
</div>
<div class="glass-panel rounded-xl p-4 transition-all-300 hover:shadow-lg hover:shadow-red-500/20">
<div class="flex items-center justify-between mb-2">
<span class="text-2xl"></span>
<span x-text="kpi.trips_trend || ''" class="text-sm text-gray-400"></span>
</div>
<p class="text-3xl font-bold text-white" x-text="kpi.circuit_trips || 0"></p>
<p class="text-sm text-gray-400 mt-1">Circuit Breakers</p>
</div>
<div class="glass-panel rounded-xl p-4 transition-all-300 hover:shadow-lg hover:shadow-cyan-500/20">
<div class="flex items-center justify-between mb-2">
<span class="text-2xl"></span>
<span x-text="kpi.latency_trend || ''" class="text-sm text-gray-400"></span>
</div>
<p class="text-3xl font-bold text-white" x-text="(kpi.avg_latency || 0).toFixed(1) + 'ms'"></p>
<p class="text-sm text-gray-400 mt-1">Avg Latency</p>
</div>
<div class="glass-panel rounded-xl p-4 transition-all-300 hover:shadow-lg hover:shadow-green-500/20">
<div class="flex items-center justify-between mb-2">
<span class="text-2xl"></span>
<span x-text="kpi.requests_trend || ''" class="text-sm text-gray-400"></span>
</div>
<p class="text-3xl font-bold text-white" x-text="kpi.requests_minute || 0"></p>
<p class="text-sm text-gray-400 mt-1">Requests/min</p>
</div>
</div>
<!-- GPU Health Scoring (Phase 3) -->
<div class="mb-6">
<div class="flex items-center justify-between mb-4">
<h2 class="text-xl font-semibold text-white flex items-center gap-2">
GPU Health Scoring
</h2>
<div class="text-sm text-gray-400">
Scoring: VRAM (40%) Temp (30%) Load (30%)
</div>
</div>
<div class="grid grid-cols-1 lg:grid-cols-2 gap-6">
<!-- GPU Score Cards -->
<div class="grid grid-cols-1 md:grid-cols-3 gap-4">
<template x-for="gpu in gpuHealth" :key="gpu.id">
<div class="glass-panel rounded-xl p-4 transition-all-300 hover:shadow-lg"
x-bind:class="{
'border-emerald-500/50': gpu.health_score < 30,
'border-amber-500/50': gpu.health_score >= 30 && gpu.health_score < 50,
'border-red-500/50': gpu.health_score >= 50
}">
<div class="flex items-center justify-between mb-3">
<div>
<p class="text-lg font-bold text-white" x-text="gpu.name"></p>
<p class="text-xs text-gray-400" x-text="gpu.model"></p>
</div>
<div x-show="gpu.is_preferred" class="px-2 py-1 bg-emerald-600 rounded-lg text-xs font-semibold">
Preferred
</div>
</div>
<div class="flex items-center justify-between mb-3">
<span class="text-4xl font-bold text-white" x-text="gpu.health_score.toFixed(1)"></span>
</div>
<div class="grid grid-cols-3 gap-2 text-xs text-gray-400 mb-3">
<div><p class="mb-1">VRAM</p><p class="text-white font-semibold" x-text="gpu.vram_pct + '%'"></p></div>
<div><p class="mb-1">Temp</p><p class="text-white font-semibold" x-text="gpu.temp + 'C'"></p></div>
<div><p class="mb-1">Load</p><p class="text-white font-semibold" x-text="gpu.load + '%'"></p></div>
</div>
<div class="health-bar">
<div class="health-fill" x-bind:style="{ width: (100 - gpu.health_score) + '%', 'background-color': gpu.health_score < 30 ? '#10b981' : (gpu.health_score < 50 ? '#f59e0b' : '#ef4444') }"></div>
</div>
</div>
</template>
</div>
<!-- Health Trend Chart -->
<div class="glass-panel rounded-xl p-4">
<h3 class="text-sm font-semibold text-gray-400 mb-3">Health Scores Over Time (1h)</h3>
<canvas id="healthTrendChart" height="200"></canvas>
</div>
</div>
</div>
<!-- Circuit Breaker Status (Phase 1) -->
<div class="mb-6">
<div class="flex items-center justify-between mb-4">
<h2 class="text-xl font-semibold text-white flex items-center gap-2">
Circuit Breaker Status
</h2>
<div class="text-sm text-gray-400">
<span class="inline-flex items-center gap-1 px-2 py-1 bg-emerald-900/50 rounded text-emerald-400 text-xs">
<span class="w-2 h-2 rounded-full bg-emerald-500"></span> Close
</span>
<span class="inline-flex items-center gap-1 px-2 py-1 bg-amber-900/50 rounded text-amber-400 text-xs ml-2">
<span class="w-2 h-2 rounded-full bg-amber-500"></span> Half-Open
</span>
<span class="inline-flex items-center gap-1 px-2 py-1 bg-red-900/50 rounded text-red-400 text-xs ml-2">
<span class="w-2 h-2 rounded-full bg-red-500"></span> Open
</span>
</div>
</div>
<div class="grid grid-cols-1 lg:grid-cols-3 gap-4 mb-6">
<template x-for="gpu in circuitBreakers" :key="gpu.name">
<div class="glass-panel rounded-xl p-4">
<div class="flex items-center justify-between mb-3">
<p class="text-lg font-semibold text-white" x-text="gpu.name"></p>
<span x-show="gpu.is_tripped" x-text="' Tripped'" x-bind:class="{ 'text-red-400': gpu.is_tripped, 'text-amber-400': !gpu.is_tripped && gpu.is_half_open }" class="text-sm"></span>
</div>
<div class="space-y-2">
<template x-for="model in gpu.models" :key="model.name">
<div class="flex items-center justify-between py-2 border-b border-gray-700/50 last:border-0">
<span class="text-sm text-gray-300" x-text="model.name"></span>
<span x-text="model.status" x-bind:class="{
'text-emerald-400 bg-emerald-900/30 px-2 py-1 rounded': model.status === 'close',
'text-amber-400 bg-amber-900/30 px-2 py-1 rounded': model.status === 'half_open',
'text-red-400 bg-red-900/30 px-2 py-1 rounded': model.status === 'open'
}" class="status-badge" x-text="model.status"></span>
</div>
</template>
</div>
<div class="mt-3 text-xs text-gray-500">
<p>Trips: <span class="text-white" x-text="gpu.trip_count"></span></p>
<p>Recovery: <span class="text-white" x-text="gpu.recovery_time || 'N/A'"></span></p>
</div>
</div>
</template>
</div>
<div class="glass-panel rounded-xl p-4">
<h3 class="text-sm font-semibold text-gray-400 mb-3">Circuit Breaker Trips (24h)</h3>
<canvas id="tripHistoryChart" height="200"></canvas>
</div>
</div>
<!-- Session Analytics (Phase 2) -->
<div class="mb-6">
<div class="flex items-center justify-between mb-4">
<h2 class="text-xl font-semibold text-white flex items-center gap-2">
Session Analytics
</h2>
<div class="flex gap-2">
<button @click="sessionTimeRange='1h'" x-bind:class="{'bg-blue-600 text-white': sessionTimeRange === '1h', 'bg-gray-700 text-gray-400': sessionTimeRange !== '1h'}" class="px-3 py-1 rounded-lg text-xs font-semibold">1H</button>
<button @click="sessionTimeRange='6h'" x-bind:class="{'bg-blue-600 text-white': sessionTimeRange === '6h', 'bg-gray-700 text-gray-400': sessionTimeRange !== '6h'}" class="px-3 py-1 rounded-lg text-xs font-semibold">6H</button>
<button @click="sessionTimeRange='24h'" x-bind:class="{'bg-blue-600 text-white': sessionTimeRange === '24h', 'bg-gray-700 text-gray-400': sessionTimeRange !== '24h'}" class="px-3 py-1 rounded-lg text-xs font-semibold">24H</button>
</div>
</div>
<div class="grid grid-cols-1 lg:grid-cols-3 gap-6">
<div class="glass-panel rounded-xl p-4">
<h3 class="text-sm font-semibold text-gray-400 mb-3">Session Distribution</h3>
<canvas id="sessionDistribution" height="250"></canvas>
</div>
<div class="glass-panel rounded-xl p-4">
<h3 class="text-sm font-semibold text-gray-400 mb-3">Peak Usage Times</h3>
<canvas id="peakUsageChart" height="250"></canvas>
</div>
<div class="glass-panel rounded-xl p-4">
<h3 class="text-sm font-semibold text-gray-400 mb-3">Concurrent Sessions</h3>
<canvas id="sessionTrend" height="250"></canvas>
</div>
</div>
</div>
<!-- System Performance -->
<div class="mb-6">
<h2 class="text-xl font-semibold text-white flex items-center gap-2 mb-4">
System Performance
</h2>
<div class="grid grid-cols-1 lg:grid-cols-2 gap-6">
<div class="glass-panel rounded-xl p-4">
<h3 class="text-sm font-semibold text-gray-400 mb-3">Latency Percentiles</h3>
<canvas id="latencyChart" height="250"></canvas>
</div>
<div class="glass-panel rounded-xl p-4">
<h3 class="text-sm font-semibold text-gray-400 mb-3">Error Rates</h3>
<canvas id="errorRates" height="250"></canvas>
</div>
</div>
</div>
<!-- Footer -->
<div class="text-center text-sm text-gray-500 pt-4 border-t border-gray-700">
<p>Inference Harness Dashboard Syslog Solution LLC Last updated: <span x-text="lastUpdate"></span></p>
<p class="mt-1 text-xs">Auto-refresh: every 10 seconds | Manual: Refresh button</p>
</div>
</div>
<!-- Alpine.js Data -->
<script>
function dashboard() {
return {
isLoading: true,
globalStatus: 'healthy',
lastUpdate: new Date().toLocaleString(),
sessionTimeRange: '1h',
refreshInterval: null,
charts: {},
kpi: { gpu_count: 0, active_sessions: 0, circuit_trips: 0, avg_latency: 0, requests_minute: 0 },
gpuHealth: [],
circuitBreakers: [],
sessionData: { distribution: {}, trend: [], peaks: {} },
systemPerf: { latency: { p50: 0, p95: 0, p99: 0 }, errorRates: {} },
init() {
console.log('Initializing Dashboard...');
this.fetchAllData();
this.startAutoRefresh();
},
async fetchAllData() {
this.isLoading = true;
try {
await Promise.all([
this.fetchGPUScores(),
this.fetchCircuitBreakers(),
this.fetchSessionAnalytics(),
this.fetchSystemPerformance()
]);
this.updateGlobalStatus();
this.lastUpdate = new Date().toLocaleString();
} catch (error) {
console.error('Data fetch failed:', error);
this.globalStatus = 'critical';
} finally {
this.isLoading = false;
}
},
async fetchGPUScores() {
try {
const metrics = await fetch('/metrics/circuit-breaker').then(r => r.json());
this.gpuHealth = [
{ id: 'gpu-vision', name: 'Qwen3.5-9B Vision', model: 'gpu-vision', health_score: metrics.gpu_vision?.gpu_health_score || 39.4, vram_pct: 45, temp: 78, load: 65, is_preferred: true },
{ id: 'deepseek-v3', name: 'DeepSeek V3', model: 'deepseek-v3', health_score: metrics.deepseek_v3?.gpu_health_score || 45.9, vram_pct: 60, temp: 82, load: 50, is_preferred: false },
{ id: 'mistral-small', name: 'Mistral Small', model: 'mistral-small', health_score: metrics.mistral_small?.gpu_health_score || 35.0, vram_pct: 30, temp: 65, load: 40, is_preferred: false }
];
console.log('GPU Health Scores loaded:', this.gpuHealth);
} catch (error) { console.error('Failed to load GPU scores:', error); }
},
async fetchCircuitBreakers() {
try {
const metrics = await fetch('/metrics/circuit-breaker').then(r => r.json());
this.circuitBreakers = Object.keys(metrics).map((gpuId) => ({
name: gpuId,
is_tripped: metrics[gpuId].is_circuit_tripped > 0,
is_half_open: metrics[gpuId].half_open_probe && !metrics[gpuId].is_circuit_tripped,
trip_count: metrics[gpuId].trip_count,
recovery_time: metrics[gpuId].last_circuit_trip ? new Date(metrics[gpuId].last_circuit_trip * 1000).toLocaleString() : null,
models: Object.keys(metrics[gpuId].models || {}).map(model => ({ name: model.replace(/_/g, ' '), status: metrics[gpuId].models[model].circuit_breaker_state }))
}));
console.log('Circuit breakers loaded:', this.circuitBreakers);
} catch (error) { console.error('Failed to load circuit breakers:', error); }
},
async fetchSessionAnalytics() {
try {
this.sessionData = {
distribution: { 'gpu-vision': 45, 'deepseek-v3': 30, 'mistral-small': 25 },
trend: Array.from({ length: 24 }, (_, i) => ({ time: `${i}:00`, sessions: Math.floor(Math.random() * 20) + 10 })),
peaks: { '09:00': 25, '14:00': 30, '18:00': 20 }
};
console.log('Session analytics loaded');
} catch (error) { console.error('Failed to load session analytics:', error); }
},
async fetchSystemPerformance() {
try {
this.systemPerf = {
latency: { p50: Math.floor(Math.random() * 50) + 100, p95: Math.floor(Math.random() * 200) + 250, p99: Math.floor(Math.random() * 500) + 400 },
errorRates: { 'gpu-vision': Math.random() * 0.01, 'deepseek-v3': Math.random() * 0.02, 'mistral-small': Math.random() * 0.015 }
};
console.log('System performance loaded');
} catch (error) { console.error('Failed to load system performance:', error); }
},
updateGlobalStatus() {
const hasCircuitTrips = this.circuitBreakers.some(gpu => gpu.is_tripped);
const hasHighLatency = this.systemPerf.latency.p99 > 1000;
if (hasCircuitTrips) this.globalStatus = 'degraded';
else if (hasHighLatency) this.globalStatus = 'degraded';
else this.globalStatus = 'healthy';
},
startAutoRefresh() {
this.refreshInterval = setInterval(() => { this.fetchAllData(); console.log('Auto-refreshing dashboard data...'); }, 10000);
},
refreshAll() { console.log('Manual refresh triggered'); this.fetchAllData(); },
async initCharts() {
try {
this.charts.healthTrend = new Chart(document.getElementById('healthTrendChart'), {
type: 'line', data: {
labels: Array.from({ length: 60 }, (_, i) => `${i}m`),
datasets: this.gpuHealth.map(gpu => ({ label: gpu.name, data: Array.from({ length: 60 }, () => gpu.health_score + (Math.random() * 10 - 5)), borderColor: this.getGPUColor(gpu.name), tension: 0.3, pointRadius: 0 }))
},
options: { responsive: true, maintainAspectRatio: false, plugins: { legend: { display: false }, tooltip: { mode: 'index', intersect: false } }, scales: { x: { grid: { color: '#374151' }, ticks: { color: '#9ca3af', font: { size: 10 } } }, y: { grid: { color: '#374151' }, ticks: { color: '#9ca3af', font: { size: 10 } }, min: 0, max: 100 } } }
});
console.log('Health trend chart initialized');
} catch (error) { console.error('Failed to initialize charts:', error); }
},
getGPUColor(name) {
const colors = { 'gpu-vision': '#3b82f6', 'deepseek-v3': '#8b5cf6', 'mistral-small': '#10b981' };
return colors[name] || '#9ca3af';
}
};
}
</script>
</body>
</html>
-2
View File
@@ -1,2 +0,0 @@
flask==3.1.*
requests==2.32.*
+35 -107
View File
@@ -3,124 +3,52 @@ version: "3.8"
services:
redis:
image: redis:7-alpine
container_name: harness-redis
restart: unless-stopped
ports:
- "127.0.0.1:6379:6379"
restart: always
networks:
- gpu-router-net
volumes:
- redis-data:/data
command: redis-server --appendonly yes --maxmemory 256mb --maxmemory-policy allkeys-lru
healthcheck:
test: ["CMD", "redis-cli", "ping"]
interval: 10s
timeout: 3s
retries: 5
postgres:
image: postgres:16-alpine
container_name: harness-postgres
restart: unless-stopped
queue-service:
build:
context: .
dockerfile: Dockerfile.queue
restart: always
networks:
- gpu-router-net
ports:
- "127.0.0.1:5432:5432"
environment:
- POSTGRES_DB=litellm
- POSTGRES_USER=litellm
- POSTGRES_PASSWORD=${POSTGRES_PASSWORD}
volumes:
- pgdata:/var/lib/postgresql/data
healthcheck:
test: ["CMD-SHELL", "pg_isready -U litellm"]
interval: 5s
timeout: 3s
retries: 5
litellm:
image: ghcr.io/docker.litellm.ai/berriai/litellm:1.99.1
command: ["--config", "/app/config.yaml", "--port", "4000"]
container_name: harness-litellm
restart: unless-stopped
ports:
- "4000:4000"
volumes:
- ./litellm_config.yaml:/app/config.yaml
- /opt/combined-ca-bundle.pem:/etc/ssl/certs/ca-certificates.crt:ro
environment:
- PROMETHEUS_EXPORTER=true
- ENFORCE_PRISMA_MIGRATION_CHECK=true
- LITELLM_MASTER_KEY=${LITELLM_MASTER_KEY}
- DATABASE_URL=postgresql://litellm:${POSTGRES_PASSWORD}@postgres:5432/litellm
- STORE_MODEL_IN_DB=True
- LITELLM_UI_USERNAME=admin
- LITELLM_UI_PASSWORD=${LITELLM_UI_PASSWORD}
- UI_USERNAME=admin
- UI_PASSWORD=${UI_PASSWORD}
- OPENAI_API_KEY=${OPENAI_API_KEY}
- PROXY_BASE_URL=https://litellm.sysloggh.net
- DOCS_URL=/docs
- ANTHROPIC_API_KEY=${ANTHROPIC_API_KEY}
- GENERIC_CLIENT_ID=FHd7bs9dP5gHad2Ki23iUL5kQvFa0GRaj3nlLnNU
- GENERIC_CLIENT_SECRET=${GENERIC_CLIENT_SECRET}
- GENERIC_AUTHORIZATION_ENDPOINT=https://auth.sysloggh.net/application/o/authorize/
- GENERIC_TOKEN_ENDPOINT=http://harness-nginx/application/o/token/
- GENERIC_USERINFO_ENDPOINT=http://harness-nginx/application/o/userinfo/
- GENERIC_SCOPE=openid email profile
- MCP_SERVER_RAHOS_URL=http://192.168.68.65:3100/mcp
- MCP_SERVER_RAHOS_TRANSPORT=http
- GENERIC_ROLE_MAPPINGS_DEFAULT_ROLE=proxy_admin
- SSL_CERT_FILE=/etc/ssl/certs/ca-certificates.crt
extra_hosts:
- "host.docker.internal:host-gateway"
- "auth.sysloggh.net:192.168.68.11"
healthcheck:
test: ["CMD", "python3", "-c", "import urllib.request; urllib.request.urlopen('http://localhost:4000/health/liveliness')"]
interval: 15s
timeout: 5s
retries: 3
start_period: 120s
- "8091:8091"
depends_on:
postgres:
condition: service_healthy
redis:
condition: service_healthy
nginx:
image: nginx:alpine
container_name: harness-nginx
restart: unless-stopped
ports:
- "80:80"
volumes:
- ./nginx/nginx.conf:/etc/nginx/nginx.conf:ro
- ./dashboard:/opt/inference-harness/dashboard:ro
extra_hosts:
- "auth.sysloggh.net:192.168.68.11"
healthcheck:
test: ["CMD", "curl", "-f", "http://127.0.0.1/health"]
interval: 30s
timeout: 15s
retries: 3
depends_on:
- litellm
- dashboard
- redis
environment:
- REDIS_HOST=redis
- REDIS_PORT=6379
dashboard:
build: ./dashboard
container_name: harness-dashboard
restart: unless-stopped
build:
context: .
dockerfile: Dockerfile.dashboard
restart: always
networks:
- gpu-router-net
ports:
- "127.0.0.1:3000:3000"
environment:
- REDIS_URL=redis://redis:6379
- GPU_SIDECARS=192.168.68.15:8090,192.168.68.8:8090,192.168.68.110:8090
healthcheck:
test: ["CMD", "python3", "-c", "import urllib.request; urllib.request.urlopen('http://localhost:3000/health')"]
interval: 15s
timeout: 5s
retries: 3
- "3001:3001"
depends_on:
- redis
gpu-dashboard:
build:
context: .
dockerfile: Dockerfile.gpu
restart: always
networks:
- gpu-router-net
ports:
- "8092:8092"
networks:
gpu-router-net:
driver: bridge
volumes:
redis-data:
pgdata:
-145
View File
@@ -1,145 +0,0 @@
version: "3.8"
services:
redis:
image: redis:7-alpine
container_name: harness-redis
restart: unless-stopped
ports:
- "127.0.0.1:6379:6379"
volumes:
- redis-data:/data
command: redis-server --appendonly yes --maxmemory 256mb --maxmemory-policy allkeys-lru
healthcheck:
test: ["CMD", "redis-cli", "ping"]
interval: 10s
timeout: 3s
retries: 5
postgres:
image: postgres:16-alpine
container_name: harness-postgres
restart: unless-stopped
ports:
- "127.0.0.1:5432:5432"
environment:
- POSTGRES_DB=litellm
- POSTGRES_USER=litellm
- POSTGRES_PASSWORD=d9fc143e3dc1a7a8e672c359fea95c5e
volumes:
- pgdata:/var/lib/postgresql/data
healthcheck:
test: ["CMD-SHELL", "pg_isready -U litellm"]
interval: 5s
timeout: 3s
retries: 5
router:
build: ./router
container_name: harness-router
restart: unless-stopped
ports:
- "127.0.0.1:9000:9000"
environment:
- REDIS_URL=redis://redis:6379
- GPU_DENSE_URL=http://192.168.68.8:8080/v1
- GPU_MOE_URL=http://192.168.68.15:8080/v1
- GPU_LIGHT_URL=http://192.168.68.110:8080/v1
- API_KEYS={"sk-syslog-local-master-key":{"tier":"enterprise","agent":"admin","deprecated":true},"sk-9e65b69a67-e54af421c1b09fb8bd4f75dacb38cb64":{"tier":"enterprise","agent":"admin"},"sk-syslog-abiba":{"tier":"enterprise","agent":"Abiba","deprecated":true},"sk-856ffb0bbb-e5aaf78b10054eca608f8fbcbd73a889":{"tier":"enterprise","agent":"Abiba"},"sk-syslog-mumuni":{"tier":"enterprise","agent":"Mumuni","deprecated":true},"sk-b57e6e042e-47573660114f3138c852c47f62da807e":{"tier":"enterprise","agent":"Mumuni"},"sk-syslog-tanko":{"tier":"enterprise","agent":"Tanko","deprecated":true},"sk-620a05e95a-e93d875476b650a4d1137249ead8eaa7":{"tier":"enterprise","agent":"Tanko"},"sk-syslog-koby":{"tier":"enterprise","agent":"Koby","deprecated":true},"sk-eb3e6fc1c0-de1bf2edf35a53cb3749a2400483fdee":{"tier":"enterprise","agent":"Koby"},"sk-syslog-kagenz0":{"tier":"enterprise","agent":"Kagenz0","deprecated":true},"sk-12b66b3392-b548aed9138aeb6f698e8e521650ed9b":{"tier":"enterprise","agent":"Kagenz0"},"sk-syslog-koonimo":{"tier":"enterprise","agent":"Koonimo","deprecated":true},"sk-680d06686c-00ee8bf9dc3c93b276af122d49a14dfe":{"tier":"enterprise","agent":"Koonimo"},"sk-starter-abc123":{"tier":"starter","agent":"test-starter","deprecated":true},"sk-55da55907a-1bd7ff344e26feda50e9ac2219697860":{"tier":"starter","agent":"test-starter"},"sk-professional-xyz789":{"tier":"professional","agent":"test-pro","deprecated":true},"sk-b5159863e6-8df3ae52fb958cfe76cc2888c8c8e676":{"tier":"professional","agent":"test-pro"}}
- ADMIN_KEY=sk-admin-ee09fffd04978b61a1569ac670c68814
healthcheck:
test: ["CMD", "python3", "-c", "import urllib.request; urllib.request.urlopen('http://localhost:9000/health')"]
interval: 30s
timeout: 15s
retries: 3
depends_on:
redis:
condition: service_healthy
litellm:
image: ghcr.io/docker.litellm.ai/berriai/litellm:1.90.0-rc.1
command: ["--config", "/app/config.yaml", "--port", "4000"]
container_name: harness-litellm
restart: unless-stopped
ports:
- "4000:4000"
volumes:
- ./litellm_config.yaml:/app/config.yaml
- /opt/combined-ca-bundle.pem:/etc/ssl/certs/ca-certificates.crt:ro
environment:
- PROMETHEUS_EXPORTER=true
- LITELLM_MASTER_KEY=sk-litellm-7f96080dd99b15c36bd4b333b58a6796
- ROUTER_API_KEY=sk-9e65b69a67-e54af421c1b09fb8bd4f75dacb38cb64
- DATABASE_URL=postgresql://litellm:d9fc143e3dc1a7a8e672c359fea95c5e@postgres:5432/litellm
- STORE_MODEL_IN_DB=True
- LITELLM_UI_USERNAME=admin
- LITELLM_UI_PASSWORD=syslog-admin-2026
- UI_USERNAME=admin
- UI_PASSWORD=syslog-admin-2026
- OPENAI_API_KEY=not-used
- PROXY_BASE_URL=https://litellm.sysloggh.net
- DOCS_URL=/docs
- ANTHROPIC_API_KEY=not-used
- GENERIC_CLIENT_ID=FHd7bs9dP5gHad2Ki23iUL5kQvFa0GRaj3nlLnNU
- GENERIC_CLIENT_SECRET=aDkXQx82duqpxc98xxqp0quzzUf1mawsnOTqj7sx1acaS7rWSt02N5ksBCi92n8ZilRavigoYME6fLakP20Ixc9H2pxnSZFiOqQLb7BPi8UtsvfxmzXklD0HJIdbKFxe
- GENERIC_AUTHORIZATION_ENDPOINT=https://auth.sysloggh.net/application/o/authorize/
- GENERIC_TOKEN_ENDPOINT=http://harness-nginx/application/o/token/
- GENERIC_USERINFO_ENDPOINT=http://harness-nginx/application/o/userinfo/
- GENERIC_SCOPE=openid email profile
- GENERIC_ROLE_MAPPINGS_DEFAULT_ROLE=proxy_admin
- SSL_CERT_FILE=/etc/ssl/certs/ca-certificates.crt
extra_hosts:
- "host.docker.internal:host-gateway"
- "auth.sysloggh.net:192.168.68.11"
healthcheck:
test: ["CMD", "python3", "-c", "import urllib.request; urllib.request.urlopen('http://localhost:4000/health/liveliness')"]
interval: 15s
timeout: 5s
retries: 3
depends_on:
postgres:
condition: service_healthy
redis:
condition: service_healthy
nginx:
image: nginx:alpine
container_name: harness-nginx
restart: unless-stopped
ports:
- "80:80"
volumes:
- ./nginx/nginx.conf:/etc/nginx/nginx.conf:ro
- ./dashboard:/opt/inference-harness/dashboard:ro
extra_hosts:
- "auth.sysloggh.net:192.168.68.11"
healthcheck:
test: ["CMD", "curl", "-f", "http://127.0.0.1/health"]
interval: 30s
timeout: 15s
retries: 3
depends_on:
- litellm
- dashboard
dashboard:
build: ./dashboard
container_name: harness-dashboard
restart: unless-stopped
ports:
- "127.0.0.1:3000:3000"
environment:
- REDIS_URL=redis://redis:6379
- GPU_SIDECARS=192.168.68.15:8090,192.168.68.8:8090,192.168.68.110:8090
healthcheck:
test: ["CMD", "python3", "-c", "import urllib.request; urllib.request.urlopen('http://localhost:3000/health')"]
interval: 15s
timeout: 5s
retries: 3
depends_on:
- redis
volumes:
redis-data:
pgdata:
+115
View File
@@ -0,0 +1,115 @@
#!/usr/bin/env python3
"""GPU metrics collector — polls sidecars + llama.cpp every 10s, writes to Workspace."""
import urllib.request, json, time, os
HOSTS = [
{"name": "amdpve", "host": "192.168.68.15", "gpu": "AMD Strix Halo", "llama_port": 8080},
{"name": "llmgpu", "host": "192.168.68.8", "gpu": "RTX 3090", "llama_port": 8080},
{"name": "ocu-llm", "host": "192.168.68.110", "gpu": "RTX 5070", "llama_port": 8080},
]
OUTPUT = "/root/hermes-workspace/public/gpu_metrics.json"
INTERVAL = 10
STALE_THRESHOLD = 30 # seconds before marking stale
DEAD_THRESHOLD = 60 # seconds before marking unreachable
last_seen = {}
def fetch_json(url, timeout=3):
try:
req = urllib.request.Request(url)
resp = urllib.request.urlopen(req, timeout=timeout)
return json.loads(resp.read().decode())
except Exception:
return None
def collect_one(h):
"""Collect GPU hardware + llama.cpp inference state for one host."""
name = h["name"]
host = h["host"]
now = time.time()
# GPU hardware from sidecar
gpu = fetch_json(f"http://{host}:8090/")
# llama.cpp inference state
llamacpp_health = fetch_json(f"http://{host}:{h['llama_port']}/health")
llamacpp_models = fetch_json(f"http://{host}:{h['llama_port']}/v1/models")
# Determine inference state
model_name = None
inference_state = "unknown"
if llamacpp_models:
models = llamacpp_models.get("data", [])
if models:
model_name = models[0].get("id")
if llamacpp_health:
status = llamacpp_health.get("status", "")
if status == "ok":
idle = llamacpp_health.get("slots_idle", 0)
processing = llamacpp_health.get("slots_processing", 0)
if idle and not processing:
inference_state = "idle"
elif processing:
inference_state = "busy"
else:
inference_state = "idle"
# Check for /slots endpoint for is_processing detail
slots = fetch_json(f"http://{host}:{h['llama_port']}/slots")
if slots and isinstance(slots, list) and len(slots) > 0:
if slots[0].get("is_processing"):
inference_state = "busy"
result = {
"host": name,
"gpu_name": h["gpu"],
"inference": {
"state": inference_state,
"model": model_name,
},
"hardware": gpu if gpu else None,
"online": gpu is not None,
"timestamp": now,
}
if gpu is not None:
last_seen[name] = now
if name in last_seen:
age = now - last_seen[name]
if age > DEAD_THRESHOLD:
result["online"] = False
elif age > STALE_THRESHOLD:
result["stale"] = True
return result
def main():
print(f"GPU collector starting, output={OUTPUT}, interval={INTERVAL}s")
os.makedirs(os.path.dirname(OUTPUT), exist_ok=True)
while True:
start = time.time()
results = [collect_one(h) for h in HOSTS]
payload = {
"updated": start,
"gpus": results,
}
with open(OUTPUT + ".tmp", "w") as f:
json.dump(payload, f)
os.rename(OUTPUT + ".tmp", OUTPUT)
elapsed = time.time() - start
sleep_for = max(0, INTERVAL - elapsed)
time.sleep(sleep_for)
if __name__ == "__main__":
main()
+183
View File
@@ -0,0 +1,183 @@
<!DOCTYPE html>
<html lang="en">
<head>
<meta charset="UTF-8">
<meta name="viewport" content="width=device-width, initial-scale=1.0">
<title>GPU Monitor</title>
<style>
* { margin: 0; padding: 0; box-sizing: border-box; }
body { background: #0d1117; color: #c9d1d9; font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', sans-serif; padding: 20px; }
h1 { font-size: 1.3em; margin-bottom: 4px; }
.topbar { display: flex; justify-content: space-between; align-items: center; margin-bottom: 20px; padding-bottom: 12px; border-bottom: 1px solid #21262d; }
.topbar .status { font-size: 0.85em; color: #8b949e; }
.topbar .status .dot { display: inline-block; width: 8px; height: 8px; border-radius: 50%; margin-right: 6px; }
.dot.green { background: #3fb950; }
.dot.yellow { background: #d2991d; }
.dot.red { background: #f85149; }
.cards { display: grid; grid-template-columns: repeat(auto-fit, minmax(320px, 1fr)); gap: 16px; }
.card { background: #161b22; border: 1px solid #21262d; border-radius: 8px; padding: 16px; }
.card.stale { opacity: 0.5; }
.card.dead { opacity: 0.3; border-color: #f85149; }
.card-header { display: flex; justify-content: space-between; align-items: center; margin-bottom: 12px; }
.card-header .name { font-weight: 600; font-size: 1.05em; }
.card-header .host { font-size: 0.8em; color: #8b949e; }
.card-header .state { font-size: 0.75em; padding: 2px 8px; border-radius: 10px; font-weight: 600; }
.state.idle { background: #1b3826; color: #3fb950; }
.state.busy { background: #3d1f1a; color: #f85149; }
.state.unknown { background: #21262d; color: #8b949e; }
.metric { margin-bottom: 10px; }
.metric-label { display: flex; justify-content: space-between; font-size: 0.82em; color: #8b949e; margin-bottom: 2px; }
.metric-label .val { color: #c9d1d9; font-weight: 500; }
.bar { height: 6px; border-radius: 3px; background: #21262d; overflow: hidden; }
.bar-fill { height: 100%; border-radius: 3px; transition: width 0.5s ease; }
.bar-fill.temp-cool { background: #3fb950; }
.bar-fill.temp-warm { background: #d2991d; }
.bar-fill.temp-hot { background: #f85149; }
.bar-fill.util { background: #58a6ff; }
.bar-fill.vram { background: #bc8cff; }
.bar-fill.power { background: #f0883e; }
.model-line { font-size: 0.82em; color: #8b949e; margin-top: 8px; padding-top: 8px; border-top: 1px solid #21262d; }
.model-line span { color: #c9d1d9; }
.error { color: #f85149; font-size: 0.85em; }
</style>
</head>
<body>
<div class="topbar">
<div>
<h1><a href="/" style="color:#58a6ff;text-decoration:none;">← Workspace</a> · GPU Monitor</h1>
<span class="status"><span class="dot green" id="status-dot"></span><span id="status-text">Loading...</span></span>
</div>
<div class="status" id="age">—</div>
</div>
<div class="cards" id="cards"></div>
<script>
const INTERVAL = 5000;
let lastFetchTime = null;
function updateClock() {
const el = document.getElementById('age');
if (!lastFetchTime) { el.textContent = '—'; return; }
const age = Math.round((Date.now() / 1000) - lastFetchTime);
el.textContent = age <= 60 ? `updated ${age}s ago` : `stale ${age}s ago`;
}
setInterval(updateClock, 1000);
const TEMP_WARN = 70, TEMP_HOT = 82;
const VRAM_WARN = 80, VRAM_HOT = 92;
function tempClass(c) { return c > TEMP_HOT ? 'temp-hot' : c > TEMP_WARN ? 'temp-warm' : 'temp-cool'; }
function vramClass(pct) { return pct > VRAM_HOT ? 'temp-hot' : pct > VRAM_WARN ? 'temp-warm' : 'temp-cool'; }
function pct(val, max) { return max ? Math.round(val / max * 100) : 0; }
function mbToGB(mb) { return mb ? (mb / 1024).toFixed(1) : '—'; }
function renderCard(g) {
const hw = g.hardware || {};
const inf = g.inference || {};
const online = g.online !== false;
const stale = g.stale === true;
let cardClass = '';
if (!online) cardClass = 'dead';
else if (stale) cardClass = 'stale';
let stateClass = inf.state || 'unknown';
let stateLabel = inf.state ? inf.state.toUpperCase() : 'UNKNOWN';
if (!online) { stateClass = 'unknown'; stateLabel = 'OFFLINE'; }
const temp = hw.temp_c;
const util = hw.gpu_util_pct;
const vramUsed = hw.vram_used_mb;
const vramTotal = hw.vram_total_mb;
const power = hw.power_w;
const powerLimit = hw.power_limit_w;
const fan = hw.fan_pct;
const vendor = hw.vendor;
let html = `<div class="card ${cardClass}">`;
html += `<div class="card-header">`;
html += `<div><div class="name">${g.gpu_name}</div><div class="host">${g.host}</div></div>`;
html += `<div class="state ${stateClass}">${stateLabel}</div>`;
html += `</div>`;
if (!online) {
html += `<div class="error">Unreachable</div>`;
} else if (hw.error) {
html += `<div class="error">${hw.error}</div>`;
} else {
// Temperature
if (temp != null) {
html += `<div class="metric"><div class="metric-label"><span>Temperature</span><span class="val">${temp}°C</span></div>`;
html += `<div class="bar"><div class="bar-fill ${tempClass(temp)}" style="width:${Math.min(temp,100)}%"></div></div></div>`;
}
// Utilization
if (util != null) {
html += `<div class="metric"><div class="metric-label"><span>GPU Utilization</span><span class="val">${util}%</span></div>`;
html += `<div class="bar"><div class="bar-fill util" style="width:${util}%"></div></div></div>`;
}
// VRAM
if (vramUsed != null && vramTotal != null) {
const vramPct = pct(vramUsed, vramTotal);
html += `<div class="metric"><div class="metric-label"><span>VRAM</span><span class="val">${mbToGB(vramUsed)} / ${mbToGB(vramTotal)} GB</span></div>`;
html += `<div class="bar"><div class="bar-fill ${vramClass(vramPct)}" style="width:${vramPct}%"></div></div></div>`;
}
// Power
if (power != null) {
const powerPct = powerLimit ? pct(power, powerLimit) : 0;
const powerText = powerLimit ? `${power}W / ${powerLimit}W` : `${power}W`;
html += `<div class="metric"><div class="metric-label"><span>Power</span><span class="val">${powerText}</span></div>`;
if (powerLimit) html += `<div class="bar"><div class="bar-fill power" style="width:${powerPct}%"></div></div>`;
html += `</div>`;
}
// Fan (NVIDIA only)
if (fan != null) {
html += `<div class="metric"><div class="metric-label"><span>Fan Speed</span><span class="val">${fan}%</span></div>`;
html += `<div class="bar"><div class="bar-fill util" style="width:${fan}%"></div></div></div>`;
}
}
// Model loaded
html += `<div class="model-line">Model: <span>${inf.model || '—'}</span></div>`;
html += `</div>`;
return html;
}
async function refresh() {
try {
const resp = await fetch('gpu_metrics.json?t=' + Date.now());
const data = await resp.json();
const gpus = data.gpus || [];
document.getElementById('cards').innerHTML = gpus.map(renderCard).join('');
// Top bar status
const online = gpus.filter(g => g.online !== false).length;
const total = gpus.length;
const dot = document.getElementById('status-dot');
const txt = document.getElementById('status-text');
if (online === total) { dot.className = 'dot green'; txt.textContent = `${online}/${total} online`; }
else if (online > 0) { dot.className = 'dot yellow'; txt.textContent = `${online}/${total} online`; }
else { dot.className = 'dot red'; txt.textContent = 'All offline'; }
// Capture fetch time for live clock
lastFetchTime = Date.now() / 1000;
} catch(e) {
document.getElementById('status-dot').className = 'dot red';
document.getElementById('status-text').textContent = 'Collector down';
}
}
// Render skeletons instantly
const SKELETONS = [
{host:'amdpve', gpu_name:'AMD Strix Halo', hardware:{}, inference:{}, online:true},
{host:'llmgpu', gpu_name:'RTX 3090', hardware:{}, inference:{}, online:true},
{host:'ocu-llm', gpu_name:'RTX 5070', hardware:{}, inference:{}, online:true},
];
document.getElementById('cards').innerHTML = SKELETONS.map(g =>
`<div class="card"><div class="card-header"><div><div class="name">${g.gpu_name}</div><div class="host">${g.host}</div></div><div class="state unknown">···</div></div><div class="model-line" style="color:#8b949e;">Loading metrics...</div></div>`
).join('');
refresh();
setInterval(refresh, INTERVAL);
</script>
</body>
</html>
+115
View File
@@ -0,0 +1,115 @@
#!/usr/bin/env python3
"""GPU metrics collector — polls sidecars + llama.cpp every 10s, writes to Workspace."""
import urllib.request, json, time, os
HOSTS = [
{"name": "amdpve", "host": "192.168.68.15", "gpu": "AMD Strix Halo", "llama_port": 8080},
{"name": "llmgpu", "host": "192.168.68.8", "gpu": "RTX 3090", "llama_port": 8080},
{"name": "ocu-llm", "host": "192.168.68.110", "gpu": "RTX 5070", "llama_port": 8080},
]
OUTPUT = "/app/public/gpu_metrics.json"
INTERVAL = 10
STALE_THRESHOLD = 30 # seconds before marking stale
DEAD_THRESHOLD = 60 # seconds before marking unreachable
last_seen = {}
def fetch_json(url, timeout=3):
try:
req = urllib.request.Request(url)
resp = urllib.request.urlopen(req, timeout=timeout)
return json.loads(resp.read().decode())
except Exception:
return None
def collect_one(h):
"""Collect GPU hardware + llama.cpp inference state for one host."""
name = h["name"]
host = h["host"]
now = time.time()
# GPU hardware from sidecar
gpu = fetch_json(f"http://{host}:8090/")
# llama.cpp inference state
llamacpp_health = fetch_json(f"http://{host}:{h['llama_port']}/health")
llamacpp_models = fetch_json(f"http://{host}:{h['llama_port']}/v1/models")
# Determine inference state
model_name = None
inference_state = "unknown"
if llamacpp_models:
models = llamacpp_models.get("data", [])
if models:
model_name = models[0].get("id")
if llamacpp_health:
status = llamacpp_health.get("status", "")
if status == "ok":
idle = llamacpp_health.get("slots_idle", 0)
processing = llamacpp_health.get("slots_processing", 0)
if idle and not processing:
inference_state = "idle"
elif processing:
inference_state = "busy"
else:
inference_state = "idle"
# Check for /slots endpoint for is_processing detail
slots = fetch_json(f"http://{host}:{h['llama_port']}/slots")
if slots and isinstance(slots, list) and len(slots) > 0:
if slots[0].get("is_processing"):
inference_state = "busy"
result = {
"host": name,
"gpu_name": h["gpu"],
"inference": {
"state": inference_state,
"model": model_name,
},
"hardware": gpu if gpu else None,
"online": gpu is not None,
"timestamp": now,
}
if gpu is not None:
last_seen[name] = now
if name in last_seen:
age = now - last_seen[name]
if age > DEAD_THRESHOLD:
result["online"] = False
elif age > STALE_THRESHOLD:
result["stale"] = True
return result
def main():
print(f"GPU collector starting, output={OUTPUT}, interval={INTERVAL}s")
os.makedirs(os.path.dirname(OUTPUT), exist_ok=True)
while True:
start = time.time()
results = [collect_one(h) for h in HOSTS]
payload = {
"updated": start,
"gpus": results,
}
with open(OUTPUT + ".tmp", "w") as f:
json.dump(payload, f)
os.rename(OUTPUT + ".tmp", OUTPUT)
elapsed = time.time() - start
sleep_for = max(0, INTERVAL - elapsed)
time.sleep(sleep_for)
if __name__ == "__main__":
main()
+14
View File
@@ -0,0 +1,14 @@
#!/bin/bash
set -e
# Start collector as background process
cd /root/hermes-workspace/public
python3 /app/collector.py &
COLLECTOR_PID=$!
echo "Collector started (PID $COLLECTOR_PID)"
echo "Serving dashboard on :8092"
# Serve the public directory (contains gpu.html + gpu_metrics.json)
cd /root/hermes-workspace/public
python3 -m http.server 8092
+122
View File
@@ -0,0 +1,122 @@
## Syslog GPU Router — Nginx Configuration (Docker-internal)
## Routes incoming agent requests to the appropriate GPU backend
## based on the X-Syslog-Model header.
upstream amdpve_pool {
## Strix Halo 395 — qwen3.6-35B-A3B (MoE) — Default workhorse
server 192.168.68.15:8080;
}
upstream llmgpu_pool {
## RTX 3090 — qwen3.5-27B (Dense) — Heavy reasoning
server 192.168.68.8:8080;
}
upstream ocu_llm_pool {
## RTX 5070 — gemma-4 (Dense 4B) — Ultra-light tasks
server 192.168.68.110:8080;
}
upstream queue_service {
## Agent queue with circuit breaker (Docker container)
server queue-service:8091;
}
upstream dashboard_service {
## Harness dashboard (Docker container)
server syslog-harness-dashboard-1:3001;
}
upstream gpu_dashboard_pool {
## GPU dashboard (Docker container)
server syslog-harness-gpu-dashboard-1:8092;
}
## ------------------------------------------------------------------
## Mapping: X-Syslog-Model header → upstream backend
## ------------------------------------------------------------------
map $http_x_syslog_model $gpu_upstream {
default amdpve_pool;
"standard" amdpve_pool;
"heavy" llmgpu_pool;
"qwen3.5-27B" llmgpu_pool;
"light" ocu_llm_pool;
"gemma-4" ocu_llm_pool;
}
## Rate limit zone — 10 req/s per IP, burst of 20
limit_req_zone $binary_remote_addr zone=perip:10m rate=10r/s;
server {
listen 80;
server_name _;
## ------------------------------------------------------------------
## Dashboard — observability UI (MUST be before / catch-all)
## ------------------------------------------------------------------
location /dashboard {
proxy_pass http://dashboard_service/;
proxy_set_header Host $host;
proxy_set_header X-Real-IP $remote_addr;
proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
}
## ------------------------------------------------------------------
## GPU Dashboard — observability UI (MUST be before / catch-all)
## ------------------------------------------------------------------
location /gpu {
proxy_pass http://gpu_dashboard_pool/;
proxy_set_header Host $host;
proxy_set_header X-Real-IP $remote_addr;
proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
proxy_set_header X-Forwarded-Proto $scheme;
}
## ------------------------------------------------------------------
## Main location — proxy to selected upstream
## ------------------------------------------------------------------
location / {
limit_req zone=perip burst=20 nodelay;
limit_req_status 503;
proxy_pass http://$gpu_upstream;
## Preserve original host and headers
proxy_set_header Host $host;
proxy_set_header X-Real-IP $remote_addr;
proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
proxy_set_header X-Forwarded-Proto $scheme;
## Pass through the model header so backends can log it
proxy_pass_header X-Syslog-Model;
## Streaming support (SSE for LLM responses)
proxy_buffering off;
proxy_cache off;
proxy_read_timeout 300s;
proxy_send_timeout 300s;
## Basic failover — retry on error or timeout
proxy_next_upstream error timeout http_502 http_503;
proxy_next_upstream_tries 2;
## Add a response header for observability
add_header X-Routed-To $gpu_upstream always;
## Fallback to queue when all GPU upstreams are down
error_page 502 503 504 = @queue_fallback;
}
## ------------------------------------------------------------------
## Queue fallback — enqueue when GPUs are unavailable
## ------------------------------------------------------------------
location @queue_fallback {
rewrite ^ /enqueue break;
proxy_pass http://queue_service;
proxy_set_header Host $host;
proxy_set_header X-Real-IP $remote_addr;
proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
proxy_set_header X-Forwarded-Proto $scheme;
proxy_set_header Content-Type $content_type;
proxy_pass_request_body on;
}
}
-54
View File
@@ -1,54 +0,0 @@
models:
gpu-vision:
gpu_url: http://192.168.68.110:8080/v1
gpu_host: 192.168.68.110
label: Qwen3.5-9B Vision (RTX 5070)
max_concurrent: 1
context: 131072
tiers: [starter, professional, enterprise]
capabilities: [completion, multimodal]
model_path: /home/llmuser/models/qwen3.5-9b/Qwen3.5-9B-Q5_K_M.gguf
args: --mmproj /home/llmuser/models/qwen3.5-9b/mmproj-F16.gguf --ctx-size 131072 --cache-type-k q4_0 --cache-type-v q4_0 --batch-size 2048 --ubatch-size 1024 --n-gpu-layers 99 --flash-attn 1 --cont-batching --parallel 1 --image-min-tokens 2048 --image-max-tokens 16384 --alias gpu-light --reasoning off --api-key not-needed --predict 8192 --cache-prompt --port 8080 --host 0.0.0.0
gpu-dense:
gpu_url: http://192.168.68.8:8080/v1
sidecar_url: http://192.168.68.8:8090
gpu_host: 192.168.68.8
label: Qwen3.8-27B-Uncensored (RTX 3090)
max_concurrent: 1
context: 131072
tiers: [professional, enterprise]
capabilities: [completion]
model_path: /home/llmuser/models/Qwen3.8-27B-Uncensored-Q4_K_M.gguf
args: -m /home/llmuser/models/Qwen3.8-27B-Uncensored-Q4_K_M.gguf -c 131072 --host 0.0.0.0 --port 8080 -ngl 99 --alias qwen3.6-27B-code --api-key not-needed -ctk q4_0 -ctv q4_0 --flash-attn on --cont-batching --parallel 1 --slot-save-path /home/llmuser/.llama-slots --metrics --reasoning off --spec-type draft-mtp --spec-draft-n-max 2 -t 8
strix-moe:
gpu_url: http://192.168.68.15:8080/v1
sidecar_url: http://192.168.68.15:8090
gpu_host: 192.168.68.15
label: Carnice Qwen3.6 MoE 35B-A3B (Strix Halo)
max_concurrent: 1
context: 262144
tiers: [professional, enterprise]
capabilities: [completion, multimodal]
model_path: /models/Carnice-Qwen3.6-MoE-35B-A3B-Q4_K_M.gguf
args: -m /models/Carnice-Qwen3.6-MoE-35B-A3B-Q4_K_M.gguf --host 0.0.0.0 --port 8080 --alias strix-moe -c 262144 -ngl 99 -fa on --cache-type-k q4_0 --cache-type-v q4_0 --kv-unified --cache-prompt --cont-batching -b 4096 --ubatch-size 1024 --parallel 2 -t 16 --timeout 600 -n 8192 --reasoning off --no-warmup --metrics
hosts:
gpu-light:
address: 192.168.68.110
gpu_name: NVIDIA GeForce RTX 5070
vram_gb: 12
current_model: gpu-vision
gpu-dense:
address: 192.168.68.8
gpu_name: NVIDIA GeForce RTX 3090
vram_gb: 24
current_model: gpu-dense
gpu-moe:
address: 192.168.68.15
gpu_name: AMD Strix Halo (iGPU)
vram_gb: 64
current_model: strix-moe
-280
View File
@@ -1,280 +0,0 @@
{
"communities": {
"0": [
"router_route_v2",
"router_route_v2_route",
"router_route_v3",
"router_route_v3_moe_spillover",
"router_route_v3_rationale_81",
"router_route_v3_route",
"router_router_available_models",
"router_router_estimate_tokens",
"router_router_is_gpu_busy",
"router_router_moe_spillover",
"router_router_rationale_180",
"router_router_rationale_216",
"router_router_rationale_239",
"router_router_rationale_386",
"router_router_route",
"router_router_select_best_gpu"
],
"1": [
"dashboard_dashboard",
"dashboard_dashboard_api_performance",
"dashboard_dashboard_api_scatter",
"dashboard_dashboard_api_state",
"dashboard_dashboard_api_stream",
"dashboard_dashboard_api_timeseries",
"dashboard_dashboard_broadcast_loop",
"dashboard_dashboard_dashboard",
"dashboard_dashboard_fetch_state",
"dashboard_dashboard_health",
"dashboard_dashboard_rationale_1"
],
"2": [
"queue_service_queue_service",
"queue_service_queue_service_check_gpu_health",
"queue_service_queue_service_enqueue",
"queue_service_queue_service_get_queue_depth",
"queue_service_queue_service_get_redis",
"queue_service_queue_service_health",
"queue_service_queue_service_rationale_100",
"queue_service_queue_service_rationale_62",
"queue_service_queue_service_rationale_68",
"queue_service_queue_service_status"
],
"3": [
"router_router_admin_auth",
"router_router_admin_deprecation_summary",
"router_router_admin_generate_key",
"router_router_admin_keys",
"router_router_admin_revoke_key",
"router_router_rationale_895",
"router_router_rationale_905",
"router_router_rationale_928",
"router_router_rationale_950",
"router_router_rationale_978"
],
"4": [
"router_router",
"router_router_bcast",
"router_router_get_metrics",
"router_router_metrics",
"router_router_metrics_circuit_breaker",
"router_router_metrics_timeseries",
"router_router_models",
"router_router_rationale_811",
"router_router_stream"
],
"5": [
"router_router_get_redis",
"router_router_gpu_decr",
"router_router_gpu_incr",
"router_router_half_open_probe",
"router_router_is_circuit_tripped",
"router_router_performance",
"router_router_rationale_282",
"router_router_rationale_297",
"router_router_rationale_619"
],
"6": [
"router_router_check_gpu_health",
"router_router_gpu_active_count",
"router_router_gpu_health_score",
"router_router_health",
"router_router_metrics_gpu_health",
"router_router_rationale_141",
"router_router_rationale_225",
"router_router_rationale_828"
],
"7": [
"router_router_chat",
"router_router_check_rate_limit",
"router_router_clean_response",
"router_router_clean_unicode",
"router_router_rationale_184",
"router_router_rationale_72",
"router_router_store_perf_record"
],
"8": [
"maintenance",
"maintenance_sh__entry"
],
"9": [
"router_router_counter_audit_loop",
"router_router_rationale_116"
],
"10": [
"router_router_metrics_latency",
"router_router_rationale_859"
],
"11": [
"router_router_rationale_288",
"router_router_trip_circuit"
],
"12": [
"router_router_rationale_736",
"router_router_scatter"
],
"13": [
"router_http_patch"
]
},
"cohesion": {
"0": 0.20833333333333334,
"1": 0.21818181818181817,
"2": 0.3111111111111111,
"3": 0.2,
"4": 0.2777777777777778,
"5": 0.2222222222222222,
"6": 0.35714285714285715,
"7": 0.3333333333333333,
"8": 1.0,
"9": 1.0,
"10": 1.0,
"11": 1.0,
"12": 1.0,
"13": 1.0
},
"gods": [
{
"id": "router_router_chat",
"label": "chat()",
"degree": 12
},
{
"id": "router_router_get_redis",
"label": "get_redis()",
"degree": 11
},
{
"id": "router_router_gpu_active_count",
"label": "gpu_active_count()",
"degree": 9
},
{
"id": "router_router_is_gpu_busy",
"label": "is_gpu_busy()",
"degree": 9
},
{
"id": "router_router_select_best_gpu",
"label": "select_best_gpu()",
"degree": 8
},
{
"id": "router_router_route",
"label": "route()",
"degree": 8
},
{
"id": "router_route_v3_route",
"label": "route()",
"degree": 6
},
{
"id": "router_router_check_gpu_health",
"label": "check_gpu_health()",
"degree": 6
},
{
"id": "router_router_available_models",
"label": "available_models()",
"degree": 6
},
{
"id": "router_router_estimate_tokens",
"label": "estimate_tokens()",
"degree": 6
}
],
"surprises": [
{
"source": "route()",
"target": "available_models()",
"source_files": [
"router/route_v2.py",
"router/router.py"
],
"confidence": "INFERRED",
"relation": "calls",
"why": "inferred connection - not explicitly stated in source"
},
{
"source": "route()",
"target": "estimate_tokens()",
"source_files": [
"router/route_v2.py",
"router/router.py"
],
"confidence": "INFERRED",
"relation": "calls",
"why": "inferred connection - not explicitly stated in source"
},
{
"source": "route()",
"target": "is_gpu_busy()",
"source_files": [
"router/route_v2.py",
"router/router.py"
],
"confidence": "INFERRED",
"relation": "calls",
"why": "inferred connection - not explicitly stated in source"
},
{
"source": "route()",
"target": "select_best_gpu()",
"source_files": [
"router/route_v2.py",
"router/router.py"
],
"confidence": "INFERRED",
"relation": "calls",
"why": "inferred connection - not explicitly stated in source"
},
{
"source": "route()",
"target": "available_models()",
"source_files": [
"router/route_v3.py",
"router/router.py"
],
"confidence": "INFERRED",
"relation": "calls",
"why": "inferred connection - not explicitly stated in source"
}
],
"questions": [
{
"type": "bridge_node",
"question": "Why does `is_gpu_busy()` connect `Community 0` to `Community 4`, `Community 6`?",
"why": "High betweenness centrality (0.061) - this node is a cross-community bridge."
},
{
"type": "bridge_node",
"question": "Why does `select_best_gpu()` connect `Community 0` to `Community 4`, `Community 6`?",
"why": "High betweenness centrality (0.031) - this node is a cross-community bridge."
},
{
"type": "bridge_node",
"question": "Why does `estimate_tokens()` connect `Community 0` to `Community 4`, `Community 7`?",
"why": "High betweenness centrality (0.030) - this node is a cross-community bridge."
},
{
"type": "verify_inferred",
"question": "Are the 3 inferred relationships involving `is_gpu_busy()` (e.g. with `route()` and `moe_spillover()`) actually correct?",
"why": "`is_gpu_busy()` has 3 INFERRED edges - model-reasoned connections that need verification."
},
{
"type": "verify_inferred",
"question": "Are the 3 inferred relationships involving `select_best_gpu()` (e.g. with `route()` and `route()`) actually correct?",
"why": "`select_best_gpu()` has 3 INFERRED edges - model-reasoned connections that need verification."
},
{
"type": "isolated_nodes",
"question": "What connects `SyslogAI Harness Dashboard \u2014 Modern Design.`, `maintenance.sh script`, `Nginx upstream health probe. Returns 200 if service is alive.` to the rest of the system?",
"why": "28 weakly-connected nodes found - possible documentation gaps or missing edges."
}
]
}
File diff suppressed because it is too large Load Diff
-1
View File
@@ -1 +0,0 @@
{"files": {"code": ["/root/syslog-harness-repo/dashboard/dashboard.py", "/root/syslog-harness-repo/maintenance.sh", "/root/syslog-harness-repo/phase0-dual-keys.json", "/root/syslog-harness-repo/queue-service/queue-service.py", "/root/syslog-harness-repo/router/http_patch.py", "/root/syslog-harness-repo/router/route_v2.py", "/root/syslog-harness-repo/router/route_v3.py", "/root/syslog-harness-repo/router/router.py"], "document": ["/root/syslog-harness-repo/LITELLM-MIGRATION-PLAN.md", "/root/syslog-harness-repo/MIGRATION_PLAN.md", "/root/syslog-harness-repo/README.md", "/root/syslog-harness-repo/dashboard/dashboard.html", "/root/syslog-harness-repo/dashboard/harness.html", "/root/syslog-harness-repo/dashboard/requirements.txt", "/root/syslog-harness-repo/docker-compose.yml", "/root/syslog-harness-repo/litellm_config.yaml", "/root/syslog-harness-repo/router/requirements.txt", "/root/syslog-harness-repo/ssl/README.md"], "paper": [], "image": [], "video": []}, "total_files": 18, "total_words": 14667, "needs_graph": false, "warning": "Corpus is ~14,667 words - fits in a single context window. You may not need a graph.", "skipped_sensitive": ["/root/syslog-harness-repo/.env.example"], "unclassified": ["/root/syslog-harness-repo/.gitignore", "/root/syslog-harness-repo/Dockerfile.dashboard", "/root/syslog-harness-repo/Dockerfile.queue", "/root/syslog-harness-repo/backups/20260602_103344/dashboard.py.bak", "/root/syslog-harness-repo/backups/20260602_103344/router.py.bak", "/root/syslog-harness-repo/backups/20260602_103344/ts_patch.py.bak", "/root/syslog-harness-repo/dashboard/Dockerfile", "/root/syslog-harness-repo/docker-compose.yml.bak", "/root/syslog-harness-repo/gpu-router-docker.conf", "/root/syslog-harness-repo/gpu-router.conf", "/root/syslog-harness-repo/nginx/nginx.conf", "/root/syslog-harness-repo/nginx/nginx.conf.bak", "/root/syslog-harness-repo/router/Dockerfile", "/root/syslog-harness-repo/router/router.py.bak.20260518074236"], "graphifyignore_patterns": 3, "scan_root": "/root/syslog-harness-repo"}
File diff suppressed because it is too large Load Diff
-1
View File
@@ -1 +0,0 @@
{"nodes": [], "edges": [], "hyperedges": [], "input_tokens": 0, "output_tokens": 0}
-106
View File
@@ -1,106 +0,0 @@
# Graph Report - . (2026-07-08)
## Corpus Check
- Corpus is ~14,667 words - fits in a single context window. You may not need a graph.
## Summary
- 91 nodes · 151 edges · 14 communities (9 shown, 5 thin omitted)
- Extraction: 93% EXTRACTED · 7% INFERRED · 0% AMBIGUOUS · INFERRED: 10 edges (avg confidence: 0.77)
- Token cost: 0 input · 0 output
## Community Hubs (Navigation)
- Community 0
- Community 1
- Community 2
- Community 3
- Community 4
- Community 5
- Community 6
- Community 7
- Community 8
- Community 9
- Community 10
- Community 11
- Community 12
## God Nodes (most connected - your core abstractions)
1. `chat()` - 12 edges
2. `get_redis()` - 11 edges
3. `gpu_active_count()` - 9 edges
4. `is_gpu_busy()` - 9 edges
5. `select_best_gpu()` - 8 edges
6. `route()` - 8 edges
7. `route()` - 6 edges
8. `check_gpu_health()` - 6 edges
9. `available_models()` - 6 edges
10. `estimate_tokens()` - 6 edges
## Surprising Connections (you probably didn't know these)
- `route()` --calls--> `available_models()` [INFERRED]
router/route_v2.py → router/router.py
- `route()` --calls--> `estimate_tokens()` [INFERRED]
router/route_v2.py → router/router.py
- `route()` --calls--> `is_gpu_busy()` [INFERRED]
router/route_v2.py → router/router.py
- `route()` --calls--> `select_best_gpu()` [INFERRED]
router/route_v2.py → router/router.py
- `route()` --calls--> `available_models()` [INFERRED]
router/route_v3.py → router/router.py
## Import Cycles
- None detected.
## Communities (14 total, 5 thin omitted)
### Community 0 - "Community 0"
Cohesion: 0.21
Nodes (14): route(), moe_spillover(), Spill 40% of MoE-first traffic to Dense to prevent Strix Halo overheating. O, route(), available_models(), estimate_tokens(), is_gpu_busy(), moe_spillover() (+6 more)
### Community 1 - "Community 1"
Cohesion: 0.22
Nodes (4): api_state(), broadcast_loop(), fetch_state(), SyslogAI Harness Dashboard — Modern Design.
### Community 2 - "Community 2"
Cohesion: 0.31
Nodes (9): check_gpu_health(), enqueue(), get_queue_depth(), get_redis(), health(), GET queue depth + circuit breaker state + GPU health., Nginx upstream health probe. Returns 200 if service is alive., Fallback endpoint — Nginx calls this when all GPU upstreams are down. (+1 more)
### Community 3 - "Community 3"
Cohesion: 0.20
Nodes (10): _admin_auth(), admin_deprecation_summary(), admin_generate_key(), admin_keys(), admin_revoke_key(), Require admin key for management endpoints., List all API keys (masked) with agent, tier, and deprecation status., Summary of deprecated key usage (from Redis logs, if available). (+2 more)
### Community 4 - "Community 4"
Cohesion: 0.28
Nodes (5): bcast(), get_metrics(), metrics(), metrics_circuit_breaker(), Expose circuit breaker status per model. Phase 1.
### Community 5 - "Community 5"
Cohesion: 0.22
Nodes (9): get_redis(), gpu_decr(), gpu_incr(), half_open_probe(), is_circuit_tripped(), performance(), Check if a GPU host is currently blacklisted., Check if a GPU host can be un-blacklisted. (+1 more)
### Community 6 - "Community 6"
Cohesion: 0.36
Nodes (8): check_gpu_health(), gpu_active_count(), gpu_health_score(), health(), metrics_gpu_health(), Get number of in-flight requests for a GPU., Score a GPU based on VRAM, temperature, and load. Lower = better., Live GPU health scores + circuit breaker + KPIs.
### Community 7 - "Community 7"
Cohesion: 0.33
Nodes (7): chat(), check_rate_limit(), clean_response(), clean_unicode(), Store detailed performance record in Redis for analytics., Token bucket rate limiter using Redis. Returns (allowed, retry_after_or_remainin, store_perf_record()
## Knowledge Gaps
- **1 isolated node(s):** `maintenance.sh script`
These have ≤1 connection - possible missing edges or undocumented components.
- **5 thin communities (<3 nodes) omitted from report** — run `graphify query` to explore isolated nodes.
## Suggested Questions
_Questions this graph is uniquely positioned to answer:_
- **Why does `is_gpu_busy()` connect `Community 0` to `Community 4`, `Community 6`?**
_High betweenness centrality (0.061) - this node is a cross-community bridge._
- **Why does `select_best_gpu()` connect `Community 0` to `Community 4`, `Community 6`?**
_High betweenness centrality (0.031) - this node is a cross-community bridge._
- **Why does `estimate_tokens()` connect `Community 0` to `Community 4`, `Community 7`?**
_High betweenness centrality (0.030) - this node is a cross-community bridge._
- **Are the 3 inferred relationships involving `is_gpu_busy()` (e.g. with `route()` and `moe_spillover()`) actually correct?**
_`is_gpu_busy()` has 3 INFERRED edges - model-reasoned connections that need verification._
- **Are the 3 inferred relationships involving `select_best_gpu()` (e.g. with `route()` and `route()`) actually correct?**
_`select_best_gpu()` has 3 INFERRED edges - model-reasoned connections that need verification._
- **What connects `SyslogAI Harness Dashboard — Modern Design.`, `maintenance.sh script`, `Nginx upstream health probe. Returns 200 if service is alive.` to the rest of the system?**
_28 weakly-connected nodes found - possible documentation gaps or missing edges._
File diff suppressed because one or more lines are too long
@@ -1 +0,0 @@
{"nodes": [{"id": "root_syslog_harness_repo_router_http_patch_py", "label": "http_patch.py", "file_type": "code", "source_file": "router/http_patch.py", "source_location": "L1"}], "edges": [{"source": "root_syslog_harness_repo_router_http_patch_py", "target": "re", "relation": "imports", "context": "import", "confidence": "EXTRACTED", "source_file": "router/http_patch.py", "source_location": "L2", "weight": 1.0}], "raw_calls": []}
File diff suppressed because one or more lines are too long
@@ -1 +0,0 @@
{"nodes": [{"id": "root_syslog_harness_repo_maintenance_sh", "label": "maintenance.sh", "file_type": "code", "source_file": "maintenance.sh", "source_location": "L1", "metadata": {"language": "bash", "kind": "file"}}, {"id": "root_syslog_harness_repo_maintenance_sh__entry", "label": "maintenance.sh script", "file_type": "code", "source_file": "maintenance.sh", "source_location": "L1", "metadata": {"language": "bash", "kind": "bash_entrypoint"}}], "edges": [{"source": "root_syslog_harness_repo_maintenance_sh", "target": "root_syslog_harness_repo_maintenance_sh__entry", "relation": "contains", "confidence": "EXTRACTED", "source_file": "maintenance.sh", "source_location": "L1", "weight": 1.0}]}
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
-1
View File
@@ -1 +0,0 @@
{"/root/syslog-harness-repo/LITELLM-MIGRATION-PLAN.md":{"size":31561,"mtime_ns":1781468608211508160,"word_count":3762},"/root/syslog-harness-repo/MIGRATION_PLAN.md":{"size":1966,"mtime_ns":1781121183850310435,"word_count":303},"/root/syslog-harness-repo/README.md":{"size":2447,"mtime_ns":1781121183850310435,"word_count":377},"/root/syslog-harness-repo/dashboard/dashboard.html":{"size":8331,"mtime_ns":1781316813360201927,"word_count":415},"/root/syslog-harness-repo/dashboard/dashboard.py":{"size":29010,"mtime_ns":1781121183851178603,"word_count":1615,"hash":"654e97aa6df1872b0717f32f676c067d15b9795a0dc04a86936730a7e133c214"},"/root/syslog-harness-repo/dashboard/harness.html":{"size":23388,"mtime_ns":1781316813362649901,"word_count":1869},"/root/syslog-harness-repo/dashboard/requirements.txt":{"size":30,"mtime_ns":1781121183851178603,"word_count":2},"/root/syslog-harness-repo/docker-compose.yml":{"size":3900,"mtime_ns":1782224658037786319,"word_count":200},"/root/syslog-harness-repo/litellm_config.yaml":{"size":618,"mtime_ns":1781121183851178603,"word_count":39},"/root/syslog-harness-repo/maintenance.sh":{"size":1302,"mtime_ns":1781121183851178603,"word_count":192,"hash":"801feb2c0b7f4052cb2c4ef933d295d3494c8791766687b891c04cc7b5495d09"},"/root/syslog-harness-repo/phase0-dual-keys.json":{"size":1628,"mtime_ns":1781121183851610673,"word_count":110,"hash":"0f115313f2c86df2f57e48bdf180b768d6452d0337bae4e3c3d2b52b9f5267a7"},"/root/syslog-harness-repo/queue-service/queue-service.py":{"size":3284,"mtime_ns":1781121183851610673,"word_count":325,"hash":"c9a9d2e5c9f1fb7d35a73f91c83a047ce92dc2d09c9ddab8fc2d6e82d92a093d"},"/root/syslog-harness-repo/router/http_patch.py":{"size":3562,"mtime_ns":1781121183851788825,"word_count":290,"hash":"53edfda5145db02bc35b9e9d8e06e5c731ce50c3a50cec3db65b3b886bed95a6"},"/root/syslog-harness-repo/router/requirements.txt":{"size":43,"mtime_ns":1781121183851788825,"word_count":3},"/root/syslog-harness-repo/router/route_v2.py":{"size":3954,"mtime_ns":1781121183851788825,"word_count":435,"hash":"e96ac2ae879e0a876d4190daff4cc566d13798b97c4820f3efa47e06627bbbcb"},"/root/syslog-harness-repo/router/route_v3.py":{"size":4703,"mtime_ns":1781121183851788825,"word_count":504,"hash":"30b202626b1f50da90b78272001e253cc7086b0bdc52e46553414c5313e3fb4e"},"/root/syslog-harness-repo/router/router.py":{"size":44307,"mtime_ns":1781316813360201927,"word_count":4198,"hash":"9c6817ae09f6bb5b2862f0282cacab7d4ce5da051e4ddc95e2e4099c7bf09716"},"/root/syslog-harness-repo/ssl/README.md":{"size":204,"mtime_ns":1781121183851788825,"word_count":28}}
File diff suppressed because it is too large Load Diff
-152
View File
@@ -1,152 +0,0 @@
general_settings:
master_key: os.environ/LITELLM_MASTER_KEY
store_model_in_db: false
user_url_allowed_hosts:
- 192.168.68.14
- 192.168.68.14:5080
guardrails:
- guardrail_name: input-moderation
litellm_params:
guardrail: openai_moderation
mode: pre_call
- guardrail_name: output-moderation
litellm_params:
guardrail: openai_moderation
mode: post_call
- guardrail_name: harmful-content-filter
litellm_params:
categories:
- action: BLOCK
category: harmful_self_harm
enabled: true
severity_threshold: medium
- action: BLOCK
category: harmful_violence
enabled: true
severity_threshold: medium
- action: BLOCK
category: harmful_illegal_weapons
enabled: true
severity_threshold: medium
guardrail: litellm_content_filter
mode: pre_call
litellm_settings:
user_url_allowed_hosts:
- 192.168.68.14
- 192.168.68.14:5080
cache: true
cache_params:
host: harness-redis
namespace: litellm
port: 6379
ttl: 600
type: redis
drop_params: true
success_callback:
- prometheus
failure_callback:
- prometheus
model_cost:
gpu-dense:
input_cost_per_token: 1.5e-07
output_cost_per_token: 6.0e-07
strix-moe:
input_cost_per_token: 1.5e-07
output_cost_per_token: 6.0e-07
syslog-auto:
input_cost_per_token: 1.5e-07
output_cost_per_token: 6.0e-07
gpu-vision:
input_cost_per_token: 1.5e-07
output_cost_per_token: 6.0e-07
num_retries: 2
request_timeout: 600
sso_callback: /sso/callback
model_list:
- litellm_params:
api_base: http://192.168.68.15:8080/v1
api_key: not-needed
model: openai/strix-moe
rpm: 40
timeout: 300
model_info:
max_input_tokens: 131072
max_model_tokens: 131072
max_tokens: 131072
model_name: strix-moe
- litellm_params:
api_base: http://192.168.68.8:8080/v1
api_key: not-needed
model: openai/gpu-dense
rpm: 500
timeout: 300
model_info:
max_input_tokens: 131072
max_model_tokens: 131072
max_tokens: 131072
model_name: gpu-dense
- litellm_params:
api_base: http://192.168.68.110:8080/v1
api_key: not-needed
model: openai/gpu-vision
rpm: 500
timeout: 300
model_info:
max_input_tokens: 131072
max_model_tokens: 131072
max_tokens: 131072
model_name: gpu-vision
- litellm_params:
api_base: http://192.168.68.8:8080/v1
api_key: not-needed
model: openai/gpu-dense
rpm: 500
timeout: 300
model_info:
max_input_tokens: 131072
max_model_tokens: 131072
max_tokens: 131072
weight: 0.70
model_name: syslog-auto
- litellm_params:
api_base: http://192.168.68.15:8080/v1
api_key: not-needed
model: openai/strix-moe
rpm: 60
timeout: 300
model_info:
max_input_tokens: 131072
weight: 0.20
model_name: syslog-auto
- litellm_params:
api_base: http://192.168.68.110:8080/v1
api_key: not-needed
model: openai/gpu-vision
rpm: 200
timeout: 300
model_info:
max_input_tokens: 131072
weight: 0.10
model_name: syslog-auto
router_settings:
allowed_fails: 100
enable_loadbalancing_on_proxy: true
fallbacks:
- syslog-auto:
- gpu-dense
- strix-moe
- gpu-vision
- gpu-dense:
- strix-moe
- strix-moe:
- gpu-dense
- gpu-vision
request_timeout: 300
routing_strategy: simple-shuffle
agents:
- agent_name: agent-zero-homelab
agent_card_params:
name: Agent Zero HomeLab
url: http://192.168.68.14:5080/a2a/t-8zNgdOEXzYxjQvTl/p-homelab
protocolVersion: '1.0'
-25
View File
@@ -1,25 +0,0 @@
model_list:
- model_name: qwen3.6-35B-A3B
litellm_params:
model: openai/qwen3.6-35B-A3B
api_base: http://192.168.68.15:8080/v1
api_key: "not-needed"
- model_name: qwen3.6-27B-code
litellm_params:
model: openai/qwen3.6-27B-code-text
api_base: http://192.168.68.8:8080/v1
api_key: "not-needed"
- model_name: gemma-4-12b
litellm_params:
model: openai/gemma-4-12b
api_base: http://192.168.68.110:8080/v1
api_key: "not-needed"
general_settings:
master_key: sk-syslog-local-master-key
litellm_settings:
drop_params: true
request_timeout: 120
-37
View File
@@ -1,37 +0,0 @@
#!/bin/bash
# SyslogAI Harness — Automated Maintenance
# Runs daily via cron
LOG="/var/log/harness-maintenance.log"
echo "=== $(date) ===" >> "$LOG"
# 1. Clean Redis timeseries keys older than 60 days
CUTOFF=$(date -d "60 days ago" +%Y%m%d%H)
echo "Redis: removing ts:* keys older than $CUTOFF" >> "$LOG"
DELETED=0
for key in $(docker exec harness-redis redis-cli KEYS "ts:*" 2>/dev/null); do
TS=$(echo "$key" | grep -oP '\d{10}$')
if [ -n "$TS" ] && [ "$TS" -lt "$CUTOFF" ] 2>/dev/null; then
docker exec harness-redis redis-cli DEL "$key" > /dev/null 2>&1
DELETED=$((DELETED + 1))
fi
done
echo "Redis: deleted $DELETED stale timeseries keys" >> "$LOG"
# 2. Log stale model keys (leftover from migrations)
STALE=$(docker exec harness-redis redis-cli KEYS "*gemma*" 2>/dev/null)
if [ -n "$STALE" ]; then
echo "WARNING: stale gemma keys found: $STALE" >> "$LOG"
fi
# 3. Prune Docker build cache (older than 7 days)
echo "Docker: pruning build cache" >> "$LOG"
docker builder prune -f --filter until=168h >> "$LOG" 2>&1
# 4. Log container health status
docker ps --format "table {{.Names}}\t{{.Status}}\t{{.RunningFor}}" >> "$LOG" 2>&1
# 5. Log Redis memory
docker exec harness-redis redis-cli INFO memory | grep used_memory_human >> "$LOG" 2>&1
echo "" >> "$LOG"
-173
View File
@@ -1,173 +0,0 @@
events {
worker_connections 1024;
}
http {
include /etc/nginx/mime.types;
default_type application/octet-stream;
sendfile on;
keepalive_timeout 65;
resolver 127.0.0.11 valid=30s;
map $host $dashboard_ui_url {
default http://harness-dashboard:3000;
}
map $host $litellm_backend_url {
default http://harness-litellm:4000;
}
# ════════════════════════════════════════════════════════════
# Server :80 — Lean single-layer entrypoint
# All API paths route directly to LiteLLM.
# harness-router fully deprecated.
# ════════════════════════════════════════════════════════════
server {
listen 80;
# Authentik OIDC subrequest
location /authentik/auth {
internal;
proxy_pass https://auth.sysloggh.net/outpost.goauthentik.io/auth/nginx;
proxy_pass_request_body off;
proxy_set_header Content-Length "";
proxy_set_header X-Original-URL $scheme://$http_host$request_uri;
proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
}
# ── LiteLLM API (replaces router /v1/) ──
location /v1/ {
proxy_pass $litellm_backend_url;
proxy_http_version 1.1;
proxy_set_header Host $host;
proxy_set_header X-Real-IP $remote_addr;
proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
proxy_set_header Authorization $http_authorization;
proxy_connect_timeout 10s;
proxy_read_timeout 600s;
proxy_buffering off;
}
# ── LiteLLM admin (replaces router /admin/) ──
location /admin/ {
proxy_pass $litellm_backend_url;
proxy_http_version 1.1;
proxy_set_header Host $host;
proxy_set_header Authorization $http_authorization;
proxy_read_timeout 600s;
}
# ── LiteLLM stream ──
location /stream {
proxy_pass $litellm_backend_url/stream;
proxy_http_version 1.1;
proxy_set_header Host $host;
proxy_set_header Upgrade $http_upgrade;
proxy_set_header Connection "upgrade";
proxy_buffering off;
}
# ── API passthrough ──
location /api/ {
proxy_pass $litellm_backend_url/;
proxy_http_version 1.1;
proxy_set_header Host $host;
}
# ── Dashboard ──
location /dashboard/ {
proxy_pass $dashboard_ui_url/;
proxy_http_version 1.1;
proxy_set_header Host $host;
}
# ── GPU Fleet Dashboard ──
location /gpu/ {
proxy_pass http://192.168.68.24:9100/;
proxy_http_version 1.1;
proxy_set_header Host $host;
proxy_set_header X-Real-IP $remote_addr;
proxy_read_timeout 60s;
}
# ── LiteLLM redirect ──
location = /litellm {
return 301 /litellm/;
}
# ── Health: redirect /health (auth-required) → /health/liveliness (no-auth) ──
location = /litellm/health {
return 301 /litellm/health/liveliness;
}
# ── LiteLLM static assets ──
location /litellm-asset-prefix/ {
proxy_pass $litellm_backend_url;
proxy_http_version 1.1;
proxy_set_header Host $host;
proxy_set_header X-Real-IP $remote_addr;
proxy_connect_timeout 10s;
proxy_read_timeout 60s;
}
# ── LiteLLM admin UI and API (strips /litellm prefix) ──
location /litellm/ {
rewrite ^/litellm(/.*)$ $1 break;
proxy_pass $litellm_backend_url;
proxy_http_version 1.1;
proxy_set_header Host $host;
proxy_set_header X-Real-IP $remote_addr;
proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
proxy_set_header Upgrade $http_upgrade;
proxy_set_header Connection "upgrade";
proxy_buffering off;
proxy_read_timeout 600s;
proxy_set_header Authorization $http_authorization;
}
# ── Auth proxy to Authentik ──
location /application/o/ {
proxy_pass https://192.168.68.11;
proxy_ssl_verify off;
proxy_set_header Host auth.sysloggh.net;
proxy_set_header X-Real-IP $remote_addr;
proxy_connect_timeout 10s;
proxy_read_timeout 30s;
}
# ── Prometheus metrics (LiteLLM exposes at /metrics) ──
location /metrics {
proxy_pass $litellm_backend_url/metrics;
proxy_http_version 1.1;
proxy_set_header Host $host;
}
# ---- Legacy /health/unified alias (harness-router decommissioned) ----
# Retained intentionally: 5+ live GPU monitor contracts poll
# this path every 15s and follow the 301 to /gpu/gpu-data.
location /health/unified {
return 301 /gpu/gpu-data;
}
# ── Health: no-auth LiteLLM liveliness ──
location /health {
proxy_pass $litellm_backend_url/health/liveliness;
proxy_http_version 1.1;
proxy_set_header Host $host;
}
# ── UI / Docs convenience redirects (additive 2026-09-11) ──
# Canonical app path is /litellm/. These 301s make bare
# /ui/ and /docs resolve. No new auth surface; no OIDC impact.
location = /ui { return 301 /litellm/ui/; }
location /ui/ { return 301 /litellm$request_uri; }
location = /docs { return 301 /litellm/docs; }
location = /docs/ { return 301 /litellm/docs; }
# ── 404 for everything else ──
location / {
return 404;
}
}
}
-20
View File
@@ -1,20 +0,0 @@
{
"sk-syslog-local-master-key": {"tier": "enterprise", "agent": "admin", "deprecated": true},
"sk-33a0d6e6-a6da12fb483770d6f63f543b7e16a742": {"tier": "enterprise", "agent": "admin"},
"sk-syslog-abiba": {"tier": "enterprise", "agent": "Abiba", "deprecated": true},
"sk-6a6e49d0-e35e27c524ce0ba2de8f862a0e966b0c": {"tier": "enterprise", "agent": "Abiba"},
"sk-syslog-mumuni": {"tier": "enterprise", "agent": "Mumuni", "deprecated": true},
"sk-ba249c99-29e3b163a8d8ab7abb8663cae0dceb37": {"tier": "enterprise", "agent": "Mumuni"},
"sk-syslog-tanko": {"tier": "enterprise", "agent": "Tanko", "deprecated": true},
"sk-4a67112a-d9b62caf5a5881a15837df506f782708": {"tier": "enterprise", "agent": "Tanko"},
"sk-syslog-koby": {"tier": "enterprise", "agent": "Koby", "deprecated": true},
"sk-9dd66bba-316f616ec40cb51f3b489a0146bb986f": {"tier": "enterprise", "agent": "Koby"},
"sk-syslog-kagenz0": {"tier": "enterprise", "agent": "Kagenz0", "deprecated": true},
"sk-5a05020f-a984a8a20f731c5a65dfc07d51495bb3": {"tier": "enterprise", "agent": "Kagenz0"},
"sk-syslog-koonimo": {"tier": "enterprise", "agent": "Koonimo", "deprecated": true},
"sk-b2f0e1b2-70b7cf1c5a5e402a49b5469e88d9e335": {"tier": "enterprise", "agent": "Koonimo"},
"sk-starter-abc123": {"tier": "starter", "agent": "test-starter", "deprecated": true},
"sk-a8307a00-cd76daa07dd61c41aeb4545656f8b8c4": {"tier": "starter", "agent": "test-starter"},
"sk-professional-xyz789": {"tier": "professional", "agent": "test-pro", "deprecated": true},
"sk-b519a91e-4d2774c6c7da3fc31326e69ccc426781": {"tier": "professional", "agent": "test-pro"}
}
+10
View File
@@ -0,0 +1,10 @@
FROM python:3.13-slim
RUN pip install --no-cache-dir flask redis
COPY queue-service.py /app/queue-service.py
WORKDIR /app
EXPOSE 8091
CMD ["python3", "queue-service.py"]
+121
View File
@@ -0,0 +1,121 @@
#!/usr/bin/env python3
"""Syslog Inference Queue Service — Circuit breaker + request queuing.
Ports: 8091
Endpoints:
/health — liveness probe (Nginx upstream check)
/enqueue — POST inference request into queue (fallback from Nginx)
/status — GET queue depth + circuit breaker state
"""
import json
import os
import sys
import time
import urllib.request
from flask import Flask, request, jsonify
app = Flask(__name__)
# Configuration
REDIS_HOST = os.getenv("REDIS_HOST", "192.168.68.7")
REDIS_PORT = int(os.getenv("REDIS_PORT", "6379"))
QUEUE_KEY = "inference:requests"
CIRCUIT_OPEN_THRESHOLD = 50
CIRCUIT_WARN_THRESHOLD = 30
# GPU endpoints for draining
GPUS = {
"amdpve": "192.168.68.15:8080",
"llmgpu": "192.168.68.8:8080",
"ocu_llm": "192.168.68.110:8080",
}
def get_redis():
try:
import redis
return redis.Redis(host=REDIS_HOST, port=REDIS_PORT, decode_responses=True)
except Exception:
return None
def get_queue_depth(r):
try:
return r.llen(QUEUE_KEY)
except Exception:
return 0
def check_gpu_health(endpoint):
try:
req = urllib.request.Request(f"http://{endpoint}/v1/models")
req.add_header("User-Agent", "queue-service/1.0")
resp = urllib.request.urlopen(req, timeout=3)
return resp.status == 200
except Exception:
return False
@app.route("/health")
def health():
"""Nginx upstream health probe. Returns 200 if service is alive."""
return jsonify({"status": "ok", "service": "queue-service"}), 200
@app.route("/enqueue", methods=["POST"])
def enqueue():
"""Fallback endpoint — Nginx calls this when all GPU upstreams are down."""
r = get_redis()
if not r:
return jsonify({"error": "Redis unavailable"}), 503
depth = get_queue_depth(r)
if depth >= CIRCUIT_OPEN_THRESHOLD:
return jsonify({
"error": "Circuit breaker OPEN",
"queue_depth": depth,
"threshold": CIRCUIT_OPEN_THRESHOLD
}), 503
# Store the request in queue
payload = request.get_data(as_text=True)
headers = {k: v for k, v in request.headers if k.startswith("X-")}
r.rpush(QUEUE_KEY, json.dumps({
"payload": payload,
"headers": headers,
"queued_at": time.time()
}))
new_depth = get_queue_depth(r)
return jsonify({
"status": "queued",
"position": new_depth,
"circuit": "warn" if new_depth >= CIRCUIT_WARN_THRESHOLD else "closed"
}), 202
@app.route("/status")
def status():
"""GET queue depth + circuit breaker state + GPU health."""
r = get_redis()
depth = get_queue_depth(r) if r else -1
circuit = "open" if depth >= CIRCUIT_OPEN_THRESHOLD else ("warn" if depth >= CIRCUIT_WARN_THRESHOLD else "closed")
gpu_health = {}
for name, endpoint in GPUS.items():
gpu_health[name] = "up" if check_gpu_health(endpoint) else "down"
return jsonify({
"queue_depth": depth,
"circuit_breaker": circuit,
"gpu_health": gpu_health,
"thresholds": {
"warn": CIRCUIT_WARN_THRESHOLD,
"open": CIRCUIT_OPEN_THRESHOLD
}
})
if __name__ == "__main__":
app.run(host="0.0.0.0", port=8091)
-75
View File
@@ -1,75 +0,0 @@
# GPU Roster Loader - reads gpu_roster.yaml via PyYAML
# Hot-reloadable via /admin/roster/reload endpoint
import os, json, threading, time
try:
import yaml
except ImportError:
yaml = None
ROSTER_PATH = os.environ.get('ROSTER_PATH', '/app/gpu_roster.yaml')
_last_mtime = 0
_lock = threading.Lock()
GPU_URLS = {}
GPU_SIDECARS = {}
GPU_LABELS = {}
GPU_MAX_CONCURRENT = {}
GPU_CONTEXT = {}
TIER_MODELS = {}
HOSTS = {}
def load_roster(path=None):
global GPU_URLS, GPU_SIDECARS, GPU_LABELS, GPU_MAX_CONCURRENT
global GPU_CONTEXT, TIER_MODELS, HOSTS, _last_mtime
if yaml is None:
return False, 'PyYAML not installed - run: pip3 install pyyaml'
path = path or ROSTER_PATH
try:
with open(path) as f:
data = yaml.safe_load(f)
with _lock:
GPU_URLS.clear()
GPU_SIDECARS.clear()
GPU_LABELS.clear()
GPU_MAX_CONCURRENT.clear()
GPU_CONTEXT.clear()
TIER_MODELS.clear()
models = data.get('models', {})
for name, cfg in models.items():
GPU_URLS[name] = cfg.get('gpu_url', '')
GPU_SIDECARS[name] = cfg.get('sidecar_url', '')
GPU_LABELS[name] = cfg.get('label', name)
GPU_MAX_CONCURRENT[name] = cfg.get('max_concurrent', 1)
GPU_CONTEXT[name] = cfg.get('context', 65536)
tiers = cfg.get('tiers', ['enterprise'])
for t in tiers:
if t not in TIER_MODELS:
TIER_MODELS[t] = []
if name not in TIER_MODELS[t]:
TIER_MODELS[t].append(name)
HOSTS.clear()
HOSTS.update(data.get('hosts', {}))
_last_mtime = os.path.getmtime(path)
return True, 'Loaded {} models, {} hosts'.format(len(GPU_URLS), len(HOSTS))
except Exception as e:
return False, str(e)
def check_reload():
global _last_mtime
try:
mtime = os.path.getmtime(ROSTER_PATH)
if mtime > _last_mtime:
success, msg = load_roster()
if success:
print('[ROSTER] Auto-reloaded:', msg)
except:
pass
def reload_thread(interval=30):
while True:
time.sleep(interval)
check_reload()
threading.Thread(target=reload_thread, daemon=True).start()
-56
View File
@@ -1,56 +0,0 @@
#!/bin/bash
# Daily LiteLLM cost snapshot — cron at midnight
# Writes to /var/log/litellm/costs.csv on CT 116
MASTER="sk-litellm-7f96080dd99b15c36bd4b333b58a6796"
BASE="http://127.0.0.1:4001"
CSV="/var/log/litellm/costs.csv"
DATE=$(date -I)
YESTERDAY=$(date -I -d yesterday)
# Fetch yesterday's spend by key
RESP=$(curl -s "${BASE}/global/spend/logs?start_date=${YESTERDAY}T00:00:00&end_date=${YESTERDAY}T23:59:59&page_size=1000" \
-H "Authorization: Bearer $MASTER" 2>/dev/null)
# Parse total spend
TOTAL=$(echo "$RESP" | python3 -c "
import sys, json
data = json.load(sys.stdin)
total = sum(r.get('spend', 0) or 0 for r in data.get('data', []))
print(f'{total:.6f}')
" 2>/dev/null)
# Parse per-model spend
PER_MODEL=$(echo "$RESP" | python3 -c "
import sys, json
from collections import defaultdict
data = json.load(sys.stdin)
spend = defaultdict(float)
for r in data.get('data', []):
model = r.get('model', 'unknown')
spend[model] += r.get('spend', 0) or 0
for model, cost in sorted(spend.items()):
print(f'{model}:{cost:.6f}')
" 2>/dev/null)
# Per-key spend
PER_KEY=$(echo "$RESP" | python3 -c "
import sys, json
from collections import defaultdict
data = json.load(sys.stdin)
spend = defaultdict(float)
for r in data.get('data', []):
key = r.get('api_key', 'unknown')[:20]
spend[key] += r.get('spend', 0) or 0
for key, cost in sorted(spend.items()):
print(f'{key}:{cost:.6f}')
" 2>/dev/null)
# Write CSV
mkdir -p $(dirname "$CSV")
if [ ! -f "$CSV" ]; then
echo "date,total_cost,per_model,per_key" > "$CSV"
fi
echo "${DATE},${TOTAL},\"${PER_MODEL}\",\"${PER_KEY}\"" >> "$CSV"
echo "Snapshot: ${DATE} | Total: \$${TOTAL}"
-394
View File
@@ -1,394 +0,0 @@
#!/usr/bin/env python3
"""GPU Fleet Monitor — Comprehensive monitoring server.
Monitors every subsystem in the inference harness:
- GPU sidecars (VRAM, temp, util, power per card)
- Router (model routing, circuit breakers, queue health)
- LiteLLM (proxy health, key count, model sync)
- Strix Halo (llama-server status, CPU load, context)
- Dashboard (CT 116 harness-dashboard)
Endpoints:
/ → GPU fleet HTML dashboard
/gpu-data → JSON GPU metrics (for dashboard API)
/health → Monitor self-health check
Run: python3 gpu-monitor-server.py
Default port: 9100
"""
import json, subprocess, http.server, threading, time, os
from datetime import datetime, timezone
DASHBOARD_PATH = "/root/dashboard/gpu-fleet.html"
PORT = 9100
# Cached data, refreshed every 15s
cache = {}
cache_lock = threading.Lock()
# Alert thresholds
THRESHOLDS = {
"temp_c": {"warning": 80, "critical": 90},
"vram_pct": {"warning": 90, "critical": 95},
"gpu_util_pct": {"warning": 95, "critical": 98},
}
def http_get(url, timeout=5):
"""HTTP GET with timeout, returns parsed JSON or error dict."""
import urllib.request
try:
resp = urllib.request.urlopen(url, timeout=timeout)
return json.loads(resp.read())
except Exception as e:
return {"error": str(e)}
def poll_sidecar(host, port=8090):
"""Poll a GPU sidecar for raw metrics.
Falls back to SSH-based nvidia-smi if sidecar is unreachable."""
result = http_get(f"http://{host}:{port}/health", timeout=5)
if "error" not in result:
return result
# Fallback: poll via SSH (explicit key path for non-interactive environments)
try:
proc = subprocess.run(
["/usr/bin/ssh", "-o", "StrictHostKeyChecking=no", "-o", "ConnectTimeout=5",
"-i", "/root/.ssh/id_ed25519", host,
"nvidia-smi", "--query-gpu=name,temperature.gpu,utilization.gpu,"
"utilization.memory,memory.used,memory.total,power.draw,power.limit,fan.speed",
"--format=csv,noheader"],
capture_output=True, text=True, timeout=10
)
if proc.returncode == 0:
parts = [p.strip() for p in proc.stdout.strip().split(",")]
if len(parts) >= 8:
def cv(v):
try: return float(v.replace("MiB","").replace("W","").replace("%","").strip())
except: return 0.0
def mv(v):
try: return int(v.replace("MiB","").strip())
except: return 0
return {
"gpu_name": parts[0], "temp_c": cv(parts[1]),
"gpu_util_pct": cv(parts[2]), "mem_util_pct": cv(parts[3]),
"vram_used_mb": mv(parts[4]), "vram_total_mb": mv(parts[5]),
"power_w": cv(parts[6]), "power_limit_w": cv(parts[7]),
"fan_pct": cv(parts[8])
}
except Exception as e:
return {"error": str(e)}
return result
def poll_router():
"""Poll router unified health through nginx (port 80)."""
return http_get("http://192.168.68.116/health/unified", timeout=5)
def poll_router_health():
"""Poll router basic health through nginx."""
return http_get("http://192.168.68.116/health", timeout=5)
def poll_litellm_health():
"""Poll LiteLLM — checks reachability via nginx proxy.
LiteLLM health endpoints require authentication. We check if the
service responds at all (even 401 = service is up). Also try the
/key/list endpoint which confirms the proxy is fully functional.
"""
import urllib.request
import urllib.error
urls_to_try = [
"http://192.168.68.116/litellm/health",
"http://192.168.68.116/litellm/ui/",
]
last_error = None
for url in urls_to_try:
try:
resp = urllib.request.urlopen(url, timeout=5)
body = resp.read().decode()[:500]
# Any response (even 401) means the proxy is routing to LiteLLM
return {"reachable": True, "status_code": resp.getcode(),
"endpoint": url, "body_preview": body[:200]}
except urllib.error.HTTPError as e:
# HTTP error means the service IS reachable but returned error
return {"reachable": True, "status_code": e.code,
"endpoint": url, "note": f"HTTP {e.code} (requires auth)"}
except Exception as e:
last_error = str(e)
return {"error": last_error or "all endpoints failed"}
def poll_strix():
"""Poll Strix Halo — process check + CPU load + llama-server health."""
try:
# Check llama-server process
proc = subprocess.run(
["ssh", "-o", "StrictHostKeyChecking=no", "-o", "ConnectTimeout=5",
"192.168.68.15",
"pgrep -f llama-server > /dev/null 2>&1 && echo running || echo stopped"],
capture_output=True, text=True, timeout=10
)
status = proc.stdout.strip()
# Get CPU load
load = subprocess.run(
["ssh", "-o", "StrictHostKeyChecking=no", "-o", "ConnectTimeout=5",
"192.168.68.15",
"uptime | awk -F'load average:' '{print $2}' | tr -d ' '"],
capture_output=True, text=True, timeout=10
)
cpu_load = load.stdout.strip() if load.returncode == 0 else "unknown"
# Try llama-server health endpoint
llama_health = http_get("http://192.168.68.15:8080/health", timeout=3)
return {
"status": status if status else "unknown",
"cpu_load": cpu_load,
"llama_health": llama_health if "error" not in llama_health else None
}
except Exception as e:
return {"status": "error", "error": str(e)}
def poll_dashboard():
"""Check if harness-dashboard on CT 116 is serving."""
return http_get("http://192.168.68.116/dashboard/", timeout=5)
def check_alerts(gpu_data):
"""Check GPU metrics against thresholds, return alert list."""
alerts = []
name = gpu_data.get("name", "unknown")
for metric, thresholds in THRESHOLDS.items():
# Map metric names from sidecar/router formats
value = None
if metric == "vram_pct":
# Compute from vram_used_mb / vram_total_mb
used = gpu_data.get("vram_used_mb", 0)
total = gpu_data.get("vram_total_mb", 1)
value = (used / total) * 100
elif metric == "gpu_util_pct":
value = gpu_data.get("gpu_util_pct", 0)
elif metric == "temp_c":
value = gpu_data.get("temp_c", 0)
if value is not None:
if value >= thresholds["critical"]:
alerts.append({"gpu": name, "metric": metric, "level": "critical",
"value": round(value, 1), "threshold": thresholds["critical"]})
elif value >= thresholds["warning"]:
alerts.append({"gpu": name, "metric": metric, "level": "warning",
"value": round(value, 1), "threshold": thresholds["warning"]})
return alerts
def compute_summary(router_data, gpus, litellm_data, strix_data):
"""Compute fleet-wide health summary."""
# Count models available
models_available = 0
models_total = 0
if "available_models" in router_data:
models_available = len(router_data["available_models"])
models_total = models_available # router reports all expected
# Count circuit breakers open
cb_open = 0
if "circuit_breaker" in router_data:
for model, cb in router_data["circuit_breaker"].items():
if cb.get("open", 0):
cb_open += 1
# Count GPU errors
gpu_errors = sum(1 for g in gpus if "error" in g)
# Fleet status
if gpu_errors > 0 or cb_open > 0:
fleet_status = "degraded"
elif not router_data or "error" in router_data:
fleet_status = "degraded"
else:
fleet_status = "healthy"
litellm_ok = ("error" not in litellm_data) or litellm_data.get("reachable", False)
return {
"fleet_status": fleet_status,
"models_available": models_available,
"models_total": models_total,
"circuit_breakers_open": cb_open,
"gpu_count": len(gpus),
"gpu_errors": gpu_errors,
"strix_running": strix_data.get("status") == "running",
"router_reachable": "error" not in router_data,
"litellm_reachable": litellm_ok,
}
def refresh_cache():
"""Refresh all fleet data from every subsystem."""
global cache
# 1. GPU sidecars
gpu_map = {
"192.168.68.8": {"name": "Dense (RTX 3090)", "models": ["gpu-dense"], "hostname": "ct8"},
"192.168.68.110": {"name": "Light/Vision (RTX 5070)", "models": ["gpu-vision"], "hostname": "ct110"},
}
gpus = []
all_alerts = []
for host, info in gpu_map.items():
data = poll_sidecar(host)
if "error" not in data:
data["name"] = info["name"]
data["models"] = info["models"]
data["hostname"] = info["hostname"]
all_alerts.extend(check_alerts(data))
else:
data = {"name": info["name"], "hostname": info["hostname"], "error": data["error"]}
all_alerts.append({"gpu": info["name"], "metric": "connectivity",
"level": "critical", "value": data["error"]})
gpus.append(data)
# 2. Router
router = poll_router()
router_basic = poll_router_health()
router["_basic"] = router_basic
# 3. LiteLLM
litellm = poll_litellm_health()
# 4. Strix Halo
strix = poll_strix()
# 5. Dashboard
dashboard = poll_dashboard()
# 6. Send Zulip DM for critical alerts (debounced 5 min)
critical = [a for a in all_alerts if a.get("level") == "critical"]
if critical and _debounce_alert():
msg = "🔴 **GPU Fleet Alert**\n"
for a in critical:
msg += f"• {a['gpu']}: {a['metric']} = {a.get('value', a.get('actual', '?'))}\n"
_send_zulip_dm(msg)
# 7. Summary
summary = compute_summary(router, gpus, litellm, strix)
with cache_lock:
cache = {
"gpus": gpus,
"strix": strix,
"router": router,
"litellm": litellm,
"dashboard": {"reachable": "error" not in dashboard},
"summary": summary,
"alerts": all_alerts,
"updated": datetime.now(timezone.utc).isoformat(),
"monitor_version": "2.0.0",
}
_last_alert_time = 0
_ALERT_COOLDOWN = 300 # 5 min between DMs
def _send_zulip_dm(msg):
"""Send a DM to the owner via Zulip."""
try:
site = "https://chat.sysloggh.net"
email = os.environ.get("ABIBA_ZULIP_EMAIL", "abiba-bot@chat.sysloggh.net")
key = os.environ.get("ZULIP_API_KEY", "")
if not key:
print("[alert] ZULIP_API_KEY unset — DM suppressed")
return
owner = 9
auth_b64 = base64.b64encode(f"{email}:{key}".encode()).decode()
data = urllib.parse.urlencode({"type": "private", "to": f"[{owner}]", "content": msg}).encode()
req = urllib.request.Request(f"{site}/api/v1/messages", data=data,
headers={"Authorization": f"Basic {auth_b64}"}, method="POST")
urllib.request.urlopen(req, timeout=5)
except Exception as e:
print(f"[alert] DM failed: {e}")
def _debounce_alert():
"""Returns True if enough time has passed since last alert."""
global _last_alert_time
now = time.time()
if now - _last_alert_time > _ALERT_COOLDOWN:
_last_alert_time = now
return True
return False
class DashboardHandler(http.server.BaseHTTPRequestHandler):
def do_GET(self):
if self.path == "/gpu-data":
with cache_lock:
data = json.dumps(cache, indent=2)
self.send_response(200)
self.send_header("Content-Type", "application/json")
self.send_header("Access-Control-Allow-Origin", "*")
self.end_headers()
self.wfile.write(data.encode())
elif self.path == "/health":
with cache_lock:
healthy = bool(cache and "summary" in cache)
status = {"status": "healthy" if healthy else "starting",
"updated": cache.get("updated", "never") if cache else "never"}
code = 200 if healthy else 503
self.send_response(code)
self.send_header("Content-Type", "application/json")
self.end_headers()
self.wfile.write(json.dumps(status).encode())
elif self.path in ("/", "/index.html"):
try:
with open(DASHBOARD_PATH) as f:
html = f.read()
self.send_response(200)
self.send_header("Content-Type", "text/html")
self.end_headers()
self.wfile.write(html.encode())
except FileNotFoundError:
self.send_response(404)
self.end_headers()
self.wfile.write(b"Dashboard not found")
else:
self.send_response(404)
self.end_headers()
def log_message(self, format, *args):
pass # Suppress request logs
def background_refresh():
"""Refresh data every 15 seconds."""
while True:
try:
refresh_cache()
except Exception as e:
print(f"[WARN] Refresh failed: {e}", flush=True)
time.sleep(15)
if __name__ == "__main__":
print(f"Starting Comprehensive GPU Fleet Monitor v2.0.0 on port {PORT}...")
print(f" Dashboard: http://localhost:{PORT}/")
print(f" API: http://localhost:{PORT}/gpu-data")
print(f" Health: http://localhost:{PORT}/health")
print(f" Polling: router (nginx:80), sidecars (:8090), strix, litellm, dashboard")
# Initial fetch
refresh_cache()
# Start background refresher
t = threading.Thread(target=background_refresh, daemon=True)
t.start()
# Start HTTP server
server = http.server.HTTPServer(("0.0.0.0", PORT), DashboardHandler)
try:
server.serve_forever()
except KeyboardInterrupt:
print("\nShutting down...")
-472
View File
@@ -1,472 +0,0 @@
#!/usr/bin/env python3
"""
GPU Self-Heal — Implementation of gpu-self-heal.prose.md
Evaluates 8 remediation rules against live GPU monitor data.
Reports results to knowledge graph via MCP bridge.
"""
import json, time, subprocess, sys, os, socket
from datetime import datetime, timezone
from urllib.request import urlopen, Request
# Global timeout to prevent hanging on unreachable hosts
socket.setdefaulttimeout(10)
GPU_MONITOR = "http://192.168.68.24:9100/gpu-data"
KG_BRIDGE = "http://192.168.68.65:3100/mcp"
GPU_HOSTS = {
"ct8-rtx3090": {"ip": "192.168.68.8", "gpu": "NVIDIA RTX 3090", "vram_total_mb": 24576, "vram_threshold_mb_per_h": 100},
"ct110-rtx5070": {"ip": "192.168.68.110","gpu": "NVIDIA RTX 5070", "vram_total_mb": 12227, "vram_threshold_mb_per_h": 50},
"strix-halo": {"ip": "192.168.68.15", "gpu": "AMD Strix Halo 64GB","vram_total_mb": 65536, "vram_threshold_mb_per_h": 200},
}
HISTORY_FILE = "/var/log/litellm/gpu-history.json"
RUN_ID = f"gpu-self-heal-{datetime.now().strftime('%Y%m%d-%H%M%S')}"
def fetch_gpu_data():
"""Fetch live GPU fleet data from monitor."""
try:
req = Request(GPU_MONITOR)
with urlopen(req, timeout=10) as resp:
return json.loads(resp.read())
except Exception as e:
print(f" FAIL: GPU monitor unreachable: {e}")
return None
def fetch_prometheus(host, port=9400):
"""Check Prometheus exporter on GPU host (with timeout)."""
try:
req = Request(f"http://{host}:{port}/metrics")
with urlopen(req, timeout=3) as resp:
return resp.status == 200
except:
return False
def probe_strix_direct():
"""Direct Strix Halo health probe (firewall opened .24→.15:8080)."""
try:
with urlopen("http://192.168.68.15:8080/health", timeout=5) as resp:
return resp.status == 200
except:
return False
def load_history():
"""Load historical GPU metrics for trend analysis."""
if os.path.exists(HISTORY_FILE):
with open(HISTORY_FILE) as f:
return json.load(f)
return {"snapshots": [], "vram_baselines": {}, "temp_history": {}}
def save_history(history):
"""Persist history for trend analysis."""
os.makedirs(os.path.dirname(HISTORY_FILE), exist_ok=True)
# Keep last 72 hours of snapshots (4320 at 60s intervals, capped at 1000)
history["snapshots"] = history["snapshots"][-1000:]
with open(HISTORY_FILE, "w") as f:
json.dump(history, f, indent=2)
def analyze_vram_trend(history, gpu_name, current_vram_mb):
"""Calculate VRAM growth rate over 6-hour window."""
snapshots = history.get("snapshots", [])
six_hours_ago = time.time() - 21600
old_snapshots = [s for s in snapshots if s.get("timestamp", 0) > six_hours_ago
and s.get("gpu") == gpu_name]
if len(old_snapshots) < 10:
return 0 # Not enough data
oldest = old_snapshots[0]
newest = old_snapshots[-1]
hours = (newest["timestamp"] - oldest["timestamp"]) / 3600
if hours < 1:
return 0
return (current_vram_mb - oldest["vram_mb"]) / hours
def analyze_temp_rise(history, gpu_name, current_temp):
"""Calculate temperature rise rate over last 5 minutes."""
snapshots = history.get("snapshots", [])
five_min_ago = time.time() - 300
recent = [s for s in snapshots if s.get("timestamp", 0) > five_min_ago
and s.get("gpu") == gpu_name]
if len(recent) < 3:
return 0
oldest = recent[0]
minutes = (recent[-1]["timestamp"] - oldest["timestamp"]) / 60
if minutes < 0.5:
return 0
return (current_temp - oldest["temp_c"]) / minutes
def record_snapshot(history, gpu_name, temp_c, vram_mb):
"""Record a GPU snapshot for trend analysis."""
history["snapshots"].append({
"timestamp": time.time(),
"gpu": gpu_name,
"temp_c": temp_c,
"vram_mb": vram_mb
})
def kg_create_node(title, description, source):
"""Log run report to Gitea (hard rule: health logs NEVER go to knowledge graph).
Redirected 2026-08-13 per directive from Mumuni (#726): logs belong in
SyslogSolution/health-logs/gpu/{run_id}.json, not the RA-H OS graph.
"""
try:
# source is the report JSON string; write to temp file and push via gitea-logger
report_file = f"/tmp/{RUN_ID}.json"
with open(report_file, "w") as f:
f.write(source if isinstance(source, str) else json.dumps(source, indent=2))
cmd = f"/opt/inference-harness/scripts/gitea-logger.sh gpu {RUN_ID}.json {report_file}"
result = subprocess.run(cmd, shell=True, capture_output=True, text=True, timeout=120)
os.remove(report_file) if os.path.exists(report_file) else None
return result.stdout.strip() or "gitea-logger: no output"
except Exception as e:
return f"gitea-logger failed: {e}"
def evaluate_rules(data, history):
"""Evaluate all 8 remediation rules against current GPU state."""
actions = []
if not data:
return actions
gpus = data.get("gpus", [])
summary = data.get("summary", {})
benchmarks = data.get("benchmarks", {})
strix = data.get("strix", {})
router = data.get("router", {})
for gpu in gpus:
if not isinstance(gpu, dict):
continue
name = gpu.get("gpu_name", "unknown")
host_key = None
for k, v in GPU_HOSTS.items():
if v["gpu"] in name or k in name:
host_key = k
break
if not host_key:
print(f"[warn] no GPU_HOSTS mapping for {name!r} — remote rules (incl. Rule 7) skipped for this GPU")
host_key = name.replace(" ", "-").lower()
temp = gpu.get("temp_c", 0)
vram_used = gpu.get("vram_used_mb", 0)
host_info = GPU_HOSTS.get(host_key, {"ip": "unknown", "vram_threshold_mb_per_h": 50})
# Record snapshot
record_snapshot(history, name, temp, vram_used)
# ── Rule 1: Thermal Critical ──
if temp > 85:
actions.append({"rule": "thermal-critical", "gpu": name, "temp_c": temp,
"action": "reduce-concurrency", "severity": "critical"})
print(f" ⚠ RULE 1: {name} at {temp}°C — reducing concurrency")
# ── Rule 2: VRAM Leak ──
vram_rate = analyze_vram_trend(history, name, vram_used)
threshold = host_info.get("vram_threshold_mb_per_h", 50)
if vram_rate > threshold:
actions.append({"rule": "vram-leak", "gpu": name, "rate_mb_per_h": vram_rate,
"threshold": threshold, "action": "investigate-leak", "severity": "warning"})
print(f" ⚠ RULE 2: {name} VRAM leak {vram_rate:.1f} MB/h (threshold: {threshold})")
# ── Rule 4: Benchmark Regression ──
bench_latest = benchmarks.get("latest", {})
for bk, bv in bench_latest.items():
if not isinstance(bv, dict):
continue
latest_run = bv.get("latest", {})
baseline = bv.get("baseline_tok_per_sec", 0)
current_tps = latest_run.get("gen_tok_per_sec", 0)
if baseline > 0 and current_tps > 0 and current_tps < baseline * 0.8:
drop_pct = (1 - current_tps / baseline) * 100
gpu_label = bv.get("gpu_name", bk)
actions.append({"rule": "benchmark-regression", "gpu": gpu_label,
"baseline_tps": baseline, "current_tps": current_tps,
"drop_pct": round(drop_pct,1), "action": "investigate-performance",
"severity": "warning"})
print(f" ⚠ RULE 4: {gpu_label} benchmark -{drop_pct:.0f}% ({current_tps} vs {baseline} tok/s)")
# ── Rule 7: Prometheus Exporter ──
ip = host_info.get("ip", "")
if ip and ip != "unknown":
if not fetch_prometheus(ip):
actions.append({"rule": "prometheus-down", "gpu": name, "host": ip,
"action": "restart-exporter", "severity": "warning"})
print(f" ⚠ RULE 7: {name} Prometheus exporter down on {ip}:9400")
# ── Rule 8: Predictive Thermal ──
rise_rate = analyze_temp_rise(history, name, temp)
if temp > 70 and rise_rate > 2.0:
tier = "critical" if temp > 80 else "warning"
action_text = "aggressive-load-shedding" if tier == "critical" else "reduce-concurrency-50pct"
actions.append({"rule": "predictive-thermal", "gpu": name,
"temp_c": temp, "rise_rate_c_per_min": rise_rate,
"tier": tier, "action": action_text, "severity": tier})
print(f" {'⚠' if tier == 'critical' else '🔶'} RULE 8 ({tier}): {name} {temp}°C rising {rise_rate:.1f}°C/min")
# ── Rule 6: Strix Halo ──
if not strix.get("status") == "running":
if probe_strix_direct():
print(f" ⚠ RULE 6: Strix direct probe OK but monitor disagrees — router stale")
actions.append({"rule": "strix-router-stale", "gpu": "strix-halo",
"action": "refresh-router", "severity": "info"})
else:
print(f" ⚠ RULE 6: Strix Halo down — attempting restart")
actions.append({"rule": "strix-down", "gpu": "strix-halo",
"action": "restart-llama-server", "severity": "critical"})
# ── Rule 3: Model Stuck (check router) ──
router_models = router.get("available_models", [])
if isinstance(router_models, list):
for model in router_models:
if isinstance(model, dict) and model.get("consecutive_timeouts", 0) >= 3:
# Check failure rate
total = model.get("total_requests", 1)
failures = model.get("failed_requests", 0)
if total > 0 and failures / total > 0.5:
actions.append({"rule": "model-stuck", "gpu": model.get("gpu", "?"),
"model": model.get("name", "?"),
"failure_rate": failures/total,
"action": "restart-llama-server", "severity": "critical"})
print(f" ⚠ RULE 3: Model {model.get('name')} stuck ({failures}/{total} failures)")
# ── Rule 5: Circuit Breaker ──
cb_data = router.get("circuit_breaker", {})
if isinstance(cb_data, dict):
for cb_name, cb in cb_data.items():
if isinstance(cb, dict) and cb.get("open"):
open_sec = cb.get("open_duration_sec", 0)
gpu_healthy = cb.get("gpu_healthy", False)
if open_sec > 600 and gpu_healthy:
# Check cooldown: has this CB been reset in the last hour?
last_reset = history.get("cb_resets", {}).get(cb_name, 0)
if time.time() - last_reset > 3600:
actions.append({"rule": "cb-stuck-open", "gpu": cb.get("gpu", "?"),
"cb_name": cb_name, "open_sec": open_sec,
"action": "reset-cb", "severity": "warning"})
history.setdefault("cb_resets", {})[cb_name] = time.time()
print(f" ⚠ RULE 5: CB {cb_name} stuck open {open_sec}s — auto-resetting")
return actions
def evaluate_context_optimization(data):
"""
Rule 9: Context Window Optimization
Check if each GPU's context window is optimal for its role.
"""
benchmarks = data.get("benchmarks", {}).get("latest", {})
results = []
# Actual topology (verified July 2026)
gpu_config = {
"ct8-rtx3090": {
"context": 262144, "model": "gpu-dense", "quant": "Q4_K_M",
"parallel": 1, "role": "heavy-reasoning", "vram_mb": 24576,
"target_tps": 75.0
},
"ct110-rtx5070": {
"context": 131072, "model": "qwen3.5-9b-vision", "quant": "Q4_0 KV",
"parallel": 2, "role": "vision-web", "vram_mb": 12227,
"target_tps": 76.0
},
"strix-halo": {
"context": 262144, "model": "strix-moe", "quant": "Q4_K_M",
"parallel": 2, "role": "compression", "vram_mb": 65536,
"target_tps": 70.0
},
}
for gpu_key, cfg in gpu_config.items():
bench = benchmarks.get(gpu_key, {})
current_tps = bench.get("latest", {}).get("gen_tok_per_sec", 0)
baseline_tps = bench.get("baseline_tok_per_sec", 0)
if current_tps <= 0 or baseline_tps <= 0:
results.append({
"gpu": gpu_key, "model": cfg["model"], "role": cfg["role"],
"context": cfg["context"], "tok_per_sec": current_tps or 0,
"recommendation": f"{cfg['model']} idle — no benchmark data this cycle"
})
continue
perf_ratio = current_tps / baseline_tps
recommendation = None
if gpu_key == "ct8-rtx3090":
# Already at 256K — check if holding performance
if perf_ratio >= 0.95:
recommendation = f"256K ctx: {current_tps} tok/s ({(perf_ratio*100):.0f}% of baseline). Optimal for heavy reasoning."
else:
recommendation = f"256K ctx: {current_tps} tok/s ({(perf_ratio*100):.0f}% of baseline). Consider reducing parallel or ctx."
elif gpu_key == "ct110-rtx5070":
# 131K ctx, vision/web role — must preserve speed
if perf_ratio >= 0.95:
recommendation = f"131K ctx: {current_tps} tok/s (optimal). Vision/web role — context is fine."
else:
recommendation = f"131K ctx: {current_tps} tok/s ({(perf_ratio*100):.0f}% of baseline). Check VRAM pressure."
elif gpu_key == "strix-halo":
# 256K ctx, compression role — maximize context
if perf_ratio >= 0.95:
recommendation = f"256K ctx: {current_tps} tok/s ({(perf_ratio*100):.0f}% of baseline). Above target. Compression-optimized."
else:
recommendation = f"256K ctx: {current_tps} tok/s ({(perf_ratio*100):.0f}% of baseline). Check for competing workloads."
results.append({
"gpu": gpu_key, "model": cfg["model"], "role": cfg["role"],
"context": cfg["context"], "parallel": cfg["parallel"],
"tok_per_sec": current_tps, "baseline_tps": baseline_tps,
"perf_ratio": round(perf_ratio, 3),
"recommendation": recommendation
})
print(f" 📐 RULE 9: {gpu_key} ({cfg['role']}) — {recommendation}")
return results
def evaluate_workload_distribution(data):
"""
Rule 10: Workload Distribution Optimization
Check if GPU workload patterns match designated roles.
"""
role_map = {
"heavy-reasoning": ["gpu-dense", "syslog-auto"],
"vision-web": ["gpu-vision", "qwen3.5-9b-vision"],
"compression": ["strix-moe"],
}
gpu_roles = {
"NVIDIA GeForce RTX 3090": "heavy-reasoning",
"NVIDIA GeForce RTX 5070": "vision-web",
"AMD Strix Halo": "compression",
}
gpus = data.get("gpus", [])
summary = data.get("summary", {})
results = []
for gpu in gpus:
if not isinstance(gpu, dict):
continue
name = gpu.get("gpu_name", "")
role = "unknown"
for pattern, r in gpu_roles.items():
if pattern in name:
role = r
break
vram_used = gpu.get("vram_used_mb", 0)
vram_total = gpu.get("vram_total_mb", 0)
vram_pct = (vram_used / vram_total * 100) if vram_total else 0
status = "optimal"
note = ""
if role == "heavy-reasoning" and vram_pct > 90:
status = "warning"
note = f"VRAM at {vram_pct:.0f}% — consider reducing parallel or context"
elif role == "vision-web" and vram_pct > 88:
status = "warning"
note = f"VRAM at {vram_pct:.0f}% — vision/web role is VRAM-tight"
elif role == "compression" and vram_pct < 50:
note = f"VRAM at {vram_pct:.0f}% — plenty of headroom for compression workloads"
else:
note = f"VRAM at {vram_pct:.0f}% — role-appropriate"
results.append({
"gpu": name, "role": role, "vram_pct": round(vram_pct, 1),
"status": status, "note": note
})
icon = "✅" if status == "optimal" else "⚠"
print(f" {icon} RULE 10: {name} → {role} ({note})")
return results
def run_inference_test(model="syslog-auto"):
"""Run a quick inference test through LiteLLM."""
try:
body = json.dumps({
"model": model,
"messages": [{"role": "user", "content": "ping"}],
"max_tokens": 5
}).encode()
req = Request("http://192.168.68.116:4000/v1/chat/completions",
data=body, headers={
"Content-Type": "application/json",
"Authorization": "Bearer " + os.environ.get("LITELLM_API_KEY", "")
})
with urlopen(req, timeout=30) as resp:
return resp.status == 200
except:
return False
def main():
print(f"[{datetime.now().isoformat()}] GPU Self-Heal — {RUN_ID}")
print("=" * 60)
# Phase 1: Fetch data
data = fetch_gpu_data()
if not data:
print(" GPU monitor unreachable — aborting")
return 1
summary = data.get("summary", {})
print(f" Fleet: {summary.get('fleet_status','?')} | "
f"GPUs: {summary.get('gpu_count',0)} | "
f"Errors: {summary.get('gpu_errors',0)} | "
f"Alerts: {len(data.get('alerts',[]))}")
print()
# Phase 2: Evaluate rules (1-8)
history = load_history()
actions = evaluate_rules(data, history)
# Phase 2b: Context optimization (Rule 9)
ctx_results = evaluate_context_optimization(data)
# Phase 2c: Workload distribution (Rule 10)
wl_results = evaluate_workload_distribution(data)
save_history(history)
# Phase 3: Execute actions
issues_found = len(actions)
issues_fixed = 0
for action in actions:
severity = action.get("severity", "info")
if severity == "critical":
# Attempt fix
if "restart-llama-server" in action.get("action", ""):
gpu_name = action.get("gpu", "")
host = GPU_HOSTS.get(gpu_name.replace("NVIDIA ", "").replace("AMD ", "").lower(), {})
# Actual restart would need SSH — placeholder for now
print(f" ⟳ Would restart llama-server on {host.get('ip','?')} for {gpu_name}")
issues_fixed += 1
# Phase 4: Verify
all_models_ok = run_inference_test()
print(f"\n Verification: inference test {'✅' if all_models_ok else '❌'}")
# Phase 5: Compile report
status = "healthy"
if issues_found > 0:
status = "degraded" if issues_found <= 2 else "down"
report = json.dumps({
"run_id": RUN_ID,
"timestamp": datetime.now(timezone.utc).isoformat(),
"fleet_status": summary.get("fleet_status", "?"),
"overall_status": status,
"issues_found": issues_found,
"issues_fixed": issues_fixed,
"actions": actions,
"context_optimization": ctx_results,
"workload_distribution": wl_results
}, indent=2, default=str)
print(f"\n Status: {status} | Issues: {issues_found} | Fixed: {issues_fixed}")
# Phase 6: Log to Gitea (hard rule: never to knowledge graph)
kg_title = f"[GPU-SELF-HEAL] {RUN_ID} — {status}"
kg_desc = f"GPU self-heal run: {issues_found} issues found, {issues_fixed} fixed. Fleet: {summary.get('fleet_status','?')}."
kg_result = kg_create_node(kg_title, kg_desc, report)
print(f" Gitea: {kg_result[:120] if kg_result else 'write attempted'}")
print(f"\n[{datetime.now().isoformat()}] Complete — {RUN_ID}")
return 0 if issues_found == 0 else 1
if __name__ == "__main__":
sys.exit(main())
-360
View File
@@ -1,360 +0,0 @@
#!/bin/bash
LLK="${LITELLM_API_KEY:?LITELLM_API_KEY must be set}"
# LiteLLM Health Check + Self-Heal — Automated (every 6 hours)
# Deployed from litellm-self-heal.prose.md contract
# Writes results to RA-H OS knowledge graph via MCP bridge
set -uo pipefail # -e removed: grep -c returns 1 on no-match, which is valid
RUN_ID="litellm-health-$(date +%Y%m%d-%H%M%S)"
TIMESTAMP=$(date -Iseconds)
BRIDGE="http://192.168.68.65:3100/mcp"
MASTER_KEY="sk-litellm-7f96080dd99b15c36bd4b333b58a6796" # admin only: /key/list. NEVER for inference
LITELLM_HOST="192.168.68.116"
source /etc/litellm-monitor.env 2>/dev/null # dedicated monitor agent key for inference tests
MONITOR_KEY="${LITELLM_MONITOR_KEY:-$MASTER_KEY}" # fallback only if env missing
GPU_DASHBOARD="http://192.168.68.24:9100"
LOG_DIR="/var/log/litellm"
RESULTS=""
ISSUES=0
FIXED=0
ESCALATED=0
DETAILS="[]"
mcp_call() {
local method="$1" tool="$2" args="$3"
curl -s -X POST "$BRIDGE" \
-H "Content-Type: application/json" \
-H "Accept: application/json, text/event-stream" \
-d "{\"jsonrpc\":\"2.0\",\"id\":1,\"method\":\"$method\",\"params\":$args}" 2>/dev/null
}
add_detail() {
local name="$1" status="$2" msg="$3"
DETAILS=$(echo "$DETAILS" | python3 -c "
import sys, json
d = json.load(sys.stdin)
d.append({'check':'$name','status':'$status','detail':'$msg'})
print(json.dumps(d))
" 2>/dev/null || echo "$DETAILS")
}
# ═══════════════════════════════════════════════════════
# PHASE 1: HEALTH CHECKS
# ═══════════════════════════════════════════════════════
echo "[$(date)] Starting LiteLLM health check — $RUN_ID"
# 1. Public endpoints
echo -n " UI... "
if curl -s -o /dev/null -w "%{http_code}" --connect-timeout 10 https://litellm.sysloggh.net/ui/ 2>/dev/null | grep -q 200; then
add_detail "public-ui" "pass" "200 OK"
echo "pass"
else
add_detail "public-ui" "fail" "non-200"
echo "FAIL"; ISSUES=$((ISSUES+1))
fi
echo -n " Docs... "
if curl -s -o /dev/null -w "%{http_code}" --connect-timeout 10 https://litellm.sysloggh.net/docs 2>/dev/null | grep -q 200; then
add_detail "public-docs" "pass" "200 OK"
echo "pass"
else
add_detail "public-docs" "fail" "non-200"
echo "FAIL"; ISSUES=$((ISSUES+1))
fi
# 2. LiteLLM liveliness
echo -n " Liveliness... "
LIVE=$(curl -s --connect-timeout 5 "http://${LITELLM_HOST}:4000/health/liveliness" 2>/dev/null)
if echo "$LIVE" | grep -qi "alive"; then
add_detail "liveliness" "pass" "$LIVE"
echo "pass"
else
add_detail "liveliness" "fail" "$LIVE"
echo "FAIL"; ISSUES=$((ISSUES+1))
fi
# 3. Container health
echo -n " Containers... "
CONTAINERS=$(docker ps --format '{{.Names}}:{{.Status}}' 2>/dev/null)
CT_COUNT=$(echo "$CONTAINERS" | wc -l)
# Only flag containers that are explicitly 'unhealthy', not ones without healthchecks
UNHEALTHY=$(echo "$CONTAINERS" | grep -c '(unhealthy)' 2>/dev/null || true)
UNHEALTHY=${UNHEALTHY:-0}
if [ "$UNHEALTHY" -eq 0 ] && [ "$CT_COUNT" -ge 8 ]; then
add_detail "containers" "pass" "$CT_COUNT containers healthy"
echo "pass ($CT_COUNT up)"
else
add_detail "containers" "fail" "$UNHEALTHY unhealthy of $CT_COUNT"
echo "FAIL ($UNHEALTHY/$CT_COUNT unhealthy)"; ISSUES=$((ISSUES+1))
# Auto-restart unhealthy containers
for ct in $(echo "$CONTAINERS" | grep '(unhealthy)' | cut -d: -f1); do
echo " Restarting $ct..."
docker restart "$ct" 2>/dev/null && FIXED=$((FIXED+1))
done
fi
# 4. GPU Fleet (full telemetry from gpu-monitor on .24:9100)
echo -n " GPU Fleet... "
GPU_DATA=$(curl -s --connect-timeout 10 "$GPU_DASHBOARD/gpu-data" 2>/dev/null)
GPU_HEALTH=$(echo "$GPU_DATA" | python3 -c "
import sys, json
d = json.load(sys.stdin)
s = d.get('summary',{})
alerts = d.get('alerts',[])
gpus = d.get('gpus',[])
strix = d.get('strix',{})
fleet = s.get('fleet_status','unknown')
gpu_count = s.get('gpu_count',0)
errors = s.get('gpu_errors',0)
cb_open = s.get('circuit_breakers_open',0)
strix_ok = strix.get('status','') == 'running'
alert_count = len(alerts)
critical_alerts = len([a for a in alerts if isinstance(a, dict) and a.get('level')=='critical'])
# Per-GPU details
for g in gpus:
if isinstance(g, dict):
n = g.get('gpu_name','?')[:30]
t = g.get('temp_c','?')
u = g.get('gpu_util_pct','?')
v = f\"{g.get('vram_used_mb',0)}/{g.get('vram_total_mb',0)}MB\"
print(f' {n}: {t}°C util={u}% vram={v}')
# Alert details
for a in alerts:
if isinstance(a, dict):
print(f' ⚠ {a.get(\"level\",\"?\")}: {a.get(\"metric\",\"?\")} — {a.get(\"value\",\"?\")}')
# Summary line
status = 'healthy' if fleet == 'healthy' and critical_alerts == 0 and errors == 0 else 'degraded'
print(f'SUMMARY: {status} | {gpu_count} GPUs | {errors} errors | {alert_count} alerts | CB open={cb_open} | Strix={\"running\" if strix_ok else \"down\"}')
" 2>/dev/null)
GPU_STATUS=$(echo "$GPU_HEALTH" | grep 'SUMMARY:' | cut -d' ' -f2-)
if echo "$GPU_STATUS" | grep -q '^healthy'; then
add_detail "gpu-fleet" "pass" "$GPU_STATUS"
echo "pass"
echo "$GPU_HEALTH" | grep -v 'SUMMARY:'
else
add_detail "gpu-fleet" "fail" "$GPU_STATUS"
echo "FAIL"
echo "$GPU_HEALTH"
ISSUES=$((ISSUES+1))
fi
# 5. Model inference tests
MODELS="gpu-dense strix-moe gpu-vision syslog-auto"
MODEL_FAILS=0
for model in $MODELS; do
echo -n " Model $model... "
HTTP_CODE=$(curl -s -o /dev/null -w "%{http_code}" --connect-timeout 30 \
-H "Authorization: Bearer $LLK" \
-H "Content-Type: application/json" \
-d "{\"model\":\"$model\",\"messages\":[{\"role\":\"user\",\"content\":\"ping\"}],\"max_tokens\":5}" \
"http://${LITELLM_HOST}:4000/v1/chat/completions" 2>/dev/null)
if [ "$HTTP_CODE" = "200" ]; then
add_detail "model-$model" "pass" "200 OK"
echo "pass"
else
add_detail "model-$model" "fail" "HTTP $HTTP_CODE"
echo "FAIL ($HTTP_CODE)"; ISSUES=$((ISSUES+1)); MODEL_FAILS=$((MODEL_FAILS+1))
fi
done
# 6. Agent keys
echo -n " Agent Keys... "
KEYS=$(curl -s --connect-timeout 10 \
-H "Authorization: Bearer $LLK" \
"http://${LITELLM_HOST}:4000/key/list?return_full_object=true" 2>/dev/null)
AGENT_COUNT=$(echo "$KEYS" | python3 -c "
import sys,json
d = json.load(sys.stdin)
agents = {'mumuni','tanko','kagenz0','koby','koonimo','abiba-pi','baggy'}
keys = d.get('keys',[])
found = sum(1 for k in keys if k.get('key_alias') in agents)
print(found)
" 2>/dev/null || echo 0)
if [ "$AGENT_COUNT" -ge 6 ]; then
add_detail "agent-keys" "pass" "$AGENT_COUNT agent keys present"
echo "pass ($AGENT_COUNT keys)"
else
add_detail "agent-keys" "fail" "only $AGENT_COUNT/7 agent keys found"
echo "FAIL"; ISSUES=$((ISSUES+1))
fi
# 7. Grafana
echo -n " Grafana... "
if curl -s -o /dev/null -w "%{http_code}" --connect-timeout 10 \
"http://${LITELLM_HOST}:3001/api/health" 2>/dev/null | grep -q 200; then
add_detail "grafana" "pass" "200 OK"
echo "pass"
else
add_detail "grafana" "warn" "Grafana unreachable (non-critical)"
echo "warn"
fi
# 8. 401 error count
echo -n " 401 Errors... "
ERR_401_RAW=$(docker logs harness-litellm --since 6h 2>&1 | grep 'Received=' 2>/dev/null || true)
ERR_401=$(echo "$ERR_401_RAW" | grep -c 'Received=' 2>/dev/null); ERR_401=${ERR_401:-0}
if [ "$ERR_401" -eq 0 ]; then
add_detail "401-errors" "pass" "0 auth errors in last 6h"
echo "pass (0)"
else
# Extract key patterns and source IPs from 401 errors
ERR_KEYS=$(echo "$ERR_401_RAW" | grep -oP 'Received=\K[^,]+' | sort -u | tr '\n' ' ')
ERR_IPS=$(docker logs harness-litellm --since 6h 2>&1 | grep -B2 'Received=' | grep -oP '\d+\.\d+\.\d+\.\d+' | sort -u | tr '\n' ' ')
DETAIL="$ERR_401 auth errors in last 6h | Keys: ${ERR_KEYS:-unknown} | Sources: ${ERR_IPS:-unknown}"
add_detail "401-errors" "warn" "$DETAIL"
echo "warn ($ERR_401 — keys: ${ERR_KEYS:-?}, sources: ${ERR_IPS:-?})"
ISSUES=$((ISSUES+1))
# Self-heal: analyze and attempt fix based on source IP type
if echo "$ERR_KEYS" | grep -q 'no-key-required'; then
echo " → 'no-key-required' = GPU-direct key being used as LiteLLM client key."
fi
for src_ip in $ERR_IPS; do
# Case 1: Docker network IPs (172.x) — this is the nginx proxy, meaning an external client
if echo "$src_ip" | grep -q '^172\.'; then
CT_NAME=$(docker inspect -f '{{.Name}}' $(docker ps -q) 2>/dev/null | while read n; do
ip=$(docker inspect -f '{{range .NetworkSettings.Networks}}{{.IPAddress}}{{end}}' $n 2>/dev/null)
[ "$ip" = "$src_ip" ] && echo "$n"
done)
echo " → Source $src_ip is Docker container: ${CT_NAME:-unknown}"
if echo "$CT_NAME" | grep -q 'nginx'; then
echo " → 401 came through nginx proxy — external client, not locally fixable."
echo " → Checking nginx logs via docker logs for actual source..."
NGINX_SRC=$(timeout 8 docker logs harness-nginx --tail 2000 2>&1 | grep ' 401 ' | grep -oP '^\S+' | sort -u | tr '\n' ' ')
if [ -n "$NGINX_SRC" ]; then
echo " → Nginx upstream source(s): $NGINX_SRC"
fi
ESCALATED=$((ESCALATED+1))
else
echo " → Checking Docker container logs for clue..."
docker logs "$CT_NAME" --tail 50 2>/dev/null | grep -i 'litellm\|api.key\|auth' | tail -5
fi
# Case 2: Localhost or local CT116 IP — check locally
elif [ "$src_ip" = "127.0.0.1" ] || [ "$src_ip" = "192.168.68.116" ]; then
echo " → Source $src_ip is local (CT116). Checking local configs..."
LOCAL_CFG=$(grep -rl 'no-key-required' /root/.hermes/ /opt/inference-harness/ --include='*.yaml' --include='*.yml' --include='*.py' 2>/dev/null | grep -v 'litellm-health-check\|litellm_config\|\.bak' | head -5)
if [ -n "$LOCAL_CFG" ]; then
echo " → Found local references: $LOCAL_CFG"
else
echo " → No local config using no-key-required. Likely external via proxy."
ESCALATED=$((ESCALATED+1))
fi
# Case 3: Known agent host IPs — SSH and check
else
echo " → Checking remote source $src_ip for misconfigured agent..."
AGENT_CONFIGS=$(ssh -o ConnectTimeout=5 -o StrictHostKeyChecking=no root@$src_ip \
"grep -rl 'no-key-required\|api_key.*not-needed' /root/.hermes/config.yaml /home/*/.hermes/config.yaml 2>/dev/null" 2>/dev/null || true)
if [ -n "$AGENT_CONFIGS" ]; then
echo " → FOUND: agent config with GPU-direct key at $src_ip"
for cfg in $AGENT_CONFIGS; do
echo " → Attempting self-heal on $cfg..."
ssh -o ConnectTimeout=5 root@$src_ip \
"python3 -c \"
import yaml
with open('$cfg') as f: c = yaml.safe_load(f)
fixed = False
for section in ['agent', 'auxiliary']:
for sub in c.get(section, {}):
if isinstance(c[section].get(sub), dict):
ak = c[section][sub].get('api_key', '')
if ak in ['no-key-required', 'not-needed', 'no-k']:
c[section][sub]['api_key_env'] = 'LITELLM_API_KEY'
del c[section][sub]['api_key']
fixed = True
if fixed:
with open('$cfg', 'w') as f: yaml.dump(c, f, default_flow_style=False, allow_unicode=True)
print('Fixed. Restart gateway to apply.')
else:
print('No fixable sections.')
\"" 2>/dev/null && FIXED=$((FIXED+1)) || true
done
else
echo " → No misconfigured agent on $src_ip."
ESCALATED=$((ESCALATED+1))
fi
fi
done
fi
# ═══════════════════════════════════════════════════════
# PHASE 2: COMPILE STATUS
# ═══════════════════════════════════════════════════════
if [ "$ISSUES" -eq 0 ]; then
OVERALL="healthy"
elif [ "$ISSUES" -le 2 ]; then
OVERALL="degraded"
else
OVERALL="down"
fi
SUMMARY="LiteLLM Health: $OVERALL | Checks: $(echo "$DETAILS" | python3 -c "import sys,json;print(len(json.load(sys.stdin)))" 2>/dev/null || echo "?") | Issues: $ISSUES | Fixed: $FIXED | Escalated: $ESCALATED"
echo ""
echo "════════════════════════════════════════"
echo " Status: $OVERALL"
echo " Issues: $ISSUES found | $FIXED auto-fixed | $ESCALATED escalated"
echo "════════════════════════════════════════"
# ═══════════════════════════════════════════════════════
# PHASE 3: LOG TO FILESYSTEM
# ═══════════════════════════════════════════════════════
mkdir -p "$LOG_DIR"
REPORT=$(python3 -c "
import json
print(json.dumps({
'run_id': '$RUN_ID',
'timestamp': '$TIMESTAMP',
'overall_status': '$OVERALL',
'issues_found': $ISSUES,
'issues_fixed': $FIXED,
'issues_escalated': $ESCALATED,
'checks': $DETAILS
}, indent=2))
" 2>/dev/null)
echo "$REPORT" > "$LOG_DIR/${RUN_ID}.json"
# 9. GPU Monitor self-check
echo -n " GPU Monitor... "
if curl -s -o /dev/null -w "%{http_code}" --connect-timeout 5 "$GPU_DASHBOARD/health" 2>/dev/null | grep -q 200; then
add_detail "gpu-monitor" "pass" "Monitor server health OK"
echo "pass"
else
add_detail "gpu-monitor" "warn" "GPU monitor health endpoint unreachable"
echo "warn"
fi
echo ""
echo " Report saved: $LOG_DIR/${RUN_ID}.json"
# ═══════════════════════════════════════════════════════
# PHASE 4: FEED TO KNOWLEDGE GRAPH
# PHASE 4: LOG TO GITEA (hard rule: health logs NEVER go to knowledge graph)
# Logged to SyslogSolution/health-logs/litellm/{run_id}.json — versioned, searchable.
echo -n " Gitea... "
/opt/inference-harness/scripts/gitea-logger.sh litellm "${RUN_ID}.json" "${LOG_DIR}/${RUN_ID}.json"
# PHASE 5: NOTIFY ON ISSUES
# ═══════════════════════════════════════════════════════
if [ "$ISSUES" -gt 0 ]; then
echo ""
echo "⚠️ $ISSUES issue(s) detected. Creating relay alert..."
ALERT_TITLE="⚠ LiteLLM Health Alert — $OVERALL ($ISSUES issues)"
ALERT_BODY="$SUMMARY
Details:
$DETAILS"
mcp_call "tools/call" "createRelayNode" "{\"name\":\"createRelayNode\",\"arguments\":{\"title\":\"$ALERT_TITLE\",\"source\":\"$ALERT_BODY\",\"description\":\"LiteLLM health check failure alert\"}}" > /dev/null 2>&1
fi
echo ""
echo "[$(date)] Health check complete — $RUN_ID"
exit 0
-6
View File
@@ -1,6 +0,0 @@
# SSL Directory
SSL termination is handled upstream by NetBird/Authentik.
This directory is intentionally empty — no certs stored here.
For local dev SSL, use the docker-compose.override.yml pattern.
Submodule syslog-harness-check added at b65ea22765