litellm_config.yaml - client-visible model_list reduced to capability names: syslog-auto, gpu-dense, strix-moe, gpu-vision - retired qwen3.6-27B-code, qwen3.8-27B-uncensored, qwen3.6-35B-udq4 (and the already-dead gemma-4-12b) - fallbacks and model_cost re-keyed to the surviving names - upstream litellm_params.model set to the capability names; the backends ignore the model field (verified HTTP 200 on all three hosts), so this needs no llama.cpp relaunch and loses no KV warmth - verified live: /v1/models advertises exactly the four names, each returns 200 through nginx gpu_roster.yaml - launch args replaced with the VERIFIED live command lines, read from the running processes via the PVE guest agent (acerpve VM 101 = .8, ocupve VM 103 = .110) and on amdpve (.15) - .8: model_path corrected to Qwen3.8-27B-Uncensored-Q4_K_M.gguf, max_concurrent 2 -> 1, context 262144 -> 131072, full arg list recorded (incl. --parallel 1 and the new --slot-save-path / --metrics added 2026-09-12) - .15: model_path corrected to Carnice-Qwen3.6-MoE-35B-A3B-Q4_K_M.gguf, full arg list recorded - keys renamed to the capability names; hosts.current_model aligned README.md / dashboard - README dense entry corrected (1 slot, 131K ctx, actual model file) - dashboard picker now uses capability names; fixed the "Gemma 4 12B" / "12B VLM" labels (the RTX 5070 serves a 9B Qwen3.5) and the gpu-vision -> gpu-light id mapping scripts/ - added the three operational monitors as tracked files (they were untracked): gpu-monitor.py, gpu-self-heal.py, litellm-health-check.sh - cleared their references to retired model names, which were causing failed calls every benchmark cycle (150 failed gemma-4-12b calls in the last 7 days); gpu-monitor.py's .110 entry also wrongly listed .8's model Intentionally NOT changed - LITELLM-MIGRATION-PLAN.md: historical planning document (June 14), not a live-state claim - backups/, graphify-out/, litellm_config.yaml.backup: historical artifacts - unrelated untracked files (router.py, docker-compose.yml.pre-1991-20260911, nginx/default.conf, dashboard/gpu-monitor.html, scripts/gitea-logger.sh): out of scope for this change
syslog-harness — Inference API Harness
CT 116 Docker stack for routing local GPU models through a unified OpenAI-compatible API.
Architecture
nginx :80 → router :9000 → GPU backends
├─ qwen3.6-35B-A3B (MoE) @ 192.168.68.15:8080 [2 slots, 262K ctx]
├─ gpu-dense (Qwen3.8-27B-U) @ 192.168.68.8:8080 [1 slot, 131K ctx]
└─ gpu-vision (VLM) @ 192.168.68.110:8080 [2 slots, 262K ctx]
Total: 6 concurrent slots
LiteLLM :8081 (fallback) | Dashboard :3000 | Redis :6379 (local)
Deploy
cd /opt/inference-harness
docker compose up -d
Endpoints
| URL | Purpose |
|---|---|
/v1/chat/completions |
Inference API (OpenAI-compatible) — API key required |
/v1/models |
Available models |
/ |
Dashboard (GPU health, routing, agents, timeseries) |
Authentication
All /v1/chat/completions requests require a valid API key via Authorization: Bearer <key>. Missing or invalid keys return 401 Unauthorized.
Agent API Keys
| Agent | Key |
|---|---|
| Abiba | sk-syslog-abiba |
| Mumuni | sk-syslog-mumuni |
| Tanko | sk-syslog-tanko |
| Koby | sk-syslog-koby |
| Kagenz0 | sk-syslog-kagenz0 |
| Koonimo | sk-syslog-koonimo |
Routing Tiers
| Tier | Trigger | Priority |
|---|---|---|
| Lightweight | No system prompt, ≤1 turn, ≤100 words | VLM → MoE → Dense |
| Simple Conv | ≤1000 tokens, ≤4 turns | VLM → MoE → Dense |
| Heavy | >4000 tokens OR >8 turns | Dense → MoE → VLM |
| Default | Everything else | MoE → VLM → Dense |
Queue
When all GPUs are saturated, requests enter a polling queue (500ms intervals) instead of returning 503 immediately. Timeout: 30s (configurable via QUEUE_TIMEOUT env or X-Queue-Timeout header).
Models
| GPU | Model | VRAM | Slots | Context | Best For | |-----|-------|------|-------| | Strix Halo | qwen3.6-35B-A3B (MoE) | 65GB | 2 | 262K | General quality | | RTX 3090 | gpu-dense (Qwen3.8-27B-Uncensored) | 24GB | 1 | 131K | Dense, general | | RTX 5070 | gpu-vision (VLM) | 12GB | 2 | 262K | Speed, vision |
Maintenance
Automated cron job runs daily at 3:00 AM UTC (/opt/inference-harness/maintenance.sh):
- Cleans Redis timeseries keys >60 days
- Prunes Docker build cache >7 days
- Logs container health and Redis memory
Logs: /var/log/harness-maintenance.log