Author SHA1 Message Date
mumuni-bot 82464d2158 README table: gpu-vision slots/ctx aligned to roster (parallel 1, 131K) 2026-09-20 15:33:31 +00:00
mumuni-bot 4c6cf66aed Merge pull request 'Align harness repo with verified live state; retire model-version names from the client surface' (#2) from fix/harness-align-20260912 into main 2026-09-20 15:32:07 +00:00
mumuni-bot afccc29034 gpu-self-heal: repair Rule-7 warning line (bad quoting/indent from prior patch) 2026-09-20 15:31:52 +00:00
mumuni-bot 502f2d17c3 health-check: route remaining bearer placeholders through $LLK env guard 2026-09-20 15:29:42 +00:00
mumuni-bot 973335d35b security+alignment pass from PR review: env-var keys (strip dead literals), Rule-7 unmapped-GPU warning, README slots vs live roster args 2026-09-20 15:27:57 +00:00
mumuni-bot b4f94bf672 security+alignment pass from PR review: env-var keys (strip dead literals), Rule-7 unmapped-GPU warning, README slots vs live roster args 2026-09-20 15:27:53 +00:00
mumuni-bot 4f3a28ff10 security+alignment pass from PR review: env-var keys (strip dead literals), Rule-7 unmapped-GPU warning, README slots vs live roster args 2026-09-20 15:27:52 +00:00
mumuni-bot f893e4f611 security+alignment pass from PR review: env-var keys (strip dead literals), Rule-7 unmapped-GPU warning, README slots vs live roster args 2026-09-20 15:27:51 +00:00
mumuni-bot 80811511e7 fix(harness): align repo with verified live state; retire model-version names from the client surface
litellm_config.yaml
- client-visible model_list reduced to capability names: syslog-auto, gpu-dense, strix-moe, gpu-vision
- retired qwen3.6-27B-code, qwen3.8-27B-uncensored, qwen3.6-35B-udq4 (and the already-dead gemma-4-12b)
- fallbacks and model_cost re-keyed to the surviving names
- upstream litellm_params.model set to the capability names; the backends ignore the model field
  (verified HTTP 200 on all three hosts), so this needs no llama.cpp relaunch and loses no KV warmth
- verified live: /v1/models advertises exactly the four names, each returns 200 through nginx

gpu_roster.yaml
- launch args replaced with the VERIFIED live command lines, read from the running processes via the
  PVE guest agent (acerpve VM 101 = .8, ocupve VM 103 = .110) and on amdpve (.15)
- .8: model_path corrected to Qwen3.8-27B-Uncensored-Q4_K_M.gguf, max_concurrent 2 -> 1,
  context 262144 -> 131072, full arg list recorded (incl. --parallel 1 and the new
  --slot-save-path / --metrics added 2026-09-12)
- .15: model_path corrected to Carnice-Qwen3.6-MoE-35B-A3B-Q4_K_M.gguf, full arg list recorded
- keys renamed to the capability names; hosts.current_model aligned

README.md / dashboard
- README dense entry corrected (1 slot, 131K ctx, actual model file)
- dashboard picker now uses capability names; fixed the "Gemma 4 12B" / "12B VLM" labels (the RTX 5070
  serves a 9B Qwen3.5) and the gpu-vision -> gpu-light id mapping

scripts/
- added the three operational monitors as tracked files (they were untracked): gpu-monitor.py,
  gpu-self-heal.py, litellm-health-check.sh
- cleared their references to retired model names, which were causing failed calls every benchmark
  cycle (150 failed gemma-4-12b calls in the last 7 days); gpu-monitor.py's .110 entry also wrongly
  listed .8's model

Intentionally NOT changed
- LITELLM-MIGRATION-PLAN.md: historical planning document (June 14), not a live-state claim
- backups/, graphify-out/, litellm_config.yaml.backup: historical artifacts
- unrelated untracked files (router.py, docker-compose.yml.pre-1991-20260911, nginx/default.conf,
  dashboard/gpu-monitor.html, scripts/gitea-logger.sh): out of scope for this change
2026-09-12 21:03:48 +00:00
8 changed files with 1268 additions and 68 deletions
+6 -6
View File
@@ -6,10 +6,10 @@ CT 116 Docker stack for routing local GPU models through a unified OpenAI-compat
```
nginx :80 → router :9000 → GPU backends
├─ qwen3.6-35B-A3B (MoE) @ 192.168.68.15:8080 [2 slots, 262K ctx]
├─ qwen3.6-27B-code (Dense) @ 192.168.68.8:8080 [2 slots, 262K ctx]
└─ gpu-vision (VLM) @ 192.168.68.110:8080 [2 slots, 262K ctx]
Total: 6 concurrent slots
├─ strix-moe (Qwen3.6-MoE-35B-A3B) @ 192.168.68.15:8080 [2 slots, 262K ctx]
├─ gpu-dense (Qwen3.8-27B-U) @ 192.168.68.8:8080 [1 slot, 131K ctx]
└─ gpu-vision (Qwen3.5-9B VLM) @ 192.168.68.110:8080 [1 slot, 131K ctx]
Total: 4 concurrent slots
LiteLLM :8081 (fallback) | Dashboard :3000 | Redis :6379 (local)
```
@@ -62,8 +62,8 @@ When all GPUs are saturated, requests enter a polling queue (500ms intervals) in
| GPU | Model | VRAM | Slots | Context | Best For |
|-----|-------|------|-------|
| Strix Halo | qwen3.6-35B-A3B (MoE) | 65GB | 2 | 262K | General quality |
| RTX 3090 | qwen3.6-27B-code (Dense) | 24GB | 2 | 262K | Code, reasoning |
| RTX 5070 | gpu-vision (VLM) | 12GB | 2 | 262K | Speed, vision |
| RTX 3090 | gpu-dense (Qwen3.8-27B-Uncensored) | 24GB | 1 | 131K | Dense, general |
| RTX 5070 | gpu-vision (VLM) | 12GB | 1 | 131K | Speed, vision |
## Maintenance
+1 -1
View File
@@ -70,7 +70,7 @@ body{background:var(--bg);color:var(--text);font-family:-apple-system,BlinkMacSy
</div>
<script>
const COLORS={'qwen3.6-35B-A3B':'#10b981','qwen3.6-27B-code':'#8b5cf6','gpu-vision':'#3b82f6'};
const COLORS={'strix-moe':'#10b981','gpu-dense':'#8b5cf6','gpu-vision':'#3b82f6'};
const HISTORY=[]; // rolling 60 sample history for chart
function Q(id){return document.getElementById(id)}
+13 -13
View File
@@ -119,9 +119,9 @@ body { background: #0b0f17; color: #bcc3cd; font-family: -apple-system, BlinkMac
<div class="d-flex gap-2">
<select id="scatter-model" onchange="loadScatter()" style="font-size:10px;background:#1e293b;color:#94a3b8;border:1px solid #334155;border-radius:4px;padding:2px 6px">
<option value="all">All Models</option>
<option value="gpu-vision">12B VLM</option>
<option value="qwen3.6-27B-code">27B Dense</option>
<option value="qwen3.6-35B-A3B">35B MoE</option>
<option value="gpu-vision">9B VLM</option>
<option value="gpu-dense">27B Dense</option>
<option value="strix-moe">35B MoE</option>
</select>
</div>
</div><div id="scatter-plot" style="height:200px;position:relative"></div><div id="scatter-legend" class="d-flex justify-content-center gap-3 mt-2 flex-wrap small"></div></div></div>
@@ -136,9 +136,9 @@ body { background: #0b0f17; color: #bcc3cd; font-family: -apple-system, BlinkMac
</div>
<script>
var MC={'gpu-vision':'#22c55e','qwen3.6-27B-code':'#f59e0b','qwen3.6-35B-A3B':'#a78bfa'};
var ML={'gpu-vision':'Gemma 4 12B','qwen3.6-27B-code':'Qwen Code','qwen3.6-35B-A3B':'Qwen MoE'};
var GL={'qwen3.6-35B-A3B':'MoE - Strix Halo','qwen3.6-27B-code':'Dense - RTX 3090','gpu-vision':'VLM - RTX 5070'};
var MC={'gpu-vision':'#22c55e','gpu-dense':'#f59e0b','strix-moe':'#a78bfa'};
var ML={'gpu-vision':'Qwen3.5-9B VLM','gpu-dense':'Qwen3.8-27B','strix-moe':'Qwen MoE'};
var GL={'strix-moe':'MoE - Strix Halo','gpu-dense':'Dense - RTX 3090','gpu-vision':'VLM - RTX 5070'};
function $(id){return document.getElementById(id);}
function render(data){
@@ -147,7 +147,7 @@ var t=Object.values(data.route_counts||{}).reduce((a,b)=>a+b,0);
var ta=0,tm=0;data.gpus.forEach(function(g){ta+=(g.active_requests||0);tm+=(g.max_concurrent||1)});
$('kpi-total').textContent=t;$('kpi-active').textContent=ta+'/'+tm;$('kpi-agents').textContent=Object.keys(data.agent_counts||{}).length;
$('update-time').textContent=new Date().toLocaleTimeString();
var ids={'qwen3.6-35B-A3B':'gpu-moe','qwen3.6-27B-code':'gpu-dense','gpu-vision':'gpu-light'};
var ids={'strix-moe':'gpu-moe','gpu-dense':'gpu-dense','gpu-vision':'gpu-vision'};
data.gpus.forEach(function(g){
var el=$(ids[g.id]);if(!el)return;
var a=g.active_requests||0,mx=g.max_concurrent||1;
@@ -179,14 +179,14 @@ var sc=pct>=100?'#ef4444':pct>=50?'#f59e0b':'#22c55e';
var circ=188.5,dash=(pct/100)*circ;
var h='<div class=\"d-inline-block position-relative mb-2\"><svg width=\"72\" height=\"72\"><circle cx=\"36\" cy=\"36\" r=\"30\" fill=\"none\" stroke=\"#1e293b\" stroke-width=\"6\"/><circle cx=\"36\" cy=\"36\" r=\"30\" fill=\"none\" stroke=\"'+sc+'\" stroke-width=\"6\" stroke-dasharray=\"'+dash+' '+(circ-dash)+'\" stroke-linecap=\"round\" transform=\"rotate(-90 36 36)\"/></svg><div style=\"position:absolute;top:50%;left:50%;transform:translate(-50%,-50%);text-align:center\"><div class=\"ring-label\" style=\"color:'+sc+'\">'+ta+'</div><div class=\"ring-sublabel\">/ '+tm+' slots</div></div></div>';
h+='<div class=\"fw-bold mb-2 small\" style=\"color:'+sc+'\">'+st+'</div>';
var lb={'qwen3.6-35B-A3B':'MoE','qwen3.6-27B-code':'Dense','gpu-vision':'VLM'};
var lb={'strix-moe':'MoE','gpu-dense':'Dense','gpu-vision':'VLM'};
data.gpus.forEach(function(g){var a=g.active_requests||0,mx=g.max_concurrent||1,gp=mx>0?Math.round(a/mx*100):0;h+='<div class=\"d-flex align-items-center gap-2 mb-1 justify-content-center\"><span class=\"small\" style=\"min-width:32px;text-align:right;font-size:10px\">'+(lb[g.id]||g.id)+'</span><div style=\"flex:1;max-width:70px;height:3px;background:#1e293b;border-radius:2px;overflow:hidden\"><div style=\"height:100%;width:'+gp+'%;background:'+sc+';border-radius:2px\"></div></div><span class=\"small\" style=\"min-width:22px;font-size:10px\">'+a+'/'+mx+'</span></div>'});
el.innerHTML=h;
}
function renderGPUMetrics(data){
var el=$('gpu-metrics-card');if(!el)return;
var lb={'qwen3.6-35B-A3B':'MoE','qwen3.6-27B-code':'Dense','gpu-vision':'VLM'};
var lb={'strix-moe':'MoE','gpu-dense':'Dense','gpu-vision':'VLM'};
var h='';data.gpus.forEach(function(g){
var nm=lb[g.id]||g.id,tp=g.temp_c||0,ut=g.gpu_util_pct||0,pw=g.power_w||0,pl=g.power_limit_w||0;
var tc=tp>85?'#ef4444':tp>70?'#f59e0b':'#22c55e',uc=ut>90?'#ef4444':ut>70?'#f59e0b':'#22c55e';
@@ -219,8 +219,8 @@ function loadPerf(){fetch('/api/performance?window='+perfWindow).then(function(r
function renderPerf(d){
var models=d.models||[],reasons=d.reasons||[],agents=d.agents||[],sum=d.summary||{};
// Latency bars: p50/p95/p99 per model
var mlab={'qwen3.6-35B-A3B':'35B MoE','qwen3.6-27B-code':'27B Dense','gpu-vision':'12B VLM'};
var mcol={'qwen3.6-35B-A3B':'#a78bfa','qwen3.6-27B-code':'#f59e0b','gpu-vision':'#22c55e'};
var mlab={'strix-moe':'35B MoE','gpu-dense':'27B Dense','gpu-vision':'9B VLM'};
var mcol={'strix-moe':'#a78bfa','gpu-dense':'#f59e0b','gpu-vision':'#22c55e'};
if(!models.length){$('perf-latency').innerHTML='<div class="text-secondary small text-center py-4">Accumulating data...</div>';return;}
var maxLat=Math.max(...models.map(function(m){return m.latency.p99||0}),1);
var latHTML=models.map(function(m){
@@ -270,8 +270,8 @@ fetch('/api/scatter?window=24&model='+m).then(function(r){return r.json()}).then
function renderScatter(d){
var pts=d.points||[],el=$('scatter-plot'),lg=$('scatter-legend');
if(!pts.length){el.innerHTML='<div class="text-secondary small text-center py-5">No data yet</div>';return;}
var mcol={'qwen3.6-35B-A3B':'#a78bfa','qwen3.6-27B-code':'#f59e0b','gpu-vision':'#22c55e','unknown':'#38bdf8'};
var mlab={'qwen3.6-35B-A3B':'35B MoE','qwen3.6-27B-code':'27B Dense','gpu-vision':'12B VLM'};
var mcol={'strix-moe':'#a78bfa','gpu-dense':'#f59e0b','gpu-vision':'#22c55e','unknown':'#38bdf8'};
var mlab={'strix-moe':'35B MoE','gpu-dense':'27B Dense','gpu-vision':'9B VLM'};
var maxX=Math.max.apply(null,pts.map(function(p){return p.prompt_tokens||0}))||1000;
var maxY=Math.max.apply(null,pts.map(function(p){return p.inference_ms||0}))||5000;
// Log scale for X axis (prompt tokens vary widely)
+13 -13
View File
@@ -8,31 +8,31 @@ models:
tiers: [starter, professional, enterprise]
capabilities: [completion, multimodal]
model_path: /home/llmuser/models/qwen3.5-9b/Qwen3.5-9B-Q5_K_M.gguf
args: --mmproj /home/llmuser/models/qwen3.5-9b/mmproj-F16.gguf --ctx-size 131072 --cache-type-k q4_0 --cache-type-v q4_0 --flash-attn 1 --parallel 1 --alias gpu-light --reasoning off --api-key not-needed --cache-prompt
args: --mmproj /home/llmuser/models/qwen3.5-9b/mmproj-F16.gguf --ctx-size 131072 --cache-type-k q4_0 --cache-type-v q4_0 --batch-size 2048 --ubatch-size 1024 --n-gpu-layers 99 --flash-attn 1 --cont-batching --parallel 1 --image-min-tokens 2048 --image-max-tokens 16384 --alias gpu-light --reasoning off --api-key not-needed --predict 8192 --cache-prompt --port 8080 --host 0.0.0.0
qwen3.6-27B-code:
gpu-dense:
gpu_url: http://192.168.68.8:8080/v1
sidecar_url: http://192.168.68.8:8090
gpu_host: 192.168.68.8
label: Qwen3.6 27B Code (RTX 3090)
max_concurrent: 2
context: 262144
label: Qwen3.8-27B-Uncensored (RTX 3090)
max_concurrent: 1
context: 131072
tiers: [professional, enterprise]
capabilities: [completion]
model_path: /root/models/Qwen3.6-27B-NEO-CODE-2T-OT-IQ4_NL.gguf
args: --cache-type-k q4_0 --cache-type-v q4_0 --flash-attn on --reasoning off --spec-type draft-mtp --spec-draft-n-max 2 -t 8
model_path: /home/llmuser/models/Qwen3.8-27B-Uncensored-Q4_K_M.gguf
args: -m /home/llmuser/models/Qwen3.8-27B-Uncensored-Q4_K_M.gguf -c 131072 --host 0.0.0.0 --port 8080 -ngl 99 --alias qwen3.6-27B-code --api-key not-needed -ctk q4_0 -ctv q4_0 --flash-attn on --cont-batching --parallel 1 --slot-save-path /home/llmuser/.llama-slots --metrics --reasoning off --spec-type draft-mtp --spec-draft-n-max 2 -t 8
qwen3.6-35B-udq4:
strix-moe:
gpu_url: http://192.168.68.15:8080/v1
sidecar_url: http://192.168.68.15:8090
gpu_host: 192.168.68.15
label: Qwen3.6 35B UD-Q4_K_M (Strix Halo)
label: Carnice Qwen3.6 MoE 35B-A3B (Strix Halo)
max_concurrent: 1
context: 262144
tiers: [professional, enterprise]
capabilities: [completion, multimodal]
model_path: /models/Qwen3.6-35B-A3B-UD-Q4_K_M.gguf
args: -c 262144 -ngl 99 --flash-attn on
model_path: /models/Carnice-Qwen3.6-MoE-35B-A3B-Q4_K_M.gguf
args: -m /models/Carnice-Qwen3.6-MoE-35B-A3B-Q4_K_M.gguf --host 0.0.0.0 --port 8080 --alias strix-moe -c 262144 -ngl 99 -fa on --cache-type-k q4_0 --cache-type-v q4_0 --kv-unified --cache-prompt --cont-batching -b 4096 --ubatch-size 1024 --parallel 2 -t 16 --timeout 600 -n 8192 --reasoning off --no-warmup --metrics
hosts:
gpu-light:
@@ -45,10 +45,10 @@ hosts:
address: 192.168.68.8
gpu_name: NVIDIA GeForce RTX 3090
vram_gb: 24
current_model: qwen3.6-27B-code
current_model: gpu-dense
gpu-moe:
address: 192.168.68.15
gpu_name: AMD Strix Halo (iGPU)
vram_gb: 64
current_model: qwen3.6-35B-udq4
current_model: strix-moe
+9 -35
View File
@@ -47,7 +47,7 @@ litellm_settings:
failure_callback:
- prometheus
model_cost:
qwen3.6-27B-code:
gpu-dense:
input_cost_per_token: 1.5e-07
output_cost_per_token: 6.0e-07
strix-moe:
@@ -56,28 +56,13 @@ litellm_settings:
syslog-auto:
input_cost_per_token: 1.5e-07
output_cost_per_token: 6.0e-07
gpu-vision:
input_cost_per_token: 1.5e-07
output_cost_per_token: 6.0e-07
num_retries: 2
request_timeout: 600
sso_callback: /sso/callback
model_list:
- litellm_params:
api_base: http://192.168.68.8:8080/v1
api_key: not-needed
model: openai/qwen3.6-27B-code
timeout: 300
model_info:
max_input_tokens: 131072
model_name: qwen3.6-27B-code
- litellm_params:
api_base: http://192.168.68.15:8080/v1
api_key: not-needed
model: openai/strix-moe
timeout: 300
model_info:
max_input_tokens: 131072
max_model_tokens: 131072
max_tokens: 131072
model_name: qwen3.6-35B-udq4
- litellm_params:
api_base: http://192.168.68.15:8080/v1
api_key: not-needed
@@ -92,7 +77,7 @@ model_list:
- litellm_params:
api_base: http://192.168.68.8:8080/v1
api_key: not-needed
model: openai/qwen3.6-27B-code
model: openai/gpu-dense
rpm: 500
timeout: 300
model_info:
@@ -114,7 +99,7 @@ model_list:
- litellm_params:
api_base: http://192.168.68.8:8080/v1
api_key: not-needed
model: openai/qwen3.6-27B-code
model: openai/gpu-dense
rpm: 500
timeout: 300
model_info:
@@ -123,17 +108,6 @@ model_list:
max_tokens: 131072
weight: 0.70
model_name: syslog-auto
- litellm_params:
api_base: http://192.168.68.8:8080/v1
api_key: not-needed
model: openai/qwen3.6-27B-code
rpm: 500
timeout: 300
model_info:
max_input_tokens: 131072
max_model_tokens: 131072
max_tokens: 131072
model_name: qwen3.8-27B-uncensored
- litellm_params:
api_base: http://192.168.68.15:8080/v1
api_key: not-needed
@@ -159,13 +133,13 @@ router_settings:
enable_loadbalancing_on_proxy: true
fallbacks:
- syslog-auto:
- qwen3.6-27B-code
- gpu-dense
- strix-moe
- gpu-vision
- qwen3.6-27B-code:
- gpu-dense:
- strix-moe
- strix-moe:
- qwen3.6-27B-code
- gpu-dense
- gpu-vision
request_timeout: 300
routing_strategy: simple-shuffle
+394
View File
@@ -0,0 +1,394 @@
#!/usr/bin/env python3
"""GPU Fleet Monitor — Comprehensive monitoring server.
Monitors every subsystem in the inference harness:
- GPU sidecars (VRAM, temp, util, power per card)
- Router (model routing, circuit breakers, queue health)
- LiteLLM (proxy health, key count, model sync)
- Strix Halo (llama-server status, CPU load, context)
- Dashboard (CT 116 harness-dashboard)
Endpoints:
/ → GPU fleet HTML dashboard
/gpu-data → JSON GPU metrics (for dashboard API)
/health → Monitor self-health check
Run: python3 gpu-monitor-server.py
Default port: 9100
"""
import json, subprocess, http.server, threading, time, os
from datetime import datetime, timezone
DASHBOARD_PATH = "/root/dashboard/gpu-fleet.html"
PORT = 9100
# Cached data, refreshed every 15s
cache = {}
cache_lock = threading.Lock()
# Alert thresholds
THRESHOLDS = {
"temp_c": {"warning": 80, "critical": 90},
"vram_pct": {"warning": 90, "critical": 95},
"gpu_util_pct": {"warning": 95, "critical": 98},
}
def http_get(url, timeout=5):
"""HTTP GET with timeout, returns parsed JSON or error dict."""
import urllib.request
try:
resp = urllib.request.urlopen(url, timeout=timeout)
return json.loads(resp.read())
except Exception as e:
return {"error": str(e)}
def poll_sidecar(host, port=8090):
"""Poll a GPU sidecar for raw metrics.
Falls back to SSH-based nvidia-smi if sidecar is unreachable."""
result = http_get(f"http://{host}:{port}/health", timeout=5)
if "error" not in result:
return result
# Fallback: poll via SSH (explicit key path for non-interactive environments)
try:
proc = subprocess.run(
["/usr/bin/ssh", "-o", "StrictHostKeyChecking=no", "-o", "ConnectTimeout=5",
"-i", "/root/.ssh/id_ed25519", host,
"nvidia-smi", "--query-gpu=name,temperature.gpu,utilization.gpu,"
"utilization.memory,memory.used,memory.total,power.draw,power.limit,fan.speed",
"--format=csv,noheader"],
capture_output=True, text=True, timeout=10
)
if proc.returncode == 0:
parts = [p.strip() for p in proc.stdout.strip().split(",")]
if len(parts) >= 8:
def cv(v):
try: return float(v.replace("MiB","").replace("W","").replace("%","").strip())
except: return 0.0
def mv(v):
try: return int(v.replace("MiB","").strip())
except: return 0
return {
"gpu_name": parts[0], "temp_c": cv(parts[1]),
"gpu_util_pct": cv(parts[2]), "mem_util_pct": cv(parts[3]),
"vram_used_mb": mv(parts[4]), "vram_total_mb": mv(parts[5]),
"power_w": cv(parts[6]), "power_limit_w": cv(parts[7]),
"fan_pct": cv(parts[8])
}
except Exception as e:
return {"error": str(e)}
return result
def poll_router():
"""Poll router unified health through nginx (port 80)."""
return http_get("http://192.168.68.116/health/unified", timeout=5)
def poll_router_health():
"""Poll router basic health through nginx."""
return http_get("http://192.168.68.116/health", timeout=5)
def poll_litellm_health():
"""Poll LiteLLM — checks reachability via nginx proxy.
LiteLLM health endpoints require authentication. We check if the
service responds at all (even 401 = service is up). Also try the
/key/list endpoint which confirms the proxy is fully functional.
"""
import urllib.request
import urllib.error
urls_to_try = [
"http://192.168.68.116/litellm/health",
"http://192.168.68.116/litellm/ui/",
]
last_error = None
for url in urls_to_try:
try:
resp = urllib.request.urlopen(url, timeout=5)
body = resp.read().decode()[:500]
# Any response (even 401) means the proxy is routing to LiteLLM
return {"reachable": True, "status_code": resp.getcode(),
"endpoint": url, "body_preview": body[:200]}
except urllib.error.HTTPError as e:
# HTTP error means the service IS reachable but returned error
return {"reachable": True, "status_code": e.code,
"endpoint": url, "note": f"HTTP {e.code} (requires auth)"}
except Exception as e:
last_error = str(e)
return {"error": last_error or "all endpoints failed"}
def poll_strix():
"""Poll Strix Halo — process check + CPU load + llama-server health."""
try:
# Check llama-server process
proc = subprocess.run(
["ssh", "-o", "StrictHostKeyChecking=no", "-o", "ConnectTimeout=5",
"192.168.68.15",
"pgrep -f llama-server > /dev/null 2>&1 && echo running || echo stopped"],
capture_output=True, text=True, timeout=10
)
status = proc.stdout.strip()
# Get CPU load
load = subprocess.run(
["ssh", "-o", "StrictHostKeyChecking=no", "-o", "ConnectTimeout=5",
"192.168.68.15",
"uptime | awk -F'load average:' '{print $2}' | tr -d ' '"],
capture_output=True, text=True, timeout=10
)
cpu_load = load.stdout.strip() if load.returncode == 0 else "unknown"
# Try llama-server health endpoint
llama_health = http_get("http://192.168.68.15:8080/health", timeout=3)
return {
"status": status if status else "unknown",
"cpu_load": cpu_load,
"llama_health": llama_health if "error" not in llama_health else None
}
except Exception as e:
return {"status": "error", "error": str(e)}
def poll_dashboard():
"""Check if harness-dashboard on CT 116 is serving."""
return http_get("http://192.168.68.116/dashboard/", timeout=5)
def check_alerts(gpu_data):
"""Check GPU metrics against thresholds, return alert list."""
alerts = []
name = gpu_data.get("name", "unknown")
for metric, thresholds in THRESHOLDS.items():
# Map metric names from sidecar/router formats
value = None
if metric == "vram_pct":
# Compute from vram_used_mb / vram_total_mb
used = gpu_data.get("vram_used_mb", 0)
total = gpu_data.get("vram_total_mb", 1)
value = (used / total) * 100
elif metric == "gpu_util_pct":
value = gpu_data.get("gpu_util_pct", 0)
elif metric == "temp_c":
value = gpu_data.get("temp_c", 0)
if value is not None:
if value >= thresholds["critical"]:
alerts.append({"gpu": name, "metric": metric, "level": "critical",
"value": round(value, 1), "threshold": thresholds["critical"]})
elif value >= thresholds["warning"]:
alerts.append({"gpu": name, "metric": metric, "level": "warning",
"value": round(value, 1), "threshold": thresholds["warning"]})
return alerts
def compute_summary(router_data, gpus, litellm_data, strix_data):
"""Compute fleet-wide health summary."""
# Count models available
models_available = 0
models_total = 0
if "available_models" in router_data:
models_available = len(router_data["available_models"])
models_total = models_available # router reports all expected
# Count circuit breakers open
cb_open = 0
if "circuit_breaker" in router_data:
for model, cb in router_data["circuit_breaker"].items():
if cb.get("open", 0):
cb_open += 1
# Count GPU errors
gpu_errors = sum(1 for g in gpus if "error" in g)
# Fleet status
if gpu_errors > 0 or cb_open > 0:
fleet_status = "degraded"
elif not router_data or "error" in router_data:
fleet_status = "degraded"
else:
fleet_status = "healthy"
litellm_ok = ("error" not in litellm_data) or litellm_data.get("reachable", False)
return {
"fleet_status": fleet_status,
"models_available": models_available,
"models_total": models_total,
"circuit_breakers_open": cb_open,
"gpu_count": len(gpus),
"gpu_errors": gpu_errors,
"strix_running": strix_data.get("status") == "running",
"router_reachable": "error" not in router_data,
"litellm_reachable": litellm_ok,
}
def refresh_cache():
"""Refresh all fleet data from every subsystem."""
global cache
# 1. GPU sidecars
gpu_map = {
"192.168.68.8": {"name": "Dense (RTX 3090)", "models": ["gpu-dense"], "hostname": "ct8"},
"192.168.68.110": {"name": "Light/Vision (RTX 5070)", "models": ["gpu-vision"], "hostname": "ct110"},
}
gpus = []
all_alerts = []
for host, info in gpu_map.items():
data = poll_sidecar(host)
if "error" not in data:
data["name"] = info["name"]
data["models"] = info["models"]
data["hostname"] = info["hostname"]
all_alerts.extend(check_alerts(data))
else:
data = {"name": info["name"], "hostname": info["hostname"], "error": data["error"]}
all_alerts.append({"gpu": info["name"], "metric": "connectivity",
"level": "critical", "value": data["error"]})
gpus.append(data)
# 2. Router
router = poll_router()
router_basic = poll_router_health()
router["_basic"] = router_basic
# 3. LiteLLM
litellm = poll_litellm_health()
# 4. Strix Halo
strix = poll_strix()
# 5. Dashboard
dashboard = poll_dashboard()
# 6. Send Zulip DM for critical alerts (debounced 5 min)
critical = [a for a in all_alerts if a.get("level") == "critical"]
if critical and _debounce_alert():
msg = "🔴 **GPU Fleet Alert**\n"
for a in critical:
msg += f"• {a['gpu']}: {a['metric']} = {a.get('value', a.get('actual', '?'))}\n"
_send_zulip_dm(msg)
# 7. Summary
summary = compute_summary(router, gpus, litellm, strix)
with cache_lock:
cache = {
"gpus": gpus,
"strix": strix,
"router": router,
"litellm": litellm,
"dashboard": {"reachable": "error" not in dashboard},
"summary": summary,
"alerts": all_alerts,
"updated": datetime.now(timezone.utc).isoformat(),
"monitor_version": "2.0.0",
}
_last_alert_time = 0
_ALERT_COOLDOWN = 300 # 5 min between DMs
def _send_zulip_dm(msg):
"""Send a DM to the owner via Zulip."""
try:
site = "https://chat.sysloggh.net"
email = os.environ.get("ABIBA_ZULIP_EMAIL", "abiba-bot@chat.sysloggh.net")
key = os.environ.get("ZULIP_API_KEY", "")
if not key:
print("[alert] ZULIP_API_KEY unset — DM suppressed")
return
owner = 9
auth_b64 = base64.b64encode(f"{email}:{key}".encode()).decode()
data = urllib.parse.urlencode({"type": "private", "to": f"[{owner}]", "content": msg}).encode()
req = urllib.request.Request(f"{site}/api/v1/messages", data=data,
headers={"Authorization": f"Basic {auth_b64}"}, method="POST")
urllib.request.urlopen(req, timeout=5)
except Exception as e:
print(f"[alert] DM failed: {e}")
def _debounce_alert():
"""Returns True if enough time has passed since last alert."""
global _last_alert_time
now = time.time()
if now - _last_alert_time > _ALERT_COOLDOWN:
_last_alert_time = now
return True
return False
class DashboardHandler(http.server.BaseHTTPRequestHandler):
def do_GET(self):
if self.path == "/gpu-data":
with cache_lock:
data = json.dumps(cache, indent=2)
self.send_response(200)
self.send_header("Content-Type", "application/json")
self.send_header("Access-Control-Allow-Origin", "*")
self.end_headers()
self.wfile.write(data.encode())
elif self.path == "/health":
with cache_lock:
healthy = bool(cache and "summary" in cache)
status = {"status": "healthy" if healthy else "starting",
"updated": cache.get("updated", "never") if cache else "never"}
code = 200 if healthy else 503
self.send_response(code)
self.send_header("Content-Type", "application/json")
self.end_headers()
self.wfile.write(json.dumps(status).encode())
elif self.path in ("/", "/index.html"):
try:
with open(DASHBOARD_PATH) as f:
html = f.read()
self.send_response(200)
self.send_header("Content-Type", "text/html")
self.end_headers()
self.wfile.write(html.encode())
except FileNotFoundError:
self.send_response(404)
self.end_headers()
self.wfile.write(b"Dashboard not found")
else:
self.send_response(404)
self.end_headers()
def log_message(self, format, *args):
pass # Suppress request logs
def background_refresh():
"""Refresh data every 15 seconds."""
while True:
try:
refresh_cache()
except Exception as e:
print(f"[WARN] Refresh failed: {e}", flush=True)
time.sleep(15)
if __name__ == "__main__":
print(f"Starting Comprehensive GPU Fleet Monitor v2.0.0 on port {PORT}...")
print(f" Dashboard: http://localhost:{PORT}/")
print(f" API: http://localhost:{PORT}/gpu-data")
print(f" Health: http://localhost:{PORT}/health")
print(f" Polling: router (nginx:80), sidecars (:8090), strix, litellm, dashboard")
# Initial fetch
refresh_cache()
# Start background refresher
t = threading.Thread(target=background_refresh, daemon=True)
t.start()
# Start HTTP server
server = http.server.HTTPServer(("0.0.0.0", PORT), DashboardHandler)
try:
server.serve_forever()
except KeyboardInterrupt:
print("\nShutting down...")
+472
View File
@@ -0,0 +1,472 @@
#!/usr/bin/env python3
"""
GPU Self-Heal — Implementation of gpu-self-heal.prose.md
Evaluates 8 remediation rules against live GPU monitor data.
Reports results to knowledge graph via MCP bridge.
"""
import json, time, subprocess, sys, os, socket
from datetime import datetime, timezone
from urllib.request import urlopen, Request
# Global timeout to prevent hanging on unreachable hosts
socket.setdefaulttimeout(10)
GPU_MONITOR = "http://192.168.68.24:9100/gpu-data"
KG_BRIDGE = "http://192.168.68.65:3100/mcp"
GPU_HOSTS = {
"ct8-rtx3090": {"ip": "192.168.68.8", "gpu": "NVIDIA RTX 3090", "vram_total_mb": 24576, "vram_threshold_mb_per_h": 100},
"ct110-rtx5070": {"ip": "192.168.68.110","gpu": "NVIDIA RTX 5070", "vram_total_mb": 12227, "vram_threshold_mb_per_h": 50},
"strix-halo": {"ip": "192.168.68.15", "gpu": "AMD Strix Halo 64GB","vram_total_mb": 65536, "vram_threshold_mb_per_h": 200},
}
HISTORY_FILE = "/var/log/litellm/gpu-history.json"
RUN_ID = f"gpu-self-heal-{datetime.now().strftime('%Y%m%d-%H%M%S')}"
def fetch_gpu_data():
"""Fetch live GPU fleet data from monitor."""
try:
req = Request(GPU_MONITOR)
with urlopen(req, timeout=10) as resp:
return json.loads(resp.read())
except Exception as e:
print(f" FAIL: GPU monitor unreachable: {e}")
return None
def fetch_prometheus(host, port=9400):
"""Check Prometheus exporter on GPU host (with timeout)."""
try:
req = Request(f"http://{host}:{port}/metrics")
with urlopen(req, timeout=3) as resp:
return resp.status == 200
except:
return False
def probe_strix_direct():
"""Direct Strix Halo health probe (firewall opened .24→.15:8080)."""
try:
with urlopen("http://192.168.68.15:8080/health", timeout=5) as resp:
return resp.status == 200
except:
return False
def load_history():
"""Load historical GPU metrics for trend analysis."""
if os.path.exists(HISTORY_FILE):
with open(HISTORY_FILE) as f:
return json.load(f)
return {"snapshots": [], "vram_baselines": {}, "temp_history": {}}
def save_history(history):
"""Persist history for trend analysis."""
os.makedirs(os.path.dirname(HISTORY_FILE), exist_ok=True)
# Keep last 72 hours of snapshots (4320 at 60s intervals, capped at 1000)
history["snapshots"] = history["snapshots"][-1000:]
with open(HISTORY_FILE, "w") as f:
json.dump(history, f, indent=2)
def analyze_vram_trend(history, gpu_name, current_vram_mb):
"""Calculate VRAM growth rate over 6-hour window."""
snapshots = history.get("snapshots", [])
six_hours_ago = time.time() - 21600
old_snapshots = [s for s in snapshots if s.get("timestamp", 0) > six_hours_ago
and s.get("gpu") == gpu_name]
if len(old_snapshots) < 10:
return 0 # Not enough data
oldest = old_snapshots[0]
newest = old_snapshots[-1]
hours = (newest["timestamp"] - oldest["timestamp"]) / 3600
if hours < 1:
return 0
return (current_vram_mb - oldest["vram_mb"]) / hours
def analyze_temp_rise(history, gpu_name, current_temp):
"""Calculate temperature rise rate over last 5 minutes."""
snapshots = history.get("snapshots", [])
five_min_ago = time.time() - 300
recent = [s for s in snapshots if s.get("timestamp", 0) > five_min_ago
and s.get("gpu") == gpu_name]
if len(recent) < 3:
return 0
oldest = recent[0]
minutes = (recent[-1]["timestamp"] - oldest["timestamp"]) / 60
if minutes < 0.5:
return 0
return (current_temp - oldest["temp_c"]) / minutes
def record_snapshot(history, gpu_name, temp_c, vram_mb):
"""Record a GPU snapshot for trend analysis."""
history["snapshots"].append({
"timestamp": time.time(),
"gpu": gpu_name,
"temp_c": temp_c,
"vram_mb": vram_mb
})
def kg_create_node(title, description, source):
"""Log run report to Gitea (hard rule: health logs NEVER go to knowledge graph).
Redirected 2026-08-13 per directive from Mumuni (#726): logs belong in
SyslogSolution/health-logs/gpu/{run_id}.json, not the RA-H OS graph.
"""
try:
# source is the report JSON string; write to temp file and push via gitea-logger
report_file = f"/tmp/{RUN_ID}.json"
with open(report_file, "w") as f:
f.write(source if isinstance(source, str) else json.dumps(source, indent=2))
cmd = f"/opt/inference-harness/scripts/gitea-logger.sh gpu {RUN_ID}.json {report_file}"
result = subprocess.run(cmd, shell=True, capture_output=True, text=True, timeout=120)
os.remove(report_file) if os.path.exists(report_file) else None
return result.stdout.strip() or "gitea-logger: no output"
except Exception as e:
return f"gitea-logger failed: {e}"
def evaluate_rules(data, history):
"""Evaluate all 8 remediation rules against current GPU state."""
actions = []
if not data:
return actions
gpus = data.get("gpus", [])
summary = data.get("summary", {})
benchmarks = data.get("benchmarks", {})
strix = data.get("strix", {})
router = data.get("router", {})
for gpu in gpus:
if not isinstance(gpu, dict):
continue
name = gpu.get("gpu_name", "unknown")
host_key = None
for k, v in GPU_HOSTS.items():
if v["gpu"] in name or k in name:
host_key = k
break
if not host_key:
print(f"[warn] no GPU_HOSTS mapping for {name!r} — remote rules (incl. Rule 7) skipped for this GPU")
host_key = name.replace(" ", "-").lower()
temp = gpu.get("temp_c", 0)
vram_used = gpu.get("vram_used_mb", 0)
host_info = GPU_HOSTS.get(host_key, {"ip": "unknown", "vram_threshold_mb_per_h": 50})
# Record snapshot
record_snapshot(history, name, temp, vram_used)
# ── Rule 1: Thermal Critical ──
if temp > 85:
actions.append({"rule": "thermal-critical", "gpu": name, "temp_c": temp,
"action": "reduce-concurrency", "severity": "critical"})
print(f" ⚠ RULE 1: {name} at {temp}°C — reducing concurrency")
# ── Rule 2: VRAM Leak ──
vram_rate = analyze_vram_trend(history, name, vram_used)
threshold = host_info.get("vram_threshold_mb_per_h", 50)
if vram_rate > threshold:
actions.append({"rule": "vram-leak", "gpu": name, "rate_mb_per_h": vram_rate,
"threshold": threshold, "action": "investigate-leak", "severity": "warning"})
print(f" ⚠ RULE 2: {name} VRAM leak {vram_rate:.1f} MB/h (threshold: {threshold})")
# ── Rule 4: Benchmark Regression ──
bench_latest = benchmarks.get("latest", {})
for bk, bv in bench_latest.items():
if not isinstance(bv, dict):
continue
latest_run = bv.get("latest", {})
baseline = bv.get("baseline_tok_per_sec", 0)
current_tps = latest_run.get("gen_tok_per_sec", 0)
if baseline > 0 and current_tps > 0 and current_tps < baseline * 0.8:
drop_pct = (1 - current_tps / baseline) * 100
gpu_label = bv.get("gpu_name", bk)
actions.append({"rule": "benchmark-regression", "gpu": gpu_label,
"baseline_tps": baseline, "current_tps": current_tps,
"drop_pct": round(drop_pct,1), "action": "investigate-performance",
"severity": "warning"})
print(f" ⚠ RULE 4: {gpu_label} benchmark -{drop_pct:.0f}% ({current_tps} vs {baseline} tok/s)")
# ── Rule 7: Prometheus Exporter ──
ip = host_info.get("ip", "")
if ip and ip != "unknown":
if not fetch_prometheus(ip):
actions.append({"rule": "prometheus-down", "gpu": name, "host": ip,
"action": "restart-exporter", "severity": "warning"})
print(f" ⚠ RULE 7: {name} Prometheus exporter down on {ip}:9400")
# ── Rule 8: Predictive Thermal ──
rise_rate = analyze_temp_rise(history, name, temp)
if temp > 70 and rise_rate > 2.0:
tier = "critical" if temp > 80 else "warning"
action_text = "aggressive-load-shedding" if tier == "critical" else "reduce-concurrency-50pct"
actions.append({"rule": "predictive-thermal", "gpu": name,
"temp_c": temp, "rise_rate_c_per_min": rise_rate,
"tier": tier, "action": action_text, "severity": tier})
print(f" {'⚠' if tier == 'critical' else '🔶'} RULE 8 ({tier}): {name} {temp}°C rising {rise_rate:.1f}°C/min")
# ── Rule 6: Strix Halo ──
if not strix.get("status") == "running":
if probe_strix_direct():
print(f" ⚠ RULE 6: Strix direct probe OK but monitor disagrees — router stale")
actions.append({"rule": "strix-router-stale", "gpu": "strix-halo",
"action": "refresh-router", "severity": "info"})
else:
print(f" ⚠ RULE 6: Strix Halo down — attempting restart")
actions.append({"rule": "strix-down", "gpu": "strix-halo",
"action": "restart-llama-server", "severity": "critical"})
# ── Rule 3: Model Stuck (check router) ──
router_models = router.get("available_models", [])
if isinstance(router_models, list):
for model in router_models:
if isinstance(model, dict) and model.get("consecutive_timeouts", 0) >= 3:
# Check failure rate
total = model.get("total_requests", 1)
failures = model.get("failed_requests", 0)
if total > 0 and failures / total > 0.5:
actions.append({"rule": "model-stuck", "gpu": model.get("gpu", "?"),
"model": model.get("name", "?"),
"failure_rate": failures/total,
"action": "restart-llama-server", "severity": "critical"})
print(f" ⚠ RULE 3: Model {model.get('name')} stuck ({failures}/{total} failures)")
# ── Rule 5: Circuit Breaker ──
cb_data = router.get("circuit_breaker", {})
if isinstance(cb_data, dict):
for cb_name, cb in cb_data.items():
if isinstance(cb, dict) and cb.get("open"):
open_sec = cb.get("open_duration_sec", 0)
gpu_healthy = cb.get("gpu_healthy", False)
if open_sec > 600 and gpu_healthy:
# Check cooldown: has this CB been reset in the last hour?
last_reset = history.get("cb_resets", {}).get(cb_name, 0)
if time.time() - last_reset > 3600:
actions.append({"rule": "cb-stuck-open", "gpu": cb.get("gpu", "?"),
"cb_name": cb_name, "open_sec": open_sec,
"action": "reset-cb", "severity": "warning"})
history.setdefault("cb_resets", {})[cb_name] = time.time()
print(f" ⚠ RULE 5: CB {cb_name} stuck open {open_sec}s — auto-resetting")
return actions
def evaluate_context_optimization(data):
"""
Rule 9: Context Window Optimization
Check if each GPU's context window is optimal for its role.
"""
benchmarks = data.get("benchmarks", {}).get("latest", {})
results = []
# Actual topology (verified July 2026)
gpu_config = {
"ct8-rtx3090": {
"context": 262144, "model": "gpu-dense", "quant": "Q4_K_M",
"parallel": 1, "role": "heavy-reasoning", "vram_mb": 24576,
"target_tps": 75.0
},
"ct110-rtx5070": {
"context": 131072, "model": "qwen3.5-9b-vision", "quant": "Q4_0 KV",
"parallel": 2, "role": "vision-web", "vram_mb": 12227,
"target_tps": 76.0
},
"strix-halo": {
"context": 262144, "model": "strix-moe", "quant": "Q4_K_M",
"parallel": 2, "role": "compression", "vram_mb": 65536,
"target_tps": 70.0
},
}
for gpu_key, cfg in gpu_config.items():
bench = benchmarks.get(gpu_key, {})
current_tps = bench.get("latest", {}).get("gen_tok_per_sec", 0)
baseline_tps = bench.get("baseline_tok_per_sec", 0)
if current_tps <= 0 or baseline_tps <= 0:
results.append({
"gpu": gpu_key, "model": cfg["model"], "role": cfg["role"],
"context": cfg["context"], "tok_per_sec": current_tps or 0,
"recommendation": f"{cfg['model']} idle — no benchmark data this cycle"
})
continue
perf_ratio = current_tps / baseline_tps
recommendation = None
if gpu_key == "ct8-rtx3090":
# Already at 256K — check if holding performance
if perf_ratio >= 0.95:
recommendation = f"256K ctx: {current_tps} tok/s ({(perf_ratio*100):.0f}% of baseline). Optimal for heavy reasoning."
else:
recommendation = f"256K ctx: {current_tps} tok/s ({(perf_ratio*100):.0f}% of baseline). Consider reducing parallel or ctx."
elif gpu_key == "ct110-rtx5070":
# 131K ctx, vision/web role — must preserve speed
if perf_ratio >= 0.95:
recommendation = f"131K ctx: {current_tps} tok/s (optimal). Vision/web role — context is fine."
else:
recommendation = f"131K ctx: {current_tps} tok/s ({(perf_ratio*100):.0f}% of baseline). Check VRAM pressure."
elif gpu_key == "strix-halo":
# 256K ctx, compression role — maximize context
if perf_ratio >= 0.95:
recommendation = f"256K ctx: {current_tps} tok/s ({(perf_ratio*100):.0f}% of baseline). Above target. Compression-optimized."
else:
recommendation = f"256K ctx: {current_tps} tok/s ({(perf_ratio*100):.0f}% of baseline). Check for competing workloads."
results.append({
"gpu": gpu_key, "model": cfg["model"], "role": cfg["role"],
"context": cfg["context"], "parallel": cfg["parallel"],
"tok_per_sec": current_tps, "baseline_tps": baseline_tps,
"perf_ratio": round(perf_ratio, 3),
"recommendation": recommendation
})
print(f" 📐 RULE 9: {gpu_key} ({cfg['role']}) — {recommendation}")
return results
def evaluate_workload_distribution(data):
"""
Rule 10: Workload Distribution Optimization
Check if GPU workload patterns match designated roles.
"""
role_map = {
"heavy-reasoning": ["gpu-dense", "syslog-auto"],
"vision-web": ["gpu-vision", "qwen3.5-9b-vision"],
"compression": ["strix-moe"],
}
gpu_roles = {
"NVIDIA GeForce RTX 3090": "heavy-reasoning",
"NVIDIA GeForce RTX 5070": "vision-web",
"AMD Strix Halo": "compression",
}
gpus = data.get("gpus", [])
summary = data.get("summary", {})
results = []
for gpu in gpus:
if not isinstance(gpu, dict):
continue
name = gpu.get("gpu_name", "")
role = "unknown"
for pattern, r in gpu_roles.items():
if pattern in name:
role = r
break
vram_used = gpu.get("vram_used_mb", 0)
vram_total = gpu.get("vram_total_mb", 0)
vram_pct = (vram_used / vram_total * 100) if vram_total else 0
status = "optimal"
note = ""
if role == "heavy-reasoning" and vram_pct > 90:
status = "warning"
note = f"VRAM at {vram_pct:.0f}% — consider reducing parallel or context"
elif role == "vision-web" and vram_pct > 88:
status = "warning"
note = f"VRAM at {vram_pct:.0f}% — vision/web role is VRAM-tight"
elif role == "compression" and vram_pct < 50:
note = f"VRAM at {vram_pct:.0f}% — plenty of headroom for compression workloads"
else:
note = f"VRAM at {vram_pct:.0f}% — role-appropriate"
results.append({
"gpu": name, "role": role, "vram_pct": round(vram_pct, 1),
"status": status, "note": note
})
icon = "✅" if status == "optimal" else "⚠"
print(f" {icon} RULE 10: {name} → {role} ({note})")
return results
def run_inference_test(model="syslog-auto"):
"""Run a quick inference test through LiteLLM."""
try:
body = json.dumps({
"model": model,
"messages": [{"role": "user", "content": "ping"}],
"max_tokens": 5
}).encode()
req = Request("http://192.168.68.116:4000/v1/chat/completions",
data=body, headers={
"Content-Type": "application/json",
"Authorization": "Bearer " + os.environ.get("LITELLM_API_KEY", "")
})
with urlopen(req, timeout=30) as resp:
return resp.status == 200
except:
return False
def main():
print(f"[{datetime.now().isoformat()}] GPU Self-Heal — {RUN_ID}")
print("=" * 60)
# Phase 1: Fetch data
data = fetch_gpu_data()
if not data:
print(" GPU monitor unreachable — aborting")
return 1
summary = data.get("summary", {})
print(f" Fleet: {summary.get('fleet_status','?')} | "
f"GPUs: {summary.get('gpu_count',0)} | "
f"Errors: {summary.get('gpu_errors',0)} | "
f"Alerts: {len(data.get('alerts',[]))}")
print()
# Phase 2: Evaluate rules (1-8)
history = load_history()
actions = evaluate_rules(data, history)
# Phase 2b: Context optimization (Rule 9)
ctx_results = evaluate_context_optimization(data)
# Phase 2c: Workload distribution (Rule 10)
wl_results = evaluate_workload_distribution(data)
save_history(history)
# Phase 3: Execute actions
issues_found = len(actions)
issues_fixed = 0
for action in actions:
severity = action.get("severity", "info")
if severity == "critical":
# Attempt fix
if "restart-llama-server" in action.get("action", ""):
gpu_name = action.get("gpu", "")
host = GPU_HOSTS.get(gpu_name.replace("NVIDIA ", "").replace("AMD ", "").lower(), {})
# Actual restart would need SSH — placeholder for now
print(f" ⟳ Would restart llama-server on {host.get('ip','?')} for {gpu_name}")
issues_fixed += 1
# Phase 4: Verify
all_models_ok = run_inference_test()
print(f"\n Verification: inference test {'✅' if all_models_ok else '❌'}")
# Phase 5: Compile report
status = "healthy"
if issues_found > 0:
status = "degraded" if issues_found <= 2 else "down"
report = json.dumps({
"run_id": RUN_ID,
"timestamp": datetime.now(timezone.utc).isoformat(),
"fleet_status": summary.get("fleet_status", "?"),
"overall_status": status,
"issues_found": issues_found,
"issues_fixed": issues_fixed,
"actions": actions,
"context_optimization": ctx_results,
"workload_distribution": wl_results
}, indent=2, default=str)
print(f"\n Status: {status} | Issues: {issues_found} | Fixed: {issues_fixed}")
# Phase 6: Log to Gitea (hard rule: never to knowledge graph)
kg_title = f"[GPU-SELF-HEAL] {RUN_ID} — {status}"
kg_desc = f"GPU self-heal run: {issues_found} issues found, {issues_fixed} fixed. Fleet: {summary.get('fleet_status','?')}."
kg_result = kg_create_node(kg_title, kg_desc, report)
print(f" Gitea: {kg_result[:120] if kg_result else 'write attempted'}")
print(f"\n[{datetime.now().isoformat()}] Complete — {RUN_ID}")
return 0 if issues_found == 0 else 1
if __name__ == "__main__":
sys.exit(main())
+360
View File
@@ -0,0 +1,360 @@
#!/bin/bash
LLK="${LITELLM_API_KEY:?LITELLM_API_KEY must be set}"
# LiteLLM Health Check + Self-Heal — Automated (every 6 hours)
# Deployed from litellm-self-heal.prose.md contract
# Writes results to RA-H OS knowledge graph via MCP bridge
set -uo pipefail # -e removed: grep -c returns 1 on no-match, which is valid
RUN_ID="litellm-health-$(date +%Y%m%d-%H%M%S)"
TIMESTAMP=$(date -Iseconds)
BRIDGE="http://192.168.68.65:3100/mcp"
MASTER_KEY="sk-litellm-7f96080dd99b15c36bd4b333b58a6796" # admin only: /key/list. NEVER for inference
LITELLM_HOST="192.168.68.116"
source /etc/litellm-monitor.env 2>/dev/null # dedicated monitor agent key for inference tests
MONITOR_KEY="${LITELLM_MONITOR_KEY:-$MASTER_KEY}" # fallback only if env missing
GPU_DASHBOARD="http://192.168.68.24:9100"
LOG_DIR="/var/log/litellm"
RESULTS=""
ISSUES=0
FIXED=0
ESCALATED=0
DETAILS="[]"
mcp_call() {
local method="$1" tool="$2" args="$3"
curl -s -X POST "$BRIDGE" \
-H "Content-Type: application/json" \
-H "Accept: application/json, text/event-stream" \
-d "{\"jsonrpc\":\"2.0\",\"id\":1,\"method\":\"$method\",\"params\":$args}" 2>/dev/null
}
add_detail() {
local name="$1" status="$2" msg="$3"
DETAILS=$(echo "$DETAILS" | python3 -c "
import sys, json
d = json.load(sys.stdin)
d.append({'check':'$name','status':'$status','detail':'$msg'})
print(json.dumps(d))
" 2>/dev/null || echo "$DETAILS")
}
# ═══════════════════════════════════════════════════════
# PHASE 1: HEALTH CHECKS
# ═══════════════════════════════════════════════════════
echo "[$(date)] Starting LiteLLM health check — $RUN_ID"
# 1. Public endpoints
echo -n " UI... "
if curl -s -o /dev/null -w "%{http_code}" --connect-timeout 10 https://litellm.sysloggh.net/ui/ 2>/dev/null | grep -q 200; then
add_detail "public-ui" "pass" "200 OK"
echo "pass"
else
add_detail "public-ui" "fail" "non-200"
echo "FAIL"; ISSUES=$((ISSUES+1))
fi
echo -n " Docs... "
if curl -s -o /dev/null -w "%{http_code}" --connect-timeout 10 https://litellm.sysloggh.net/docs 2>/dev/null | grep -q 200; then
add_detail "public-docs" "pass" "200 OK"
echo "pass"
else
add_detail "public-docs" "fail" "non-200"
echo "FAIL"; ISSUES=$((ISSUES+1))
fi
# 2. LiteLLM liveliness
echo -n " Liveliness... "
LIVE=$(curl -s --connect-timeout 5 "http://${LITELLM_HOST}:4000/health/liveliness" 2>/dev/null)
if echo "$LIVE" | grep -qi "alive"; then
add_detail "liveliness" "pass" "$LIVE"
echo "pass"
else
add_detail "liveliness" "fail" "$LIVE"
echo "FAIL"; ISSUES=$((ISSUES+1))
fi
# 3. Container health
echo -n " Containers... "
CONTAINERS=$(docker ps --format '{{.Names}}:{{.Status}}' 2>/dev/null)
CT_COUNT=$(echo "$CONTAINERS" | wc -l)
# Only flag containers that are explicitly 'unhealthy', not ones without healthchecks
UNHEALTHY=$(echo "$CONTAINERS" | grep -c '(unhealthy)' 2>/dev/null || true)
UNHEALTHY=${UNHEALTHY:-0}
if [ "$UNHEALTHY" -eq 0 ] && [ "$CT_COUNT" -ge 8 ]; then
add_detail "containers" "pass" "$CT_COUNT containers healthy"
echo "pass ($CT_COUNT up)"
else
add_detail "containers" "fail" "$UNHEALTHY unhealthy of $CT_COUNT"
echo "FAIL ($UNHEALTHY/$CT_COUNT unhealthy)"; ISSUES=$((ISSUES+1))
# Auto-restart unhealthy containers
for ct in $(echo "$CONTAINERS" | grep '(unhealthy)' | cut -d: -f1); do
echo " Restarting $ct..."
docker restart "$ct" 2>/dev/null && FIXED=$((FIXED+1))
done
fi
# 4. GPU Fleet (full telemetry from gpu-monitor on .24:9100)
echo -n " GPU Fleet... "
GPU_DATA=$(curl -s --connect-timeout 10 "$GPU_DASHBOARD/gpu-data" 2>/dev/null)
GPU_HEALTH=$(echo "$GPU_DATA" | python3 -c "
import sys, json
d = json.load(sys.stdin)
s = d.get('summary',{})
alerts = d.get('alerts',[])
gpus = d.get('gpus',[])
strix = d.get('strix',{})
fleet = s.get('fleet_status','unknown')
gpu_count = s.get('gpu_count',0)
errors = s.get('gpu_errors',0)
cb_open = s.get('circuit_breakers_open',0)
strix_ok = strix.get('status','') == 'running'
alert_count = len(alerts)
critical_alerts = len([a for a in alerts if isinstance(a, dict) and a.get('level')=='critical'])
# Per-GPU details
for g in gpus:
if isinstance(g, dict):
n = g.get('gpu_name','?')[:30]
t = g.get('temp_c','?')
u = g.get('gpu_util_pct','?')
v = f\"{g.get('vram_used_mb',0)}/{g.get('vram_total_mb',0)}MB\"
print(f' {n}: {t}°C util={u}% vram={v}')
# Alert details
for a in alerts:
if isinstance(a, dict):
print(f' ⚠ {a.get(\"level\",\"?\")}: {a.get(\"metric\",\"?\")} — {a.get(\"value\",\"?\")}')
# Summary line
status = 'healthy' if fleet == 'healthy' and critical_alerts == 0 and errors == 0 else 'degraded'
print(f'SUMMARY: {status} | {gpu_count} GPUs | {errors} errors | {alert_count} alerts | CB open={cb_open} | Strix={\"running\" if strix_ok else \"down\"}')
" 2>/dev/null)
GPU_STATUS=$(echo "$GPU_HEALTH" | grep 'SUMMARY:' | cut -d' ' -f2-)
if echo "$GPU_STATUS" | grep -q '^healthy'; then
add_detail "gpu-fleet" "pass" "$GPU_STATUS"
echo "pass"
echo "$GPU_HEALTH" | grep -v 'SUMMARY:'
else
add_detail "gpu-fleet" "fail" "$GPU_STATUS"
echo "FAIL"
echo "$GPU_HEALTH"
ISSUES=$((ISSUES+1))
fi
# 5. Model inference tests
MODELS="gpu-dense strix-moe gpu-vision syslog-auto"
MODEL_FAILS=0
for model in $MODELS; do
echo -n " Model $model... "
HTTP_CODE=$(curl -s -o /dev/null -w "%{http_code}" --connect-timeout 30 \
-H "Authorization: Bearer $LLK" \
-H "Content-Type: application/json" \
-d "{\"model\":\"$model\",\"messages\":[{\"role\":\"user\",\"content\":\"ping\"}],\"max_tokens\":5}" \
"http://${LITELLM_HOST}:4000/v1/chat/completions" 2>/dev/null)
if [ "$HTTP_CODE" = "200" ]; then
add_detail "model-$model" "pass" "200 OK"
echo "pass"
else
add_detail "model-$model" "fail" "HTTP $HTTP_CODE"
echo "FAIL ($HTTP_CODE)"; ISSUES=$((ISSUES+1)); MODEL_FAILS=$((MODEL_FAILS+1))
fi
done
# 6. Agent keys
echo -n " Agent Keys... "
KEYS=$(curl -s --connect-timeout 10 \
-H "Authorization: Bearer $LLK" \
"http://${LITELLM_HOST}:4000/key/list?return_full_object=true" 2>/dev/null)
AGENT_COUNT=$(echo "$KEYS" | python3 -c "
import sys,json
d = json.load(sys.stdin)
agents = {'mumuni','tanko','kagenz0','koby','koonimo','abiba-pi','baggy'}
keys = d.get('keys',[])
found = sum(1 for k in keys if k.get('key_alias') in agents)
print(found)
" 2>/dev/null || echo 0)
if [ "$AGENT_COUNT" -ge 6 ]; then
add_detail "agent-keys" "pass" "$AGENT_COUNT agent keys present"
echo "pass ($AGENT_COUNT keys)"
else
add_detail "agent-keys" "fail" "only $AGENT_COUNT/7 agent keys found"
echo "FAIL"; ISSUES=$((ISSUES+1))
fi
# 7. Grafana
echo -n " Grafana... "
if curl -s -o /dev/null -w "%{http_code}" --connect-timeout 10 \
"http://${LITELLM_HOST}:3001/api/health" 2>/dev/null | grep -q 200; then
add_detail "grafana" "pass" "200 OK"
echo "pass"
else
add_detail "grafana" "warn" "Grafana unreachable (non-critical)"
echo "warn"
fi
# 8. 401 error count
echo -n " 401 Errors... "
ERR_401_RAW=$(docker logs harness-litellm --since 6h 2>&1 | grep 'Received=' 2>/dev/null || true)
ERR_401=$(echo "$ERR_401_RAW" | grep -c 'Received=' 2>/dev/null); ERR_401=${ERR_401:-0}
if [ "$ERR_401" -eq 0 ]; then
add_detail "401-errors" "pass" "0 auth errors in last 6h"
echo "pass (0)"
else
# Extract key patterns and source IPs from 401 errors
ERR_KEYS=$(echo "$ERR_401_RAW" | grep -oP 'Received=\K[^,]+' | sort -u | tr '\n' ' ')
ERR_IPS=$(docker logs harness-litellm --since 6h 2>&1 | grep -B2 'Received=' | grep -oP '\d+\.\d+\.\d+\.\d+' | sort -u | tr '\n' ' ')
DETAIL="$ERR_401 auth errors in last 6h | Keys: ${ERR_KEYS:-unknown} | Sources: ${ERR_IPS:-unknown}"
add_detail "401-errors" "warn" "$DETAIL"
echo "warn ($ERR_401 — keys: ${ERR_KEYS:-?}, sources: ${ERR_IPS:-?})"
ISSUES=$((ISSUES+1))
# Self-heal: analyze and attempt fix based on source IP type
if echo "$ERR_KEYS" | grep -q 'no-key-required'; then
echo " → 'no-key-required' = GPU-direct key being used as LiteLLM client key."
fi
for src_ip in $ERR_IPS; do
# Case 1: Docker network IPs (172.x) — this is the nginx proxy, meaning an external client
if echo "$src_ip" | grep -q '^172\.'; then
CT_NAME=$(docker inspect -f '{{.Name}}' $(docker ps -q) 2>/dev/null | while read n; do
ip=$(docker inspect -f '{{range .NetworkSettings.Networks}}{{.IPAddress}}{{end}}' $n 2>/dev/null)
[ "$ip" = "$src_ip" ] && echo "$n"
done)
echo " → Source $src_ip is Docker container: ${CT_NAME:-unknown}"
if echo "$CT_NAME" | grep -q 'nginx'; then
echo " → 401 came through nginx proxy — external client, not locally fixable."
echo " → Checking nginx logs via docker logs for actual source..."
NGINX_SRC=$(timeout 8 docker logs harness-nginx --tail 2000 2>&1 | grep ' 401 ' | grep -oP '^\S+' | sort -u | tr '\n' ' ')
if [ -n "$NGINX_SRC" ]; then
echo " → Nginx upstream source(s): $NGINX_SRC"
fi
ESCALATED=$((ESCALATED+1))
else
echo " → Checking Docker container logs for clue..."
docker logs "$CT_NAME" --tail 50 2>/dev/null | grep -i 'litellm\|api.key\|auth' | tail -5
fi
# Case 2: Localhost or local CT116 IP — check locally
elif [ "$src_ip" = "127.0.0.1" ] || [ "$src_ip" = "192.168.68.116" ]; then
echo " → Source $src_ip is local (CT116). Checking local configs..."
LOCAL_CFG=$(grep -rl 'no-key-required' /root/.hermes/ /opt/inference-harness/ --include='*.yaml' --include='*.yml' --include='*.py' 2>/dev/null | grep -v 'litellm-health-check\|litellm_config\|\.bak' | head -5)
if [ -n "$LOCAL_CFG" ]; then
echo " → Found local references: $LOCAL_CFG"
else
echo " → No local config using no-key-required. Likely external via proxy."
ESCALATED=$((ESCALATED+1))
fi
# Case 3: Known agent host IPs — SSH and check
else
echo " → Checking remote source $src_ip for misconfigured agent..."
AGENT_CONFIGS=$(ssh -o ConnectTimeout=5 -o StrictHostKeyChecking=no root@$src_ip \
"grep -rl 'no-key-required\|api_key.*not-needed' /root/.hermes/config.yaml /home/*/.hermes/config.yaml 2>/dev/null" 2>/dev/null || true)
if [ -n "$AGENT_CONFIGS" ]; then
echo " → FOUND: agent config with GPU-direct key at $src_ip"
for cfg in $AGENT_CONFIGS; do
echo " → Attempting self-heal on $cfg..."
ssh -o ConnectTimeout=5 root@$src_ip \
"python3 -c \"
import yaml
with open('$cfg') as f: c = yaml.safe_load(f)
fixed = False
for section in ['agent', 'auxiliary']:
for sub in c.get(section, {}):
if isinstance(c[section].get(sub), dict):
ak = c[section][sub].get('api_key', '')
if ak in ['no-key-required', 'not-needed', 'no-k']:
c[section][sub]['api_key_env'] = 'LITELLM_API_KEY'
del c[section][sub]['api_key']
fixed = True
if fixed:
with open('$cfg', 'w') as f: yaml.dump(c, f, default_flow_style=False, allow_unicode=True)
print('Fixed. Restart gateway to apply.')
else:
print('No fixable sections.')
\"" 2>/dev/null && FIXED=$((FIXED+1)) || true
done
else
echo " → No misconfigured agent on $src_ip."
ESCALATED=$((ESCALATED+1))
fi
fi
done
fi
# ═══════════════════════════════════════════════════════
# PHASE 2: COMPILE STATUS
# ═══════════════════════════════════════════════════════
if [ "$ISSUES" -eq 0 ]; then
OVERALL="healthy"
elif [ "$ISSUES" -le 2 ]; then
OVERALL="degraded"
else
OVERALL="down"
fi
SUMMARY="LiteLLM Health: $OVERALL | Checks: $(echo "$DETAILS" | python3 -c "import sys,json;print(len(json.load(sys.stdin)))" 2>/dev/null || echo "?") | Issues: $ISSUES | Fixed: $FIXED | Escalated: $ESCALATED"
echo ""
echo "════════════════════════════════════════"
echo " Status: $OVERALL"
echo " Issues: $ISSUES found | $FIXED auto-fixed | $ESCALATED escalated"
echo "════════════════════════════════════════"
# ═══════════════════════════════════════════════════════
# PHASE 3: LOG TO FILESYSTEM
# ═══════════════════════════════════════════════════════
mkdir -p "$LOG_DIR"
REPORT=$(python3 -c "
import json
print(json.dumps({
'run_id': '$RUN_ID',
'timestamp': '$TIMESTAMP',
'overall_status': '$OVERALL',
'issues_found': $ISSUES,
'issues_fixed': $FIXED,
'issues_escalated': $ESCALATED,
'checks': $DETAILS
}, indent=2))
" 2>/dev/null)
echo "$REPORT" > "$LOG_DIR/${RUN_ID}.json"
# 9. GPU Monitor self-check
echo -n " GPU Monitor... "
if curl -s -o /dev/null -w "%{http_code}" --connect-timeout 5 "$GPU_DASHBOARD/health" 2>/dev/null | grep -q 200; then
add_detail "gpu-monitor" "pass" "Monitor server health OK"
echo "pass"
else
add_detail "gpu-monitor" "warn" "GPU monitor health endpoint unreachable"
echo "warn"
fi
echo ""
echo " Report saved: $LOG_DIR/${RUN_ID}.json"
# ═══════════════════════════════════════════════════════
# PHASE 4: FEED TO KNOWLEDGE GRAPH
# PHASE 4: LOG TO GITEA (hard rule: health logs NEVER go to knowledge graph)
# Logged to SyslogSolution/health-logs/litellm/{run_id}.json — versioned, searchable.
echo -n " Gitea... "
/opt/inference-harness/scripts/gitea-logger.sh litellm "${RUN_ID}.json" "${LOG_DIR}/${RUN_ID}.json"
# PHASE 5: NOTIFY ON ISSUES
# ═══════════════════════════════════════════════════════
if [ "$ISSUES" -gt 0 ]; then
echo ""
echo "⚠️ $ISSUES issue(s) detected. Creating relay alert..."
ALERT_TITLE="⚠ LiteLLM Health Alert — $OVERALL ($ISSUES issues)"
ALERT_BODY="$SUMMARY
Details:
$DETAILS"
mcp_call "tools/call" "createRelayNode" "{\"name\":\"createRelayNode\",\"arguments\":{\"title\":\"$ALERT_TITLE\",\"source\":\"$ALERT_BODY\",\"description\":\"LiteLLM health check failure alert\"}}" > /dev/null 2>&1
fi
echo ""
echo "[$(date)] Health check complete — $RUN_ID"
exit 0