Testing the consolidated litellm-self-heal contract against live infra found:
- IP drift: pi host is .24, not .65 (confused with ra-h-os bridge)
- Fixed gpu-monitor.prose.md: 192.168.68.65 → 192.168.68.24 (3 locations)
- Fixed gpu-fleet.prose.md: .65 → .24 in config files + agent keys tables
- Container count: 10 running on CT 116, not 8 as contract stated
- Added harness-docker-stats + harness-pve-exporter to litellm-self-heal
- Updated container health check list
Verified: GPU monitor restarted on .24:9100, ornith model inference working,
all 10 containers healthy.
Manages the GPU inference fleet across all hosts. Handles model deployment, registration, health checks, LiteLLM sync, agent key management, GPU saturation watchdog, Prometheus/Grafana monitoring, and self-healing. Current as of 2026-07-08: context reduced to 128K on NVIDIA GPUs, parallel 2 on all GPUs, LiteLLM timeouts tuned (gemma 25→120s, qwen 40→90s), router fully deprecated — nginx routes /v1 → LiteLLM directly.
abiba
on model add/remove
on GPU health degradation
on agent key rotation
on router restart (roster must be loaded)
Maintains
gpu_roster: { models: map, hosts: map } — Single source of truth for all GPU models
llama-server daemon (Vulkan, Strix Halo). Note: llama-server.service and llama-server@.service are masked on .15 to prevent port 8080 collisions.
Prometheus & Grafana
Component
URL
Details
Grafana
http://192.168.68.116:3001/
admin / syslog-grafana-2026
GPU Dashboard
http://192.168.68.116:3001/d/gpu-fleet
Gauges + time series
Prometheus
http://192.168.68.116:9090/ (internal)
5 scrape targets
GPU Exporters
:9400/metrics on .8, .110, .15
NVIDIA/AMD GPU metrics
Router Exporter
:9401/metrics on .24
Router + LiteLLM metrics
Known Issues & Watch Points
Router startup race: Compose router.py doesn't call load_roster(). Reload thread sleeps 30s first.
Fix: trigger roster reload via SSH after restart, or rebuild image with startup load_roster().
LiteLLM /metrics: Requires auth. Prometheus uses /health/liveliness as workaround.
VRAM (2026-07-08): RTX 3090 at 20.3/24GB (83%), RTX 5070 at 9.4/12.2GB (77%), Strix Halo at 24.4/64GB (35%). Context reduced from 256K→128K on NVIDIA GPUs freed ~3.3GB (.8) and ~1.5GB (.110).
All GPUs at --parallel 2 (2026-07-08): Fleet serves 6 concurrent requests (was 3). 2× throughput.
Strix Halo GPU: Vulkan is the working backend (ROCm/HIP path abandoned — HSA runtime blocked on Debian 13). Build at /root/llama.cpp/build-vk/, commit 4fc4ec5 (2026-07-01), ggml 0.15.3 shared-lib arch. Mesa RADV 25.0.7, KHR_coopmat fast path active. ~70 tok/s gen, 532 tok/s prompt. Service: ornith-server.service on port 8080, 256K context, flash-attn + q8 KV.
Port conflict detection (2026-07-05): All 3 GPU wrappers now detect ghost processes squatting port 8080 before starting. .8 and .110 use inline pre-start check in llama-wrapper.sh; .15 uses /usr/local/bin/port-cleanup.sh ExecStartPre. Replaces the blanket pkill -9 -x llama-server on .15 which would kill ALL llama-server instances regardless of port. Ghost detection was the root cause of .8 crash-looping for 27+ restarts (stale pid 25836 squatting 8080 after OOM kill).
Strix Halo thermal safeguard (2026-07-02): ornith-server.service has -n 8192 (hard generation cap per request). Without it, --predict defaults to -1 (infinity) — a runaway request from .123 (Mumuni) decoded 39,868 tokens over 24 min, pushing Tctl to 98°C (crit 89.8°C) and throttling 70→29 t/s. The cap bounds worst-case generation to ~5 min. Do NOT remove -n without a replacement ceiling. Sustained load hits ~84°C even at 92s; the APU is fanless/low-flow. Clients MUST also set max_tokens.
Port 8080 firewall: amdpve iptables restricts 8080 to 192.168.68.116 (LiteLLM/router host) only. All inbound connections are from .116 (LiteLLM proxied via nginx). Localhost curls hang (SYN dropped). Always test from .116.
Router sidecar fallback: router.pycheck_gpu_health() now probes GPU /health directly when sidecar at :8090 is absent. Sidecar JSON exporters not deployed on any GPU host — router relies on GPU-direct fallback.
Router GPU_MOE_URL bug (fixed 2026-07-01): docker-compose had GPU_MOE_URL=.110:8080 (gemma host) instead of .15:8080 (amdpve). Corrected.
Alert migration: All alerts now go to #agent-hub topics (alerts-gpu, alerts-pm2, alerts-infra) instead of DMs. Cross-agent visibility enabled.
tok/s benchmarks: Measured every 5 min via LiteLLM proxy. Baselines tracked with 30%/50% degradation thresholds.
NetBird 502: Tanko routes through NetBird for litellm.sysloggh.net. Use direct IP if NetBird down.
GPU Inference Benchmarks (Current)
GPU
Model
Gen tok/s
Prompt tok/s
Baseline
Samples
RTX 3090 (.8)
qwen3.6-27B-code
75
305
74
6
RTX 5070 (.110)
gemma-4-12b
75
323
75
6
Strix Halo (.15)
ornith-1.0-35b
70
532
70
6
Benchmarks run through LiteLLM proxy (192.168.68.116:4001) every 5 minutes.
Degradation alerts fire at 30% (warning) and 50% (critical) below baseline.
History stored at /root/data/toks-history.json with 7-day rolling window.
Note (2026-07-01): Strix Halo prompt tok/s jumped 209→532 after Vulkan rebuild (cooperative-matrix fast path now active on GFX1151). Baseline may need re-calibration.