- gpu-fleet: documented parallel=2/batch=4096 optimizations for RTX 3090/5070 - gpu-fleet: added Strix Halo ROCm status (installed, blocked by HSA runtime on Debian 13) - gpu-fleet: documented alert migration from DM → #agent-hub stream topics - gpu-monitor: added Alert Delivery section with stream topic conventions - All GPU alerts now go to #agent-hub for cross-agent visibility
11 KiB
11 KiB
kind, name, description, agent, triggers
| kind | name | description | agent | triggers | ||||
|---|---|---|---|---|---|---|---|---|
| responsibility | gpu-fleet | Manages the GPU inference fleet across all hosts. Handles model deployment, registration, health checks, LiteLLM sync, agent key management, GPU saturation watchdog, Prometheus/Grafana monitoring, and self-healing. Current as of 2026-06-30 post-migration: nginx routes /v1 → LiteLLM, router runs behind LiteLLM on :9000 (internal-only). | abiba |
|
Maintains
- gpu_roster: { models: map, hosts: map } — Single source of truth for all GPU models
- router: { status: "healthy", roster_loaded: bool, models: array }
- litellm: { status: "healthy", keys: array, models: array }
- agent_keys: { agent: api_key } — All agent API keys registered in LiteLLM DB
- health: { gpus: array, circuit_breakers: array } — Fleet-wide health state
- monitor: { status: "running", version: "2.0.0" } — GPU monitor server on pi (:9100)
- watchdog: { status: "running" } — GPU saturation watchdog (restarts stuck llama-server)
- benchmarks: { tok_per_sec: map, baseline: map, history: array } — Inference speed benchmarks tracked over time
- grafana: { status: "running", dashboards: ["gpu-fleet"] } — Grafana on CT 116 (:3001)
- prometheus: { status: "running", targets: 5 } — Scrapes GPU :9400 exporters + LiteLLM
Fleet Topology (Current — June 2026)
┌──────────────────────────────────────────────────────────────────┐
│ CT 116 (192.168.68.116) — Inference Harness Host │
│ │
│ nginx:80 (entrypoint) │
│ ├─ /v1/* → harness-litellm:4000 (API requests) │
│ ├─ /admin/* → harness-litellm:4000 (admin endpoints) │
│ ├─ /dashboard/ → harness-dashboard:3000 (harness UI) │
│ ├─ /litellm/* → harness-litellm:4000 (LiteLLM UI + API) │
│ ├─ /grafana/* → harness-grafana:3000 (dashboards) │
│ ├─ /health/* → harness-litellm:4000 (health probes) │
│ └─ /gpu/* → 192.168.68.24:9100 (fleet monitor) │
│ │
│ Containers: │
│ ┌──────────┐ ┌──────────┐ ┌──────────┐ ┌──────────┐ │
│ │ LiteLLM │ │ Router │ │Dashboard │ │ Grafana │ │
│ │ :4000 │─▶│ :9000 │ │ :3000 │ │ :3000 │ │
│ │ keys+sync│ │internal │ │ harness │ │ Prometheus│ │
│ │ fallback │ │only! │ │ UI │ │ data src │ │
│ └──────────┘ └───┬──────┘ └──────────┘ └──────────┘ │
│ │ │
│ ┌──────────┐ ┌──────────┐ ┌──────────┐ │
│ │PostgreSQL│ │ Redis │ │Prometheus│ │
│ │ :5432 │ │ :6379 │ │ :9090 │ │
│ └──────────┘ └──────────┘ └──────────┘ │
└────────────────────┼─────────────────────────────────────────────┘
│
┌───────────────┼───────────────┬──────────────────┐
│ │ │ │
┌────▼─────┐ ┌──────▼──────┐ ┌────▼──────┐ ┌───────▼──────┐
│ CT 8 │ │ CT 110 │ │ CT 15 │ │ pi (.24) │
│ RTX 3090 │ │ RTX 5070 │ │ Strix Halo│ │ GPU Monitor │
│ 24GB │ │ 12GB │ │ 64GB UMA │ │ :9100 │
│ │ │ │ │ llama-srv │ │ Watchdog │
│ qwen3.6 │ │ gemma-4-12b │ │ ornith35B │ │ Prometheus │
│ 27B-code │ │ :8090 │ │ :8080 │ │ exporter │
│ :8090 │ │ :9400 (exp) │ │ :9400(exp)│ │ :9401 │
│ :9400 │ └─────────────┘ └───────────┘ └──────────────┘
└──────────┘
Current Model Assignments
| Model | GPU | Host | VRAM | Status |
|---|---|---|---|---|
| qwen3.6-27B-code | RTX 3090 | .8 (llm-gpu) | 23.4/24GB | ✅ healthy |
| gemma-4-12b | RTX 5070 | .110 (ocu-llm) | 11.2/12.2GB | ✅ healthy |
| ornith-1.0-35b | Strix Halo CPU | .15 (amdpve) | 28.2/64GB | ✅ healthy |
Operations
add-model
- Download model file from Hugging Face or source
- Check disk space + GPU VRAM compatibility
- Start llama-server via systemd service on GPU host
- Add to
/opt/inference-harness/gpu_roster.yaml(or router env vars) - Hot-reload or restart router
- Add model to LiteLLM config (
model_list+fallbacks) - Generate agent keys for new model access via
/key/generate - Add Prometheus scrape target for the new GPU exporter
- Verify end-to-end: LiteLLM → Router → Model
remove-model
- Drain active requests (wait for active=0)
- Remove from LiteLLM config
- Remove from router config
- Stop llama-server (systemd)
- Remove Prometheus scrape target
- Cleanup model files (optional)
heal
- Check all GPUs via router internal
:9000/health/unified - Check LiteLLM health via nginx
:80/litellm/health/liveliness - Reset stuck circuit breakers if idle (Redis)
- Restart dead llama-server instances via SSH
- Flush Redis active counters if stale
- Verify GPU monitor server is running on pi (:9100)
- Verify watchdog is running on pi
- Restart router if roster not loaded (check logs for STARTUP ROSTER)
- Reload roster via
POST :9000/admin/roster/reloadif available
sync-keys
- List all agent keys in LiteLLM DB via
GET /key/list - Compare against expected agent list: [tanko, mumuni, abiba, koby, koonimo, kagenz0]
- Generate missing keys via
POST /key/generatewith unlimited budget - Update agent configs —
/etc/environmentLITELLM_API_KEY - Send Zulip DM to agents that can't be reached via SSH
- Verify each key with test request through full chain
- Document keys in knowledge graph
list
Show full fleet status: GPUs, models, VRAM, active requests, circuit breakers, keys
Agent Keys (LiteLLM DB — Current 2026-06-30)
| Agent | Key | Host | Access |
|---|---|---|---|
| Tanko | sk-3ZsdWJbbNSo9zSnJN2OsJw |
.122 | SSH jerome |
| Mumuni | sk-XY2aUfvy2BIs6kp1ZPh6VA |
.123 | SSH root |
| Abiba | sk-kxbPgbnV2Zkn6TKQbAEcEg |
.24 | local |
| Koby | sk-wb86PvcZXWkjShAIBxrpDg |
? | Zulip DM |
| Koonimo | sk-Prx_vRXYb2VEzDmhPCbNeg |
.114 (baggy) | no SSH |
| Kagenz0 | sk-Dh4CDkaHebMLEp8qqq20qA |
? | no SSH |
Key update procedure: Update /etc/environment → LITELLM_API_KEY=sk-... → restart Hermes.
If no SSH access, send Zulip DM via abiba-bot.
Configuration Files
| File | Host | Purpose |
|---|---|---|
/opt/inference-harness/docker-compose.yml |
CT 116 | All containers (router, litellm, nginx, postgres, redis, dashboard) |
/opt/inference-harness/litellm_config.yaml |
CT 116 | LiteLLM proxy config (models, fallbacks, timeouts) |
/opt/inference-harness/router/router.py |
CT 116 | Router source (builds via compose) |
/etc/nginx/nginx.conf |
CT 116 (nginx container) | Routes /v1→LiteLLM, /dashboard/, /litellm/, /grafana/, /health |
/opt/monitoring/prometheus.yml |
CT 116 | Prometheus scrape config (5 targets) |
/root/scripts/gpu-monitor-server.py |
pi (.24) | GPU fleet monitor v2.1.0 (with benchmarks) |
/root/scripts/gpu_benchmark.py |
pi (.24) | GPU inference benchmark module (tok/s tracking) |
/root/scripts/gpu-saturation-watchdog.py |
pi (.24) | Auto-restart stuck llama-server |
/root/dashboard/gpu-fleet.html |
pi (.24) | Live HTML dashboard |
/etc/systemd/system/llama-server.service |
.8, .110, .15 | llama-server daemons |
Prometheus & Grafana
| Component | URL | Details |
|---|---|---|
| Grafana | http://192.168.68.116:3001/ |
admin / syslog-grafana-2026 |
| GPU Dashboard | http://192.168.68.116:3001/d/gpu-fleet |
Gauges + time series |
| Prometheus | http://192.168.68.116:9090/ (internal) |
5 scrape targets |
| GPU Exporters | :9400/metrics on .8, .110, .15 |
NVIDIA/AMD GPU metrics |
| Router Exporter | :9401/metrics on .24 |
Router + LiteLLM metrics |
Known Issues & Watch Points
- Router startup race: Compose router.py doesn't call load_roster(). Reload thread sleeps 30s first. Fix: trigger roster reload via SSH after restart, or rebuild image with startup load_roster().
- LiteLLM /metrics: Requires auth. Prometheus uses
/health/livelinessas workaround. - VRAM: RTX 3090 at ~9.6/24GB (60% free), RTX 5070 at ~6.5/12GB (46% free). Healthy.
- RTX 3090 optimization:
--parallel 2 --batch-size 4096— doubled concurrent throughput. - RTX 5070 optimization: Removed draft model freed ~2GB VRAM for vision tasks.
--parallel 2. - Strix Halo GPU: ROCm 6.3.1 installed, HIP compiler works via clang++ + device libs. Blocked: HSA runtime driver needs
rocm-llvm→libstdc++-11-devnot on Debian 13. Fixes: Docker with ROCm, or reinstall to Ubuntu 24.04. - Alert migration: All alerts now go to
#agent-hubtopics (alerts-gpu,alerts-pm2,alerts-infra) instead of DMs. Cross-agent visibility enabled. - tok/s benchmarks: Measured every 5 min via LiteLLM proxy. Baselines tracked with 30%/50% degradation thresholds.
- NetBird 502: Tanko routes through NetBird for litellm.sysloggh.net. Use direct IP if NetBird down.
GPU Inference Benchmarks (Current)
| GPU | Model | Gen tok/s | Prompt tok/s | Baseline | Samples |
|---|---|---|---|---|---|
| RTX 3090 (.8) | gemma-4-12b | 75 | 305 | 74 | 6 |
| RTX 5070 (.110) | qwen3.6-27B-code | 75 | 323 | 75 | 6 |
| Strix Halo (.15) | ornith-1.0-35b | 70 | 209 | 70 | 6 |
Benchmarks run through LiteLLM proxy (192.168.68.116:4001) every 5 minutes.
Degradation alerts fire at 30% (warning) and 50% (critical) below baseline.
History stored at /root/data/toks-history.json with 7-day rolling window.