--- kind: responsibility name: gpu-fleet description: > Manages the GPU inference fleet across all hosts. Handles model deployment, registration, health checks, LiteLLM sync, agent key management, GPU saturation watchdog, Prometheus/Grafana monitoring, and self-healing. UPDATED 2026-07-15: Stable role-based aliases introduced: strix-moe, gpu-dense, gpu-vision (gpu-light was superseded by gpu-vision on 2026-09-12). These never change — only the underlying model does. Strix Halo: strix-moe → Carnice-Qwen3.6-MoE-35B-A3B-Q4_K_M.gguf (22GB, 256K ctx). RTX 5070: gpu-vision — IQ4_NL + MTP draft (~122 tok/s, 2x faster). UPDATED 2026-07-17: NVIDIA host context reduced from 256K to 128K for stability. Strix Halo model: Carnice-Qwen3.6-MoE-35B-A3B-Q4_K_M.gguf (alias strix-moe, 256K context). Strix Halo runs 256K (n_ctx 262144, --kv-unified); RTX 3090 and RTX 5070 remain at 128K. For >128K on NVIDIA hosts → fall back to external providers (deepseek). VRAM headroom improved: RTX 3090 ~70%, RTX 5070 ~65%. agent: abiba triggers: - on model add/remove - on GPU health degradation - on agent key rotation - on harness container restart (LiteLLM reloads its model list) --- ## Maintains - gpu_roster: { models: map, hosts: map } — GPU host/model roster; the authoritative alias/weight/fallback registry is CT 116 `litellm_config.yaml` - router: { status: "healthy", roster_loaded: bool, models: array } - litellm: { status: "healthy", keys: array, models: array } - agent_keys: { agent: api_key } — All agent API keys registered in LiteLLM DB - health: { gpus: array, circuit_breakers: array } — Fleet-wide health state - monitor: { status: "running", version: "2.0.0" } — GPU monitor server on pi (:9100) - watchdog: { status: "running" } — GPU saturation watchdog (restarts stuck llama-server) - benchmarks: { tok_per_sec: map, baseline: map, history: array } — Inference speed benchmarks tracked over time - grafana: { status: "running", dashboards: ["gpu-fleet"] } — Grafana on CT 116 (:3001) - prometheus: { status: "running", targets: 5 } — Scrapes GPU :9400 exporters + LiteLLM - port_conflict_detection: { status: "active" } — All 3 GPU wrappers detect ghost processes before binding ## Fleet Topology (Current — 2026-09-11, router decommissioned) ``` ┌──────────────────────────────────────────────────────────────────┐ │ CT 116 (192.168.68.116) — Inference Harness Host │ │ │ │ nginx:80 (entrypoint) │ │ ├─ /v1/* → harness-litellm:4000 (API requests) │ │ ├─ /admin/* → harness-litellm:4000 (admin endpoints) │ │ ├─ /dashboard/ → harness-dashboard:3000 (harness UI) │ │ ├─ /litellm/* → harness-litellm:4000 (LiteLLM UI + API) │ │ ├─ /health/* → harness-litellm:4000 (health probes) │ │ └─ /gpu/* → 192.168.68.24:9100 (fleet monitor) │ │ │ │ Containers: │ │ ┌──────────┐ ┌──────────┐ ┌──────────┐ │ │ │ LiteLLM │ │Dashboard │ │ Grafana │ │ │ │ :4000 │ │ :3000 │ │ :3001 │ │ │ │ keys+sync│ │ harness │ │Prometheus│ │ │ │ fallback │ │ UI │ │ data src │ │ │ └──────────┘ └──────────┘ └──────────┘ │ │ │ │ ┌──────────┐ ┌──────────┐ ┌──────────┐ │ │ │PostgreSQL│ │ Redis │ │Prometheus│ │ │ │ :5432 │ │ :6379 │ │ :9090 │ │ │ └──────────┘ └──────────┘ └──────────┘ │ └───────┼──────────────────────────────────────────────────────────┘ │ ┌─┴─────────────┬───────────────┬───────────────┐ │ │ │ │ ┌─────▼──────┐ ┌─────▼──────┐ ┌─────▼──────┐ ┌─────▼──────┐ │ CT 8 │ │ CT 110 │ │ CT 15 │ │ pi (.24) │ │ RTX 3090 │ │ RTX 5070 │ │ Strix Halo │ │ GPU Monitor│ │ 24GB │ │ 12GB │ │ 64GB UMA │ │ :9100 │ │ :8080 │ │ :9400 exp │ │ :9400 exp │ │ :9401 │ └────────────┘ └────────────┘ └────────────┘ └────────────┘ ``` ## Stable Role-Based Aliases (Introduced 2026-07-15) Agent configs, cron jobs, and workflows MUST use these aliases, never model-specific names. When a model is swapped on a GPU, ONLY the infrastructure layer changes — agent configs are untouched. | Alias | Serves | Where | Kind | |-------|--------|-------|------| | `gpu-dense` | heavy reasoning | RTX 3090 (192.168.68.8) | direct alias | | `gpu-vision` | vision / web extract / light tasks | RTX 5070 (192.168.68.110) | direct alias AND `syslog-auto` pool member | | `strix-moe` | compression (MoE) | Strix Halo (192.168.68.15) | direct alias | | `syslog-auto` | balanced default | weighted pool across the three GPU hosts | pool router | Single source of truth for models, aliases, rpm caps, weights and fallback chains: CT 116 `/opt/inference-harness/litellm_config.yaml`. Do not duplicate those values in contracts — read them there. **No backward compatibility**: Old model-specific names (qwen3.6-27B-code, qwen3.5-9b-it) are retired as of 2026-09-12 and no longer resolve; do not use them in agent configs. Only the stable aliases survive model swaps. `gemma-4-12b` is retired and returns 400 `Invalid model name`. ## Routing Configuration (LiteLLM — July 2026) Model, alias, rpm/weight and fallback values are owned by CT 116 `/opt/inference-harness/litellm_config.yaml` (see § Stable Role-Based Aliases above). Note: All syslog-auto entries route directly to GPUs with `api_key: not-needed`. The router (port 9000) was decommissioned 2026-09-11 and is NOT in the inference path. ## Operations ### add-model 1. Download model file from Hugging Face or source 2. Check disk space + GPU VRAM compatibility 3. Start llama-server via systemd service on GPU host 4. Add to `/opt/inference-harness/gpu_roster.yaml` (or router env vars) 5. Hot-reload or restart router 6. Add model to LiteLLM config (`model_list` + `fallbacks`) 7. Generate agent keys for new model access via `/key/generate` 8. Add Prometheus scrape target for the new GPU exporter 9. Verify end-to-end: LiteLLM → Router → Model ### remove-model 1. Drain active requests (wait for active=0) 2. Remove from LiteLLM config 3. Remove from router config 4. Stop llama-server (systemd) 5. Remove Prometheus scrape target 6. Cleanup model files (optional) ### heal 1. Check all GPUs via gpu-monitor `{{gpu_dashboard_url}}/gpu-data` 2. Check LiteLLM health via nginx `:80/litellm/health/liveliness` 3. Reset stuck circuit breakers if idle (Redis) 4. Restart dead llama-server instances via SSH 5. Flush Redis active counters if stale 6. Verify GPU monitor server is running on pi (:9100) 7. Verify watchdog is running on pi 8. Restart router if roster not loaded (check logs for STARTUP ROSTER) 9. Router roster reload — REMOVED 2026-09-11 (router decommissioned; LiteLLM fallbacks handle routing) ### sync-keys 1. List all agent keys in LiteLLM DB via `GET /key/list` 2. Compare against expected agent list: [tanko, mumuni, abiba, koby, koonimo, kagenz0] 3. Generate missing keys via `POST /key/generate` with unlimited budget 4. Update Infisical vault: `infisical secrets set LITELLM_API_KEY= --project=agents --env=production` 5. Send Zulip DM to agents that can't be reached via SSH (provide vault login instructions) 6. Verify each key with test request through full chain 7. Document keys in knowledge graph ### list Show full fleet status: GPUs, models, VRAM, context windows, parallel slots, active requests, circuit breakers, keys ### health-check 1. Check GPU hardware: nvidia-smi (.8, .110) + amdgpu sysfs (.15 via /sys/class/drm/card0/device/) 2. Check llama-server processes: `ps aux | grep llama-server` on all 3 hosts 3. Check LiteLLM: `curl http://192.168.68.116/health` (expect "I'm alive!") 4. Check LiteLLM models: `curl -H "Authorization: Bearer $MASTER_KEY" http://192.168.68.116/v1/models` 5. Check LiteLLM timeouts: `grep -n 'timeout:' /opt/inference-harness/litellm_config.yaml` — read the live values from the authority config; do not assert them from this contract. 6. Check AMD metrics: `curl http://192.168.68.15:9400/metrics` (Radeon 8060S, util%, VRAM, temp, power) 7. Check port conflicts: verify only one llama-server on :8080 per host 8. Verify agent keys: 9 keys in LiteLLM DB (`GET /key/list`) ## Agent Keys (LiteLLM DB — Current 2026-07-11) Keys stored in Infisical vault (project=agents, env=production, secret=LITELLM_API_KEY). Agent gateways inject keys at runtime via `infisical run --` wrapper. Plaintext keys removed from this contract post-vault-migration. | Agent | CT | IP | LiteLLM Alias | Key Source | Access | |-------|-----|-----|---------------|------------|--------| | Tanko | 112 | .122 | `tanko` | Infisical vault | SSH jerome | | Mumuni | 105 (kagentz) | .14 | `mumuni` | Infisical vault | SSH root | | Abiba | 100 | .24 | `abiba-pi` | Infisical vault | local (pi agent) | | Koby | 111 | ? | `koby` | Infisical vault | Zulip DM | | Koonimo | 113 | ? | `koonimo` | Infisical vault (migrated 2026-07-11) | no SSH | | Kagenz0 | 105 | ? | `kagenz0` | Infisical vault | no SSH | > **Note**: CT hostnames differ from agent identities. CT111=tdunna runs koby; CT113=baggy runs koonimo. **Key update procedure**: Update Infisical vault → `infisical secrets set LITELLM_API_KEY=sk-... --project=agents --env=production` → restart agent gateway. Agent picks up new key via `infisical run --` wrapper at startup. If no SSH access, send Zulip DM via abiba-bot with vault update instructions. ## Configuration Files | File | Host | Purpose | |------|------|---------| | `/opt/inference-harness/docker-compose.yml` | CT 116 | All containers (router, litellm, nginx, postgres, redis, dashboard) | | `/opt/inference-harness/litellm_config.yaml` | CT 116 | LiteLLM proxy config (models, fallbacks, timeouts) | | `/opt/inference-harness/router/router.py` | CT 116 | Router source (builds via compose) | | `/etc/nginx/nginx.conf` | CT 116 (nginx container) | Routes /v1→LiteLLM, /dashboard/, /litellm/, /health | | `/opt/monitoring/prometheus.yml` | CT 116 | Prometheus scrape config (5 targets) | | `/root/scripts/gpu-monitor-server.py` | pi (.24) | GPU fleet monitor v2.1.0 (with benchmarks) | | `/root/scripts/gpu_benchmark.py` | pi (.24) | GPU inference benchmark module (tok/s tracking) | | `/root/scripts/gpu-saturation-watchdog.py` | pi (.24) | Auto-restart stuck llama-server | | `/root/dashboard/gpu-fleet.html` | pi (.24) | Live HTML dashboard | | `/etc/systemd/system/llama-server.service` | .8, .110 | llama-server daemons (Nvidia GPUs) | | `/etc/systemd/system/strix-server.service` | .15 (amdpve) | llama-server daemon (Vulkan, Strix Halo) running Carnice-Qwen3.6-MoE-35B-A3B-Q4_K_M.gguf (256K ctx). Note: `llama-server.service` and `llama-server@.service` are **masked** on .15 to prevent port 8080 collisions. | ## Prometheus & Grafana | Component | URL | Details | |-----------|-----|---------| | Grafana | `http://192.168.68.116:3001/` | admin / vault (`GRAFANA_ADMIN_PASSWORD`) | | GPU Dashboard | `http://192.168.68.116:3001/d/gpu-fleet` | Gauges + time series | | Prometheus | `http://192.168.68.116:9090/` (internal) | 5 scrape targets | | GPU Exporters | `:9400/metrics` on .8, .110, .15 | NVIDIA/AMD GPU metrics | | Router Exporter | `:9401/metrics` on .24 | Router + LiteLLM metrics | ## Known Issues & Watch Points - **Router startup race**: Compose router.py doesn't call load_roster(). Reload thread sleeps 30s first. Fix: trigger roster reload via SSH after restart, or rebuild image with startup load_roster(). - **LiteLLM /metrics**: Requires auth. Prometheus uses `/health/liveliness` as workaround. - **Strix Halo GPU**: Vulkan is the working backend (ROCm/HIP path abandoned — HSA runtime blocked on Debian 13). Build at `/root/llama.cpp/build-vk/`, commit `4fc4ec5` (2026-07-01), ggml 0.15.3 shared-lib arch. Mesa RADV 25.0.7, KHR_coopmat fast path active. ~70 tok/s gen, 532 tok/s prompt. Service: `strix-server.service` on port 8080, model: `Carnice-Qwen3.6-MoE-35B-A3B-Q4_K_M.gguf`, alias `strix-moe`, 256K context (n_ctx 262144), --parallel 2 --kv-unified, flash-attn + q4 KV, multimodal (mmproj loaded). - **Port conflict detection (2026-07-05)**: All 3 GPU wrappers now detect ghost processes squatting port 8080 before starting. `.8` and `.110` use inline pre-start check in `llama-wrapper.sh`; `.15` uses `/usr/local/bin/port-cleanup.sh` ExecStartPre. Replaces the blanket `pkill -9 -x llama-server` on .15 which would kill ALL llama-server instances regardless of port. Ghost detection was the root cause of .8 crash-looping for 27+ restarts (stale pid 25836 squatting 8080 after OOM kill). - **Strix Halo thermal safeguard (2026-07-02)**: `strix-server.service` has `-n 8192` (hard generation cap per request). Without it, `--predict` defaults to -1 (infinity) — a runaway request from .123 (old Mumuni CT114 — now inside Abiba CT100 at .24) decoded 39,868 tokens over 24 min, pushing Tctl to 98°C (crit 89.8°C) and throttling 70→29 t/s. The cap bounds worst-case generation to ~5 min. Do NOT remove `-n` without a replacement ceiling. Sustained load hits ~84°C even at 92s; the APU is fanless/low-flow. Clients MUST also set `max_tokens`. - **Port 8080 firewall**: amdpve iptables restricts 8080 to 192.168.68.116 (LiteLLM/router host) only. All inbound connections are from .116 (LiteLLM proxied via nginx). Localhost curls hang (SYN dropped). Always test from .116. - **Router sidecar fallback**: `router.py` `check_gpu_health()` now probes GPU `/health` directly when sidecar at :8090 is absent. Sidecar JSON exporters not deployed on any GPU host — router relies on GPU-direct fallback. - **Router GPU_MOE_URL bug (fixed 2026-07-01)**: docker-compose had `GPU_MOE_URL=.110:8080` (gemma host) instead of `.15:8080` (amdpve). Corrected. - **Alert migration**: All alerts now go to `#agent-hub` topics (`alerts-gpu`, `alerts-pm2`, `alerts-infra`) instead of DMs. Cross-agent visibility enabled. - **tok/s benchmarks**: Measured every 5 min via LiteLLM proxy. Baselines tracked with 30%/50% degradation thresholds. - **NetBird 502**: Tanko routes through NetBird for litellm.sysloggh.net. Use direct IP if NetBird down. - **Alias-retirement sweep (2026-09-12)**: `gemma-4-12b`, `gpu-light` and `crew-auto` are retired and replaced by `gpu-vision` / no cap respectively. The agent-facing templates (`hermes-config-template.prose.md`, `hermes-agent-baseline.prose.md`), `litellm-api-keys.prose.md`, `gpu-self-heal.prose.md`, `hermes-key-enforcement.prose.md`, `inference-optimization.prose.md`, `litellm-client-timeouts.prose.md` and the executable `audit-hermes-config.py` were all updated to the live canonical alias in the same change. **koby's config on .129 still names `gpu-light` (and `gemma-4-E4B`); .129 is report-only, so that is recorded for its owner and NOT edited here.** ## GPU Inference Benchmarks (Current) | GPU | Model | Gen tok/s | Prompt tok/s | Baseline | Context | |-----|-------|-----------|--------------|----------|---------| | RTX 3090 (.8) | Qwen3.8-27B-Uncensored-Q4_K_M | **TBD** | — | — | **128K** | | Strix Halo (.15) | Carnice-Qwen3.6-MoE-35B-A3B-Q4_K_M.gguf (strix-moe) | **65** | 140 | — | **256K** | Benchmarks from 2026-07-17. Strix Halo model: Carnice-Qwen3.6-MoE-35B-A3B-Q4_K_M.gguf (alias strix-moe), n_ctx 262144, --parallel 2 --kv-unified. RTX 5070 MTP provides 2.7x speedup over pre-upgrade 70 tok/s. GPU contexts: RTX 3090 (.8) and RTX 5070 (.110) at 128K; Strix Halo (.15) at 256K (2026-09-12). Benchmarks run through LiteLLM proxy (192.168.68.116:4000) every 5 minutes. Degradation alerts fire at 30% (warning) and 50% (critical) below baseline. History stored at `/root/data/toks-history.json` with 7-day rolling window. **Note (2026-07-01)**: Strix Halo prompt tok/s jumped 209→532 after Vulkan rebuild (cooperative-matrix fast path now active on GFX1151). Baseline may need re-calibration. ## Agent Config Implications (2026-07-15) ### Stable Aliases — CRITICAL All agent configs MUST use stable role-based aliases, never model-specific names: - `compression.model: syslog-auto` - `auxiliary.vision.model: gpu-vision` - `delegation.model: gpu-dense` - `auxiliary.web_extract.model: gpu-vision` When the underlying model is swapped, only the LiteLLM config changes — agent configs are untouched. ### Context Windows - RTX 3090: **128K** (reduced from 256K 2026-07-17) | RTX 5070: **128K** (reduced from 256K) | Strix Halo: **256K** (2026-09-12) - **Agents via `syslog-auto`**: 128K ceiling — the pool's safe floor (NVIDIA hosts are 128K). For >128K workloads, use external providers (deepseek) - Compression threshold 0.65 (audit Rule 9): fires at ~85K (~43K headroom before 128K ceiling) - **Pi agents (Abiba)**: `compaction.reserveTokens: 52739` (≈60% of 128K) - Mumuni compression model alias: `syslog-auto` ### Mumuni Agent Profile Mumuni (kagentz CT105, 192.168.68.14 — migrated from CT100 2026-08-29) is the primary business assistant. This profile is the reference for all agent configs. The compression values below are the current required values per template Rules 7/9 and `audit-hermes-config.py`; whether Mumuni's LIVE config currently complies is a separate operational question. | Setting | Value | Notes | |---------|-------|-------| | `model.default` | `syslog-auto` | Balanced default (pool router) | | `model.provider` | `custom:litellm` | LiteLLM on CT116 | | `compression.model` | `syslog-auto` | Rule 7: auto-routing, prevents Strix Halo overload | | `aux.compression.model` | `syslog-auto` | Must match `compression.model` (Rule 7) | | `aux.vision.model` | `gpu-vision` | Vision tasks (RTX 5070) | | `aux.web_extract.model` | `gpu-vision` | Web extraction | | `delegation.model` | `gpu-dense` | Sub-agent reasoning (RTX 3090) | | `context.max_context_window` | 131072 (128K) | Conservative `syslog-auto` pool floor (NVIDIA hosts 128K; Strix Halo 256K) | | `compression.threshold` | 0.65 | Rule 9: triggers at ~85K for a 128K window | | `compression.target_ratio` | 0.3 | Compresses to ~38K | | `compression.protect_last_n` | 40 | Preserves last 40 messages | | `memory.memory_char_limit` | 800 | Brief memory entries | | `personalities` | `creative` | Creative assistant personality | | Platforms | cli, homeassistant, signal, telegram, zulip | All Hermes platforms | ### Agent Update Status (2026-07-15) | Agent | Host | Status | |-------|------|--------| | **Mumuni** | CT105 (.14) | ✅ Updated to stable aliases | | **Tanko** | CT112 (.122) | ✅ Updated to stable aliases | | **Koby** | CT111 (.129) | ❌ SSH unreachable — needs Zulip DM | | **Koonimo** | CT113 | ❌ SSH unreachable — needs Zulip DM | | **Kagenz0** | CT105 | ❌ SSH unreachable — needs Zulip DM | All Hermes clients MUST set `max_tokens: 4096` — first line of defense before server-side `-n 8192` cap.