20 KiB
kind, name, description, agent, triggers
| kind | name | description | agent | triggers | ||||
|---|---|---|---|---|---|---|---|---|
| responsibility | gpu-fleet | Manages the GPU inference fleet across all hosts. Handles model deployment, registration, health checks, LiteLLM sync, agent key management, GPU saturation watchdog, Prometheus/Grafana monitoring, and self-healing. UPDATED 2026-07-15: Stable role-based aliases introduced: strix-moe, gpu-dense, gpu-vision (gpu-light was superseded by gpu-vision on 2026-09-12). These never change — only the underlying model does. Strix Halo: strix-moe → Carnice-Qwen3.6-MoE-35B-A3B-Q4_K_M.gguf (22GB, 256K ctx). RTX 5070: gpu-vision — IQ4_NL + MTP draft (~122 tok/s, 2x faster). UPDATED 2026-07-17: NVIDIA host context reduced from 256K to 128K for stability. Strix Halo model: Carnice-Qwen3.6-MoE-35B-A3B-Q4_K_M.gguf (alias strix-moe, 256K context). Strix Halo runs 256K (n_ctx 262144, --kv-unified); RTX 3090 and RTX 5070 remain at 128K. For >128K on NVIDIA hosts → fall back to external providers (deepseek). VRAM headroom improved: RTX 3090 ~70%, RTX 5070 ~65%. | abiba |
|
Maintains
- gpu_roster: { models: map, hosts: map } — GPU host/model roster; the authoritative alias/weight/fallback registry is CT 116
litellm_config.yaml - router: { status: "healthy", roster_loaded: bool, models: array }
- litellm: { status: "healthy", keys: array, models: array }
- agent_keys: { agent: api_key } — All agent API keys registered in LiteLLM DB
- health: { gpus: array, circuit_breakers: array } — Fleet-wide health state
- monitor: { status: "running", version: "2.0.0" } — GPU monitor server on pi (:9100)
- watchdog: { status: "running" } — GPU saturation watchdog (restarts stuck llama-server)
- benchmarks: { tok_per_sec: map, baseline: map, history: array } — Inference speed benchmarks tracked over time
- grafana: { status: "running", dashboards: ["gpu-fleet"] } — Grafana on CT 116 (:3001)
- prometheus: { status: "running", targets: 5 } — Scrapes GPU :9400 exporters + LiteLLM
- port_conflict_detection: { status: "active" } — All 3 GPU wrappers detect ghost processes before binding
Fleet Topology (Current — 2026-09-11, router decommissioned)
┌──────────────────────────────────────────────────────────────────┐
│ CT 116 (192.168.68.116) — Inference Harness Host │
│ │
│ nginx:80 (entrypoint) │
│ ├─ /v1/* → harness-litellm:4000 (API requests) │
│ ├─ /admin/* → harness-litellm:4000 (admin endpoints) │
│ ├─ /dashboard/ → harness-dashboard:3000 (harness UI) │
│ ├─ /litellm/* → harness-litellm:4000 (LiteLLM UI + API) │
│ ├─ /health/* → harness-litellm:4000 (health probes) │
│ └─ /gpu/* → 192.168.68.24:9100 (fleet monitor) │
│ │
│ Containers: │
│ ┌──────────┐ ┌──────────┐ ┌──────────┐ │
│ │ LiteLLM │ │Dashboard │ │ Grafana │ │
│ │ :4000 │ │ :3000 │ │ :3001 │ │
│ │ keys+sync│ │ harness │ │Prometheus│ │
│ │ fallback │ │ UI │ │ data src │ │
│ └──────────┘ └──────────┘ └──────────┘ │
│ │
│ ┌──────────┐ ┌──────────┐ ┌──────────┐ │
│ │PostgreSQL│ │ Redis │ │Prometheus│ │
│ │ :5432 │ │ :6379 │ │ :9090 │ │
│ └──────────┘ └──────────┘ └──────────┘ │
└───────┼──────────────────────────────────────────────────────────┘
│
┌─┴─────────────┬───────────────┬───────────────┐
│ │ │ │
┌─────▼──────┐ ┌─────▼──────┐ ┌─────▼──────┐ ┌─────▼──────┐
│ CT 8 │ │ CT 110 │ │ CT 15 │ │ pi (.24) │
│ RTX 3090 │ │ RTX 5070 │ │ Strix Halo │ │ GPU Monitor│
│ 24GB │ │ 12GB │ │ 64GB UMA │ │ :9100 │
│ :8080 │ │ :9400 exp │ │ :9400 exp │ │ :9401 │
└────────────┘ └────────────┘ └────────────┘ └────────────┘
Stable Role-Based Aliases (Introduced 2026-07-15)
Agent configs, cron jobs, and workflows MUST use these aliases, never model-specific names. When a model is swapped on a GPU, ONLY the infrastructure layer changes — agent configs are untouched.
| Alias | Serves | Where | Kind |
|---|---|---|---|
gpu-dense |
heavy reasoning | RTX 3090 (192.168.68.8) | direct alias |
gpu-vision |
vision / web extract / light tasks | RTX 5070 (192.168.68.110) | direct alias AND syslog-auto pool member |
strix-moe |
compression (MoE) | Strix Halo (192.168.68.15) | direct alias |
syslog-auto |
balanced default | weighted pool across the three GPU hosts | pool router |
Single source of truth for models, aliases, rpm caps, weights and fallback chains:
CT 116 /opt/inference-harness/litellm_config.yaml. Do not duplicate those values in
contracts — read them there.
No backward compatibility: Old model-specific names (qwen3.6-27B-code, qwen3.5-9b-it) are retired as of 2026-09-12
but are deprecated for agent configs. Only the stable aliases survive model swaps. gemma-4-12b
is retired and returns 400 Invalid model name.
Routing Configuration (LiteLLM — July 2026)
Model, alias, rpm/weight and fallback values are owned by CT 116
/opt/inference-harness/litellm_config.yaml (see § Stable Role-Based Aliases above).
Note: All syslog-auto entries route directly to GPUs with api_key: not-needed. The router (port 9000) was decommissioned 2026-09-11 and is NOT in the inference path.
Operations
add-model
- Download model file from Hugging Face or source
- Check disk space + GPU VRAM compatibility
- Start llama-server via systemd service on GPU host
- Add to
/opt/inference-harness/gpu_roster.yaml(or router env vars) - Hot-reload or restart router
- Add model to LiteLLM config (
model_list+fallbacks) - Generate agent keys for new model access via
/key/generate - Add Prometheus scrape target for the new GPU exporter
- Verify end-to-end: LiteLLM → Router → Model
remove-model
- Drain active requests (wait for active=0)
- Remove from LiteLLM config
- Remove from router config
- Stop llama-server (systemd)
- Remove Prometheus scrape target
- Cleanup model files (optional)
heal
- Check all GPUs via gpu-monitor
{{gpu_dashboard_url}}/gpu-data - Check LiteLLM health via nginx
:80/litellm/health/liveliness - Reset stuck circuit breakers if idle (Redis)
- Restart dead llama-server instances via SSH
- Flush Redis active counters if stale
- Verify GPU monitor server is running on pi (:9100)
- Verify watchdog is running on pi
- Restart router if roster not loaded (check logs for STARTUP ROSTER)
- Router roster reload — REMOVED 2026-09-11 (router decommissioned; LiteLLM fallbacks handle routing)
sync-keys
- List all agent keys in LiteLLM DB via
GET /key/list - Compare against expected agent list: [tanko, mumuni, abiba, koby, koonimo, kagenz0]
- Generate missing keys via
POST /key/generatewith unlimited budget - Update Infisical vault:
infisical secrets set LITELLM_API_KEY=<key> --project=agents --env=production - Send Zulip DM to agents that can't be reached via SSH (provide vault login instructions)
- Verify each key with test request through full chain
- Document keys in knowledge graph
list
Show full fleet status: GPUs, models, VRAM, context windows, parallel slots, active requests, circuit breakers, keys
health-check
- Check GPU hardware: nvidia-smi (.8, .110) + amdgpu sysfs (.15 via /sys/class/drm/card0/device/)
- Check llama-server processes:
ps aux | grep llama-serveron all 3 hosts - Check LiteLLM:
curl http://192.168.68.116/health(expect "I'm alive!") - Check LiteLLM models:
curl -H "Authorization: Bearer $MASTER_KEY" http://192.168.68.116/v1/models - Check LiteLLM timeouts:
grep -n 'timeout:' /opt/inference-harness/litellm_config.yaml— read the live values from the authority config; do not assert them from this contract. - Check AMD metrics:
curl http://192.168.68.15:9400/metrics(Radeon 8060S, util%, VRAM, temp, power) - Check port conflicts: verify only one llama-server on :8080 per host
- Verify agent keys: 9 keys in LiteLLM DB (
GET /key/list)
Agent Keys (LiteLLM DB — Current 2026-07-11)
Keys stored in Infisical vault (project=agents, env=production, secret=LITELLM_API_KEY).
Agent gateways inject keys at runtime via infisical run -- wrapper.
Plaintext keys removed from this contract post-vault-migration.
| Agent | CT | IP | LiteLLM Alias | Key Source | Access |
|---|---|---|---|---|---|
| Tanko | 112 | .122 | tanko |
Infisical vault | SSH jerome |
| Mumuni | 105 (kagentz) | .14 | mumuni |
Infisical vault | SSH root |
| Abiba | 100 | .24 | abiba-pi |
Infisical vault | local (pi agent) |
| Koby | 111 | ? | koby |
Infisical vault | Zulip DM |
| Koonimo | 113 | ? | koonimo |
Infisical vault (migrated 2026-07-11) | no SSH |
| Kagenz0 | 105 | ? | kagenz0 |
Infisical vault | no SSH |
Note
: CT hostnames differ from agent identities. CT111=tdunna runs koby; CT113=baggy runs koonimo.
Key update procedure: Update Infisical vault → infisical secrets set LITELLM_API_KEY=sk-... --project=agents --env=production → restart agent gateway. Agent picks up new key via infisical run -- wrapper at startup.
If no SSH access, send Zulip DM via abiba-bot with vault update instructions.
Configuration Files
| File | Host | Purpose |
|---|---|---|
/opt/inference-harness/docker-compose.yml |
CT 116 | All containers (router, litellm, nginx, postgres, redis, dashboard) |
/opt/inference-harness/litellm_config.yaml |
CT 116 | LiteLLM proxy config (models, fallbacks, timeouts) |
/opt/inference-harness/router/router.py |
CT 116 | Router source (builds via compose) |
/etc/nginx/nginx.conf |
CT 116 (nginx container) | Routes /v1→LiteLLM, /dashboard/, /litellm/, /health |
/opt/monitoring/prometheus.yml |
CT 116 | Prometheus scrape config (5 targets) |
/root/scripts/gpu-monitor-server.py |
pi (.24) | GPU fleet monitor v2.1.0 (with benchmarks) |
/root/scripts/gpu_benchmark.py |
pi (.24) | GPU inference benchmark module (tok/s tracking) |
/root/scripts/gpu-saturation-watchdog.py |
pi (.24) | Auto-restart stuck llama-server |
/root/dashboard/gpu-fleet.html |
pi (.24) | Live HTML dashboard |
/etc/systemd/system/llama-server.service |
.8, .110 | llama-server daemons (Nvidia GPUs) |
/etc/systemd/system/strix-server.service |
.15 (amdpve) | llama-server daemon (Vulkan, Strix Halo) running Carnice-Qwen3.6-MoE-35B-A3B-Q4_K_M.gguf (256K ctx). Note: llama-server.service and llama-server@.service are masked on .15 to prevent port 8080 collisions. |
Prometheus & Grafana
| Component | URL | Details |
|---|---|---|
| Grafana | http://192.168.68.116:3001/ |
admin / vault (GRAFANA_ADMIN_PASSWORD) |
| GPU Dashboard | http://192.168.68.116:3001/d/gpu-fleet |
Gauges + time series |
| Prometheus | http://192.168.68.116:9090/ (internal) |
5 scrape targets |
| GPU Exporters | :9400/metrics on .8, .110, .15 |
NVIDIA/AMD GPU metrics |
| Router Exporter | :9401/metrics on .24 |
Router + LiteLLM metrics |
Known Issues & Watch Points
- Router startup race: Compose router.py doesn't call load_roster(). Reload thread sleeps 30s first. Fix: trigger roster reload via SSH after restart, or rebuild image with startup load_roster().
- LiteLLM /metrics: Requires auth. Prometheus uses
/health/livelinessas workaround. - Strix Halo GPU: Vulkan is the working backend (ROCm/HIP path abandoned — HSA runtime blocked on Debian 13). Build at
/root/llama.cpp/build-vk/, commit4fc4ec5(2026-07-01), ggml 0.15.3 shared-lib arch. Mesa RADV 25.0.7, KHR_coopmat fast path active. ~70 tok/s gen, 532 tok/s prompt. Service:strix-server.serviceon port 8080, model:Carnice-Qwen3.6-MoE-35B-A3B-Q4_K_M.gguf, aliasstrix-moe, 256K context (n_ctx 262144), --parallel 2 --kv-unified, flash-attn + q4 KV, multimodal (mmproj loaded). - Port conflict detection (2026-07-05): All 3 GPU wrappers now detect ghost processes squatting port 8080 before starting.
.8and.110use inline pre-start check inllama-wrapper.sh;.15uses/usr/local/bin/port-cleanup.shExecStartPre. Replaces the blanketpkill -9 -x llama-serveron .15 which would kill ALL llama-server instances regardless of port. Ghost detection was the root cause of .8 crash-looping for 27+ restarts (stale pid 25836 squatting 8080 after OOM kill). - Strix Halo thermal safeguard (2026-07-02):
strix-server.servicehas-n 8192(hard generation cap per request). Without it,--predictdefaults to -1 (infinity) — a runaway request from .123 (old Mumuni CT114 — now inside Abiba CT100 at .24) decoded 39,868 tokens over 24 min, pushing Tctl to 98°C (crit 89.8°C) and throttling 70→29 t/s. The cap bounds worst-case generation to ~5 min. Do NOT remove-nwithout a replacement ceiling. Sustained load hits ~84°C even at 92s; the APU is fanless/low-flow. Clients MUST also setmax_tokens. - Port 8080 firewall: amdpve iptables restricts 8080 to 192.168.68.116 (LiteLLM/router host) only. All inbound connections are from .116 (LiteLLM proxied via nginx). Localhost curls hang (SYN dropped). Always test from .116.
- Router sidecar fallback:
router.pycheck_gpu_health()now probes GPU/healthdirectly when sidecar at :8090 is absent. Sidecar JSON exporters not deployed on any GPU host — router relies on GPU-direct fallback. - Router GPU_MOE_URL bug (fixed 2026-07-01): docker-compose had
GPU_MOE_URL=.110:8080(gemma host) instead of.15:8080(amdpve). Corrected. - Alert migration: All alerts now go to
#agent-hubtopics (alerts-gpu,alerts-pm2,alerts-infra) instead of DMs. Cross-agent visibility enabled. - tok/s benchmarks: Measured every 5 min via LiteLLM proxy. Baselines tracked with 30%/50% degradation thresholds.
- NetBird 502: Tanko routes through NetBird for litellm.sysloggh.net. Use direct IP if NetBird down.
- Alias-retirement sweep (2026-09-12):
gemma-4-12b,gpu-lightandcrew-autoare retired and replaced bygpu-vision/ no cap respectively. The agent-facing templates (hermes-config-template.prose.md,hermes-agent-baseline.prose.md),litellm-api-keys.prose.md,gpu-self-heal.prose.md,hermes-key-enforcement.prose.md,inference-optimization.prose.md,litellm-client-timeouts.prose.mdand the executableaudit-hermes-config.pywere all updated to the live canonical alias in the same change. koby's config on .129 still namesgpu-light(andgemma-4-E4B); .129 is report-only, so that is recorded for its owner and NOT edited here.
GPU Inference Benchmarks (Current)
| GPU | Model | Gen tok/s | Prompt tok/s | Baseline | Context |
|---|---|---|---|---|---|
| RTX 3090 (.8) | Qwen3.8-27B-Uncensored-Q4_K_M | TBD | — | — | 128K |
| Strix Halo (.15) | Carnice-Qwen3.6-MoE-35B-A3B-Q4_K_M.gguf (strix-moe) | 65 | 140 | — | 256K |
Benchmarks from 2026-07-17. Strix Halo model: Carnice-Qwen3.6-MoE-35B-A3B-Q4_K_M.gguf (alias strix-moe), n_ctx 262144, --parallel 2 --kv-unified. RTX 5070 MTP provides 2.7x speedup over pre-upgrade 70 tok/s. GPU contexts: RTX 3090 (.8) and RTX 5070 (.110) at 128K; Strix Halo (.15) at 256K (2026-09-12).
Benchmarks run through LiteLLM proxy (192.168.68.116:4000) every 5 minutes.
Degradation alerts fire at 30% (warning) and 50% (critical) below baseline.
History stored at /root/data/toks-history.json with 7-day rolling window.
Note (2026-07-01): Strix Halo prompt tok/s jumped 209→532 after Vulkan rebuild (cooperative-matrix fast path now active on GFX1151). Baseline may need re-calibration.
Agent Config Implications (2026-07-15)
Stable Aliases — CRITICAL
All agent configs MUST use stable role-based aliases, never model-specific names:
compression.model: syslog-autoauxiliary.vision.model: gpu-visiondelegation.model: gpu-denseauxiliary.web_extract.model: gpu-vision
When the underlying model is swapped, only the LiteLLM config changes — agent configs are untouched.
Context Windows
- RTX 3090: 128K (reduced from 256K 2026-07-17) | RTX 5070: 128K (reduced from 256K) | Strix Halo: 256K (2026-09-12)
- Agents via
syslog-auto: 128K ceiling — the pool's safe floor (NVIDIA hosts are 128K). For >128K workloads, use external providers (deepseek) - Compression threshold 0.65 (audit Rule 9): fires at ~85K (~43K headroom before 128K ceiling)
- Pi agents (Abiba):
compaction.reserveTokens: 52739(≈60% of 128K) - Mumuni compression model alias:
syslog-auto
Mumuni Agent Profile
Mumuni (kagentz CT105, 192.168.68.14 — migrated from CT100 2026-08-29) is the primary business assistant. This profile is the reference for all agent configs. The compression values below are the current required values per template Rules 7/9 and audit-hermes-config.py; whether Mumuni's LIVE config currently complies is a separate operational question.
| Setting | Value | Notes |
|---|---|---|
model.default |
syslog-auto |
Balanced default (pool router) |
model.provider |
custom:litellm |
LiteLLM on CT116 |
compression.model |
syslog-auto |
Rule 7: auto-routing, prevents Strix Halo overload |
aux.compression.model |
syslog-auto |
Must match compression.model (Rule 7) |
aux.vision.model |
gpu-vision |
Vision tasks (RTX 5070) |
aux.web_extract.model |
gpu-vision |
Web extraction |
delegation.model |
gpu-dense |
Sub-agent reasoning (RTX 3090) |
context.max_context_window |
131072 (128K) | Conservative syslog-auto pool floor (NVIDIA hosts 128K; Strix Halo 256K) |
compression.threshold |
0.65 | Rule 9: triggers at ~85K for a 128K window |
compression.target_ratio |
0.3 | Compresses to ~38K |
compression.protect_last_n |
40 | Preserves last 40 messages |
memory.memory_char_limit |
800 | Brief memory entries |
personalities |
creative |
Creative assistant personality |
| Platforms | cli, homeassistant, signal, telegram, zulip | All Hermes platforms |
Agent Update Status (2026-07-15)
| Agent | Host | Status |
|---|---|---|
| Mumuni | CT105 (.14) | ✅ Updated to stable aliases |
| Tanko | CT112 (.122) | ✅ Updated to stable aliases |
| Koby | CT111 (.129) | ❌ SSH unreachable — needs Zulip DM |
| Koonimo | CT113 | ❌ SSH unreachable — needs Zulip DM |
| Kagenz0 | CT105 | ❌ SSH unreachable — needs Zulip DM |
All Hermes clients MUST set max_tokens: 4096 — first line of defense before server-side -n 8192 cap.