From 76fc2595ea91450797f09e88b7e5d8d5c0927391 Mon Sep 17 00:00:00 2001 From: root Date: Sun, 28 Jun 2026 16:24:40 +0000 Subject: [PATCH] feat(gpu-fleet): GPU fleet management prose contract Declarative GPU roster replaces hardcoded router.py model configs. Operations: add-model, remove-model, heal, sync-keys, list. Documents agent key architecture with LITELLM_API_KEY env var. Covers all 3 GPU hosts (RTX 5070, RTX 3090, Strix Halo). --- gpu-fleet.prose.md | 111 +++++++++++++++++++++++++++++++++++++++++++++ 1 file changed, 111 insertions(+) create mode 100644 gpu-fleet.prose.md diff --git a/gpu-fleet.prose.md b/gpu-fleet.prose.md new file mode 100644 index 0000000..cbc0e4d --- /dev/null +++ b/gpu-fleet.prose.md @@ -0,0 +1,111 @@ +--- +kind: responsibility +name: gpu-fleet +description: > + Manages the GPU inference fleet across all hosts. Handles model deployment, + registration, hot-reload, health checks, LiteLLM sync, agent key management, + and stale model cleanup. Replaces hardcoded router.py config with declarative + gpu_roster.yaml. +agent: abiba +triggers: + - on model add/remove + - on GPU health degradation + - on agent key rotation +--- + +## Maintains + +- gpu_roster: { models: map, hosts: map } — Single source of truth for all GPU models +- router: { status: "healthy", models: array } +- litellm: { status: "healthy", models: array } +- agent_keys: { agent: api_key } — All agent API keys registered in LiteLLM DB +- health: { gpus: array, circuit_breakers: array } + +## Operations + +### add-model +1. Download model file from Hugging Face or source +2. Check disk space + GPU VRAM compatibility +3. Build/update llama.cpp if needed (cmake, Vulkan/CUDA flags) +4. Start llama-server via systemd service +5. Add to `gpu_roster.yaml` +6. Hot-reload router: `POST /admin/roster/reload` +7. Sync LiteLLM config: add model to model_list + fallbacks +8. Generate/update agent keys for new model access +9. Verify end-to-end: LiteLLM → Router → Model + +### remove-model +1. Drain active requests (wait for active=0) +2. Update `gpu_roster.yaml` — remove entry +3. Hot-reload router +4. Remove from LiteLLM config +5. Stop llama-server (systemd) +6. Cleanup model files (optional — keep if model might return) + +### heal +1. Check all GPUs via router `/health/unified` +2. Reset stuck circuit breakers if idle +3. Restart dead llama-server instances +4. Flush Redis counters if stale +5. Sync LiteLLM if configs diverge + +### sync-keys +1. List all agent keys in LiteLLM DB via `/key/list` +2. Compare against expected agent list +3. Generate missing keys via `/key/generate` +4. Update agent configs to use `api_key_env: LITELLM_API_KEY` +5. Set `LITELLM_API_KEY` in system-wide env + startup wrapper +6. Restart gateways + +### list +Show full fleet status: GPUs, models, VRAM, active requests, circuit breakers + +## Configuration Files + +| File | Host | Purpose | +|------|------|---------| +| `gpu_roster.yaml` | CT 116 (`/opt/inference-harness/`) | Model-to-GPU mapping | +| `litellm_config.yaml` | CT 116 | LiteLLM proxy config | +| `roster_loader.py` | CT 116 (router container) | Hot-reloadable roster parser | +| `/etc/systemd/system/moe-server.service` | .15 | Ornith-1.0-35B server | +| Agent configs | .122, .123, .114 | `api_key_env: LITELLM_API_KEY` | + +## GPU Fleet Topology + +``` +┌──────────────────────────────────────────────────┐ +│ LiteLLM (:4001) │ +│ Config: /opt/inference-harness/litellm_config.yaml │ +└──────────────┬───────────────────────────────────┘ + │ +┌──────────────▼───────────────────────────────────┐ +│ Router (:9000) │ +│ Roster: /app/gpu_roster.yaml (hot-reload) │ +│ Redis: slot booking + circuit breakers │ +└──────┬──────────┬──────────┬─────────────────────┘ + │ │ │ + ┌────▼───┐ ┌───▼────┐ ┌──▼──────┐ + │ gemma │ │ qwen │ │ ornith │ + │ 4-12b │ │ 27B │ │ 1.0-35B │ + │ .110 │ │ .8 │ │ .15 │ + │ 5070 │ │ 3090 │ │ Strix │ + └────────┘ └────────┘ └─────────┘ +``` + +## Agent Key Architecture + +All agents use `api_key_env: LITELLM_API_KEY` in Hermes configs: + +| Agent | Host | Config | Env Source | +|-------|------|--------|------------| +| Mumuni | .123 | 7 refs + 6 subagent profiles | `/etc/environment` + wrapper | +| Tanko | .122 | 4 refs | `/etc/environment` + wrapper | +| Koonimo | .114 | 5 refs | `/etc/environment` + wrapper | + +To rotate all keys: update `LITELLM_API_KEY` value, restart gateways. + +## Sync Protocol + +1. Edit `gpu_roster.yaml` → `POST /admin/roster/reload` (hot, no restart) +2. Edit `litellm_config.yaml` → `docker restart harness-litellm` +3. Agent key changes → generate via `POST /key/generate` with master key