feat(gpu-fleet): GPU fleet management prose contract

Declarative GPU roster replaces hardcoded router.py model configs.
Operations: add-model, remove-model, heal, sync-keys, list.
Documents agent key architecture with LITELLM_API_KEY env var.
Covers all 3 GPU hosts (RTX 5070, RTX 3090, Strix Halo).
This commit is contained in:
root
2026-06-28 16:24:40 +00:00
parent 8a190fa803
commit 76fc2595ea
+111
View File
@@ -0,0 +1,111 @@
---
kind: responsibility
name: gpu-fleet
description: >
Manages the GPU inference fleet across all hosts. Handles model deployment,
registration, hot-reload, health checks, LiteLLM sync, agent key management,
and stale model cleanup. Replaces hardcoded router.py config with declarative
gpu_roster.yaml.
agent: abiba
triggers:
- on model add/remove
- on GPU health degradation
- on agent key rotation
---
## Maintains
- gpu_roster: { models: map, hosts: map } — Single source of truth for all GPU models
- router: { status: "healthy", models: array }
- litellm: { status: "healthy", models: array }
- agent_keys: { agent: api_key } — All agent API keys registered in LiteLLM DB
- health: { gpus: array, circuit_breakers: array }
## Operations
### add-model
1. Download model file from Hugging Face or source
2. Check disk space + GPU VRAM compatibility
3. Build/update llama.cpp if needed (cmake, Vulkan/CUDA flags)
4. Start llama-server via systemd service
5. Add to `gpu_roster.yaml`
6. Hot-reload router: `POST /admin/roster/reload`
7. Sync LiteLLM config: add model to model_list + fallbacks
8. Generate/update agent keys for new model access
9. Verify end-to-end: LiteLLM → Router → Model
### remove-model
1. Drain active requests (wait for active=0)
2. Update `gpu_roster.yaml` — remove entry
3. Hot-reload router
4. Remove from LiteLLM config
5. Stop llama-server (systemd)
6. Cleanup model files (optional — keep if model might return)
### heal
1. Check all GPUs via router `/health/unified`
2. Reset stuck circuit breakers if idle
3. Restart dead llama-server instances
4. Flush Redis counters if stale
5. Sync LiteLLM if configs diverge
### sync-keys
1. List all agent keys in LiteLLM DB via `/key/list`
2. Compare against expected agent list
3. Generate missing keys via `/key/generate`
4. Update agent configs to use `api_key_env: LITELLM_API_KEY`
5. Set `LITELLM_API_KEY` in system-wide env + startup wrapper
6. Restart gateways
### list
Show full fleet status: GPUs, models, VRAM, active requests, circuit breakers
## Configuration Files
| File | Host | Purpose |
|------|------|---------|
| `gpu_roster.yaml` | CT 116 (`/opt/inference-harness/`) | Model-to-GPU mapping |
| `litellm_config.yaml` | CT 116 | LiteLLM proxy config |
| `roster_loader.py` | CT 116 (router container) | Hot-reloadable roster parser |
| `/etc/systemd/system/moe-server.service` | .15 | Ornith-1.0-35B server |
| Agent configs | .122, .123, .114 | `api_key_env: LITELLM_API_KEY` |
## GPU Fleet Topology
```
┌──────────────────────────────────────────────────┐
│ LiteLLM (:4001) │
│ Config: /opt/inference-harness/litellm_config.yaml │
└──────────────┬───────────────────────────────────┘
┌──────────────▼───────────────────────────────────┐
│ Router (:9000) │
│ Roster: /app/gpu_roster.yaml (hot-reload) │
│ Redis: slot booking + circuit breakers │
└──────┬──────────┬──────────┬─────────────────────┘
│ │ │
┌────▼───┐ ┌───▼────┐ ┌──▼──────┐
│ gemma │ │ qwen │ │ ornith │
│ 4-12b │ │ 27B │ │ 1.0-35B │
│ .110 │ │ .8 │ │ .15 │
│ 5070 │ │ 3090 │ │ Strix │
└────────┘ └────────┘ └─────────┘
```
## Agent Key Architecture
All agents use `api_key_env: LITELLM_API_KEY` in Hermes configs:
| Agent | Host | Config | Env Source |
|-------|------|--------|------------|
| Mumuni | .123 | 7 refs + 6 subagent profiles | `/etc/environment` + wrapper |
| Tanko | .122 | 4 refs | `/etc/environment` + wrapper |
| Koonimo | .114 | 5 refs | `/etc/environment` + wrapper |
To rotate all keys: update `LITELLM_API_KEY` value, restart gateways.
## Sync Protocol
1. Edit `gpu_roster.yaml``POST /admin/roster/reload` (hot, no restart)
2. Edit `litellm_config.yaml``docker restart harness-litellm`
3. Agent key changes → generate via `POST /key/generate` with master key