feat: tighten contracts with GPU fleet topology, model routing, Strix Halo details
litellm-health: added GPU fleet topology table, model routing map, 3-tier GPU checks (reachability, VRAM, temp, test inference) litellm-self-heal: added GPU fleet reference, rules for GPU unreachable and model not responding Infrastructure topology now fully documented: - amdpve: Strix Halo CPU (35B, 16 threads, 262K ctx, llama-server) - llm-gpu: RTX 3090 24GB (Dense tier) - ocu-llm: RTX 5070 12GB (MoE + Light tiers)
This commit is contained in:
+29
-2
@@ -12,6 +12,7 @@ description: >
|
|||||||
- public_url: string — The public LiteLLM URL (default: "https://litellm.sysloggh.net")
|
- public_url: string — The public LiteLLM URL (default: "https://litellm.sysloggh.net")
|
||||||
- backend_host: string — Internal CT host to check containers (default: "192.168.68.116")
|
- backend_host: string — Internal CT host to check containers (default: "192.168.68.116")
|
||||||
- auth_host: string — Authentik server for OIDC (default: "192.168.68.11")
|
- auth_host: string — Authentik server for OIDC (default: "192.168.68.11")
|
||||||
|
- gpu_hosts: array — GPU inference hosts to check (default: ["192.168.68.8", "192.168.68.110", "192.168.68.15"])
|
||||||
|
|
||||||
## Returns
|
## Returns
|
||||||
|
|
||||||
@@ -34,6 +35,23 @@ description: >
|
|||||||
- If /health/unified reports any non-healthy components, status reflects it
|
- If /health/unified reports any non-healthy components, status reflects it
|
||||||
- Checks cover at minimum: admin UI, API docs, OIDC, containers, unified health, nginx proxy
|
- Checks cover at minimum: admin UI, API docs, OIDC, containers, unified health, nginx proxy
|
||||||
|
|
||||||
|
## GPU Fleet Topology
|
||||||
|
|
||||||
|
| Host | IP | Hardware | Models Served | Engine |
|
||||||
|
|------|-----|----------|---------------|--------|
|
||||||
|
| llm-gpu | 192.168.68.8 | NVIDIA RTX 3090 (24 GB) | gemma-4-12b (Dense) | llama-server Docker |
|
||||||
|
| ocu-llm | 192.168.68.110 | NVIDIA RTX 5070 (12 GB) | qwen3.6-27B-code (MoE), all Light tier | llama-server Docker |
|
||||||
|
| amdpve | 192.168.68.15 | AMD Strix Halo (CPU-only, -ngl 0) | ornith-1.0-35b (35B) | llama-server bare-metal |
|
||||||
|
|
||||||
|
## Model Routing (LiteLLM → Router → GPU)
|
||||||
|
|
||||||
|
| Model Request | Router Routes To |
|
||||||
|
|--------------|-----------------|
|
||||||
|
| `gemma-4-12b` | `GPU_DENSE_URL` → 192.168.68.8:8080 |
|
||||||
|
| `qwen3.6-27B-code` | `GPU_MOE_URL` → 192.168.68.110:8080 |
|
||||||
|
| `ornith-1.0-35b` | Router routes to 192.168.68.15:8080 |
|
||||||
|
| `syslog-auto` | Auto-routed by router tier logic |
|
||||||
|
|
||||||
## Execution
|
## Execution
|
||||||
|
|
||||||
1. **Read parameters** — Use provided values or defaults
|
1. **Read parameters** — Use provided values or defaults
|
||||||
@@ -50,7 +68,16 @@ description: >
|
|||||||
5. **Check backend container health**:
|
5. **Check backend container health**:
|
||||||
- SSH to {{backend_host}} → `docker ps` → verify all 6 containers are healthy
|
- SSH to {{backend_host}} → `docker ps` → verify all 6 containers are healthy
|
||||||
- Containers: harness-litellm, harness-nginx, harness-router, harness-postgres, harness-redis, harness-dashboard
|
- Containers: harness-litellm, harness-nginx, harness-router, harness-postgres, harness-redis, harness-dashboard
|
||||||
6. **Check OIDC auth endpoint**:
|
6. **Check GPU fleet health**:
|
||||||
|
- For each {{gpu_hosts}}:
|
||||||
|
- SSH or nvidia-smi check → verify GPU is reachable
|
||||||
|
- Check VRAM usage (warn if >90%)
|
||||||
|
- Check temperature (warn if >80°C)
|
||||||
|
- Verify model responds to a small test prompt
|
||||||
|
7. **Check model inference** — Send test prompt to each model via LiteLLM API:
|
||||||
|
- POST /v1/chat/completions with model name → expect 200 + valid response
|
||||||
|
- Record token counts for baseline performance
|
||||||
|
8. **Check OIDC auth endpoint**:
|
||||||
- Verify auth.sysloggh.net resolves to {{auth_host}}
|
- Verify auth.sysloggh.net resolves to {{auth_host}}
|
||||||
- GET https://auth.sysloggh.net/ → expect login page
|
- GET https://auth.sysloggh.net/ → expect login page
|
||||||
7. **Compile and report** — Determine overall_status from individual check results
|
9. **Compile and report** — Determine overall_status from individual check results
|
||||||
|
|||||||
@@ -66,6 +66,14 @@ logged as a `[REPORT]` knowledge graph node for long-term trending.
|
|||||||
If a fix required another agent's help (e.g., Authentik restart), a relay
|
If a fix required another agent's help (e.g., Authentik restart), a relay
|
||||||
message is sent to the responsible agent with full context.
|
message is sent to the responsible agent with full context.
|
||||||
|
|
||||||
|
## GPU Fleet
|
||||||
|
|
||||||
|
| Host | IP | Hardware | Models | Engine |
|
||||||
|
|------|-----|----------|--------|--------|
|
||||||
|
| amdpve | 192.168.68.15 | Strix Halo CPU | ornith-1.0-35b (16t, 262K ctx) | llama-server :8080 |
|
||||||
|
| llm-gpu | 192.168.68.8 | RTX 3090 24GB | gemma-4-12b Dense | llama-server :8080 |
|
||||||
|
| ocu-llm | 192.168.68.110 | RTX 5070 12GB | qwen3.6-27B-code MoE+Light | llama-server :8080 |
|
||||||
|
|
||||||
## Remediation Rules
|
## Remediation Rules
|
||||||
|
|
||||||
### Rule 1: Container Not Healthy
|
### Rule 1: Container Not Healthy
|
||||||
@@ -82,6 +90,16 @@ Detect → immediate escalate (external dependency)
|
|||||||
### Rule 4: Docs 404
|
### Rule 4: Docs 404
|
||||||
Detect → fix DOCS_URL env var → verify → log
|
Detect → fix DOCS_URL env var → verify → log
|
||||||
|
|
||||||
|
### Rule 5: GPU Unreachable
|
||||||
|
Detect → SSH to GPU host fails or model not responding
|
||||||
|
Fix → restart llama-server on host via SSH
|
||||||
|
Escalate → immediately (downtime affects inference)
|
||||||
|
|
||||||
|
### Rule 6: Model Not Responding
|
||||||
|
Detect → test prompt via LiteLLM returns non-200
|
||||||
|
Fix → restart llama-server on GPU host → verify
|
||||||
|
Escalate → after 2 failed restarts
|
||||||
|
|
||||||
## Execution
|
## Execution
|
||||||
|
|
||||||
1. Run health check
|
1. Run health check
|
||||||
|
|||||||
Reference in New Issue
Block a user