Merged PR #50: fix/gpu-dense-docs
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 0s
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 0s
This commit is contained in:
@@ -1,4 +1,6 @@
|
||||
---
|
||||
report_only_agents:
|
||||
- koby # ⛔ KOBY IS NEVER REPAIRED (Rule 17, 2026-08-17) — detect + report, never fix on .129
|
||||
kind: responsibility
|
||||
name: gpu-self-heal
|
||||
description: >
|
||||
@@ -16,6 +18,7 @@ depends_on:
|
||||
- gpu-monitor.prose.md (live data source on .24:9100)
|
||||
- gpu-fleet.prose.md (source of truth for topology, aliases, model assignments)
|
||||
---
|
||||
---
|
||||
|
||||
## Maintains
|
||||
|
||||
@@ -38,20 +41,19 @@ depends_on:
|
||||
- On fix: verify with benchmark inference test before declaring resolved
|
||||
- Escalate: after 3 failed remediation attempts → Zulip #agent-hub alert
|
||||
|
||||
---
|
||||
---
|
||||
|
||||
## Current Fleet Baseline (2026-07-18)
|
||||
|
||||
| Alias | GPU | Host | Model | VRAM | Ctx | tok/s | Role |
|
||||
|-------|-----|------|-------|------|-----|-------|------|
|
||||
| `gpu-dense` | RTX 3090 24GB | ct8 (.8:8080) | Qwen3.8-27B-Uncensored-Q4_K_M (alias qwen3.6-27B-code) | ~16.8/24.6GB | 128K | — | Heavy reasoning, code gen |
|
||||
| `gpu-light` | RTX 5070 12GB | ct110 (.110:8080) | HauhauCS Gemma4-12B QAT Q4_K_M + MTP draft | 10.1/12.2GB (83%) | 128K | 169.6 | Vision, web extract, light tasks |
|
||||
| `strix-moe` | Strix Halo 64GB | ct15 (.15:8080) | qwen3.6-35B-udq4 | ~10/64GB (16%) | 128K | 62.9 | Compression, summarization, long docs |
|
||||
|
||||
Key notes:
|
||||
- All models use direct GPU routing via LiteLLM (`api_key: not-needed`). Router (port 9000) is deprecated and NOT in the inference path.
|
||||
- Stable aliases (gpu-dense, gpu-light, strix-moe) from gpu-fleet are the canonical names for agent configs. Model-specific names still work but are deprecated.
|
||||
- RTX 5070 tok/s is 2.3x faster than RTX 3090 for its model — gpu-light is the fastest endpoint. Route vision/web/light work there first.
|
||||
- RTX 5070 tok/s is ~145 for Qwen3.5-9B — gpu-light is the fastest endpoint. Route vision/web/light work there first. NOTE: Qwen3.5-9B is multimodal (image+text), NOT text-only like gemma-4-12b was.
|
||||
- Strix Halo is 62.9 tok/s (89% of 70.5 baseline) — below optimal but stable. Check for competing workloads.
|
||||
- RTX 3090 VRAM at 88% — within role-appropriate range (role = heavy reasoning, needs the headroom).
|
||||
- RTX 5070 VRAM at 83% — role-appropriate for vision/web (smaller batch sizes).
|
||||
@@ -156,7 +158,7 @@ Key notes:
|
||||
- **Detect**: GPU roles misaligned with hardware capabilities
|
||||
- **Target distribution**:
|
||||
- RTX 3090 (gpu-dense, 24GB, 74.9 tok/s) → Heavy reasoning, code gen, long conversations (slowest per-token but largest context capacity). Weight: 0.55 (LiteLLM).
|
||||
- RTX 5070 (gpu-light, 12GB, 169.6 tok/s) → Vision/image, web search, lightweight tasks (2.3x faster than 3090 per token). Weight: 0.15 (LiteLLM).
|
||||
- RTX 5070 (gpu-light, 12GB, ~145 tok/s) → Vision (image+text), web search, lightweight tasks. Weight: 0.15 (LiteLLM).
|
||||
- Strix Halo (strix-moe, 64GB, 62.9 tok/s) → Context compression, summarization, long docs (MoE model). Weight: 0.30 (LiteLLM).
|
||||
- **Note**: RTX 5070 is the fastest endpoint per token. Route high-volume, low-complexity work there first.
|
||||
- **Fix**:
|
||||
@@ -166,6 +168,7 @@ Key notes:
|
||||
- **Verify**: Each GPU's request pattern matches its designated role within 24h
|
||||
- **Escalate**: If role mismatch persists >48h → agent alias audit needed
|
||||
|
||||
---
|
||||
---
|
||||
|
||||
## Execution
|
||||
@@ -247,6 +250,7 @@ call update-gpu-health
|
||||
}
|
||||
```
|
||||
|
||||
---
|
||||
---
|
||||
|
||||
## Reporting
|
||||
@@ -263,6 +267,7 @@ Pushed to `SyslogSolution/health-logs/gpu/{run_id}.json` — versioned, searchab
|
||||
- Per-GPU tok/s trend over 7 days
|
||||
- Regression alerts if any GPU degrades >10% week-over-week
|
||||
|
||||
---
|
||||
---
|
||||
|
||||
## Design Decisions (Verified 2026-07-12, Reaffirmed 2026-07-18)
|
||||
|
||||
Reference in New Issue
Block a user