- hermes-config-template.prose.md: Add Rule 17 — Koby is never repaired, full stop
- hermes-agent-baseline.prose.md: Document Koby report-only posture
- contract-registry.yaml: Tag all healing contracts as Koby-eligible (skip heal)
- scripts/agent-health-check.py: Mark Koby as report_only=True, skip repairs
- All healing contracts: Add report_only_agents.koby marker
Captain-approved ship via no-mistakes. PR auto-merges green.
GPU fleet self-healing — detects anomalies, applies remediation, tracks benchmarks, and predicts failures before they happen. Extends gpu-monitor (v2.1.0) with active remediation rules, Prometheus metrics consumption, VRAM trend analysis, and predictive alerting. UPDATED 2026-07-18: Model assignments synced to 2026-07-17 swaps. Router (port 9000) references replaced with direct GPU routing. Benchmark baselines refreshed to live values. Prometheus exporters removed — not deployed; fall back to direct sidecar probes. Stable role-based aliases (strix-moe, gpu-dense, gpu-light) from gpu-fleet.
abiba
gpu-monitor.prose.md (live data source on .24:9100)
gpu-fleet.prose.md (source of truth for topology, aliases, model assignments)
gpu-monitor:function — Live fleet data from localhost:9100/gpu-data
Direct sidecar probe access to all GPU hosts (:8080/health)
SSH access to GPU hosts for restart operations
Continuity
Self-driven: check every 60 seconds against GPU monitor data
Also wakes on gpu-fleet health degradation
On fix: verify with benchmark inference test before declaring resolved
Escalate: after 3 failed remediation attempts → Zulip #agent-hub alert
Current Fleet Baseline (2026-07-18)
Alias
GPU
Host
Model
VRAM
Ctx
tok/s
Role
gpu-dense
RTX 3090 24GB
ct8 (.8:8080)
ThinkingCap Qwen3.6-27B Q4_K_M + MTP + vision
21.6/24.6GB (88%)
128K
74.9
Heavy reasoning, code gen
gpu-light
RTX 5070 12GB
ct110 (.110:8080)
HauhauCS Gemma4-12B QAT Q4_K_M + MTP draft
10.1/12.2GB (83%)
128K
169.6
Vision, web extract, light tasks
strix-moe
Strix Halo 64GB
ct15 (.15:8080)
qwen3.6-35B-udq4
~10/64GB (16%)
128K
62.9
Compression, summarization, long docs
Key notes:
All models use direct GPU routing via LiteLLM (api_key: not-needed). Router (port 9000) is deprecated and NOT in the inference path.
Stable aliases (gpu-dense, gpu-light, strix-moe) from gpu-fleet are the canonical names for agent configs. Model-specific names still work but are deprecated.
RTX 5070 tok/s is 2.3x faster than RTX 3090 for its model — gpu-light is the fastest endpoint. Route vision/web/light work there first.
Strix Halo is 62.9 tok/s (89% of 70.5 baseline) — below optimal but stable. Check for competing workloads.
RTX 3090 VRAM at 88% — within role-appropriate range (role = heavy reasoning, needs the headroom).
RTX 5070 VRAM at 83% — role-appropriate for vision/web (smaller batch sizes).
Remediation Rules
Rule 1: GPU Temperature Critical (>85°C for >2 min)
Detect: Any GPU temp >85°C sustained for 2+ consecutive polls
Fix:
Reduce inference concurrency on that GPU (load-side cooling only — NO fan control)
Redirect new requests to cooler GPUs via LiteLLM fallback chains (gemma → qwen, qwen → gemma)
If all GPUs hot, alert about cooling infrastructure
Detect: gen_tok_per_sec drops >20% below baseline over 3+ benchmarks
RTX 3090 baseline: 74.8 tok/s → alert at <59.8 tok/s
RTX 5070 baseline: 165.2 tok/s → alert at <132.2 tok/s
Strix Halo baseline: 70.5 tok/s → alert at <56.4 tok/s
Fix:
Check GPU utilization — if >90%, other process is competing
Check power limit — if throttled, restore to max
Check thermal — if hot, apply Rule 1
Verify: Benchmark returns to within 10% of baseline
Escalate after: persistent regression → possible hardware degradation
Rule 5: Circuit Breaker Stuck Open
Detect: Circuit breaker open >10 minutes with GPU reporting healthy
Note: Router (port 9000) is deprecated. If circuit breakers are reported by gpu-monitor, they come from LiteLLM's internal tracking, not the old router.
Fix:
Verify GPU /health returns 200 on direct port (:8080)
If GPU healthy, alert but do NOT reset via router API (deprecated)
Fan control: ❌ NO auto fan control. Load-side cooling only (reduce concurrency, redirect).
Model restart: ✅ Only if >50% failure rate over 60s + 30s grace period. Not on single stuck request.
Strix direct access: ✅ Open firewall .15:8080 → .24 for direct health probe + restart.
VRAM thresholds: Tiered — 300MB/h (RTX 3090), 300MB/h (RTX 5070), 200MB/h (Strix). Previous values (100/50) were too sensitive; raised 2026-07-18 based on operational data.
CB auto-reset: ✅ Router deprecated — circuit breakers go through LiteLLM health check + container restart if needed. No per-GPU auto-reset.
Benchmark baseline: Rolling 30-day average, recalculated weekly. Current baselines live in gpu-monitor.
Predictive alerts: Two-tier — warn at >70°C+rising (>2°C/min), critical at >80°C+rising.
Prometheus: ❌ Not deployed. Use direct sidecar probes (:8080/health) and gpu-monitor API. Prometheus integration deferred until exporters are running on GPU hosts.
Lessons Learned (2026-07-12, Updated 2026-07-18)
L1: API Key Standardization Is Critical
All GPU llama-servers MUST use the same api-key as the LiteLLM config.
RTX 5070 had --api-key sk-loc...5678 while LiteLLM sent not-needed.
This caused cascading 401 → fallback → timeout → 401 loops.
Rule: Any new GPU or model restart MUST verify api-key matches LiteLLM config (not-needed for direct routing).
L2: Fallback Chain Cascading Failures
When one model returns 401 (auth) and another is slow (timeout), the fallback
chain creates an infinite loop.
Rule: If a model returns 401 (auth error), do NOT fall back to it again.
Mark it as permanently failed for this request.
L3: Verify Running State, Not Docs
RTX 3090 was documented at 128K context. Running at 128K (verified 2026-07-18).
Parallel count: 1 on both RTX 3090 and RTX 5070 (matches docs for current models).
Rule: Before making decisions, check /proc/PID/cmdline on GPU hosts.
L4: Infisical Is Not Always Available
Keep a local .env fallback for LITELLM_API_KEY.
Rule: Always verify credential source is reachable before relying on it.
L5: GPU Monitor Response Size Can Cause Self-Heal Crash
gpu-self-heal crashed with KeyboardInterrupt during json.loads() of 20MB response.
Root cause: router poll returns accumulated data → cache balloons.
Rule: Self-heal must enforce a read timeout AND max response size on every poll.
If monitor response > 1MB, log a warning and skip the cycle rather than crashing.
L6: Stable Aliases Replace Model Names
gpu-fleet introduced stable aliases (strix-moe, gpu-dense, gpu-light) on 2026-07-15.
Self-heal must use aliases for reporting and alerting, not model-specific names.
Rule: All alert messages and KG nodes use the stable alias as the GPU identifier.