Files
prose-contracts/gpu-self-heal.prose.md
T
root 226f2ad55e
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 9s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 10s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s
fix: reconcile the health-logs self-heal contracts, and add a dead-man's-switch
relay-785 / self-heal-contract-stale-script-pointers-20260924. gpu/ had been
silent since 2026-09-14T18:02:03Z (12 days) and pm2/ had never posted a run.

FINDING - the gpu-self-heal executor was never lost, only its schedule was.
The brief concluded the mechanism was gone. It is not: on CT 116
/opt/inference-harness/scripts/gpu-self-heal.py exists (mtime 2026-09-21 11:43),
is self-posting via gitea-logger.sh, and runs cleanly - I ran it and it posted
health-logs/gpu/gpu-self-heal-20260926-144618.json. What was missing was the
cron entry, dropped during a CT 116 /etc/cron.d rework on 2026-09-21. So this is
a schedule restoration, not a resurrection.

DECISIONS
* gpu-self-heal -> RESTORE. Schedule re-added as /etc/cron.d/gpu-self-heal
   on CT 116, matching the observed historical cadence (:02 past
  0/6/12/18). Proven with a REAL unattended cron fire (temporary */1 entry,
  removed after): gpu-self-heal-20260926-144802.json committed 14:48:03Z.
  The contract now names the executor, schedule, log and posting, which it
  previously did not.
* pm2 health-logs posting -> RETIRE the claim. pm2-self-heal.prose.md required
  appending to health-logs/pm2/ as a 'hard rule', but scripts/pm2-self-heal.sh
  contains no Gitea or push code and never did; the directory has held only its
  init commit since 2026-07-28. Corrected to point at the contract-runner's
  durable per-run logs and failure note instead of adding a second, redundant
  posting path.

DEAD-MAN'S-SWITCH - scripts/health-log-freshness.py + contract. Absence of logs
must raise an alarm, and the producer cannot raise it, so this runs on CT 100,
a different host from the producers, and fails when the newest health-logs/gpu
entry is older than 12h (litellm 18h). Verified it would have caught the real
gap: evaluated at 2026-09-20 the newest gpu/ entry was 126h old against a 12h
limit -> STALE.

Bugs found and fixed while testing, each caught by a test that bit:
* the documented HEALTH_LOG_MAX_AGE_* override was never implemented;
* a Gitea password held under GITEA_TOKEN was sent as an API token -> HTTP 401;
  now every candidate auth is tried and the first that works is used;
* the ~/.git-credentials fallback filtered on GITEA_URL's host, which missed
  whenever GITEA_URL pointed at the internal IP -> zero candidates.

Exit codes: 0 fresh, 1 stale-or-unreadable, 2 cannot run. A directory that
cannot be read is a failure, never a skip.

Evidence: live PASS; stale override names the directory and producer; no
credential -> exit 2; whole thing runs green through contract-run.sh.
prose-lint: PASSED.
2026-09-26 14:50:52 +00:00

16 KiB

report_only_agents, kind, name, description, agent, depends_on
report_only_agents kind name description agent depends_on
koby
responsibility gpu-self-heal GPU fleet self-healing — detects anomalies, applies remediation, tracks benchmarks, and predicts failures before they happen. Extends gpu-monitor (v2.1.0) with active remediation rules, Prometheus metrics consumption, VRAM trend analysis, and predictive alerting. UPDATED 2026-07-18: Model assignments synced to 2026-07-17 swaps. Router (port 9000) DECOMMISSIONED 2026-09-11; references replaced with direct GPU routing. Benchmark baselines refreshed to live values. Prometheus exporters removed — not deployed; fall back to direct sidecar probes. Stable role-based aliases (strix-moe, gpu-dense, gpu-vision) from gpu-fleet. abiba
gpu-monitor.prose.md (live data source on .24:9100)
gpu-fleet.prose.md (source of truth for topology, aliases, model assignments)

Maintains

  • gpu-health: { status: "healthy"|"degraded"|"down", issues: array, actions: array }
  • gpu-self-heal-log: array of { timestamp, gpu, issue, action, result } — audit trail
  • benchmark-regression: { gpu, baseline_tok_sec, current_tok_sec, trend, alerts }
  • vram-trend: { gpu, current_mb, rate_mb_per_hour, projected_full_in_hours }
  • circuit-breaker-status: { gpu, open, auto_reset_attempted, last_reset }

Requires

  • gpu-monitor:function — Live fleet data from localhost:9100/gpu-data
  • Direct sidecar probe access to all GPU hosts (:8080/health)
  • SSH access to GPU hosts for restart operations

Continuity

  • Self-driven: check every 60 seconds against GPU monitor data
  • Also wakes on gpu-fleet health degradation
  • On fix: verify with benchmark inference test before declaring resolved
  • Escalate: after 3 failed remediation attempts → Zulip #agent-hub alert


Current Fleet Baseline (2026-07-18)

Alias GPU Host Model VRAM Ctx tok/s Role
strix-moe Strix Halo 64GB ct15 (.15:8080) Carnice-Qwen3.6-MoE-35B-A3B-Q4_K_M.gguf ~10/64GB (16%) 256K 62.9 Compression, summarization, long docs

Key notes:

  • All models use direct GPU routing via LiteLLM (api_key: not-needed). Router (port 9000) was decommissioned 2026-09-11 and is NOT in the inference path.
  • Stable aliases (gpu-dense, gpu-vision, strix-moe) from gpu-fleet are the canonical names for agent configs — use these, not model-specific names. The retired names gpu-light and gemma-4-12b were superseded by gpu-vision on 2026-09-12 and no longer resolve (400 Invalid model name).
  • The RTX 5070 is the fastest endpoint per token — gpu-vision is its canonical alias. Route vision/web/light work there first. The RTX 5070 model is multimodal (image+text).
  • Strix Halo is 62.9 tok/s (89% of 70.5 baseline) — below optimal but stable. Check for competing workloads.
  • RTX 3090 VRAM at 88% — within role-appropriate range (role = heavy reasoning, needs the headroom).
  • RTX 5070 VRAM at 83% — role-appropriate for vision/web (smaller batch sizes).

Remediation Rules

Rule 1: GPU Temperature Critical (>85°C for >2 min)

  • Detect: Any GPU temp >85°C sustained for 2+ consecutive polls
  • Fix:
    1. Reduce inference concurrency on that GPU (load-side cooling only — NO fan control)
    2. Redirect new requests to cooler GPUs via LiteLLM fallback chains (gpu-vision → gpu-dense, gpu-dense → gpu-vision)
    3. If all GPUs hot, alert about cooling infrastructure
  • Verify: Temp drops below 80°C within 5 minutes
  • Escalate after: 3 verification failures → Zulip alert

Rule 2: VRAM Leak Detection (tiered by GPU capacity)

  • Detect: VRAM growing at sustained rate over 6+ hour window
    • RTX 3090 (24GB): ≥300MB/hour
    • RTX 5070 (12GB): ≥300MB/hour
    • Strix Halo (64GB UMA): ≥200MB/hour
  • Fix:
    1. Log VRAM snapshot with process list (nvidia-smi/rocm-smi + ps aux)
    2. If llama-server is the growth source → restart with memory cap flag
    3. If unknown process → kill and alert
  • Verify: VRAM growth rate drops below threshold
  • Escalate after: persistent leak after restart → hardware investigation

Rule 3: Model Inference Timeout / GPU Stuck

  • Detect: >50% failure rate over 60s window + 30s grace period (not just single stuck request)
  • Fix:
    1. Restart llama-server on affected GPU host
    2. Wait 15s for model to reload
    3. Run benchmark inference test
  • Verify: Model returns 200 with <30s response, failure rate drops to 0%
  • Escalate after: 3 restarts in 1 hour → GPU hardware check

Rule 4: Benchmark Regression (>20% drop)

  • Detect: gen_tok_per_sec drops >20% below baseline over 3+ benchmarks
    • RTX 3090 baseline: 74.8 tok/s → alert at <59.8 tok/s
    • RTX 5070 baseline: 165.2 tok/s → alert at <132.2 tok/s
    • Strix Halo baseline: 70.5 tok/s → alert at <56.4 tok/s
  • Fix:
    1. Check GPU utilization — if >90%, other process is competing
    2. Check power limit — if throttled, restore to max
    3. Check thermal — if hot, apply Rule 1
  • Verify: Benchmark returns to within 10% of baseline
  • Escalate after: persistent regression → possible hardware degradation

Rule 5: Circuit Breaker Stuck Open

  • Detect: Circuit breaker open >10 minutes with GPU reporting healthy
  • Note: Router (port 9000) was decommissioned 2026-09-11. If circuit breakers are reported by gpu-monitor, they come from LiteLLM's internal tracking, not the old router.
  • Fix:
    1. Verify GPU /health returns 200 on direct port (:8080)
    2. If GPU healthy, alert but do NOT reset via router API (decommissioned 2026-09-11)
    3. Check LiteLLM health directly: http://192.168.68.116/litellm/health/liveliness
    4. Restart LiteLLM container on CT 116 if circuit breakers are stuck
  • Verify: LiteLLM returns healthy, circuit breaker clears within 60s
  • Escalate after: LiteLLM restart doesn't clear → human investigation

Rule 6: Strix Halo Unreachable

  • Detect: Strix not responding — probe .15:8080 directly (firewall opened .24→.15)
  • Fix:
    1. SSH to .15 → check llama-server process
    2. Restart llama-server if not running
    3. Verify through both direct probe AND LiteLLM health
  • Verify: Direct health probe returns 200, LiteLLM reports model healthy
  • Escalate: If host .15 itself is unreachable → infrastructure alert

Rule 7: GPU Data Source Unreachable (replaces old Prometheus rule)

  • Detect: gpu-monitor endpoint (localhost:9100/gpu-data) or sidecar port (:8080) on any GPU unreachable for >2 polls
  • Fix:
    1. If gpu-monitor is down: restart systemd service gpu-monitor.service on this host
    2. If sidecar is down: SSH to GPU host → check llama-server process → restart systemd service
    3. Fall back to direct nvidia-smi/rocm-smi probe via SSH if all API paths fail
  • Verify: gpu-monitor returns healthy + all sidecars reachable
  • Escalate after: 3 failed restarts → networking issue

Rule 8: Predictive Thermal Warning (two-tier)

  • Detect:
    • Tier 1 (warning): temp >70°C AND rising >2°C/min → reduce concurrency, no alert
    • Tier 2 (critical): temp >80°C AND still rising → full alert + aggressive load shedding
  • Fix:
    • Tier 1: silently reduce parallel requests to that GPU by 50%
    • Tier 2: redirect all new requests away, alert #agent-hub, apply Rule 1 logic
  • Verify: Temp rise rate drops below 1°C/min (Tier 1) or temp drops below 80°C (Tier 2)
  • Escalate: If Tier 2 triggers and temp still rising after 5 min → possible hardware failure

Rule 9: Context Window Optimization

  • Detect: Benchmark tok/s vs baseline for each GPU at its current context
    • RTX 3090 (128K ctx, ThinkingCap): baseline 74.8 tok/s — currently at 74.9 (100%)
    • RTX 5070 (128K ctx, HauhauCS QAT): baseline 165.2 tok/s — currently at 169.6 (103%)
    • Strix Halo (256K ctx, Carnice-Qwen3.6-MoE-35B-A3B-Q4_K_M.gguf): baseline 70.5 tok/s — currently at 62.9 (89%)
  • Fix:
    • If tok/s > baseline → context has headroom, consider increasing
    • If tok/s < 90% baseline → reduce context by 25% and retest
    • If tok/s within 10% of baseline → optimal, no change
    • Strix Halo at 89% of baseline → MONITOR but do not reduce yet (recent model swap may still be settling)
  • Verify: Re-benchmark after context change, confirm within 10% of target
  • Escalate: If context can't be adjusted without significant perf loss

Rule 10: Workload Distribution Optimization (updated 2026-07-18)

  • Detect: GPU roles misaligned with hardware capabilities
  • Target distribution:
    • RTX 3090 (gpu-dense, 24GB, 74.9 tok/s) → Heavy reasoning, code gen, long conversations (slowest per-token but largest context capacity).
    • RTX 5070 (gpu-vision, 12GB, ~145 tok/s) → Vision (image+text), web search, lightweight tasks.
    • Strix Halo (strix-moe, 64GB, 62.9 tok/s) → Context compression, summarization, long docs (MoE model).
  • Note: RTX 5070 is the fastest endpoint per token. Route high-volume, low-complexity work there first.
  • Weights are not restated here — the live syslog-auto pool weights and rpm caps live in CT 116 /opt/inference-harness/litellm_config.yaml, the single source of truth.
  • Fix:
    • Alert if any GPU is handling workload outside its designated role
    • Recommend agent alias updates to match workload to GPU role (use stable aliases: gpu-dense, gpu-vision, strix-moe)
    • Track per-GPU request distribution via LiteLLM spend logs
  • Verify: Each GPU's request pattern matches its designated role within 24h
  • Escalate: If role mismatch persists >48h → agent alias audit needed


Execution

Executor, schedule and dead-man's-switch

This contract is implemented by a real script, scheduled on the inference host:

Executor /opt/inference-harness/scripts/gpu-self-heal.py on CT 116
Schedule /etc/cron.d/gpu-self-heal on CT 116 — 2 */6 * * *
Log /var/log/litellm/gpu-self-heal.log
Posting the script calls gitea-logger.sh gpu {RUN_ID}.json <report> → SyslogSolution/health-logs/gpu/{RUN_ID}.json

Dead-man's-switch: absence of logs must raise an alarm, because that is how this went silent for 12 days. That alarm cannot live on the producer — a stopped job cannot report that it stopped — so it lives off-host as the health-log-freshness contract on CT 100, which fails when health-logs/gpu/ is older than 12 h. A failed run is also visible in the log above, but only the off-host check catches a missing run.

History (2026-09-26, relay-785): the executor was never lost — only its schedule was, dropped during a CT 116 /etc/cron.d rework on 2026-09-21. The gpu/ log was silent from 2026-09-14T18:02:03Z. The schedule was restored and the off-host freshness check added; do not treat either as optional.

-- Phase 1: Fetch live GPU data
let fleet = call gpu-monitor
  endpoint: "http://localhost:9100/gpu-data"

-- Phase 2: Evaluate each GPU against remediation rules
let actions = []
for gpu in fleet.gpus:
  -- Rule 1: Thermal critical
  if gpu.temp_c > 85 and sustained_for(gpu, 120):
    push actions apply-thermal-fix(gpu)
  
  -- Rule 2: VRAM leak
  let vram_rate = calculate-vram-trend(gpu, hours=6)
  if vram_rate > 50:
    push actions apply-vram-fix(gpu, vram_rate)
  
  -- Rule 4: Benchmark regression
  let bench = fleet.benchmarks[gpu.hostname]
  if bench.current_tok_s < bench.baseline_tok_s * 0.8:
    push actions apply-benchmark-fix(gpu, bench)

-- Rule 3: Model stuck
for model in fleet.summary.available_models:
  if model.consecutive_timeouts >= 3:
    push actions apply-model-restart(model)

-- Rule 5: Circuit breaker check via LiteLLM (router deprecated)
if fleet.summary.circuit_breakers_open > 0:
  push actions check-litellm-circuit-breakers()

-- Rule 6: Strix Halo
if not fleet.strix.running and pingable("192.168.68.15"):
  push actions apply-strix-restart()

-- Rule 7: GPU data source
if not fleet.gpus or len(fleet.gpus) < 2:
  push actions check-gpu-monitor-service()

-- Rule 8: Predictive thermal
for gpu in fleet.gpus:
  let rise_rate = calculate-temp-rise(gpu, minutes=5)
  if rise_rate > 2.0 and gpu.temp_c < 80:
    push actions apply-proactive-cooling(gpu)

-- Phase 3: Execute actions, verify, log
for action in actions:
  let result = execute-with-verify(action)
  log-to-kg(action, result)
  if result.failed:
    escalate-if-needed(action)

-- Phase 4: Update health state
call update-gpu-health
  gpus: fleet.gpus
  actions: actions
  status: derive-overall-status(fleet, actions)

-- Wait 60s and repeat

Audit Trail Format

{
  "run_id": "gpu-self-heal-20260718-001",
  "timestamp": "2026-07-18T08:00:00Z",
  "gpu": "ct8-rtx3090",
  "issue": "thermal-critical",
  "detected": { "temp_c": 87, "duration_s": 180 },
  "action": "load-shedding",
  "result": "resolved",
  "verification": { "temp_c": 76, "after_s": 300 },
  "escalated": false
}


Reporting

1. Gitea Log (not knowledge graph — hard rule)

Pushed to SyslogSolution/health-logs/gpu/{run_id}.json — versioned, searchable, not in graph.

2. Zulip Alerts (#agent-hub → alerts-gpu)

  • issues_fixed > 0 → "🛠 GPU Self-Heal — resolved"
  • issues_escalated > 0 → "⚠ GPU Self-Heal — needs attention"
  • Every 100th clean cycle → "✅ GPU Fleet: All Clear"

3. Weekly Benchmark Report

  • Per-GPU tok/s trend over 7 days
  • Regression alerts if any GPU degrades >10% week-over-week


Design Decisions (Verified 2026-07-12, Reaffirmed 2026-07-18)

  1. Fan control: ❌ NO auto fan control. Load-side cooling only (reduce concurrency, redirect).
  2. Model restart: ✅ Only if >50% failure rate over 60s + 30s grace period. Not on single stuck request.
  3. Strix direct access: ✅ Open firewall .15:8080 → .24 for direct health probe + restart.
  4. VRAM thresholds: Tiered — 300MB/h (RTX 3090), 300MB/h (RTX 5070), 200MB/h (Strix). Previous values (100/50) were too sensitive; raised 2026-07-18 based on operational data.
  5. CB auto-reset: ✅ Router decommissioned 2026-09-11 — circuit breakers go through LiteLLM health check + container restart if needed. No per-GPU auto-reset.
  6. Benchmark baseline: Rolling 30-day average, recalculated weekly. Current baselines live in gpu-monitor.
  7. Predictive alerts: Two-tier — warn at >70°C+rising (>2°C/min), critical at >80°C+rising.
  8. Prometheus: ❌ Not deployed. Use direct sidecar probes (:8080/health) and gpu-monitor API. Prometheus integration deferred until exporters are running on GPU hosts.

Lessons Learned (2026-07-12, Updated 2026-07-18)

L1: API Key Standardization Is Critical

  • All GPU llama-servers MUST use the same api-key as the LiteLLM config.
  • RTX 5070 had --api-key sk-loc...5678 while LiteLLM sent not-needed. This caused cascading 401 → fallback → timeout → 401 loops.
  • Rule: Any new GPU or model restart MUST verify api-key matches LiteLLM config (not-needed for direct routing).

L2: Fallback Chain Cascading Failures

  • When one model returns 401 (auth) and another is slow (timeout), the fallback chain creates an infinite loop.
  • Rule: If a model returns 401 (auth error), do NOT fall back to it again. Mark it as permanently failed for this request.

L3: Verify Running State, Not Docs

  • RTX 3090 was documented at 128K context. Running at 128K (verified 2026-07-18).
  • Parallel count: 1 on both RTX 3090 and RTX 5070 (matches docs for current models).
  • Rule: Before making decisions, check /proc/PID/cmdline on GPU hosts.

L4: Infisical Is Not Always Available

  • Keep a local .env fallback for LITELLM_API_KEY.
  • Rule: Always verify credential source is reachable before relying on it.

L5: GPU Monitor Response Size Can Cause Self-Heal Crash

  • gpu-self-heal crashed with KeyboardInterrupt during json.loads() of 20MB response.
  • Root cause: router poll returns accumulated data → cache balloons.
  • Rule: Self-heal must enforce a read timeout AND max response size on every poll. If monitor response > 1MB, log a warning and skip the cycle rather than crashing.

L6: Stable Aliases Replace Model Names

  • gpu-fleet introduced stable aliases (strix-moe, gpu-dense, gpu-light) on 2026-07-15; gpu-light was superseded by gpu-vision on 2026-09-12.
  • Self-heal must use aliases for reporting and alerting, not model-specific names.
  • Rule: All alert messages and KG nodes use the stable alias as the GPU identifier.