Compare commits

..
Author SHA1 Message Date
root aebc98ead6 fix: remove trailing whitespace from litellm-api-keys.prose.md
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 2s
2026-07-18 17:48:52 +00:00
root 17a77e6b3f fix: fleet config issues from 2026-07-18 relay review
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
- hermes-agent-baseline: add gpu-dense and gpu-light to models list
- hermes-config-template: fix pgrep traps (exclude infisical wrapper) + add
  vault empty-key guard documentation in Rule 13
- litellm-api-keys: fix pgrep pattern in auditable check
- scripts/agent-health-check: fix pgrep to exclude infisical bash wrapper

Addresses issues found during relay inbox resolution session:
1. pgrep -f 'hermes_cli.main gateway run' matches both the real python
   gateway and the infisical bash wrapper, causing false health readings
2. infisical vault stores empty key silently — no guard/monitoring
3. gpu-dense/gpu-light stable aliases missing from baseline config
2026-07-18 17:39:46 +00:00
7 changed files with 106 additions and 194 deletions
+14 -26
View File
@@ -9,9 +9,9 @@ description: >
gpu-dense, gpu-light. These never change — only the underlying model does. gpu-dense, gpu-light. These never change — only the underlying model does.
Strix Halo: strix-moe → unsloth/Qwen3.6-35B-A3B-MTP (UD-Q4_K_M, 22GB). Strix Halo: strix-moe → unsloth/Qwen3.6-35B-A3B-MTP (UD-Q4_K_M, 22GB).
RTX 5070: gemma-4-12b Q4_K_M → IQ4_NL + MTP draft (122 tok/s, 2x faster). RTX 5070: gemma-4-12b Q4_K_M → IQ4_NL + MTP draft (122 tok/s, 2x faster).
UPDATED 2026-07-17: Strix Halo model swapped to Genesis Hermes V3 APEX (LuffyTheFox, 24GB, uncensored, UPDATED 2026-07-17: Context reduced fleet-wide from 256K to 128K for stability.
Strix Halo model swapped to Genesis Hermes V3 APEX (LuffyTheFox, 24GB, uncensored,
Hermes agent fine-tune, tensor repair, multimodal with mmproj). Hermes agent fine-tune, tensor repair, multimodal with mmproj).
RTX 5070 swapped to HauhauCS Gemma4-12B QAT Uncensored Balanced (Q4_K_M, 87 tok/s, 0/465 refusals).
Instability observed near 100K at 256K. 128K is the stable ceiling. Instability observed near 100K at 256K. 128K is the stable ceiling.
For larger context needs → fall back to external providers (deepseek). For larger context needs → fall back to external providers (deepseek).
VRAM headroom improved: RTX 3090 ~70%, RTX 5070 ~65%. VRAM headroom improved: RTX 3090 ~70%, RTX 5070 ~65%.
@@ -87,32 +87,20 @@ When a model is swapped on a GPU, ONLY the infrastructure layer changes — agen
| Alias | GPU | Current Model | Will Route To | | Alias | GPU | Current Model | Will Route To |
|-------|-----|---------------|---------------| |-------|-----|---------------|---------------|
| `strix-moe` | Strix Halo (.15) | qwen3.6-35B-udq4 | Whatever runs on Strix Halo | | `strix-moe` | Strix Halo (.15) | qwen3.6-35B-udq4 | Whatever runs on Strix Halo |
| `gpu-dense` | RTX 3090 (.8) | qwen3.6-27B-code (ThinkingCap) | Whatever runs on RTX 3090 | | `gpu-dense` | RTX 3090 (.8) | qwen3.6-27B-code | Whatever runs on RTX 3090 |
| `gpu-light` | RTX 5070 (.110) | gemma-4-12b | Whatever runs on RTX 5070 | | `gpu-light` | RTX 5070 (.110) | gemma-4-12b | Whatever runs on RTX 5070 |
**Backward compatibility**: Old model-specific names (qwen3.6-27B-code, gemma-4-12b, qwen3.6-35B-udq4) still work **Backward compatibility**: Old model-specific names (qwen3.6-27B-code, gemma-4-12b, qwen3.6-35B-udq4) still work
but are deprecated for agent configs. Only the stable aliases survive model swaps. but are deprecated for agent configs. Only the stable aliases survive model swaps.
## Current Model Assignments (2026-07-17) ## Current Model Assignments (2026-07-15)
| Model | GPU | Host | VRAM | Ctx | KV Cache | Parallel | Batch/Ubatch | Status | | Model | GPU | Host | VRAM | Ctx | KV Cache | Parallel | Batch/Ubatch | Status |
|-------|-----|------|------|-----|----------|----------|-------------|--------| |-------|-----|------|------|-----|----------|----------|-------------|--------|
| qwen3.6-27B-code (ThinkingCap) | RTX 3090 | .8 (llm-gpu) | ~20.9/24.6GB (85%) | **128K** | turbo4 | 1 | default | ✅ 68 tok/s | | qwen3.6-27B-code (MTP) | RTX 3090 | .8 (llm-gpu) | ~17/24.6GB (70%) | **128K** | turbo4 | 2 | default | ✅ 63 tok/s |
| gemma-4-12b (HauhauCS QAT) | RTX 5070 | .110 (ocu-llm) | ~10.0/12.2GB (82%) | 128K | q4_0 | 1 | 2048/1024 | ✅ 87 tok/s | | gemma-4-12b | RTX 5070 | .110 (ocu-llm) | ~7.8/12.2GB (65%) | 128K | q4_0 | 2 | 2048/1024 | ✅ healthy |
| Genesis Hermes V3 APEX | Strix Halo Vulkan | .15 (amdpve) | ~10GB/64GB | 128K | q4_0 | 1 | 4096/1024 | ✅ 65 tok/s | | Genesis Hermes V3 APEX | Strix Halo Vulkan | .15 (amdpve) | ~10GB/64GB | 128K | q4_0 | 1 | 4096/1024 | ✅ 65 tok/s |
> **RTX 5070 model swap (2026-07-17)**: Switched from `gemma-4-12b-it-IQ4_NL` (Unsloth, 191 tok/s)
> to `HauhauCS/Gemma4-12B-QAT-Uncensored-HauhauCS-Balanced` (Q4_K_M QAT, 87 tok/s).
> Trade: 54% slower generation for QAT quality, 0/465 refusals, and agent-optimized tuning.
> MTP draft also swapped: Q8_0 (444MB) → tuned draft (242MB), saving 200MB VRAM.
> Role unchanged: gpu-light (vision, web extract, light auxiliary tasks).
> **RTX 3090 model swap (2026-07-17)**: Switched from `Qwopus3.6-27B-v2-MTP-Q4_K_M` (63 tok/s)
> to `bottlecapai/ThinkingCap-Qwen3.6-27B` (Q4_K_M QAT, 68 tok/s).
> RL-finetuned: 50% fewer thinking tokens, MMLU-Pro 0.85 vs 0.83 base.
> Self-spec MTP REQUIRED (crashes without it on turboquant build).
> Added vision via mmproj (0.9GB). VRAM 85%.
## Routing Configuration (LiteLLM — July 2026) ## Routing Configuration (LiteLLM — July 2026)
### syslog-auto Weighted Pool (Direct GPU — bypasses router) ### syslog-auto Weighted Pool (Direct GPU — bypasses router)
@@ -130,16 +118,16 @@ Note: All syslog-auto entries route directly to GPUs with `api_key: not-needed`.
| Model | RPM Cap | Notes | | Model | RPM Cap | Notes |
|-------|---------|-------| |-------|---------|-------|
| strix-moe (Hermes V3) | 40 | Tight cap — prevents Strix overload | | strix-moe (Hermes V3) | 40 | Tight cap — prevents Strix overload |
| qwen3.6-27B-code | 500 | ThinkingCap Q4_K_M + MTP self-spec + vision, 68 tok/s | | qwen3.6-27B-code | 500 | High cap — primary workhorse |
| gemma-4-12b | 500 | HauhauCS QAT Uncensored Balanced + MTP, 87 tok/s | | gemma-4-12b | 500 | High cap — IQ4_NL+MTP, 122 tok/s |
### Stable Aliases (for agent configs — never change) ### Stable Aliases (for agent configs — never change)
| Alias | RPM Cap | Routes To | Purpose | | Alias | RPM Cap | Routes To | Purpose |
|-------|---------|-----------|---------| |-------|---------|-----------|---------|
| `strix-moe` | 40 | Strix Halo | Compression tasks (MoE models) | | `strix-moe` | 40 | Strix Halo | Compression tasks (MoE models) |
| `gpu-dense` | 500 | RTX 3090 (ThinkingCap) | Heavy reasoning, code gen, delegation | | `gpu-dense` | 500 | RTX 3090 | Heavy reasoning |
| `gpu-light` | 500 | RTX 5070 (HauhauCS QAT) | Vision, web extract, light tasks | | `gpu-light` | 500 | RTX 5070 | Vision, web extract, light tasks |
### Fallback Chains ### Fallback Chains
- gemma → qwen - gemma → qwen
@@ -261,8 +249,8 @@ If no SSH access, send Zulip DM via abiba-bot with vault update instructions.
- **LiteLLM /metrics**: Requires auth. Prometheus uses `/health/liveliness` as workaround. - **LiteLLM /metrics**: Requires auth. Prometheus uses `/health/liveliness` as workaround.
- **VRAM (2026-07-15)**: RTX 3090 at ~17/24.6GB (~70%) with **128K context** (reduced from 256K 2026-07-17). RTX 5070 at ~7.8/12.2GB (~65%) with 128K context + MTP. Strix Halo at ~7GB/64GB. - **VRAM (2026-07-15)**: RTX 3090 at ~17/24.6GB (~70%) with **128K context** (reduced from 256K 2026-07-17). RTX 5070 at ~7.8/12.2GB (~65%) with 128K context + MTP. Strix Halo at ~7GB/64GB.
- **RTX 3090 runs `--parallel 2`** with MTP draft (spec-type draft-mtp, spec-draft-n-max 2). - **RTX 3090 runs `--parallel 2`** with MTP draft (spec-type draft-mtp, spec-draft-n-max 2).
- **RTX 3090 config**: `-c 131072 -ctk turbo4 -ctv turbo4 --parallel 1 --flash-attn on --cont-batching --spec-type draft-mtp --spec-draft-n-max 4`. ThinkingCap Qwen3.6-27B Q4_K_M (15.7GB) + mmproj (0.9GB) + MTP self-spec. VRAM: ~85%. Service: `/home/llmuser/llama-wrapper.sh`. ⚠️ MTP REQUIRED for stability on this turboquant build — model segfaults without `--spec-type draft-mtp`. Outputs reasoning_content (hidden from agent, improves answer quality). - **RTX 3090 config**: `-c 131072 -ctk turbo4 -ctv turbo4 --parallel 2 --flash-attn on --cont-batching --spec-type draft-mtp`. Context reduced to 128K (2026-07-17, was 256K). VRAM: ~70%. Service: `/home/llmuser/llama-wrapper.sh`.
- **RTX 5070 config (2026-07-17)**: HauhauCS Gemma4-12B QAT Uncensored Balanced (Q4_K_M) + tuned MTP draft (242MB) at 128K context, single slot. Gen speed: 87 tok/s (vs 191 IQ4_NL). VRAM: ~10.0/12.2GB (~82%). Service: `/home/llmuser/llama-wrapper.sh`. Model: `Gemma4-12B-QAT-Uncensored-HauhauCS-Balanced-Q4_K_M.gguf`, MTP: `mtp-gemma-4-12B-it.gguf`, mmproj: `mmproj-Gemma4-12B-QAT-Uncensored-HauhauCS-Balanced-BF16.gguf`. Recommended sampling: temp 0.6, top_k 64, top_p 0.9, min_p 0.05, repeat_penalty 1.1. - **RTX 5070 config (2026-07-15)**: Switched to IQ4_NL + MTP draft (Q8_0) at 128K context. Gen speed: 122 tok/s. VRAM: ~7.8/12.2GB (~65%). Service: `/home/llmuser/llama-wrapper.sh`. Config: `--model gemma-4-12b-it-IQ4_NL.gguf --spec-draft-model gemma-4-12b-it-Q8_0-MTP.gguf --spec-type draft-mtp --spec-draft-n-max 4 --ctx-size 131072`.
- **LiteLLM timeout tuning (verified 2026-07-16 against `/opt/inference-harness/litellm_config.yaml` on CT 116)**: gemma-4-12b 120s, qwen3.6-27B-code 300s, qwen3.6-35B-udq4 300s, strix-moe 300s, syslog-auto routes all 300s. Nginx proxy_read_timeout: 600s. Global request_timeout: 300s. - **LiteLLM timeout tuning (verified 2026-07-16 against `/opt/inference-harness/litellm_config.yaml` on CT 116)**: gemma-4-12b 120s, qwen3.6-27B-code 300s, qwen3.6-35B-udq4 300s, strix-moe 300s, syslog-auto routes all 300s. Nginx proxy_read_timeout: 600s. Global request_timeout: 300s.
- **Strix Halo GPU**: Vulkan is the working backend (ROCm/HIP path abandoned — HSA runtime blocked on Debian 13). Build at `/root/llama.cpp/build-vk/`, commit `4fc4ec5` (2026-07-01), ggml 0.15.3 shared-lib arch. Mesa RADV 25.0.7, KHR_coopmat fast path active. ~70 tok/s gen, 532 tok/s prompt. Service: `strix-server.service` on port 8080, model: `LuffyTheFox/Qwen3.6-35B-A3B-Uncensored-Genesis-Hermes-V3-GGUF` (APEX quant), alias `strix-moe`, 128K context, flash-attn + q4 KV, multimodal (mmproj loaded). Hermes agent fine-tune, tensor repair (SSM layers fixed via SVD), uncensored (0/465 refusals). - **Strix Halo GPU**: Vulkan is the working backend (ROCm/HIP path abandoned — HSA runtime blocked on Debian 13). Build at `/root/llama.cpp/build-vk/`, commit `4fc4ec5` (2026-07-01), ggml 0.15.3 shared-lib arch. Mesa RADV 25.0.7, KHR_coopmat fast path active. ~70 tok/s gen, 532 tok/s prompt. Service: `strix-server.service` on port 8080, model: `LuffyTheFox/Qwen3.6-35B-A3B-Uncensored-Genesis-Hermes-V3-GGUF` (APEX quant), alias `strix-moe`, 128K context, flash-attn + q4 KV, multimodal (mmproj loaded). Hermes agent fine-tune, tensor repair (SSM layers fixed via SVD), uncensored (0/465 refusals).
- **Port conflict detection (2026-07-05)**: All 3 GPU wrappers now detect ghost processes squatting port 8080 before starting. `.8` and `.110` use inline pre-start check in `llama-wrapper.sh`; `.15` uses `/usr/local/bin/port-cleanup.sh` ExecStartPre. Replaces the blanket `pkill -9 -x llama-server` on .15 which would kill ALL llama-server instances regardless of port. Ghost detection was the root cause of .8 crash-looping for 27+ restarts (stale pid 25836 squatting 8080 after OOM kill). - **Port conflict detection (2026-07-05)**: All 3 GPU wrappers now detect ghost processes squatting port 8080 before starting. `.8` and `.110` use inline pre-start check in `llama-wrapper.sh`; `.15` uses `/usr/local/bin/port-cleanup.sh` ExecStartPre. Replaces the blanket `pkill -9 -x llama-server` on .15 which would kill ALL llama-server instances regardless of port. Ghost detection was the root cause of .8 crash-looping for 27+ restarts (stale pid 25836 squatting 8080 after OOM kill).
@@ -278,8 +266,8 @@ If no SSH access, send Zulip DM via abiba-bot with vault update instructions.
| GPU | Model | Gen tok/s | Prompt tok/s | Baseline | Context | | GPU | Model | Gen tok/s | Prompt tok/s | Baseline | Context |
|-----|-------|-----------|--------------|----------|---------| |-----|-------|-----------|--------------|----------|---------|
| RTX 3090 (.8) | ThinkingCap-Qwen3.6-27B Q4_K_M | **68** | — | — | **128K** | | RTX 3090 (.8) | qwen3.6-27B-code (MTP) | **63** | — | — | **128K** |
| RTX 5070 (.110) | HauhauCS QAT Uncensored Balanced | **87** | — | — | **128K** | | RTX 5070 (.110) | gemma-4-12b (IQ4_NL+MTP) | **191** | — | — | **128K** |
| Strix Halo (.15) | Genesis Hermes V3 APEX | **65** | 140 | — | **128K** | | Strix Halo (.15) | Genesis Hermes V3 APEX | **65** | 140 | — | **128K** |
Benchmarks from 2026-07-17. Strix Halo swapped to Genesis Hermes V3 APEX (LuffyTheFox). RTX 5070 MTP provides 2.7x speedup over pre-upgrade 70 tok/s. Benchmarks from 2026-07-17. Strix Halo swapped to Genesis Hermes V3 APEX (LuffyTheFox). RTX 5070 MTP provides 2.7x speedup over pre-upgrade 70 tok/s.
+69 -93
View File
@@ -6,15 +6,10 @@ description: >
benchmarks, and predicts failures before they happen. Extends gpu-monitor benchmarks, and predicts failures before they happen. Extends gpu-monitor
(v2.1.0) with active remediation rules, Prometheus metrics consumption, (v2.1.0) with active remediation rules, Prometheus metrics consumption,
VRAM trend analysis, and predictive alerting. VRAM trend analysis, and predictive alerting.
UPDATED 2026-07-18: Model assignments synced to 2026-07-17 swaps.
Router (port 9000) references replaced with direct GPU routing.
Benchmark baselines refreshed to live values.
Prometheus exporters removed — not deployed; fall back to direct sidecar probes.
Stable role-based aliases (strix-moe, gpu-dense, gpu-light) from gpu-fleet.
agent: abiba agent: abiba
depends_on: depends_on:
- gpu-monitor.prose.md (live data source on .24:9100) - gpu-monitor.prose.md (live data source on .24:9100)
- gpu-fleet.prose.md (source of truth for topology, aliases, model assignments) - gpu-fleet.prose.md (source of truth for topology)
--- ---
## Maintains ## Maintains
@@ -27,8 +22,8 @@ depends_on:
## Requires ## Requires
- gpu-monitor:function — Live fleet data from localhost:9100/gpu-data - gpu-monitor:function — Live fleet data from .24:9100/gpu-data
- Direct sidecar probe access to all GPU hosts (:8080/health) - Prometheus exporters on all 3 GPUs (:9400/metrics)
- SSH access to GPU hosts for restart operations - SSH access to GPU hosts for restart operations
## Continuity ## Continuity
@@ -40,37 +35,21 @@ depends_on:
--- ---
## Current Fleet Baseline (2026-07-18)
| Alias | GPU | Host | Model | VRAM | Ctx | tok/s | Role |
|-------|-----|------|-------|------|-----|-------|------|
| `gpu-dense` | RTX 3090 24GB | ct8 (.8:8080) | ThinkingCap Qwen3.6-27B Q4_K_M + MTP + vision | 21.6/24.6GB (88%) | 128K | 74.9 | Heavy reasoning, code gen |
| `gpu-light` | RTX 5070 12GB | ct110 (.110:8080) | HauhauCS Gemma4-12B QAT Q4_K_M + MTP draft | 10.1/12.2GB (83%) | 128K | 169.6 | Vision, web extract, light tasks |
| `strix-moe` | Strix Halo 64GB | ct15 (.15:8080) | Genesis Hermes V3 APEX (LuffyTheFox, 24GB) | ~10/64GB (16%) | 128K | 62.9 | Compression, summarization, long docs |
Key notes:
- All models use direct GPU routing via LiteLLM (`api_key: not-needed`). Router (port 9000) is deprecated and NOT in the inference path.
- Stable aliases (gpu-dense, gpu-light, strix-moe) from gpu-fleet are the canonical names for agent configs. Model-specific names still work but are deprecated.
- RTX 5070 tok/s is 2.3x faster than RTX 3090 for its model — gpu-light is the fastest endpoint. Route vision/web/light work there first.
- Strix Halo is 62.9 tok/s (89% of 70.5 baseline) — below optimal but stable. Check for competing workloads.
- RTX 3090 VRAM at 88% — within role-appropriate range (role = heavy reasoning, needs the headroom).
- RTX 5070 VRAM at 83% — role-appropriate for vision/web (smaller batch sizes).
## Remediation Rules ## Remediation Rules
### Rule 1: GPU Temperature Critical (>85°C for >2 min) ### Rule 1: GPU Temperature Critical (>85°C for >2 min)
- **Detect**: Any GPU temp >85°C sustained for 2+ consecutive polls - **Detect**: Any GPU temp >85°C sustained for 2+ consecutive polls
- **Fix**: - **Fix**:
1. Reduce inference concurrency on that GPU (load-side cooling only — NO fan control) 1. Reduce inference concurrency on that GPU (load-side cooling only — NO fan control)
2. Redirect new requests to cooler GPUs via LiteLLM fallback chains (gemma → qwen, qwen → gemma) 2. Redirect new requests to cooler GPUs via LiteLLM fallback chains
3. If all GPUs hot, alert about cooling infrastructure 3. If all GPUs hot, alert about cooling infrastructure
- **Verify**: Temp drops below 80°C within 5 minutes - **Verify**: Temp drops below 80°C within 5 minutes
- **Escalate after**: 3 verification failures → Zulip alert - **Escalate after**: 3 verification failures → Zulip alert
### Rule 2: VRAM Leak Detection (tiered by GPU capacity) ### Rule 2: VRAM Leak Detection (tiered by GPU capacity)
- **Detect**: VRAM growing at sustained rate over 6+ hour window - **Detect**: VRAM growing at sustained rate over 6+ hour window
- RTX 3090 (24GB): ≥300MB/hour - RTX 3090 (24GB): ≥100MB/hour
- RTX 5070 (12GB): ≥300MB/hour - RTX 5070 (12GB): ≥50MB/hour
- Strix Halo (64GB UMA): ≥200MB/hour - Strix Halo (64GB UMA): ≥200MB/hour
- **Fix**: - **Fix**:
1. Log VRAM snapshot with process list (nvidia-smi/rocm-smi + ps aux) 1. Log VRAM snapshot with process list (nvidia-smi/rocm-smi + ps aux)
@@ -90,9 +69,6 @@ Key notes:
### Rule 4: Benchmark Regression (>20% drop) ### Rule 4: Benchmark Regression (>20% drop)
- **Detect**: gen_tok_per_sec drops >20% below baseline over 3+ benchmarks - **Detect**: gen_tok_per_sec drops >20% below baseline over 3+ benchmarks
- RTX 3090 baseline: 74.8 tok/s → alert at <59.8 tok/s
- RTX 5070 baseline: 165.2 tok/s → alert at <132.2 tok/s
- Strix Halo baseline: 70.5 tok/s → alert at <56.4 tok/s
- **Fix**: - **Fix**:
1. Check GPU utilization — if >90%, other process is competing 1. Check GPU utilization — if >90%, other process is competing
2. Check power limit — if throttled, restore to max 2. Check power limit — if throttled, restore to max
@@ -102,31 +78,32 @@ Key notes:
### Rule 5: Circuit Breaker Stuck Open ### Rule 5: Circuit Breaker Stuck Open
- **Detect**: Circuit breaker open >10 minutes with GPU reporting healthy - **Detect**: Circuit breaker open >10 minutes with GPU reporting healthy
- **Note**: Router (port 9000) is deprecated. If circuit breakers are reported by gpu-monitor, they come from LiteLLM's internal tracking, not the old router.
- **Fix**: - **Fix**:
1. Verify GPU /health returns 200 on direct port (:8080) 1. Verify GPU /health returns 200
2. If GPU healthy, alert but do NOT reset via router API (deprecated) 2. If GPU healthy, send 1 test inference
3. Check LiteLLM health directly: http://192.168.68.116/litellm/health/liveliness 3. If test succeeds → reset circuit breaker via router API
4. Restart LiteLLM container on CT 116 if circuit breakers are stuck 4. 60s cooldown — if CB re-opens immediately, it was legitimate, do NOT re-reset
- **Verify**: LiteLLM returns healthy, circuit breaker clears within 60s 5. Max 1 auto-reset per GPU per hour
- **Escalate after**: LiteLLM restart doesn't clear → human investigation - **Verify**: CB closes, inference succeeds, CB stays closed for 60s+
- **Escalate after**: CB won't close after reset → router issue
### Rule 6: Strix Halo Unreachable ### Rule 6: Strix Halo Unreachable
- **Detect**: Strix not responding — probe .15:8080 directly (firewall opened .24→.15) - **Detect**: Strix not responding — probe .15:8080 directly (firewall opened .24→.15)
- **Fix**: - **Fix**:
1. SSH to .15 → check llama-server process 1. SSH to .15 → check llama-server process
2. Restart llama-server if not running 2. Restart llama-server if not running
3. Verify through both direct probe AND LiteLLM health 3. Verify through both direct probe AND router
- **Verify**: Direct health probe returns 200, LiteLLM reports model healthy - **Verify**: Direct health probe returns 200, router reports Strix healthy
- **Escalate**: If host .15 itself is unreachable → infrastructure alert - **Escalate**: If host .15 itself is unreachable → infrastructure alert
### Rule 7: GPU Data Source Unreachable (replaces old Prometheus rule) ### Rule 7: Prometheus Exporter Down
- **Detect**: gpu-monitor endpoint (localhost:9100/gpu-data) or sidecar port (:8080) on any GPU unreachable for >2 polls - **Detect**: Any GPU :9400/metrics unreachable for >2 polls
- **Fix**: - **Fix**:
1. If gpu-monitor is down: restart systemd service `gpu-monitor.service` on this host 1. SSH to GPU host → check prometheus-exporter process
2. If sidecar is down: SSH to GPU host → check llama-server process → restart systemd service 2. Restart exporter if dead
3. Fall back to direct nvidia-smi/rocm-smi probe via SSH if all API paths fail 3. While exporter is down, fall back to nvidia-smi/rocm-smi direct probes
- **Verify**: gpu-monitor returns healthy + all sidecars reachable 4. If exporter is running but unreachable → check firewall/host networking
- **Verify**: :9400/metrics returns 200
- **Escalate after**: 3 failed restarts → networking issue - **Escalate after**: 3 failed restarts → networking issue
### Rule 8: Predictive Thermal Warning (two-tier) ### Rule 8: Predictive Thermal Warning (two-tier)
@@ -140,31 +117,29 @@ Key notes:
- **Escalate**: If Tier 2 triggers and temp still rising after 5 min → possible hardware failure - **Escalate**: If Tier 2 triggers and temp still rising after 5 min → possible hardware failure
### Rule 9: Context Window Optimization ### Rule 9: Context Window Optimization
- **Detect**: Benchmark tok/s vs baseline for each GPU at current context (all 128K) - **Detect**: Benchmark tok/s vs baseline for each GPU at current context
- RTX 3090 (128K ctx, ThinkingCap): baseline 74.8 tok/s — currently at 74.9 (100%) - RTX 3090 (128K ctx, qwen3.6-27B-code): target 75+ tok/s — currently at baseline
- RTX 5070 (128K ctx, HauhauCS QAT): baseline 165.2 tok/s — currently at 169.6 (103%) - RTX 5070 (128K ctx, gemma-4-12b): target 76+ tok/s — optimal for vision/web role
- Strix Halo (128K ctx, Genesis Hermes V3): baseline 70.5 tok/s — currently at 62.9 (89%) - Strix Halo (128K ctx, strix-moe / qwen3.6-35B-udq4): target 70+ tok/s — currently above baseline
- **Fix**: - **Fix**:
- If tok/s > baseline → context has headroom, consider increasing - If tok/s > baseline → context has headroom, consider increasing
- If tok/s < 90% baseline → reduce context by 25% and retest - If tok/s < 90% baseline → reduce context by 25% and retest
- If tok/s within 10% of baseline → optimal, no change - If tok/s within 10% of baseline → optimal, no change
- Strix Halo at 89% of baseline → MONITOR but do not reduce yet (recent model swap may still be settling)
- **Verify**: Re-benchmark after context change, confirm within 10% of target - **Verify**: Re-benchmark after context change, confirm within 10% of target
- **Escalate**: If context can't be adjusted without significant perf loss - **Escalate**: If context can't be adjusted without significant perf loss
### Rule 10: Workload Distribution Optimization (updated 2026-07-18) ### Rule 10: Workload Distribution Optimization
- **Detect**: GPU roles misaligned with hardware capabilities - **Detect**: GPU roles misaligned with hardware capabilities
- **Target distribution**: - **Target distribution**:
- RTX 3090 (gpu-dense, 24GB, 74.9 tok/s) → Heavy reasoning, code gen, long conversations (slowest per-token but largest context capacity). Weight: 0.55 (LiteLLM). - RTX 3090 (24GB, 128K, 75 tok/s) → Heavy reasoning, code gen, long conversations
- RTX 5070 (gpu-light, 12GB, 169.6 tok/s) → Vision/image, web search, lightweight tasks (2.3x faster than 3090 per token). Weight: 0.15 (LiteLLM). - RTX 5070 (12GB, 128K, 76 tok/s) → Vision/image, web search, quick lightweight tasks
- Strix Halo (strix-moe, 64GB, 62.9 tok/s) → Context compression, summarization, long docs (MoE model). Weight: 0.30 (LiteLLM). - Strix Halo (64GB, 128K, 72 tok/s) → Context compression, summarization, long docs
- **Note**: RTX 5070 is the fastest endpoint per token. Route high-volume, low-complexity work there first.
- **Fix**: - **Fix**:
- Alert if any GPU is handling workload outside its designated role - Alert if any GPU is handling workload outside its designated role
- Recommend agent alias updates to match workload to GPU role (use stable aliases: gpu-dense, gpu-light, strix-moe) - Recommend Hermes agent profile updates to match workload to GPU
- Track per-GPU request distribution via LiteLLM spend logs - Track per-GPU request distribution via LiteLLM spend logs
- **Verify**: Each GPU's request pattern matches its designated role within 24h - **Verify**: Each GPU's request pattern matches its designated role within 24h
- **Escalate**: If role mismatch persists >48h → agent alias audit needed - **Escalate**: If role mismatch persists >48h → agent profile audit needed
--- ---
@@ -173,7 +148,7 @@ Key notes:
```prose ```prose
-- Phase 1: Fetch live GPU data -- Phase 1: Fetch live GPU data
let fleet = call gpu-monitor let fleet = call gpu-monitor
endpoint: "http://localhost:9100/gpu-data" endpoint: "http://192.168.68.24:9100/gpu-data"
-- Phase 2: Evaluate each GPU against remediation rules -- Phase 2: Evaluate each GPU against remediation rules
let actions = [] let actions = []
@@ -189,25 +164,27 @@ for gpu in fleet.gpus:
-- Rule 4: Benchmark regression -- Rule 4: Benchmark regression
let bench = fleet.benchmarks[gpu.hostname] let bench = fleet.benchmarks[gpu.hostname]
if bench.current_tok_s < bench.baseline_tok_s * 0.8: if bench.current_tok_sec < bench.baseline_tok_sec * 0.8:
push actions apply-benchmark-fix(gpu, bench) push actions apply-benchmark-fix(gpu, bench)
-- Rule 3: Model stuck -- Rule 3: Model stuck
for model in fleet.summary.available_models: for model in fleet.router.available_models:
if model.consecutive_timeouts >= 3: if model.consecutive_timeouts >= 3:
push actions apply-model-restart(model) push actions apply-model-restart(model)
-- Rule 5: Circuit breaker check via LiteLLM (router deprecated) -- Rule 5: Circuit breaker
if fleet.summary.circuit_breakers_open > 0: for cb in fleet.router.circuit_breaker:
push actions check-litellm-circuit-breakers() if cb.open and cb.open_duration > 600 and gpu_is_healthy(cb.gpu):
push actions apply-cb-reset(cb)
-- Rule 6: Strix Halo -- Rule 6: Strix Halo
if not fleet.strix.running and pingable("192.168.68.15"): if not fleet.strix.running and pingable("192.168.68.15"):
push actions apply-strix-restart() push actions apply-strix-restart()
-- Rule 7: GPU data source -- Rule 7: Prometheus exporters
if not fleet.gpus or len(fleet.gpus) < 2: for gpu in fleet.gpus:
push actions check-gpu-monitor-service() if not prometheus_reachable(gpu.hostname, 9400):
push actions apply-exporter-restart(gpu)
-- Rule 8: Predictive thermal -- Rule 8: Predictive thermal
for gpu in fleet.gpus: for gpu in fleet.gpus:
@@ -235,12 +212,12 @@ call update-gpu-health
```json ```json
{ {
"run_id": "gpu-self-heal-20260718-001", "run_id": "gpu-self-heal-20260712-001",
"timestamp": "2026-07-18T08:00:00Z", "timestamp": "2026-07-12T16:00:00Z",
"gpu": "ct8-rtx3090", "gpu": "ct8-rtx3090",
"issue": "thermal-critical", "issue": "thermal-critical",
"detected": { "temp_c": 87, "duration_s": 180 }, "detected": { "temp_c": 87, "duration_s": 180 },
"action": "load-shedding", "action": "set-fan-100pct",
"result": "resolved", "result": "resolved",
"verification": { "temp_c": 76, "after_s": 300 }, "verification": { "temp_c": 76, "after_s": 300 },
"escalated": false "escalated": false
@@ -259,53 +236,52 @@ Every action logged as `[GPU-SELF-HEAL] <run_id>` node with full audit trail.
- `issues_escalated > 0` → "⚠ GPU Self-Heal — <gpu> needs attention" - `issues_escalated > 0` → "⚠ GPU Self-Heal — <gpu> needs attention"
- Every 100th clean cycle → "✅ GPU Fleet: All Clear" - Every 100th clean cycle → "✅ GPU Fleet: All Clear"
### 3. Weekly Benchmark Report ### 3. Prometheus/Grafana Integration
- GPU self-heal actions exposed as Prometheus counter metrics
- Dashboard panel: "GPU Interventions (24h)" showing count/type/result
### 4. Weekly Benchmark Report
- Per-GPU tok/s trend over 7 days - Per-GPU tok/s trend over 7 days
- Regression alerts if any GPU degrades >10% week-over-week - Regression alerts if any GPU degrades >10% week-over-week
--- ---
## Design Decisions (Verified 2026-07-12, Reaffirmed 2026-07-18) ## Design Decisions (Grilled & Confirmed 2026-07-12)
1. **Fan control**: ❌ NO auto fan control. Load-side cooling only (reduce concurrency, redirect). 1. **Fan control**: ❌ NO auto fan control. Load-side cooling only (reduce concurrency, redirect).
2. **Model restart**: ✅ Only if >50% failure rate over 60s + 30s grace period. Not on single stuck request. 2. **Model restart**: ✅ Only if >50% failure rate over 60s + 30s grace period. Not on single stuck request.
3. **Strix direct access**: ✅ Open firewall .15:8080 → .24 for direct health probe + restart. 3. **Strix direct access**: ✅ Open firewall .15:8080 → .24 for direct health probe + restart.
4. **VRAM thresholds**: Tiered — **300MB/h** (RTX 3090), **300MB/h** (RTX 5070), 200MB/h (Strix). Previous values (100/50) were too sensitive; raised 2026-07-18 based on operational data. 4. **VRAM thresholds**: Tiered — 100MB/h (RTX 3090), 50MB/h (RTX 5070), 200MB/h (Strix).
5. **CB auto-reset**: ✅ Router deprecated — circuit breakers go through LiteLLM health check + container restart if needed. No per-GPU auto-reset. 5. **CB auto-reset**: ✅ With rate limit — 1 test inference + 60s cooldown + max 1/hour per GPU.
6. **Benchmark baseline**: Rolling 30-day average, recalculated weekly. Current baselines live in gpu-monitor. 6. **Benchmark baseline**: Rolling 30-day average, recalculated weekly. Original baseline kept in Grafana.
7. **Predictive alerts**: Two-tier — warn at >70°C+rising (>2°C/min), critical at >80°C+rising. 7. **Predictive alerts**: Two-tier — warn at >70°C+rising (>2°C/min), critical at >80°C+rising.
8. **Prometheus**: ❌ Not deployed. Use direct sidecar probes (:8080/health) and gpu-monitor API. Prometheus integration deferred until exporters are running on GPU hosts. 8. **Prometheus**: Primary source. Fall back to nvidia-smi/rocm-smi direct probes if exporter down.
## Lessons Learned (2026-07-12, Updated 2026-07-18) ## Lessons Learned (2026-07-12)
### L1: API Key Standardization Is Critical ### L1: API Key Standardization Is Critical
- All GPU llama-servers MUST use the same api-key as the LiteLLM config. - All GPU llama-servers MUST use the same api-key as the LiteLLM config.
- RTX 5070 had `--api-key sk-loc...5678` while LiteLLM sent `not-needed`. - RTX 5070 had `--api-key sk-loc...5678` while LiteLLM sent `not-needed`.
This caused cascading 401 → fallback → timeout → 401 loops. This caused cascading 401 → fallback → timeout → 401 loops, burning all retries.
- **Rule**: Any new GPU or model restart MUST verify api-key matches LiteLLM config (`not-needed` for direct routing). - **Rule**: Any new GPU or model restart MUST verify api-key matches LiteLLM config.
### L2: Fallback Chain Cascading Failures ### L2: Fallback Chain Cascading Failures
- When one model returns 401 (auth) and another is slow (timeout), the fallback - When one model returns 401 (auth) and another is slow (timeout), the fallback
chain creates an infinite loop. chain creates an infinite loop: gemma 401 → qwen timeout → gemma 401 → ...
- **Rule**: If a model returns 401 (auth error), do NOT fall back to it again. - **Rule**: If a model returns 401 (auth error), do NOT fall back to it again.
Mark it as permanently failed for this request. Mark it as permanently failed for this request.
### L3: Verify Running State, Not Docs ### L3: Verify Running State, Not Docs
- RTX 3090 was documented at 128K context. Running at 128K (verified 2026-07-18). - RTX 3090 was documented at 128K context. Actually running at 256K.
- Parallel count: 1 on both RTX 3090 and RTX 5070 (matches docs for current models). - Parallel count wrong (docs said 2, actual is 1 on RTX 3090).
- **Rule**: Before making decisions, check `/proc/PID/cmdline` on GPU hosts. - **Rule**: Before making decisions, check `/proc/PID/cmdline` on GPU hosts.
### L4: Infisical Is Not Always Available ### L4: Infisical Is Not Always Available
- Keep a local `.env` fallback for `LITELLM_API_KEY`. - Tanko's Infisical service token was 404 — gateway ran without API key for hours.
- **Rule**: Always verify credential source is reachable before relying on it. - **Rule**: Always keep a local `.env` fallback for `LITELLM_API_KEY`.
- Contract hermes-config-template Rule 3 updated.
### L5: GPU Monitor Response Size Can Cause Self-Heal Crash ### L5: Zulip Event Queue Can Silently Die
- gpu-self-heal crashed with KeyboardInterrupt during json.loads() of 20MB response. - Mumuni's queue accumulated 41 errors/reconnects then stopped polling.
- Root cause: router poll returns accumulated data → cache balloons. Gateway was running but ignoring all messages.
- **Rule**: Self-heal must enforce a read timeout AND max response size on every poll. - **Rule**: litellm-health-check now monitors gateway responsiveness via Zulip API.
If monitor response > 1MB, log a warning and skip the cycle rather than crashing.
### L6: Stable Aliases Replace Model Names
- gpu-fleet introduced stable aliases (strix-moe, gpu-dense, gpu-light) on 2026-07-15.
- Self-heal must use aliases for reporting and alerting, not model-specific names.
- **Rule**: All alert messages and KG nodes use the stable alias as the GPU identifier.
+2
View File
@@ -193,6 +193,8 @@ Key is injected via `infisical run --` wrapper at PM2 startup:
"models": [ "models": [
{ "id": "syslog-auto" }, { "id": "syslog-auto" },
{ "id": "strix-moe" }, { "id": "strix-moe" },
{ "id": "gpu-dense" },
{ "id": "gpu-light" },
{ "id": "qwen3.6-27B-code" }, { "id": "qwen3.6-27B-code" },
{ "id": "gemma-4-12b" } { "id": "gemma-4-12b" }
] ]
+13 -2
View File
@@ -325,7 +325,8 @@ verify ALL FOUR of these against the live config. They are the only root causes
One-line agent health check (run on the agent host): One-line agent health check (run on the agent host):
```bash ```bash
PID=$(pgrep -f "python -m hermes_cli.main gateway run" | head -1) # Use grep -v infisical to avoid matching the bash wrapper that contains the same string
PID=$(pgrep -f "python -m hermes_cli.main gateway run" | grep -v infisical | head -1)
cat /proc/$PID/environ | tr '\0' '\n' | grep ^LITELLM_API_KEY= | sed 's/=.*/<set>/' cat /proc/$PID/environ | tr '\0' '\n' | grep ^LITELLM_API_KEY= | sed 's/=.*/<set>/'
curl -s -o /dev/null -w 'key_health: %{http_code}\n' -H "Authorization: Bearer $(cat /proc/$PID/environ | tr '\0' '\n' | grep ^LITELLM_API_KEY= | cut -d= -f2)" http://192.168.68.116/v1/models curl -s -o /dev/null -w 'key_health: %{http_code}\n' -H "Authorization: Bearer $(cat /proc/$PID/environ | tr '\0' '\n' | grep ^LITELLM_API_KEY= | cut -d= -f2)" http://192.168.68.116/v1/models
``` ```
@@ -351,9 +352,19 @@ directly (no infisical). Apply with `systemctl daemon-reload && systemctl restar
The wrapper sources `~/.hermes/.env` then exports `LITELLM_API_KEY="$<AGENT>_LITELLM_API_KEY"`. The wrapper sources `~/.hermes/.env` then exports `LITELLM_API_KEY="$<AGENT>_LITELLM_API_KEY"`.
See litellm-api-keys.prose.md § Machine Identity for Vault Writes for vault sync. See litellm-api-keys.prose.md § Machine Identity for Vault Writes for vault sync.
**⚠️ Vault empty-key guard:** If the vault stores the secret as an empty string,
the wrapper will inject an empty key and the gateway will silently get 401 errors
on all LiteLLM requests (triggering silent DeepSeek fallback). The `.env` fallback
is present but the vault takes precedence when the secret key exists (even if empty).
**Fix:** The wrapper MUST validate the key length after injection. If LITELLM_API_KEY
is empty or shorter than 20 chars, log a warning and either fail with a clear error
message or fall back to the `.env` value before starting the gateway.
**Verification (all agents):** **Verification (all agents):**
```bash ```bash
GP=$(pgrep -f "python -m hermes_cli.main gateway run" | head -1) # Use grep -v infisical to avoid matching the bash wrapper that contains the same string
GP=$(pgrep -f "python -m hermes_cli.main gateway run" | grep -v infisical | head -1)
K=$(cat /proc/$GP/environ | tr '\0' '\n' | grep '^LITELLM_API_KEY=' | cut -d= -f2) K=$(cat /proc/$GP/environ | tr '\0' '\n' | grep '^LITELLM_API_KEY=' | cut -d= -f2)
curl -s -o /dev/null -w '%{http_code}' -H "Authorization: Bearer $K" http://192.168.68.116/v1/models # must be 200 curl -s -o /dev/null -w '%{http_code}' -H "Authorization: Bearer $K" http://192.168.68.116/v1/models # must be 200
``` ```
+1 -1
View File
@@ -171,7 +171,7 @@ through its agent wrapper.
- **Survives gateway crash**: the wrapper's `while true` + systemd `Restart=always` revive the gateway. Two-layer defense. - **Survives gateway crash**: the wrapper's `while true` + systemd `Restart=always` revive the gateway. Two-layer defense.
- **Survives Hermes updates**: systemd drop-in overrides unit file ExecStart — `hermes gateway install` cannot break the vault injection. - **Survives Hermes updates**: systemd drop-in overrides unit file ExecStart — `hermes gateway install` cannot break the vault injection.
- **Survives reboot**: systemd user service + `loginctl enable-linger` ensures gateway starts at boot without a login session. - **Survives reboot**: systemd user service + `loginctl enable-linger` ensures gateway starts at boot without a login session.
- **Auditable**: `cat /proc/$(pgrep hermes_cli)/environ` shows all injected keys; `infisical secrets` shows the vault source. - **Auditable**: `cat /proc/$(pgrep -f 'python.*hermes_cli.main.gateway.run' | grep -v infisical | head -1)/environ` shows all injected keys (note: pipe through grep -v infisical to avoid matching the bash wrapper); `infisical secrets` shows the vault source.
### Migration status (2026-07-17) ### Migration status (2026-07-17)
+2 -2
View File
@@ -184,8 +184,8 @@ def check_agents():
print(f"{name} (CT {ct}): cannot SSH — skip liveness check") print(f"{name} (CT {ct}): cannot SSH — skip liveness check")
continue continue
# Gateway process # Gateway process (exclude the infisical bash wrapper that contains the same string)
pid = ssh(host, "pgrep -f 'hermes_cli.main gateway run' | head -1", user=user) pid = ssh(host, "pgrep -f 'hermes_cli.main gateway run' | grep -v infisical | head -1", user=user)
if not pid: if not pid:
print(f"{name}: GATEWAY NOT RUNNING") print(f"{name}: GATEWAY NOT RUNNING")
FAIL.append(f"gateway-down:{name}") FAIL.append(f"gateway-down:{name}")
-65
View File
@@ -409,68 +409,3 @@ Backup v2 before starting: `cp index.js index.js.v2-backup-$(date +%Y%m%d-%H%M%S
| Queue expiry handling | Crash | Auto re-register | | Queue expiry handling | Crash | Auto re-register |
| Busy worker deadlock | Router death | Worker SIGKILL + error DM | | Busy worker deadlock | Router death | Worker SIGKILL + error DM |
| PM2 restart exhaustion | Yes (max_restarts=10) | No (max_restarts=100 + watchdog) | | PM2 restart exhaustion | Yes (max_restarts=10) | No (max_restarts=100 + watchdog) |
---
## Incident Log — 2026-07-18 Fleet-Wide Audit
### Fleet State After Audit
| Agent | Platform | Zulip State | Issues Found | Fix Applied |
|-------|----------|-------------|--------------|-------------|
| **Abiba** | pi (CT 100) | ✅ Connected | API key missing from Infisical injection; poll timeout noise | Added .env fallback; AbortError treated as empty poll (no retry); poll timeout 65s→90s |
| **Tanko** | Hermes (CT 112) | ✅ Connected | Gateway disconnected since Jul 11; watchdog restart didn't re-establish Zulip | Full gateway restart (kill wrapper, let infisical-gateway.sh respawn) |
| **Mumuni** | Hermes (CT 114) | ✅ Connected | No issues found | None needed |
### Key Fixes Applied
**1. Abiba — Credential Fallback (L4 Pattern)**
- Root cause: `zulip.api_key` in config.yaml is `""` (expected from Infisical). Infisical vault `ABIBA_ZULIP_API_KEY` wasn't being injected into the process environment.
- Fix: Added `.env` file fallback at `/root/.pi/agent/extensions/zulip/.env` with known-working key, sourced before the Infisical `exec`.
- Lesson: Per L4 from gpu-self-heal, Infisical is not always available — always keep a local `.env` fallback.
**2. Abiba — Poll Timeout Handling**
- Root cause: Zulip long-poll uses `AbortSignal.timeout(65000)`. Zulip's default `event_queue_longpoll_timeout_seconds` can exceed 65s. When the signal fires, an `AbortError` is thrown and caught by the circuit breaker as a failure.
- Fix: Caught `AbortError` inside `poll()` and return empty array (no events) instead of throwing. Extended timeout to 90s to match Zulip server default.
- Reference: [Zulip Events System — long-poll timeout](https://zulip.readthedocs.io/en/11.6/subsystems/events-system.html)
**3. Tanko — Gateway Restart**
- Root cause: Gateway process was running but Zulip platform stayed in "disconnected" state since Jul 11, 2026. The wrapper script (`infisical-gateway.sh`) restarts on crash but the gateway wasn't re-establishing Zulip on restart.
- Fix: Killed gateway PID to trigger wrapper restart. New gateway (PID 331991) established Zulip connection successfully.
### Fleet-Wide Zulip Health Metrics (as of 2026-07-18)
| Metric | Value |
|--------|-------|
| Zulip server | ✅ HTTP 200 |
| Agents connected | 3/3 (Abiba, Tanko, Mumuni) |
| Abiba circuit breaker | CLOSED (0 failures) |
| Abiba uptime | 2D (post-restart) |
| Tanko gateway uptime | Ongoing |
| Mumuni gateway uptime | Ongoing |
| Watchdog status | ✅ Online (2D uptime) |
### Hermes Agent Zulip Plugin Improvements
Based on the audit, improvements that should be ported to all Hermes Zulip adapters:
1. **Circuit breaker pattern** — Already in Abiba's pi extension. Hermes adapters should add the same CLOSED→OPEN→HALF_OPEN state machine with exponential backoff.
2. **Credential fallback** — All Hermes agents use Infisical for credentials. Add `.env` local fallback per L4 pattern for `ZULIP_API_KEY`.
3. **Queue re-registration** — Handle `BAD_EVENT_QUEUE_ID` with automatic re-registration instead of gateway restart.
4. **Supervisor watchdog** — Hermes uses PM2 which auto-restarts on crash, but has no health-check watchdog. Add lightweight external health checks.
5. **Streaming** — All agents have `streaming: true` in their zulip config. Verify `edit_message()` is implemented in each adapter.
### Abiba pi Zulip Extension v2 — Implemented Resilience Summary
| Feature | Status | Notes |
|---------|--------|-------|
| Circuit breaker | ✅ | CLOSED→OPEN→HALF_OPEN; 50% failure threshold; 30s reset timeout |
| Retry with jitter | ✅ | 2 attempts, 200ms base, 50-100% jitter |
| Queue lifecycle | ✅ | 10min idle_queue_timeout; BAD_EVENT_QUEUE_ID handling |
| Crash prevention | ✅ | uncaughtException + unhandledRejection recovery |
| Worker timeout | ✅ | 5min busy timeout → SIGKILL + error DM |
| Health endpoint | ✅ | :9200 with circuit breaker metrics |
| Echo prevention | ✅ | Dynamic bot user resolution |
| Poll timeout (AbortError) | ✅ v2.1 | Normal timeout returns [] instead of error |
| Credential fallback | ✅ v2.1 | .env file before Infisical exec |
| Provider auto-fix | ✅ | Detects reasoning_content models, switches to compatible |