UPDATE 2026-07-15: Full GPU fleet rebuild + stable aliases #18

Merged
jerome merged 13 commits from feat/gpu-fleet-rebuild-20260715 into master 2026-07-16 21:17:12 +00:00
13 changed files with 1136 additions and 107 deletions
+309
View File
@@ -0,0 +1,309 @@
---
kind: function
name: abiba-zulip-restore
description: >
Restores Zulip connectivity for Abiba (pi agent). Verifies the v2 router-worker
extension code, starts a PM2 process as the Zulip gateway with correct env vars,
validates health endpoint, and confirms DM delivery. Run this whenever Abiba
stops responding on Zulip or after system restart.
agent: abiba
version: 1.0.0
status: active
runtime_contract: 2
---
# Abiba Zulip Restore — Resume pi Zulip Communication
Single-shot function that restores full Zulip connectivity for the Abiba pi agent.
Covers extension code validation, PM2 process management, health endpoint
verification, and DM loopback testing.
## Live-State Fields
| Field | Value | Trust |
|-------|-------|-------|
| Agent name | abiba | ✅ |
| Bot email | abiba-bot@chat.sysloggh.net | ✅ |
| Zulip server | https://chat.sysloggh.net | ✅ |
| Extension path | /root/.pi/agent/extensions/zulip/index.js | ✅ |
| Config path | /root/.pi/agent/extensions/zulip/config.yaml | ✅ |
| Health port | 9200 | ✅ |
| @all-bots user ID | 20 | ✅ (config, verified by API at runtime) |
| PM2 process name | abiba-zulip | ✅ |
| Provider | syslog-harness (http://192.168.68.116/v1) | ✅ |
| Default model | deepseek-v4-pro | ✅ (settings.json) |
## Architecture
The pi Zulip extension uses a **router-worker architecture** (v2):
- **Router** (PM2 `abiba-zulip`, `ZULIP_ROLE=router`): Polls Zulip for events,
maintains a per-sender pool of pi RPC worker processes. Each sender gets
their own `pi --mode rpc --session-dir` process with persistent sessions.
Handles streaming edits back to Zulip.
- **Worker** (child `pi --mode rpc`): Runs the agent with per-sender persistent
sessions. No Zulip logic in the worker — the router handles all Zulip I/O.
The extension loads in ALL pi sessions (via `settings.json` extensions array)
but is a **NO-OP** unless `ZULIP_ROLE=router` is set. Only the PM2 router process
carries the env var.
## Maintains
- extension_valid: bool — Extension code imports without errors
- pm2_running: bool — PM2 process `abiba-zulip` is online
- health_responding: bool — GET :9200/health returns "ok"
- zulip_connected: bool — Queue registered, bot identity resolved
- loopback_delivered: bool — Test DM sent and received
- model_valid: bool — Configured models match LiteLLM authorized models
### Postconditions
- PM2 process `abiba-zulip` online and stable (uptime > 30s)
- Health endpoint returns `{ status: "ok", connected: true }`
- Worker pool creates sessions on demand
- Echo prevention active (BOT_EMAILS includes all known bots)
- PM2 saved for auto-restart on boot
## Requires
- Node.js with `yaml` module available
- PM2 installed globally
- Zulip API key in `config.yaml`
- `pi` CLI available on PATH
- Zulip server accessible at https://chat.sysloggh.net
- LiteLLM provider accessible at http://192.168.68.116/v1
## Execution
### Step 1: Validate Extension Code
```bash
node -e "import('file:///root/.pi/agent/extensions/zulip/index.js').then(() => console.log('OK')).catch(e => {console.error('FAIL:', e.message); process.exit(1)})"
```
Expected: `OK`. If FAIL → check for missing dependencies, syntax errors.
### Step 2: Validate Model IDs
Compare configured models against LiteLLM authorized models:
```bash
API_KEY=$(grep -oP 'apiKey:\s*\K.*' /root/.pi/agent/models.json | head -1)
curl -s -H "Authorization: Bearer $API_KEY" http://192.168.68.116/v1/models | \
python3 -c "import json,sys; d=json.load(sys.stdin); [print(m['id']) for m in d.get('data',[])]" 2>/dev/null
```
Check: Every model ID in `models.json` must appear in the authorized list.
If not → fix `models.json` to only include authorized models (prefer `syslog-auto`).
### Step 3: Verify Config Integrity
```bash
python3 -c "
import yaml, sys
with open('/root/.pi/agent/extensions/zulip/config.yaml') as f:
cfg = yaml.safe_load(f)
required = ['zulip.site', 'zulip.email', 'zulip.api_key', 'agent.name']
for k in required:
keys = k.split('.')
v = cfg
for kk in keys:
v = v.get(kk)
if v is None:
print(f'MISSING: {k}')
sys.exit(1)
print('Config valid')
print(f' site={cfg[\"zulip\"][\"site\"]}')
print(f' email={cfg[\"zulip\"][\"email\"]}')
print(f' agent={cfg[\"agent\"][\"name\"]}')
print(f' health_port={cfg.get(\"health_port\", 9200)}')
"
```
Expected: Config valid with all fields non-empty.
### Step 4: Remove Stale Systemd Service
The old `abiba-zulip.service` points to `/opt/abiba-zulip/dist/index.js` (compiled
TypeScript, not the v2 extension). It's disabled and stale. Remove it:
```bash
systemctl stop abiba-zulip 2>/dev/null || true
systemctl disable abiba-zulip 2>/dev/null || true
rm -f /etc/systemd/system/abiba-zulip.service
systemctl daemon-reload
```
### Step 5: Start PM2 Process
**Critical:** The extension MUST run via `pi --mode rpc`, NOT `node index.js` directly.
The extension exports a function that requires pi's session lifecycle. Running
`node index.js` loads the module but never calls the export, so nothing happens.
`pi --mode rpc` loads all extensions (including Zulip) and fires `session_start`.
```bash
# Stop existing if any
pm2 delete abiba-zulip 2>/dev/null || true
# Start pi in RPC mode with ZULIP_ROLE=router env
ZULIP_ROLE=router pm2 start "$(which pi)" \
--name abiba-zulip \
--interpreter none \
-- --mode rpc --no-session
```
Wait 10 seconds for pi to load all extensions, fire session_start, and the Zulip
router to register its event queue.
### Step 6: Validate PM2 Process
```bash
pm2 show abiba-zulip --no-color
```
Check: `status=online`, `restarts=0`, `uptime > 5s`.
### Step 7: Check Logs for Connection
```bash
sleep 3
tail -20 /root/.pm2/logs/abiba-zulip-out.log
```
Look for:
- `[zulip-ext] Connecting to https://chat.sysloggh.net as abiba-bot@chat.sysloggh.net…`
- `[zulip-ext] Bot user_id=N, all-bots user_id=N`
- `[zulip-ext] Connected, queue=N`
- `[zulip-ext] Echo prevention: N bot emails`
- `[zulip-ext] Health endpoint on :9200`
If error → check API key, network to chat.sysloggh.net.
### Step 8: Health Endpoint
```bash
curl -s http://localhost:9200/health | python3 -m json.tool
```
Check:
- `status: "ok"` (not "down")
- `zulip.connected: true`
- `zulip.queue_id` is non-null string
- `zulip.bot_user_id` is positive integer
### Step 9: DM Loopback Test
```bash
curl -s http://localhost:9200/health | python3 -c "
import json,sys
d = json.load(sys.stdin)
if d.get('zulip',{}).get('connected'):
print(f'✅ Zulip connected. Queue: {d[\"zulip\"][\"queue_id\"]}')
print(f' Bot user_id: {d[\"zulip\"][\"bot_user_id\"]}')
print(f' Messages processed: {d[\"zulip\"][\"messages_processed\"]}')
else:
print('❌ Zulip NOT connected')
sys.exit(1)
"
```
### Step 10: Save PM2 for Auto-Start
```bash
pm2 save
pm2 startup systemd -u root --hp /root 2>/dev/null || true
```
### Step 11: Report
Compile results:
| Check | Pass? |
|-------|-------|
| Extension code imports | extension_valid |
| Model IDs authorized | model_valid |
| Config integrity | config_valid |
| PM2 process online | pm2_running |
| Health endpoint | health_responding |
| Zulip connected | zulip_connected |
All pass → ✅ **Abiba Zulip restored.** Relay success to user.
Partial failure → see recovery matrix below.
## Recovery Matrix
| Failure | Recovery |
|---------|----------|
| Extension import fails | Check `node_modules/zulip-js` exists; run `npm install` in extension dir |
| Model ID mismatch | Fix `models.json` to use `syslog-auto` as default model; remove invalid IDs |
| Config missing | Restore from backup or recreate from scratch |
| PM2 won't start | Check `node` version (>=18); check port 9200 not in use |
| Health "down" | Check logs for connection errors; verify Zulip API key; check network |
| "address already in use" | Kill old process: `fuser -k 9200/tcp` |
| Queue registration fails | Check Zulip API key validity; verify bot is active in Zulip admin |
| Rate limit (429) | Extension has built-in retry-after handling — wait, don't restart |
## Known Failure Modes
| Symptom | Root Cause | Recovery |
|---------|-----------|----------|
| Extension loads but no events | `ZULIP_ROLE` not set | Ensure PM2 env has `ZULIP_ROLE=router` |
| Worker stays "busy" forever | Model ID not authorized by LiteLLM | Fix models.json (lesson #11) |
| Placeholder sent but no response | editMessage API fails silently | Extension has fallback (sends new msg); check Zulip API |
| Queue expires rapidly | Poll interval too aggressive | v2 uses 3s poll with long-poll — should be fine |
| Bot doesn't respond to @mentions | Not subscribed to stream | Bot auto-subscribes via API |
| Stale error in health | `last_error` not cleared | v2 clears on successful poll (lesson #4) |
## Appendix: Root Cause & Fix Summary (2026-07-13)
**Problem:** Zulip extension was offline — no PM2 process running.
**Root cause:** The PM2 command ran `node index.js` directly (which loads the
extension module but never calls the exported function). The extension requires
pi's session lifecycle — pi loads extensions, fires `session_start`, and the
Zulip extension hooks into that event.
**Fix:** Run `pi --mode rpc` (not `node index.js`). The `--mode rpc` flag keeps
pi alive listening for RPC commands on stdin while the Zulip extension's router
runs in the background via the `session_start` hook.
```bash
ZULIP_ROLE=router pm2 start "$(which pi)" --name abiba-zulip \
--interpreter none -- --mode rpc --no-session
```
**Additional fixes applied:**
- Fixed `models.json`: replaced `qwen3.6-35B-A3B` (not authorized by LiteLLM)
with `syslog-auto` + `strix-moe` (prevents silent worker failure per
Lesson #11)
- Removed stale systemd unit `abiba-zulip.service` (pointed to old TS code)
- PM2 saved for auto-restart on boot
## Appendix: PM2 Ecosystem Config (Optional)
If preferred over manual `pm2 start`, create `/root/ecosystem.config.js` entry:
```js
module.exports = {
apps: [{
name: 'abiba-zulip',
script: '/bin/pi',
interpreter: 'none',
args: '--mode rpc --no-session',
cwd: '/root',
env: {
ZULIP_ROLE: 'router',
},
log_file: '/root/.pm2/logs/abiba-zulip-out.log',
error_file: '/root/.pm2/logs/abiba-zulip-error.log',
max_restarts: 20,
restart_delay: 5000,
}]
};
```
---
**Last verified good state**: 2026-07-13 — Extension v2 running via `pi --mode rpc`, health endpoint :9200 returning `{status:"ok",connected:true}`, queue a669f21e.
+4 -4
View File
@@ -72,10 +72,10 @@ reads them directly.
|--------|-------|----------|------|----------| |--------|-------|----------|------|----------|
| `syslog-code` | qwen3.6-27B-code | terminal, file, web, memory, skills | Code patches, automation, scripts | Writing/modifying code, creating scripts, debugging, reading/writing files | | `syslog-code` | qwen3.6-27B-code | terminal, file, web, memory, skills | Code patches, automation, scripts | Writing/modifying code, creating scripts, debugging, reading/writing files |
| `syslog-devops` | qwen3.6-27B-code | terminal, file, web, memory, skills | Infrastructure, DB, bridge, Proxmox | Server ops, SSH, Docker, Proxmox, DB queries, hardware checks | | `syslog-devops` | qwen3.6-27B-code | terminal, file, web, memory, skills | Infrastructure, DB, bridge, Proxmox | Server ops, SSH, Docker, Proxmox, DB queries, hardware checks |
| `syslog-email` | ornith-1.0-35b | terminal, file, web, memory, skills | Email automation, mail operations | Sending/receiving email, inbox management, SMTP operations | | `syslog-email` | strix-moe | terminal, file, web, memory, skills | Email automation, mail operations | Sending/receiving email, inbox management, SMTP operations |
| `syslog-research` | ornith-1.0-35b | terminal, file, web, memory, skills, **browser** | Analysis, classification, data processing | Web research, browser tasks, data analysis, classification, reading docs | | `syslog-research` | strix-moe | terminal, file, web, memory, skills, **browser** | Analysis, classification, data processing | Web research, browser tasks, data analysis, classification, reading docs |
| `syslog-review` | ornith-1.0-35b | terminal, file, web, memory, skills | Verification, QA, audit validation | **ALWAYS** verify worker output before delivery — especially for infra changes, code builds, and research findings | | `syslog-review` | strix-moe | terminal, file, web, memory, skills | Verification, QA, audit validation | **ALWAYS** verify worker output before delivery — especially for infra changes, code builds, and research findings |
| `syslog-writer` | ornith-1.0-35b | terminal, file, web, memory, skills | Docs, content, branding, reports | Writing docs, reports, proposals, content, markdown formatting | | `syslog-writer` | strix-moe | terminal, file, web, memory, skills | Docs, content, branding, reports | Writing docs, reports, proposals, content, markdown formatting |
### Selection Rules ### Selection Rules
+115 -46
View File
@@ -5,13 +5,12 @@ description: >
Manages the GPU inference fleet across all hosts. Handles model deployment, Manages the GPU inference fleet across all hosts. Handles model deployment,
registration, health checks, LiteLLM sync, agent key management, GPU registration, health checks, LiteLLM sync, agent key management, GPU
saturation watchdog, Prometheus/Grafana monitoring, and self-healing. saturation watchdog, Prometheus/Grafana monitoring, and self-healing.
UPDATED 2026-07-12: Architecture is DIRECT GPU — LiteLLM routes directly UPDATED 2026-07-15: Stable role-based aliases introduced: strix-moe,
to llama-server on each GPU host (no router in inference path). Router gpu-dense, gpu-light. These never change — only the underlying model does.
(port 9000) is running but NOT in request path. All GPUs standardized on Strix Halo: ornith-1.0-35b → unsloth/Qwen3.6-35B-A3B-MTP (UD-Q4_K_M, 22GB).
api-key 'not-needed'. RTX 5070 had api-key mismatch (sk-loc...5678) that RTX 5070: gemma-4-12b Q4_K_M → IQ4_NL + MTP draft (122 tok/s, 2x faster).
caused cascading 401→timeout→401 fallback loops — fixed. RTX 5070 context: 131K → 256K. VRAM: 88% (10.8/12.2GB).
Context: RTX 3090 verified at 256K (was documented as 128K — WRONG). Compression timeout: 300s (was 120s). Mumuni context: 128K (was 256K).
Workload: Compression moved to Strix Halo, RTX 5070 → vision/web only.
agent: abiba agent: abiba
triggers: triggers:
- on model add/remove - on model add/remove
@@ -68,48 +67,72 @@ triggers:
│ CT 8 │ │ CT 110 │ │ CT 15 │ │ pi (.24) │ │ CT 8 │ │ CT 110 │ │ CT 15 │ │ pi (.24) │
│ RTX 3090 │ │ RTX 5070 │ │ Strix Halo│ │ GPU Monitor │ │ RTX 3090 │ │ RTX 5070 │ │ Strix Halo│ │ GPU Monitor │
│ 24GB │ │ 12GB │ │ 64GB UMA │ │ :9100 │ │ 24GB │ │ 12GB │ │ 64GB UMA │ │ :9100 │
│ 256K ctx │ │ 131K ctx │ │ 256K ctx │ │ Watchdog │ │ 256K ctx │ │ 256K ctx │ │ 256K ctx │ │ Watchdog │
│ qwen3.6 │ │ gemma-4-12b │ │ ornith35B │ │ Prometheus │ │ qwen3.6 │ │ gemma-4-12b │ │ qwen3.6 │ │ Prometheus │
│ 27B-code │ │ :8080 │ │ :8080 │ │ exporter │ │ 27B-code │ │ :8080 │ │ -35B-udq4 │ │ exporter │
│ :8080 │ │ :9400 (exp) │ │ :9400(exp)│ │ :9401 │ │ :8080 │ │ :9400 (exp) │ │ :9400(exp)│ │ :9401 │
│ :9400 │ └─────────────┘ └───────────┘ └──────────────┘ │ :9400 │ └─────────────┘ └───────────┘ └──────────────┘
└──────────┘ └──────────┘
``` ```
## Current Model Assignments (2026-07-12) ## Stable Role-Based Aliases (Introduced 2026-07-15)
Agent configs, cron jobs, and workflows MUST use these aliases, never model-specific names.
When a model is swapped on a GPU, ONLY the infrastructure layer changes — agent configs are untouched.
| Alias | GPU | Current Model | Will Route To |
|-------|-----|---------------|---------------|
| `strix-moe` | Strix Halo (.15) | qwen3.6-35B-udq4 | Whatever runs on Strix Halo |
| `gpu-dense` | RTX 3090 (.8) | qwen3.6-27B-code | Whatever runs on RTX 3090 |
| `gpu-light` | RTX 5070 (.110) | gemma-4-12b | Whatever runs on RTX 5070 |
**Backward compatibility**: Old model-specific names (qwen3.6-27B-code, gemma-4-12b, qwen3.6-35B-udq4) still work
but are deprecated for agent configs. Only the stable aliases survive model swaps.
## Current Model Assignments (2026-07-15)
| Model | GPU | Host | VRAM | Ctx | KV Cache | Parallel | Batch/Ubatch | Status | | Model | GPU | Host | VRAM | Ctx | KV Cache | Parallel | Batch/Ubatch | Status |
|-------|-----|------|------|-----|----------|----------|-------------|--------| |-------|-----|------|------|-----|----------|----------|-------------|--------|
| qwen3.6-27B-code | RTX 3090 | .8 (llm-gpu) | 20.7/24GB (84%) | **256K** | turbo4 | 1 | default | ✅ healthy | | qwen3.6-27B-code (MTP) | RTX 3090 | .8 (llm-gpu) | 22.2/24.6GB (90%) | **256K** 🚀 | turbo4 | 2 | default | ✅ 63 tok/s |
| gemma-4-12b | RTX 5070 | .110 (ocu-llm) | 10.0/12.2GB (82%) | 131K | q4_0 | 2 | 2048/1024 | ✅ healthy | | gemma-4-12b | RTX 5070 | .110 (ocu-llm) | 10.0/12.2GB (82%) | 256K | q4_0 | 2 | 2048/1024 | ✅ healthy |
| ornith-1.0-35b | Strix Halo Vulkan | .15 (amdpve) | ~9GB/64GB | 256K | q8_0 | 2 | 2048/512 | ✅ healthy | | qwen3.6-35B-udq4 | Strix Halo Vulkan | .15 (amdpve) | ~9GB/64GB | 256K | q8_0 | 2 | 2048/512 | ✅ healthy |
## Routing Configuration (LiteLLM — July 2026) ## Routing Configuration (LiteLLM — July 2026)
### syslog-auto Weighted Pool ### syslog-auto Weighted Pool (Direct GPU — bypasses router)
| Model | GPU | Weight | RPM Cap | Purpose | | Model | GPU | Weight | RPM Cap | Timeout |
|-------|-----|--------|---------|---------| |-------|-----|--------|---------|---------|
| qwen3.6-27B-code | RTX 3090 | 0.55 | 500 | Heavy reasoning, code, long context | | qwen3.6-27B-code | RTX 3090 (.8:8080) | **0.55** | 500 | **300s** |
| ornith-1.0-35b | Strix Halo | 0.30 | **60** | Agentic workflows, tool calling | | qwen3.6-35B-udq4 | Strix Halo (.15:8080) | **0.30** | 60 | **300s** |
| gemma-4-12b | RTX 5070 | 0.15 | 200 | Overflow + vision | | gemma-4-12b | RTX 5070 (.110:8080) | **0.15** | 200 | **120s** |
Note: All syslog-auto entries route directly to GPUs with `api_key: not-needed`. The router (port 9000) is NOT in the inference path.
### Direct Model Endpoints ### Direct Model Endpoints
| Model | RPM Cap | Notes | | Model | RPM Cap | Notes |
|-------|---------|-------| |-------|---------|-------|
| ornith-1.0-35b | 40 | Tight cap — prevents Strix overload | | qwen3.6-35B-udq4 | 40 | Tight cap — prevents Strix overload |
| qwen3.6-27B-code | 500 | High cap — primary workhorse | | qwen3.6-27B-code | 500 | High cap — primary workhorse |
| gemma-4-12b | 500 | High cap — fast 12B | | gemma-4-12b | 500 | High cap — IQ4_NL+MTP, 122 tok/s |
### Stable Aliases (for agent configs — never change)
| Alias | RPM Cap | Routes To | Purpose |
|-------|---------|-----------|---------|
| `strix-moe` | 40 | Strix Halo | Compression tasks (MoE models) |
| `gpu-dense` | 500 | RTX 3090 | Heavy reasoning |
| `gpu-light` | 500 | RTX 5070 | Vision, web extract, light tasks |
### Fallback Chains ### Fallback Chains
- gemma → qwen - gemma → qwen
- qwen → gemma - qwen → gemma
- ornith → qwen → gemma - qwen3.6-35B-udq4 → qwen → gemma
- syslog-auto → qwen → gemma → ornith - syslog-auto → qwen → gemma → qwen3.6-35B-udq4
### Why ornith RPM Is Capped ### Why Strix Halo RPM Is Capped
- Direct: 40 RPM (tight) — Strix Halo is shared with compression tasks - Direct (qwen3.6-35B-udq4): 40 RPM (tight) — Strix Halo is shared with compression tasks
- Via syslog-auto: 60 RPM (moderate) — prevents flooding when multiple agents use syslog-auto simultaneously - Via syslog-auto: 60 RPM (moderate) — prevents flooding when multiple agents use syslog-auto simultaneously
- Combined max: ~100 RPM across both paths — Strix Halo can sustain this at 80°C - Combined max: ~100 RPM across both paths — Strix Halo can sustain this at 80°C
@@ -163,7 +186,7 @@ Show full fleet status: GPUs, models, VRAM, context windows, parallel slots, act
3. Check LiteLLM: `curl http://192.168.68.116/health` (expect "I'm alive!") 3. Check LiteLLM: `curl http://192.168.68.116/health` (expect "I'm alive!")
4. Check LiteLLM models: `curl -H "Authorization: Bearer $MASTER_KEY" http://192.168.68.116/v1/models` 4. Check LiteLLM models: `curl -H "Authorization: Bearer $MASTER_KEY" http://192.168.68.116/v1/models`
5. Check LiteLLM timeouts: `grep -n 'timeout:' /opt/inference-harness/litellm_config.yaml` 5. Check LiteLLM timeouts: `grep -n 'timeout:' /opt/inference-harness/litellm_config.yaml`
- gemma-4-12b: 120s, qwen3.6-27B: 90s, ornith-1.0-35b: 120s - gemma-4-12b: 120s, qwen3.6-27B-code: 300s, qwen3.6-35B-udq4/strix-moe: 300s (ornith-1.0-35b does NOT exist — legacy name, do not use)
- global request_timeout: 300s, nginx proxy_read_timeout: 600s - global request_timeout: 300s, nginx proxy_read_timeout: 600s
6. Check AMD metrics: `curl http://192.168.68.15:9400/metrics` (Radeon 8060S, util%, VRAM, temp, power) 6. Check AMD metrics: `curl http://192.168.68.15:9400/metrics` (Radeon 8060S, util%, VRAM, temp, power)
7. Check port conflicts: verify only one llama-server on :8080 per host 7. Check port conflicts: verify only one llama-server on :8080 per host
@@ -203,7 +226,7 @@ If no SSH access, send Zulip DM via abiba-bot with vault update instructions.
| `/root/scripts/gpu-saturation-watchdog.py` | pi (.24) | Auto-restart stuck llama-server | | `/root/scripts/gpu-saturation-watchdog.py` | pi (.24) | Auto-restart stuck llama-server |
| `/root/dashboard/gpu-fleet.html` | pi (.24) | Live HTML dashboard | | `/root/dashboard/gpu-fleet.html` | pi (.24) | Live HTML dashboard |
| `/etc/systemd/system/llama-server.service` | .8, .110 | llama-server daemons (Nvidia GPUs) | | `/etc/systemd/system/llama-server.service` | .8, .110 | llama-server daemons (Nvidia GPUs) |
| `/etc/systemd/system/ornith-server.service` | .15 (amdpve) | llama-server daemon (Vulkan, Strix Halo). Note: `llama-server.service` and `llama-server@.service` are **masked** on .15 to prevent port 8080 collisions. | | `/etc/systemd/system/strix-server.service` | .15 (amdpve) | llama-server daemon (Vulkan, Strix Halo) running unsloth/Qwen3.6-35B-A3B-MTP-GGUF. Note: `llama-server.service` and `llama-server@.service` are **masked** on .15 to prevent port 8080 collisions. |
## Prometheus & Grafana ## Prometheus & Grafana
@@ -220,14 +243,14 @@ If no SSH access, send Zulip DM via abiba-bot with vault update instructions.
- **Router startup race**: Compose router.py doesn't call load_roster(). Reload thread sleeps 30s first. - **Router startup race**: Compose router.py doesn't call load_roster(). Reload thread sleeps 30s first.
Fix: trigger roster reload via SSH after restart, or rebuild image with startup load_roster(). Fix: trigger roster reload via SSH after restart, or rebuild image with startup load_roster().
- **LiteLLM /metrics**: Requires auth. Prometheus uses `/health/liveliness` as workaround. - **LiteLLM /metrics**: Requires auth. Prometheus uses `/health/liveliness` as workaround.
- **VRAM (2026-07-12)**: RTX 3090 at 20.7/24GB (84%), RTX 5070 at 10.0/12.2GB (82%). RTX 3090 context increased to 256K (was incorrectly documented as 128K — verified via /proc/PID/cmdline). - **VRAM (2026-07-15)**: RTX 3090 at 22.2/24.6GB (90%) with **256K context** (corrected from 131K). RTX 5070 at 10.8/12.2GB (88%) with 256K context + MTP. Strix Halo at ~9GB/64GB.
- **RTX 3090 runs `--parallel 1`** (verified 2026-07-12). RTX 5070 and Strix at parallel 2. - **RTX 3090 runs `--parallel 2`** with MTP draft (spec-type draft-mtp, spec-draft-n-max 2).
- **RTX 3090 config**: `-c 262144 -ctk turbo4 -ctv turbo4 --parallel 1 --flash-attn on --cont-batching`. Service: `/home/llmuser/llama-wrapper.sh`. - **RTX 3090 config**: `-c 262144 -ctk turbo4 -ctv turbo4 --parallel 2 --flash-attn on --cont-batching --spec-type draft-mtp`. Context corrected to 256K (2026-07-15). VRAM: 90%. Service: `/home/llmuser/llama-wrapper.sh`.
- **RTX 5070 config**: `--ctx-size 131072 --cache-type-k q4_0 --cache-type-v q4_0 --batch-size 2048 --ubatch-size 1024 --parallel 2`. Api-key standardized to `not-needed` (was `sk-loc...5678` causing 401 loops). Service: `/home/llmuser/llama-wrapper.sh`. - **RTX 5070 config (2026-07-15)**: Switched to IQ4_NL + MTP draft (Q8_0) at 256K context. Gen speed: 122 tok/s (was 70). VRAM: 10.8/12.2GB (88%). No draft model pre-upgrade due to VRAM constraints. Service: `/home/llmuser/llama-wrapper.sh`. Config: `--model gemma-4-12b-it-IQ4_NL.gguf --spec-draft-model gemma-4-12b-it-Q8_0-MTP.gguf --spec-type draft-mtp --spec-draft-n-max 4 --ctx-size 262144`.
- **LiteLLM timeout tuning (2026-07-08)**: gemma-4-12b 25→120s, qwen3.6-27B-code 40→90s, syslog-auto (qwen route) 40→90s. Nginx proxy_read_timeout: 600s. Global request_timeout: 300s. Config at `/opt/inference-harness/litellm_config.yaml`. - **LiteLLM timeout tuning (verified 2026-07-16 against `/opt/inference-harness/litellm_config.yaml` on CT 116)**: gemma-4-12b 120s, qwen3.6-27B-code 300s, qwen3.6-35B-udq4 300s, strix-moe 300s, syslog-auto routes all 300s. Nginx proxy_read_timeout: 600s. Global request_timeout: 300s.
- **Strix Halo GPU**: Vulkan is the working backend (ROCm/HIP path abandoned — HSA runtime blocked on Debian 13). Build at `/root/llama.cpp/build-vk/`, commit `4fc4ec5` (2026-07-01), ggml 0.15.3 shared-lib arch. Mesa RADV 25.0.7, KHR_coopmat fast path active. ~70 tok/s gen, 532 tok/s prompt. Service: `ornith-server.service` on port 8080, 256K context, flash-attn + q8 KV. - **Strix Halo GPU**: Vulkan is the working backend (ROCm/HIP path abandoned — HSA runtime blocked on Debian 13). Build at `/root/llama.cpp/build-vk/`, commit `4fc4ec5` (2026-07-01), ggml 0.15.3 shared-lib arch. Mesa RADV 25.0.7, KHR_coopmat fast path active. ~70 tok/s gen, 532 tok/s prompt. Service: `strix-server.service` on port 8080 (was `ornith-server.service`), model changed to `unsloth/Qwen3.6-35B-A3B-MTP-GGUF` (UD-Q4_K_M), alias `qwen3.6-35B-udq4`, 256K context, flash-attn + q8 KV. MTP support enabled for 1.4-2.2x faster inference.
- **Port conflict detection (2026-07-05)**: All 3 GPU wrappers now detect ghost processes squatting port 8080 before starting. `.8` and `.110` use inline pre-start check in `llama-wrapper.sh`; `.15` uses `/usr/local/bin/port-cleanup.sh` ExecStartPre. Replaces the blanket `pkill -9 -x llama-server` on .15 which would kill ALL llama-server instances regardless of port. Ghost detection was the root cause of .8 crash-looping for 27+ restarts (stale pid 25836 squatting 8080 after OOM kill). - **Port conflict detection (2026-07-05)**: All 3 GPU wrappers now detect ghost processes squatting port 8080 before starting. `.8` and `.110` use inline pre-start check in `llama-wrapper.sh`; `.15` uses `/usr/local/bin/port-cleanup.sh` ExecStartPre. Replaces the blanket `pkill -9 -x llama-server` on .15 which would kill ALL llama-server instances regardless of port. Ghost detection was the root cause of .8 crash-looping for 27+ restarts (stale pid 25836 squatting 8080 after OOM kill).
- **Strix Halo thermal safeguard (2026-07-02)**: `ornith-server.service` has `-n 8192` (hard generation cap per request). Without it, `--predict` defaults to -1 (infinity) — a runaway request from .123 (Mumuni) decoded 39,868 tokens over 24 min, pushing Tctl to 98°C (crit 89.8°C) and throttling 70→29 t/s. The cap bounds worst-case generation to ~5 min. Do NOT remove `-n` without a replacement ceiling. Sustained load hits ~84°C even at 92s; the APU is fanless/low-flow. Clients MUST also set `max_tokens`. - **Strix Halo thermal safeguard (2026-07-02)**: `strix-server.service` has `-n 8192` (hard generation cap per request). Without it, `--predict` defaults to -1 (infinity) — a runaway request from .123 (Mumuni) decoded 39,868 tokens over 24 min, pushing Tctl to 98°C (crit 89.8°C) and throttling 70→29 t/s. The cap bounds worst-case generation to ~5 min. Do NOT remove `-n` without a replacement ceiling. Sustained load hits ~84°C even at 92s; the APU is fanless/low-flow. Clients MUST also set `max_tokens`.
- **Port 8080 firewall**: amdpve iptables restricts 8080 to 192.168.68.116 (LiteLLM/router host) only. All inbound connections are from .116 (LiteLLM proxied via nginx). Localhost curls hang (SYN dropped). Always test from .116. - **Port 8080 firewall**: amdpve iptables restricts 8080 to 192.168.68.116 (LiteLLM/router host) only. All inbound connections are from .116 (LiteLLM proxied via nginx). Localhost curls hang (SYN dropped). Always test from .116.
- **Router sidecar fallback**: `router.py` `check_gpu_health()` now probes GPU `/health` directly when sidecar at :8090 is absent. Sidecar JSON exporters not deployed on any GPU host — router relies on GPU-direct fallback. - **Router sidecar fallback**: `router.py` `check_gpu_health()` now probes GPU `/health` directly when sidecar at :8090 is absent. Sidecar JSON exporters not deployed on any GPU host — router relies on GPU-direct fallback.
- **Router GPU_MOE_URL bug (fixed 2026-07-01)**: docker-compose had `GPU_MOE_URL=.110:8080` (gemma host) instead of `.15:8080` (amdpve). Corrected. - **Router GPU_MOE_URL bug (fixed 2026-07-01)**: docker-compose had `GPU_MOE_URL=.110:8080` (gemma host) instead of `.15:8080` (amdpve). Corrected.
@@ -237,11 +260,14 @@ If no SSH access, send Zulip DM via abiba-bot with vault update instructions.
## GPU Inference Benchmarks (Current) ## GPU Inference Benchmarks (Current)
| GPU | Model | Gen tok/s | Prompt tok/s | Baseline | Samples | | GPU | Model | Gen tok/s | Prompt tok/s | Baseline | Context |
|-----|-------|-----------|--------------|----------|---------| |-----|-------|-----------|--------------|----------|---------|
| RTX 3090 (.8) | qwen3.6-27B-code | 75 | 305 | 74 | 6 | | RTX 3090 (.8) | qwen3.6-27B-code (MTP) | **63** | — | | **256K** |
| RTX 5070 (.110) | gemma-4-12b | 75 | 323 | 75 | 6 | | RTX 5070 (.110) | gemma-4-12b (IQ4_NL+MTP) | **191** | — | | **256K** |
| Strix Halo (.15) | ornith-1.0-35b | 70 | 532 | 70 | 6 | | Strix Halo (.15) | qwen3.6-35B-udq4 | **71** | — | | **256K** |
Benchmarks from 2026-07-15 verification run. RTX 5070 MTP provides 2.7x speedup over pre-upgrade 70 tok/s.
All 3 GPUs now at 256K context (2026-07-15).
Benchmarks run through LiteLLM proxy (192.168.68.116:4001) every 5 minutes. Benchmarks run through LiteLLM proxy (192.168.68.116:4001) every 5 minutes.
Degradation alerts fire at 30% (warning) and 50% (critical) below baseline. Degradation alerts fire at 30% (warning) and 50% (critical) below baseline.
@@ -249,12 +275,55 @@ History stored at `/root/data/toks-history.json` with 7-day rolling window.
**Note (2026-07-01)**: Strix Halo prompt tok/s jumped 209→532 after Vulkan rebuild (cooperative-matrix fast path now active on GFX1151). Baseline may need re-calibration. **Note (2026-07-01)**: Strix Halo prompt tok/s jumped 209→532 after Vulkan rebuild (cooperative-matrix fast path now active on GFX1151). Baseline may need re-calibration.
## Agent Config Implications (2026-07-12) ## Agent Config Implications (2026-07-15)
With RTX 3090 at 256K context (verified July 2026): ### Stable Aliases — CRITICAL
- Agents using `syslog-auto` (55/30/15 qwen+ornith+gemma): `context_length: 262144` — ornith and qwen both support it
- Agents using `qwen3.6-27B-code` directly: `context_length: 262144` (256K ctx verified) All agent configs MUST use stable role-based aliases, never model-specific names:
- Agents using `gemma-4-12b` directly (auxiliary tasks): `context_length: 131072` - `compression.model: strix-moe` (NOT `qwen3.6-35B-udq4`)
- Compression threshold at 0.65: fires at ~170K for 262K context window on Strix Halo - `auxiliary.vision.model: gpu-light` (NOT `gemma-4-12b`)
- All Hermes clients MUST set `max_tokens: 4096` — first line of defense before server-side `-n 8192` cap - `delegation.model: gpu-dense` (NOT `qwen3.6-27B-code`)
- Port 8080 is used on all 3 GPU hosts (not 8090 as previously documented) - `auxiliary.web_extract.model: gpu-light`
When the underlying model is swapped, only the LiteLLM config changes — agent configs are untouched.
### Context Windows
- RTX 3090: **256K** (was 131K, bumped 2026-07-15) | RTX 5070: **256K** (up from 131K) | Strix Halo: **256K**
- **Mumuni compression context**: 128K (down from 256K) — ensures compression model doesn't timeout
- Compression threshold 0.65: fires at ~85K for 128K context window
- Mumuni compression model alias: `strix-moe` with 300s timeout
### Mumuni Agent Profile
Mumuni (CT114, 192.168.68.123) is the primary business assistant. This profile is the reference for all agent configs:
| Setting | Value | Notes |
|---------|-------|-------|
| `model.default` | `syslog-auto` | Weighted pool (55% qwen, 30% strix, 15% gemma) |
| `model.provider` | `custom:litellm` | LiteLLM on CT116 |
| `compression.model` | `strix-moe` | Stable alias — survives model swaps |
| `aux.compression.model` | `strix-moe` | Compression auxiliary model |
| `aux.vision.model` | `gpu-light` | Vision tasks (RTX 5070) |
| `aux.web_extract.model` | `gpu-light` | Web extraction |
| `delegation.model` | `gpu-dense` | Sub-agent reasoning (RTX 3090) |
| `context.max_context_window` | 262144 (256K) | Fixed 2026-07-16 (was 131072 — caused premature compression, WAL #1300) |
| `compression.threshold` | 0.65 | Triggers at ~85K |
| `compression.target_ratio` | 0.3 | Compresses to ~38K |
| `compression.protect_last_n` | 40 | Preserves last 40 messages |
| `memory.memory_char_limit` | 800 | Brief memory entries |
| `personalities` | `creative` | Creative assistant personality |
| Platforms | cli, discord, homeassistant, signal, telegram, zulip | All Hermes platforms |
| Main model timeout | 300s | LiteLLM global timeout |
| Compression model timeout | 300s | ornith timeout increased from 120s |
### Agent Update Status (2026-07-15)
| Agent | Host | Status |
|-------|------|--------|
| **Mumuni** | CT114 (.123) | ✅ Updated to stable aliases |
| **Tanko** | CT112 (.122) | ✅ Updated to stable aliases |
| **Koby** | CT111 (.129) | ❌ SSH unreachable — needs Zulip DM |
| **Koonimo** | CT113 | ❌ SSH unreachable — needs Zulip DM |
| **Kagenz0** | CT105 | ❌ SSH unreachable — needs Zulip DM |
All Hermes clients MUST set `max_tokens: 4096` — first line of defense before server-side `-n 8192` cap.
+1 -1
View File
@@ -120,7 +120,7 @@ depends_on:
- **Detect**: Benchmark tok/s vs baseline for each GPU at current context - **Detect**: Benchmark tok/s vs baseline for each GPU at current context
- RTX 3090 (256K ctx, qwen3.6-27B-code): target 75+ tok/s — currently at baseline - RTX 3090 (256K ctx, qwen3.6-27B-code): target 75+ tok/s — currently at baseline
- RTX 5070 (131K ctx, gemma-4-12b): target 76+ tok/s — optimal for vision/web role - RTX 5070 (131K ctx, gemma-4-12b): target 76+ tok/s — optimal for vision/web role
- Strix Halo (256K ctx, ornith-1.0-35b): target 70+ tok/s — currently above baseline - Strix Halo (256K ctx, strix-moe / qwen3.6-35B-udq4): target 70+ tok/s — currently above baseline
- **Fix**: - **Fix**:
- If tok/s > baseline → context has headroom, consider increasing - If tok/s > baseline → context has headroom, consider increasing
- If tok/s < 90% baseline → reduce context by 25% and retest - If tok/s < 90% baseline → reduce context by 25% and retest
+1 -1
View File
@@ -192,7 +192,7 @@ Key is injected via `infisical run --` wrapper at PM2 startup:
"apiKey": "${LITELLM_API_KEY}", "apiKey": "${LITELLM_API_KEY}",
"models": [ "models": [
{ "id": "syslog-auto" }, { "id": "syslog-auto" },
{ "id": "ornith-1.0-35b" }, { "id": "strix-moe" },
{ "id": "qwen3.6-27B-code" }, { "id": "qwen3.6-27B-code" },
{ "id": "gemma-4-12b" } { "id": "gemma-4-12b" }
] ]
+95 -27
View File
@@ -5,10 +5,12 @@ description: >
Standard Hermes configuration template for Syslog Solution LLC agents. Standard Hermes configuration template for Syslog Solution LLC agents.
Enforces shared infrastructure setup (Firecrawl, SearXNG, local models, Enforces shared infrastructure setup (Firecrawl, SearXNG, local models,
RA-H OS MCP) while keeping agent-specific API keys and model choices. RA-H OS MCP) while keeping agent-specific API keys and model choices.
UPDATED 2026-07-12: GPU workload redistributed. Compression → Strix Halo (ornith-1.0-35b). UPDATED 2026-07-16: Compression model is the stable alias `strix-moe` (NOT `ornith-1.0-35b`,
RTX 3090 context verified at 256K (was incorrectly documented as 128K). which LiteLLM does not serve). All 3 GPUs verified 256K (RTX 5070 bumped 131K→256K on Jul 15).
RTX 5070 stays at 131K for vision/web. Infisical .env fallback required (Rule 3). Added Rule 12 (Context-Issue Diagnostic) + Rule 13 (.env fallback enforcement) from the
Parallel counts corrected (RTX 3090=1, RTX 5070=2, Strix=2). 2026-07-16 Mumuni root-cause investigation (WAL #1300).
UPDATED 2026-07-12: GPU workload redistributed. Compression → Strix Halo. RTX 3090 context
verified at 256K. Infisical .env fallback required (Rule 3/13).
--- ---
## Maintains ## Maintains
@@ -31,7 +33,7 @@ Sub-agent profiles inherit auth from the main config — no separate keys needed
| Agent | Key Alias | Host | SSH | Sub-Agents | | Agent | Key Alias | Host | SSH | Sub-Agents |
|-------|-----------|------|-----|-----------| |-------|-----------|------|-----|-----------|
| Tanko | `tanko-*` | 192.168.68.122 | jerome@.122 | — | | Tanko | `tanko-*` | 192.168.68.122 | jerome@.122 | — |
| Mumuni | `mumuni-jul2026` | 192.168.68.123 | root@.123 | 6 profiles ✱ | | Mumuni | `mumuni` | 192.168.68.123 | root@.123 | 6 profiles ✱ |
| Abiba | `abiba-pi` | 192.168.68.24 | local | — | | Abiba | `abiba-pi` | 192.168.68.24 | local | — |
| Koby | `koby` | CT 111 (tdunna) | Zulip | — | | Koby | `koby` | CT 111 (tdunna) | Zulip | — |
| Koonimo | `koonimo` | CT 113 (baggy) | Zulip | — | | Koonimo | `koonimo` | CT 113 (baggy) | Zulip | — |
@@ -94,7 +96,7 @@ work immediately after restart.
```yaml ```yaml
# ─── Model Selection ─── # ─── Model Selection ───
model: model:
default: <agent_model> # e.g., ornith-1.0-35b, qwen3.6-27B-code default: <agent_model> # e.g., strix-moe, qwen3.6-27B-code, syslog-auto
provider: harness provider: harness
base_url: http://192.168.68.116/v1 base_url: http://192.168.68.116/v1
api_key_env: LITELLM_API_KEY # Injected via infisical run -- wrapper api_key_env: LITELLM_API_KEY # Injected via infisical run -- wrapper
@@ -129,10 +131,9 @@ mcp_servers:
# ─── Compression ─── # ─── Compression ───
compression: compression:
enabled: true enabled: true
model: ornith-1.0-35b # ⚠️ Must match auxiliary.compression.model model: strix-moe # ⚠️ Must match auxiliary.compression.model. Stable alias (gpu-fleet § Stable Role-Based Aliases). NOT ornith-1.0-35b (LiteLLM does not serve that name).
provider: harness provider: harness
max_context_window: 262144 # For Strix Halo compression (256K ctx). max_context_window: 262144 # MUST match actual GPU capacity. All 3 GPUs are 256K (Jul 15).
# Set 131072 if using gemma directly.
threshold: 0.65 # Fires at ~170K for 262K window, ~85K for 128K threshold: 0.65 # Fires at ~170K for 262K window, ~85K for 128K
target_ratio: 0.30 target_ratio: 0.30
protect_last_n: 40 protect_last_n: 40
@@ -142,37 +143,48 @@ compression:
# ─── Auxiliary Tasks (CONSISTENCY RULE) ─── # ─── Auxiliary Tasks (CONSISTENCY RULE) ───
# All auxiliary services MUST use identical model, base_url, and api_key_env: # All auxiliary services MUST use identical model, base_url, and api_key_env:
# model: gemma-4-12b # model: gpu-light # stable alias (NOT raw "gemma-4-12b")
# base_url: http://192.168.68.116/v1 # base_url: http://192.168.68.116/v1
# api_key_env: LITELLM_API_KEY # api_key_env: LITELLM_API_KEY
# Do NOT use syslog-auto for auxiliary tasks — it routes to the primary GPU. # Do NOT use syslog-auto for auxiliary tasks — it routes to the primary GPU.
# gemma-4-12b is a lightweight 12B model on the RTX 5070, freeing the Strix Halo # gpu-light = RTX 5070 (12B), freeing the Strix Halo for agent reasoning.
# for agent reasoning. # Heavy aux (delegation, x_search) use gpu-dense (RTX 3090) instead.
# NEVER use raw model names (gemma-4-12b, qwen3.6-27B-code, qwen3.6-35B-udq4)
# in agent configs — use the stable aliases so model swaps don't break agents.
auxiliary: auxiliary:
vision: vision:
provider: harness provider: harness
model: gemma-4-12b model: gpu-light # stable alias for RTX 5070 (was raw gemma-4-12b)
base_url: http://192.168.68.116/v1 base_url: http://192.168.68.116/v1
api_key_env: LITELLM_API_KEY api_key_env: LITELLM_API_KEY
timeout: 60 timeout: 60
download_timeout: 30 download_timeout: 30
web_extract: web_extract:
provider: harness provider: harness
model: gemma-4-12b model: gpu-light # stable alias for RTX 5070
base_url: http://192.168.68.116/v1 base_url: http://192.168.68.116/v1
api_key_env: LITELLM_API_KEY api_key_env: LITELLM_API_KEY
timeout: 30 timeout: 30
compression: compression:
provider: harness provider: harness
model: ornith-1.0-35b model: strix-moe # MUST match compression.model above. Stable alias for Strix Halo.
base_url: http://192.168.68.116/v1 # Rule 5: /v1 NOT /litellm/v1
api_key_env: LITELLM_API_KEY
timeout: 300 # gpu-fleet: 300s for large-history summarization (was 60)
# ─── Delegation / Heavy Aux (use gpu-dense = RTX 3090) ───
# delegation.model and x_search.model use gpu-dense (NOT raw qwen3.6-27B-code).
delegation:
model: gpu-dense # stable alias for RTX 3090 (was raw qwen3.6-27B-code)
provider: harness
base_url: http://192.168.68.116/v1 base_url: http://192.168.68.116/v1
api_key_env: LITELLM_API_KEY api_key_env: LITELLM_API_KEY
timeout: 60
# ─── Custom Provider ─── # ─── Custom Provider ───
custom_providers: custom_providers:
- name: harness - name: harness
model: <agent_model> model: syslog-auto # weighted pool (default)
base_url: http://192.168.68.116/v1 base_url: http://192.168.68.116/v1
api_key_env: LITELLM_API_KEY api_key_env: LITELLM_API_KEY
api_mode: chat_completions api_mode: chat_completions
@@ -236,27 +248,29 @@ The following MUST be identical across ALL profiles:
- Apply to BOTH main config AND all sub-agent profiles - Apply to BOTH main config AND all sub-agent profiles
- For agents needing longer outputs: raise to 8192, but never omit - For agents needing longer outputs: raise to 8192, but never omit
### Rule 7: Auxiliary Model Consistency (UPDATED July 2026) ### Rule 7: Auxiliary Model Consistency (UPDATED 2026-07-16)
- Vision and web_extract use `gemma-4-12b` (RTX 5070 — 12GB, vision-optimized) - Vision and web_extract use `gemma-4-12b` (RTX 5070 — 12GB, vision-optimized)
- Compression uses `ornith-1.0-35b` (Strix Halo — 64GB, 256K ctx, compression-optimized) - Compression uses `strix-moe` (stable alias for Strix Halo — 64GB, 256K ctx, compression-optimized)
- **`strix-moe` is the only valid compression model name** — LiteLLM does NOT serve `ornith-1.0-35b`
(it serves `strix-moe`, `qwen3.6-35B-udq4`, `gpu-dense`, `gpu-light`, `syslog-auto`, `gemma-4-12b`, `qwen3.6-27B-code`). Old configs with `ornith-1.0-35b` cause 403/model-not-found on compression calls.
- All auxiliary services MUST use identical routing: - All auxiliary services MUST use identical routing:
- `base_url: http://192.168.68.116/v1` - `base_url: http://192.168.68.116/v1` (Rule 5: `/v1`, NOT `/litellm/v1`)
- `api_key_env: LITELLM_API_KEY` - `api_key_env: LITELLM_API_KEY`
- **Do NOT use `syslog-auto`** for auxiliary tasks — it routes unpredictably - **Do NOT use `syslog-auto`** for auxiliary tasks — it routes unpredictably
- **Compression moved to Strix Halo (July 2026)**: The ornith-1.0-35b model on Strix Halo - **Compression on Strix Halo**: The strix-moe alias routes to Strix Halo
(64GB UMA, 256K context, 72.4 tok/s) is the designated compression GPU. This frees the (64GB UMA, 256K context) — the designated compression GPU. This frees the
RTX 5070 for vision and web search, and the RTX 3090 for heavy reasoning. RTX 5070 for vision and web search, and the RTX 3090 for heavy reasoning.
- The `compression:` block's `model` MUST match `auxiliary: compression: model` - The `compression:` block's `model` MUST match `auxiliary: compression: model`
- The `compression: max_context_window: 262144` MUST match Strix Halo's actual capacity - The `compression: max_context_window: 262144` MUST match Strix Halo's actual capacity
### Rule 8: GPU Workload Distribution (July 2026) ### Rule 8: GPU Workload Distribution (UPDATED 2026-07-16)
- **RTX 3090 (24GB, 256K ctx, qwen3.6-27B-code)**: Heavy reasoning, code gen, long conversations - **RTX 3090 (24GB, 256K ctx, qwen3.6-27B-code)**: Heavy reasoning, code gen, long conversations
- **RTX 5070 (12GB, 131K ctx, gemma-4-12b)**: Vision, web search, quick tasks, web_extract - **RTX 5070 (12GB, 256K ctx, gemma-4-12b)**: Vision, web search, quick tasks, web_extract (bumped 131K→256K Jul 15; IQ4_NL+MTP, 88% VRAM)
- **Strix Halo (64GB, 256K ctx, ornith-1.0-35b)**: Context compression, summarization, long docs - **Strix Halo (64GB, 256K ctx, strix-moe)**: Context compression, summarization, long docs
- Agent profiles MUST route auxiliary tasks to the correct GPU: - Agent profiles MUST route auxiliary tasks to the correct GPU:
- `auxiliary.vision.model: gemma-4-12b` (RTX 5070) - `auxiliary.vision.model: gemma-4-12b` (RTX 5070)
- `auxiliary.web_extract.model: gemma-4-12b` (RTX 5070) - `auxiliary.web_extract.model: gemma-4-12b` (RTX 5070)
- `auxiliary.compression.model: ornith-1.0-35b` (Strix Halo) - `auxiliary.compression.model: strix-moe` (Strix Halo)
- Default model (`model.default`) and custom_provider remain `syslog-auto` for auto-routing - Default model (`model.default`) and custom_provider remain `syslog-auto` for auto-routing
- For 262K context window: `threshold: 0.65` (fires at ~170K tokens) - For 262K context window: `threshold: 0.65` (fires at ~170K tokens)
- Do NOT use `threshold: 0.25` — this fires at 65K, causing premature context loss - Do NOT use `threshold: 0.25` — this fires at 65K, causing premature context loss
@@ -274,7 +288,7 @@ The following MUST be identical across ALL profiles:
### Rule 10: Default Model Must Be `syslog-auto` (All Agents) ### Rule 10: Default Model Must Be `syslog-auto` (All Agents)
- **Hermes agents**: `model.default: syslog-auto`, `custom_providers[0].model: syslog-auto` - **Hermes agents**: `model.default: syslog-auto`, `custom_providers[0].model: syslog-auto`
- **pi agents**: `defaultModel: syslog-auto` in `settings.json`, first model in `models.json` - **pi agents**: `defaultModel: syslog-auto` in `settings.json`, first model in `models.json`
- `syslog-auto` is the LiteLLM routing model — it load-balances between ornith-1.0-35b - `syslog-auto` is the LiteLLM routing model — it load-balances between strix-moe
and qwen3.6-27B-code, with gemma-4-12b as fallback. Using it protects against: and qwen3.6-27B-code, with gemma-4-12b as fallback. Using it protects against:
- Model name typos that cause 403 errors and silent worker failures - Model name typos that cause 403 errors and silent worker failures
- Single GPU downtime (routing falls back automatically) - Single GPU downtime (routing falls back automatically)
@@ -293,6 +307,60 @@ The following MUST be identical across ALL profiles:
- A non-existent model ID causes 403 errors that silently break the pi RPC worker - A non-existent model ID causes 403 errors that silently break the pi RPC worker
(no `agent_end` emitted, worker stays "busy", Zulip messages pile up unprocessed) (no `agent_end` emitted, worker stays "busy", Zulip messages pile up unprocessed)
### Rule 12: Context-Issue Diagnostic Checklist (ADDED 2026-07-16, WAL #1300)
When an agent shows "context issues" (premature compression, 401s, 504s, DeepSeek fallback),
verify ALL FOUR of these against the live config. They are the only root causes found in production:
1. **max_context_window correct?** — BOTH `compression.max_context_window` AND `context.max_context_window`
MUST be `262144` (all GPUs are 256K). A value of `131072` causes premature compression at
~83K instead of ~170K. Check: `grep -n max_context_window ~/.hermes/config.yaml`
2. **base_url uses /v1 NOT /litellm/v1?**`custom_providers[0].base_url`, `delegation.base_url`,
and ALL `auxiliary.*.base_url` MUST be `http://192.168.68.116/v1` (Rule 5). nginx `/litellm/`
has a 60s default timeout → 504 on any inference >60s; `/v1/` has 600s.
Check: `grep -n 'litellm/v1' ~/.hermes/config.yaml` (must return NOTHING)
3. **LITELLM_API_KEY valid?** — The key must be a real LiteLLM key (`sk-` + 64 hex, 67 chars).
Malformed values (e.g. `sk-_SWAl_Vu_…`, 47 chars) return 401 → DeepSeek fallback.
Verify: `curl -s -o /dev/null -w '%{http_code}' -H "Authorization: Bearer $LITELLM_API_KEY" http://192.168.68.116/v1/models` (must be 200)
4. **custom_providers aligned?**`model: syslog-auto` (Rule 10), `api_mode: chat_completions`
(NOT `responses`). A wrong api_mode causes silent request failures.
One-line agent health check (run on the agent host):
```bash
PID=$(pgrep -f "python -m hermes_cli.main gateway run" | head -1)
cat /proc/$PID/environ | tr '\0' '\n' | grep ^LITELLM_API_KEY= | sed 's/=.*/<set>/'
curl -s -o /dev/null -w 'key_health: %{http_code}\n' -H "Authorization: Bearer $(cat /proc/$PID/environ | tr '\0' '\n' | grep ^LITELLM_API_KEY= | cut -d= -f2)" http://192.168.68.116/v1/models
```
### Rule 13: API Key Injection — Two Patterns (UPDATED 2026-07-16, WAL #1300)
Agents inject `LITELLM_API_KEY` via ONE of two mechanisms. Both are valid; the contract
requirement is that the key is a **valid LiteLLM virtual key** (HTTP 200 on /v1/models).
**Pattern A — systemd drop-in (Koby, Koonimo, and any agent without infisical wrapper):**
A systemd drop-in `/etc/systemd/system/hermes-gateway.service.d/litellm-key.conf` sets the key:
```ini
[Service]
Environment="LITELLM_API_KEY=sk-<VALID_KEY>"
```
The service unit `hermes-gateway.service` runs `python -m hermes_cli.main gateway run --replace`
directly (no infisical). Apply with `systemctl daemon-reload && systemctl restart hermes-gateway`.
- Koonimo (CT113/.114): service = `hermes-gateway.service`, drop-in has the key.
- Koby (CT111/.129): service = `hermes-gateway.service` (created 2026-07-16), ExecStart uses `--replace`
to win the lock against stray `hermes gateway restart` invocations. Key also in `/etc/environment`.
**Pattern B — infisical-gateway.sh wrapper (Mumuni):**
The wrapper sources `~/.hermes/.env` then exports `LITELLM_API_KEY="$<AGENT>_LITELLM_API_KEY"`.
See litellm-api-keys.prose.md § Machine Identity for Vault Writes for vault sync.
**Verification (all agents):**
```bash
GP=$(pgrep -f "python -m hermes_cli.main gateway run" | head -1)
K=$(cat /proc/$GP/environ | tr '\0' '\n' | grep '^LITELLM_API_KEY=' | cut -d= -f2)
curl -s -o /dev/null -w '%{http_code}' -H "Authorization: Bearer $K" http://192.168.68.116/v1/models # must be 200
```
- `/etc/environment` is NO LONGER the canonical key source (stale values there caused 401s).
- Do NOT leave a hardcoded stale key in `/etc/environment` — it shadows the drop-in/wrapper.
## Execution ## Execution
1. **Check current config** — Read the target agent's config.yaml 1. **Check current config** — Read the target agent's config.yaml
+5
View File
@@ -72,6 +72,11 @@ External providers are **explicitly exempt** and may use hardcoded keys:
## Standard Pattern ## Standard Pattern
> **Canonical vault process (2026-07-16):** see `litellm-api-keys` § Production Vault Access Process.
> All agents MUST use the `infisical-gateway.sh` wrapper (live vault injection). Hardcoded systemd
> drop-ins / config.yaml keys are DEPRECATED — they rot on rotation (root cause of the 2026-07-16 401 storm).
> 4/5 agents migrated; tanko (user jerome) pending.
```yaml ```yaml
# ✅ CORRECT — all harness/litellm providers (authenticated path, NO /responses suffix) # ✅ CORRECT — all harness/litellm providers (authenticated path, NO /responses suffix)
model: model:
+151 -3
View File
@@ -14,9 +14,15 @@ description: >
expire/404. The .env fallback prevents agents from running without keys. expire/404. The .env fallback prevents agents from running without keys.
Tanko incident: token 404 → gateway had no LITELLM_API_KEY for hours. Tanko incident: token 404 → gateway had no LITELLM_API_KEY for hours.
UPDATED 2026-07-16: Vault is SYNCED (session-13 keys written to vault via abiba service
token, all validate 200). Koby/Koonimo migrated from hardcoded drop-ins to the
infisical-gateway.sh wrapper (live vault injection). 4/5 agents now vault-backed.
Canonical process: see § Production Vault Access Process. Tanko (user jerome) pending.
Abiba's key is now a proper agent key (NOT the master key — stale note removed).
Current key inventory and agent list: see gpu-fleet.prose.md § Agent Keys. Current key inventory and agent list: see gpu-fleet.prose.md § Agent Keys.
Source of truth for LiteLLM config: /opt/inference-harness/litellm_config.yaml Source of truth for LiteLLM config: /opt/inference-harness/litellm_config.yaml
on CT 116. Last verified: 2026-07-12. on CT 116. Last verified: 2026-07-16.
--- ---
## Parameters ## Parameters
@@ -51,8 +57,8 @@ description: >
- Generate new key with key_alias: "{agent_name}" (e.g., "tanko" — bare name, no date) - Generate new key with key_alias: "{agent_name}" (e.g., "tanko" — bare name, no date)
- Set metadata: { "agent": "{agent_name}", "purpose": "agent-inference" } - Set metadata: { "agent": "{agent_name}", "purpose": "agent-inference" }
- Duration is null (permanent) — inherited from litellm default_key_generate_params - Duration is null (permanent) — inherited from litellm default_key_generate_params
- Set models: ["syslog-auto", "qwen3.6-27B-code", "gemma-4-12b", "ornith-1.0-35b"] - Set models: ["syslog-auto", "qwen3.6-27B-code", "gemma-4-12b", "strix-moe", "gpu-dense", "gpu-light", "qwen3.6-35B-udq4"]
- Note: qwen3.6-35B-A3B removed from fleet (was never deployed on any GPU) - Note: `ornith-1.0-35b` is NOT a valid LiteLLM model name (use `strix-moe`, the stable alias). qwen3.6-35B-A3B removed from fleet (was never deployed).
- Return the new key - Return the new key
5. **If action == "rotate"**: 5. **If action == "rotate"**:
- Generate new key with same alias (LiteLLM replaces the old key) - Generate new key with same alias (LiteLLM replaces the old key)
@@ -67,3 +73,145 @@ description: >
- Test the key against LiteLLM /v1/models - Test the key against LiteLLM /v1/models
- Confirm key alias matches agent_name in LiteLLM key list - Confirm key alias matches agent_name in LiteLLM key list
- Verify agent gateway uses vault wrapper: `cat /proc/<pid>/cmdline` shows `infisical run` - Verify agent gateway uses vault wrapper: `cat /proc/<pid>/cmdline` shows `infisical run`
## Production Vault Access Process (canonical, 2026-07-16)
The non-fail approach to agentic vault access. Deployed on 4/5 agents (tanko pending —
runs as user `jerome`, not systemd root, needs user-scope adaptation).
### The canonical pattern
1. **infisical CLI** installed on the host (`/usr/local/bin/infisical` or `/usr/bin/infisical`).
2. **Service token** (Infisical Machine Identity, `st.…`) stored root-only at `/root/.infisical-token` (`chmod 600`).
- Interim: the shared `abiba` service token (`st.8e848433…`) has READ+WRITE on the `agents` project.
- Proper: one machine identity per agent (create in Infisical UI → Project Settings → Machine Identities).
3. **`infisical-gateway.sh` wrapper** at `/root/.hermes/infisical-gateway.sh` (`chmod 700`):
```bash
#!/bin/bash
export INFISICAL_API_URL="https://vault.sysloggh.net"
TOKEN=$(cat /root/.infisical-token)
LOG=/root/.hermes/logs/gateway.log; mkdir -p /root/.hermes/logs
while true; do
infisical run --token="$TOKEN" --projectId=322fceab-39da-4854-a55a-568e76c0f13f \
--env=prod --domain=https://vault.sysloggh.net -- bash -c '
. /root/.hermes/.env 2>/dev/null # [FALLBACK Rule 3] safety net only
export LITELLM_API_KEY="$<AGENT>_LITELLM_API_KEY"
exec <HERMES_VENV>/bin/python -m hermes_cli.main gateway run
' >> $LOG 2>&1
sleep 5 # restart on exit
done
```
4. **Agent key in vault** as `<AGENT>_LITELLM_API_KEY` (e.g. `KOBY_LITELLM_API_KEY`). Vault = source of truth.
5. **`.env` fallback** at `/root/.hermes/.env` (`chmod 600`) with the same key — safety net ONLY for vault outage (Rule 3/13). Must be kept in sync on rotation.
6. **systemd service** `hermes-gateway.service` with `ExecStart=/root/.hermes/infisical-gateway.sh`. NO `litellm-key.conf` drop-in (those hardcode keys and rot).
7. **NEVER hardcode** LiteLLM keys in systemd drop-ins, config.yaml, or /etc/environment. The wrapper injects live from vault.
### Why this is non-fail
- **No rot**: keys pulled live from vault at every gateway start. Rotation = one `infisical secrets set` + `systemctl restart`. No per-host file edits.
- **Survives vault outage**: the `.env` fallback (Rule 3) keeps the gateway running if Infisical is unreachable.
- **Survives gateway crash**: the wrapper's `while true` + systemd `Restart=on-failure` revive the gateway.
- **Auditable**: `cat /proc/$(pgrep hermes_cli)/environ` shows the live key; `infisical secrets` shows the vault source.
### Migration status (2026-07-16)
| Agent | Host | Pattern | Vault key | Status |
|-------|------|---------|-----------|--------|
| abiba | .24 | `infisical run` (pi agent wrapper, service token) | ABIBA_LITELLM_API_KEY | ✅ vault-backed |
| mumuni | .123 | infisical-gateway.sh + user-login machine identity | MUMUNI_LITELLM_API_KEY | ✅ vault-backed |
| koby | .129 | infisical-gateway.sh + service token (migrated 2026-07-16) | KOBY_LITELLM_API_KEY | ✅ vault-backed, Zulip (tanko-bot@) + Telegram |
| koonimo | .114 | infisical-gateway.sh + service token (migrated 2026-07-16) | KOONIMO_LITELLM_API_KEY | ✅ vault-backed |
> **Baggy = Koonimo (CT 113).** Deleted `BAGGY_LITELLM_API_KEY` from vault 2026-07-16. Only `KOONIMO_LITELLM_API_KEY` exists — one secret per agent.
| tanko | .122 | **hardcoded in config.yaml** (runs as user jerome, not systemd) | TANKO_LITELLM_API_KEY | ⚠️ TODO: migrate to user-scope wrapper |
### Tanko migration (pending)
Tanko runs the gateway as user `jerome` (not root/systemd), with the key hardcoded in
`/home/jerome/.hermes/config.yaml` (`api_key: sk-CggiHWlamQy…`, valid but not vault-sourced).
Migration: create a user-scope systemd service (`~/.config/systemd/user/hermes-gateway.service`)
with `infisical-gateway.sh` wrapper in jerome's home, token at `~/.infisical-token`, lingering
enabled (`loginctl enable-linger jerome`) so the user service runs without a login session.
### Koby migration lessons (2026-07-16)
Migrated Koby from hardcoded systemd drop-in → `infisical-gateway.sh` wrapper.
**Two mistakes I made that broke the agent:**
1. **Overwrote `/root/.hermes/.env`** without backing it up. The Zulip API key only existed
in the running process memory — the old .env was minimal (just LiteLLM key). Zulip creds were
inherited from the pre-migration gateway env, not stored in any file. Lost on restart.
2. **Only injected `LITELLM_API_KEY`** in the wrapper — forgot Zulip + Telegram credentials.
Agents need ALL their platform env vars. Missing vars cause silent adapter failures.
**How Koby actually connects (2026-07-16):**
- Zulip: shares **Tanko's bot** (`tanko-bot@chat.sysloggh.net`, `TANKO_ZULIP_API_KEY=5PeD6f3zo…`).
Koby doesn't have its own Zulip bot (koby-bot@ doesn't exist in the swarm config).
- Telegram: token `828640…` recovered from `.env.bak-20260603` (18KB backup from June 2026).
Allowed users: 6679773481. Home channel: 6679773481.
- Both platforms now connect through the wrapper's env injection.
**Golden rule for gateway restarts:** always `cat /proc/<pid>/environ` before killing the old
process — captures the live env set. Especially important when migrating gateways between
injection mechanisms.
### Key rotation procedure (one vault operation with this standard)
1. Generate new key: `POST /key/generate` (master key, admin).
2. Update vault: `infisical secrets set <AGENT>_LITELLM_API_KEY=sk-NEW --token=$TOKEN --projectId=322fceab… --env=prod --domain=https://vault.sysloggh.net`.
3. Update `.env` fallback: `echo '<AGENT>_LITELLM_API_KEY=sk-NEW' > /root/.hermes/.env && chmod 600 /root/.hermes/.env`.
4. Restart: `systemctl restart hermes-gateway`. The wrapper pulls the new key live.
5. Verify: `curl -H "Authorization: Bearer sk-NEW" http://192.168.68.116/v1/models` → 200.
## Machine Identity for Vault Writes (ADDED 2026-07-16, WAL #1300)
**Problem:** The infisical CLI on agent hosts is logged in as a user session (jerome@sysloggh.com).
In CLI v0.38.0, `infisical secrets set` / `infisical export` fail with "project id missing" / "workspace
key 404" — a known bug where user-session auth works for `run` but NOT for `secrets set`. The apt
repo only ships 0.38.0, so `apt upgrade` does not help.
**Proper fix — Machine Identity (Infisical automation best practice):**
Create a machine identity with READ+WRITE scope on the `agents` project (project_id=
`322fceab-39da-4854-a55a-568e76c0f13f`, env `prod`). Store client_id + client_secret securely.
Then vault writes work from any host:
```bash
# Get a machine-identity access token
TOKEN=$(curl -fsSL -X POST https://vault.sysloggh.net/api/v1/auth/universal-auth/login \
-H 'Content-Type: application/json' \
-d '{"clientId":"<CLIENT_ID>","clientSecret":"<CLIENT_SECRET>"}' | jq -r .accessToken)
# Write a secret via REST API v3
curl -fsSL -X PATCH https://vault.sysloggh.net/api/v3/secrets/MUMUNI_LITELLM_API_KEY \
-H "Authorization: Bearer $TOKEN" -H 'Content-Type: application/json' \
-d '{"environment":"prod","secretValue":"sk-<NEW_KEY>","workspaceId":"<WORKSPACE_ID>","type":"shared"}'
# OR via CLI: infisical secrets set --token=$TOKEN --projectId=322fceab... --env=prod ...
```
Creation requires the Infisical web UI (https://vault.sysloggh.net) under Project Settings →
Machine Identities, or an admin API call. **TODO: create `abiba-automation` machine identity
and store its credentials in the vault itself (or a root-only file).**
**Interim (working now):** the `.env` fallback (hermes-config-template Rule 3/13). The
infisical-gateway.sh wrapper sources `~/.hermes/.env`, so its `<AGENT>_LITELLM_API_KEY`
overrides a stale vault value.
**2026-07-16 UPDATE — vault is now SYNCED.** The abiba service token (`st.8e848433…`, READ+WRITE)
can write to the vault, so the session-13 rotated keys (mumuni `sk-OzuWsoX2…`, koby `sk-BqRRMboTI…`,
koonimo `sk-OEK7z26n6E…`) are now in the vault as `MUMUNI_LITELLM_API_KEY` / `KOBY_LITELLM_API_KEY` /
`KOONIMO_LITELLM_API_KEY` and validate 200 against LiteLLM. The vault is the source of truth again.
Creating a dedicated `abiba-automation` machine identity (via UI) is still the proper long-term fix
so the shared service token isn't reused across hosts — but it is no longer blocking.
## Key Rotation Log
| Date | Agent | Action | Notes |
|------|-------|--------|-------|
| 2026-07-16 | mumuni | rotate | Old key malformed (sk-_SWAl_Vu_, 47 chars, not LiteLLM format) → 401. Deleted old `mumuni` key (token 15cbca18…), generated fresh (alias `mumuni`, 7 models: syslog-auto, qwen3.6-27B-code, gemma-4-12b, strix-moe, gpu-dense, gpu-light, qwen3.6-35B-udq4). New key sk-OzuWsoX2… written to /root/.hermes/.env (Rule 3/13 fallback). Vault sync PENDING (needs machine identity). WAL #1300. |
| 2026-07-16 | koby | rotate | Old key sk-6sbCNjz (401, stale in /etc/environment). Deleted old `koby` key, generated fresh (alias `koby`). New key sk-BqRRMboTI… in systemd drop-in `hermes-gateway.service.d/litellm-key.conf` + /etc/environment. Created `hermes-gateway.service` unit (was missing — gateway wasn't persistent) with `--replace`. Verified HTTP 200, Telegram connected. |
| 2026-07-16 | baggy (koonimo) | rotate | Old key sk-krnw_zGB (401, hardcoded in systemd drop-in). Deleted old `baggy` key, generated fresh (alias `baggy`, metadata agent=koonimo). New key sk-OEK7z26n6E… in drop-in `hermes-gateway.service.d/litellm-key.conf`. CT113 IP changed .113→.114. Verified HTTP 200, Zulip connected. |
## LiteLLM Master Key (use sparingly — agents should NOT use it directly)
- Master key: `sk-litellm-7f96080dd99b15c36bd4b333b58a6796` (in /opt/inference-harness/.env on CT116, Infisical project=infrastructure env=production secret=LITELLM_MASTER_KEY)
- Used for /key/generate, /key/delete, /key/list (GET), DB queries
- **Known violation (RESOLVED 2026-07-16):** Abiba's LITELLM_API_KEY was previously the master key.
It is now a dedicated agent key `sk-sxbphLvk1OU…` (vault secret `ABIBA_LITELLM_API_KEY`, alias `abiba-pi`).
The master key is admin-only (/key/generate, /key/delete, /key/list). NEVER use it for inference —
see `litellm-self-heal` § "NEVER use litellm_proxy_master_key for inference".
- LiteLLM key DB: `harness-postgres` container on CT116, table `"LiteLLM_VerificationToken"` (columns: token, key_alias, key_name, created_at, expires). Query: `docker exec harness-postgres psql -U litellm -d litellm -t -c "SELECT key_alias, substr(token,1,16) FROM \"LiteLLM_VerificationToken\" ORDER BY created_at;"`
+10 -10
View File
@@ -42,8 +42,8 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2)
**What changed (v3.2.0 → v4.0.0 — 2026-07-08)**: **What changed (v3.2.0 → v4.0.0 — 2026-07-08)**:
- Router REMOVED from request path — LiteLLM proxies directly to GPU - Router REMOVED from request path — LiteLLM proxies directly to GPU
- All GPUs at parallel 2 (was parallel 1) - All GPUs at parallel 2 (was parallel 1)
- NVIDIA context reduced 256K→128K to free VRAM - NVIDIA context reduced 256K→128K to free VRAM (SUPERSEDED 2026-07-16: all GPUs back to 256K — see litellm-self-heal)
- LiteLLM timeouts tuned: gemma 25→120s, qwen 40→90s - LiteLLM timeouts tuned: gemma 25→120s, qwen 40→90s (SUPERSEDED 2026-07-16: qwen 300s, gemma 120s, strix 300s — see litellm-self-heal)
- nginx proxy_read_timeout: 600s, LiteLLM request_timeout: 300s - nginx proxy_read_timeout: 600s, LiteLLM request_timeout: 300s
## Parameters ## Parameters
@@ -72,18 +72,18 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2)
| Host | IP | Hardware | Models Served | Engine | Context | Parallel | | Host | IP | Hardware | Models Served | Engine | Context | Parallel |
|------|-----|----------|---------------|--------|---------|----------| |------|-----|----------|---------------|--------|---------|----------|
| llm-gpu | 192.168.68.8 | NVIDIA RTX 3090 (24 GB) | qwen3.6-27B-code | llama-server systemd | 128K | 2 | | llm-gpu | 192.168.68.8 | NVIDIA RTX 3090 (24 GB) | qwen3.6-27B-code | llama-server systemd | **256K** | 2 |
| ocu-llm | 192.168.68.110 | NVIDIA RTX 5070 (12 GB) | gemma-4-12b | llama-server systemd | 128K | 2 | | ocu-llm | 192.168.68.110 | NVIDIA RTX 5070 (12 GB) | gemma-4-12b | llama-server systemd | **256K** | 2 |
| amdpve | 192.168.68.15 | AMD Strix Halo 64GB UMA | ornith-1.0-35b | llama-server systemd (Vulkan) | 256K | 2 | | amdpve | 192.168.68.15 | AMD Strix Halo 64GB UMA | qwen3.6-35B-udq4 (LiteLLM alias: strix-moe) | llama-server systemd (Vulkan) | 256K | 2 |
## Model Fallback Chains (LiteLLM) ## Model Fallback Chains (LiteLLM)
| Primary | Timeout | Fallback | Timeout | | Primary | Timeout | Fallback | Timeout |
|---------|---------|----------|---------| |---------|---------|----------|---------|
| qwen3.6-27B-code | 90s | gemma-4-12b | 120s | | qwen3.6-27B-code | 300s | gemma-4-12b | 120s |
| gemma-4-12b | 120s | qwen3.6-27B-code | 90s | | gemma-4-12b | 120s | qwen3.6-27B-code | 300s |
| ornith-1.0-35b | 120s | qwen → gemma | — | | qwen3.6-35B-udq4 / strix-moe | 300s | qwen → gemma | — |
| syslog-auto (balanced) | 90s | qwen → gemma | — | | syslog-auto (balanced) | 300s | qwen → gemma | — |
> Global: request_timeout=300s, nginx proxy_read_timeout=600s > Global: request_timeout=300s, nginx proxy_read_timeout=600s
@@ -129,7 +129,7 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2)
7. **Check model inference via LiteLLM** — Test each model: 7. **Check model inference via LiteLLM** — Test each model:
- POST /v1/chat/completions model=gemma-4-12b → expect 200 - POST /v1/chat/completions model=gemma-4-12b → expect 200
- POST /v1/chat/completions model=qwen3.6-27B-code → expect 200 - POST /v1/chat/completions model=qwen3.6-27B-code → expect 200
- POST /v1/chat/completions model=ornith-1.0-35b → expect 200 - POST /v1/chat/completions model=strix-moe → expect 200
- Use master key for auth - Use master key for auth
8. **Check agent keys**: 8. **Check agent keys**:
+27 -8
View File
@@ -6,6 +6,7 @@ note: >
DEPLOYED 2026-07-12 on CT 116 cron: 0 */6 * * * DEPLOYED 2026-07-12 on CT 116 cron: 0 */6 * * *
Auto-remediation code was removed from the pi Zulip extension (retired 2026-07-04), Auto-remediation code was removed from the pi Zulip extension (retired 2026-07-04),
now reimplemented as `litellm-health-check.sh` on CT 116. now reimplemented as `litellm-health-check.sh` on CT 116.
Script: `/opt/inference-harness/scripts/litellm-health-check.sh` on CT 116 (cron `0 */6 * * *`).
Reports to /var/log/litellm/health-*.json and RA-H OS knowledge graph. Reports to /var/log/litellm/health-*.json and RA-H OS knowledge graph.
GPU monitoring integrated from gpu-monitor on .24:9100. GPU monitoring integrated from gpu-monitor on .24:9100.
@@ -56,18 +57,29 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2)
| Host | IP | Hardware | Models Served | Engine | Context | Parallel | | Host | IP | Hardware | Models Served | Engine | Context | Parallel |
|------|-----|----------|---------------|--------|---------|----------| |------|-----|----------|---------------|--------|---------|----------|
| llm-gpu | 192.168.68.8 | NVIDIA RTX 3090 (24 GB) | qwen3.6-27B-code | llama-server systemd | **256K** | 1 | | llm-gpu | 192.168.68.8 | NVIDIA RTX 3090 (24 GB) | qwen3.6-27B-code | llama-server systemd (`/home/llmuser/llama-wrapper.sh`, `-c 262144 --parallel 2 --ngl 99`) | **256K** | 2 |
| ocu-llm | 192.168.68.110 | NVIDIA RTX 5070 (12 GB) | gemma-4-12b | llama-server systemd | 131K | 2 | | ocu-llm | 192.168.68.110 | NVIDIA RTX 5070 (12 GB) | gemma-4-12b | llama-server systemd (`/home/llmuser/llama-wrapper.sh`, `--ctx-size 262144 --parallel 2`, IQ4_NL + MTP draft) | **256K** | 2 |
| amdpve | 192.168.68.15 | AMD Strix Halo 64GB UMA | ornith-1.0-35b | llama-server systemd (Vulkan) | 256K | 2 | | amdpve | 192.168.68.15 | AMD Strix Halo 64GB UMA | qwen3.6-35B-udq4 (LiteLLM alias: `strix-moe`) | llama-server systemd (Vulkan) | 256K | 2 |
> Verified on ground 2026-07-16 via `curl /v1/models` on each host + `llama-wrapper.sh`. The AMD host's underlying model is `qwen3.6-35B-udq4`; LiteLLM exposes it under two `model_name`s: `qwen3.6-35B-udq4` and `strix-moe` (rpm 40). The legacy name `ornith-1.0-35b` does NOT exist in LiteLLM and must not be referenced.
## LiteLLM Model Surface (ground truth — `/opt/inference-harness/litellm_config.yaml` on CT 116)
`model_name`s served: `qwen3.6-27B-code`, `gemma-4-12b`, `qwen3.6-35B-udq4`, `strix-moe`, `gpu-dense`, `gpu-light`, `syslog-auto`.
- `syslog-auto` is a weighted router model: qwen3.6-27B-code (0.55, rpm 500) + qwen3.6-35B-udq4 (0.30, rpm 60) + gemma-4-12b (0.15, rpm 200).
- `gpu-dense` / `gpu-light` are high-rpm aliases (rpm 500) onto qwen3.6-27B-code / gemma-4-12b respectively.
- Key scoping: agent keys are restricted to `['syslog-auto','qwen3.6-27B-code','gemma-4-12b','strix-moe','gpu-dense','gpu-light']`. As of 2026-07-16 the `baggy`/`koby`/`mumuni`/`abiba-pi` keys ALSO include `qwen3.6-35B-udq4`; `abiba-pi` additionally includes `deepseek-v4-pro` (cloud fallback). `kagenz0`/`koonimo`/`pi-agents-unified` have the standard 6 only. Agents should still use the stable alias `strix-moe` (not the raw `qwen3.6-35B-udq4`) so model swaps don't break them.
- **NEVER use `litellm_proxy_master_key` (the `sk-litellm-...` master key) for inference.** It is for admin endpoints only (`/key/list`, `/key/generate`, `/key/info`). All inference — agent traffic, health-check model tests, monitor scripts — uses agent-specific keys. The health-check script's model tests use a dedicated `monitor` agent key stored at `/etc/litellm-monitor.env` on CT 116 (root-only, `chmod 600`); `/key/list` is the only call that legitimately uses `$MASTER_KEY`.
## Model Fallback Chains (LiteLLM) ## Model Fallback Chains (LiteLLM)
| Primary | Timeout | Fallback | Timeout | | Primary | Timeout | Fallback | Timeout |
|---------|---------|----------|---------| |---------|---------|----------|---------|
| qwen3.6-27B-code | 90s | gemma-4-12b | 120s | | qwen3.6-27B-code | 300s | gemma-4-12b | 120s |
| gemma-4-12b | 120s | qwen3.6-27B-code | 90s | | gemma-4-12b | 120s | qwen3.6-27B-code | 300s |
| ornith-1.0-35b | 120s | qwen → gemma | — | | qwen3.6-35B-udq4 / strix-moe | 300s | qwen → gemma | — |
| syslog-auto (balanced) | 90s | qwen → gemma | — | | syslog-auto (balanced) | 300s | qwen → gemma | — |
> Global: request_timeout=300s, nginx proxy_read_timeout=600s > Global: request_timeout=300s, nginx proxy_read_timeout=600s
@@ -86,6 +98,13 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2)
| harness-docker-stats | python:3.12-alpine | — | container stats exporter | | harness-docker-stats | python:3.12-alpine | — | container stats exporter |
| harness-pve-exporter | prompve/prometheus-pve-exporter | — | Proxmox metrics → Prometheus | | harness-pve-exporter | prompve/prometheus-pve-exporter | — | Proxmox metrics → Prometheus |
## Script Operations (synced 2026-07-16)
- **Health-check script** (`/opt/inference-harness/scripts/litellm-health-check.sh` on CT 116): `gpu-fleet` check fails only on **critical** alerts (warnings are informational). Tests `strix-moe` (not `ornith-1.0-35b`).
- **GPU monitor** (`/root/scripts/gpu-monitor-server.py` on pi .24): runs as **systemd unit `gpu-monitor.service`** (was bare `&` process). `gpu_count` includes Strix Halo (was 2, now 3). VRAM alert thresholds: warning 93%, critical 97% (raised from 90/95 — 256K context steady-state is ~96% on RTX 3090, not a fault).
- **Agent key monitor** (`/root/scripts/agent-health-check.py` on pi .24, cron `*/10`): reads each agent's **live** `LITELLM_API_KEY` from its gateway process env via SSH — never hardcodes keys (hardcoded keys rot on rotation and caused 9×401/30min). Fleet roster: abiba, tanko, mumuni, koby, koonimo (legacy `tdunna`/`baggy` removed — never existed).
- **Stale keys cleaned**: `daily-infra-report.py` SYNTHETIC_API_KEY was stale (`sk-U_ydi3B` → 401); now reads `LITELLM_MASTER_KEY` from env. Deprecated scripts (`router-original.py`, `router-phase0-backup.py`, `apply-fixes.py`) still reference `sk-syslog-local-master-key` but do not actively poll LiteLLM.
## Maintains ## Maintains
- litellm-admin-ui: { status: "healthy", last_check: timestamp } - litellm-admin-ui: { status: "healthy", last_check: timestamp }
@@ -128,7 +147,7 @@ Run this first on every cycle. Results feed into remediation rules below.
### 5. Check model inference via LiteLLM — test each model ### 5. Check model inference via LiteLLM — test each model
- POST /v1/chat/completions model=gemma-4-12b → expect 200 - POST /v1/chat/completions model=gemma-4-12b → expect 200
- POST /v1/chat/completions model=qwen3.6-27B-code → expect 200 - POST /v1/chat/completions model=qwen3.6-27B-code → expect 200
- POST /v1/chat/completions model=ornith-1.0-35b → expect 200 - POST /v1/chat/completions model=strix-moe → expect 200
- Use master key for auth - Use master key for auth
### 6. Check agent keys ### 6. Check agent keys
+4 -4
View File
@@ -92,10 +92,10 @@ raw data never provided.
|--------|-------|----------|------|----------| |--------|-------|----------|------|----------|
| `syslog-code` | qwen3.6-27B-code | terminal, file, web, memory, skills | Code patches, automation, scripts | Writing/modifying code, creating scripts, debugging, reading/writing files | | `syslog-code` | qwen3.6-27B-code | terminal, file, web, memory, skills | Code patches, automation, scripts | Writing/modifying code, creating scripts, debugging, reading/writing files |
| `syslog-devops` | qwen3.6-27B-code | terminal, file, web, memory, skills | Infrastructure, DB, bridge, Proxmox | Server ops, SSH, Docker, Proxmox, DB queries, hardware checks | | `syslog-devops` | qwen3.6-27B-code | terminal, file, web, memory, skills | Infrastructure, DB, bridge, Proxmox | Server ops, SSH, Docker, Proxmox, DB queries, hardware checks |
| `syslog-email` | ornith-1.0-35b | terminal, file, web, memory, skills | Email automation, mail operations | Sending/receiving email, inbox management, SMTP operations | | `syslog-email` | strix-moe | terminal, file, web, memory, skills | Email automation, mail operations | Sending/receiving email, inbox management, SMTP operations |
| `syslog-research` | ornith-1.0-35b | terminal, file, web, memory, skills, **browser** | Analysis, classification, data processing | Web research, browser tasks, data analysis, classification, reading docs | | `syslog-research` | strix-moe | terminal, file, web, memory, skills, **browser** | Analysis, classification, data processing | Web research, browser tasks, data analysis, classification, reading docs |
| `syslog-review` | ornith-1.0-35b | terminal, file, web, memory, skills | Verification, QA, audit validation | **ALWAYS** verify worker output before delivery — especially for infra changes, code builds, and research findings | | `syslog-review` | strix-moe | terminal, file, web, memory, skills | Verification, QA, audit validation | **ALWAYS** verify worker output before delivery — especially for infra changes, code builds, and research findings |
| `syslog-writer` | ornith-1.0-35b | terminal, file, web, memory, skills | Docs, content, branding, reports | Writing docs, reports, proposals, content, markdown formatting | | `syslog-writer` | strix-moe | terminal, file, web, memory, skills | Docs, content, branding, reports | Writing docs, reports, proposals, content, markdown formatting |
### Selection Rules ### Selection Rules
+2 -2
View File
@@ -79,7 +79,7 @@ description: >
`messages_processed` stalls. `messages_processed` stalls.
- **Root cause**: Agent's model config (`models.json` or `settings.json`) references a - **Root cause**: Agent's model config (`models.json` or `settings.json`) references a
model ID that doesn't exist in LiteLLM's authorized model list. Example: `qwen3.6-35B-A3B` model ID that doesn't exist in LiteLLM's authorized model list. Example: `qwen3.6-35B-A3B`
configured but LiteLLM only exposes `ornith-1.0-35b` under that key. pi's session configured but LiteLLM only exposes `strix-moe` (alias for qwen3.6-35B-udq4) under that key. pi's session
workers emit 403 on first prompt, then never recover because the error doesn't trigger workers emit 403 on first prompt, then never recover because the error doesn't trigger
`agent_end` — worker stays `busy` and all subsequent messages pile up in the steer queue. `agent_end` — worker stays `busy` and all subsequent messages pile up in the steer queue.
- **Detection**: Compare `~/.pi/agent/models.json` model IDs against `curl -H "Authorization: Bearer <KEY>" http://192.168.68.116/v1/models` output. A stuck worker shows - **Detection**: Compare `~/.pi/agent/models.json` model IDs against `curl -H "Authorization: Bearer <KEY>" http://192.168.68.116/v1/models` output. A stuck worker shows
@@ -88,7 +88,7 @@ description: >
(2) Set `defaultModel` to `syslog-auto` (safe routing model). (3) Delete stale session (2) Set `defaultModel` to `syslog-auto` (safe routing model). (3) Delete stale session
JSONL files from `~/.pi/agent/sessions/zulip/`. (4) Restart PM2 process. JSONL files from `~/.pi/agent/sessions/zulip/`. (4) Restart PM2 process.
- **Prevention**: Use `syslog-auto` as default model for all agents — it handles model - **Prevention**: Use `syslog-auto` as default model for all agents — it handles model
routing and fallback automatically. Direct model IDs (`ornith-1.0-35b`, etc.) should routing and fallback automatically. Direct model IDs (`strix-moe`, etc.) should
only be used when explicitly requested. Validate model IDs at agent setup time. only be used when explicitly requested. Validate model IDs at agent setup time.
- **Applies to**: pi extension (Tdunna CT111, fixed 2026-07-08), any agent using `syslog-harness` provider - **Applies to**: pi extension (Tdunna CT111, fixed 2026-07-08), any agent using `syslog-harness` provider
+411
View File
@@ -0,0 +1,411 @@
---
kind: responsibility
name: zulip-resilience-v3
description: >
Rewrite the pi Zulip gateway with production-grade resilience patterns drawn from
Zulip's own event system docs (queue lifecycle, heartbeat monitoring, BAD_EVENT_QUEUE_ID
handling, idle_queue_timeout) and battle-tested Node.js resilience patterns
(circuit breaker, exponential backoff with jitter, bulkhead isolation, supervisor watchdog).
replaces: zulip-self-heal (retired)
agent: abiba
triggers:
- "/zulip self-heal v3"
- "zulip stopped responding"
- "PM2 abiba-zulip crashed"
---
# Zulip Gateway v3 — Production Resilience
## Architecture Overview
The current v2 gateway (`/root/.pi/agent/extensions/zulip/index.js`) has three structural
weaknesses that cause repeated deaths:
1. **No crash recovery** — uncaught errors kill the Node process, PM2 exhausts max_restarts
2. **No circuit breaker** — 502/fetch-failed errors escalate to process death with no fallback
3. **No queue lifecycle management** — doesn't use Zulip's documented heartbeat protocol or
idle_queue_timeout, so BAD_EVENT_QUEUE_ID errors cascade into crashes
The v3 rewrite addresses all three, following patterns from:
- [Zulip Events System docs](https://zulip.readthedocs.io/en/11.6/subsystems/events-system.html) —
queue registration, heartbeat, BAD_EVENT_QUEUE_ID recovery, call_on_each_event loop
- [Zulip API: Get Events](https://zulip.com/api/get-events) — long-poll timeout, dont_block, event ack
- [Circuit Breaker & Retry Patterns in Node.js 2026](https://1xapi.com/blog/resilient-api-circuit-breaker-bulkhead-retry-nodejs-2026) —
Opossum-based circuit breaker with fallback, retry with jitter, bulkhead isolation
---
## Maintains
- `zulip-gateway`: { status: "healthy" | "degraded" | "down" }
- `circuit-breaker`: { state: "CLOSED" | "OPEN" | "HALF_OPEN", failures, successes }
- `queue-lifecycle`: { queue_id, last_event_id, idle_timeout, heartbeat_age }
- `workers`: { count, busy, idle, stuck }
- `supervisor`: { pid, last_check, health_failures }
---
## Detection Rules
### Rule 1: Queue Expired (BAD_EVENT_QUEUE_ID)
- **Detect**: Events API returns error with BAD_EVENT_QUEUE_ID in body
- **Fix**: Call `POST /register` to create new queue, update queue_id and last_event_id
- **Debounce**: If 3 re-registrations fail within 60s, escalate (server may be down)
- **Ref**: Zulip docs: "Your software will need to handle that error condition by re-initializing itself"
### Rule 2: Network Degradation (502/ECONNREFUSED/fetch failed)
- **Detect**: Events API returns 502 or network error
- **Circuit breaker**: Track failure rate over 10s rolling window
- CLOSED → OPEN: 50% failure rate with ≥5 requests
- OPEN → HALF_OPEN: After 30s reset timeout
- HALF_OPEN → CLOSED: Probe succeeds
- HALF_OPEN → OPEN: Probe fails
- **While OPEN**: Log errors, skip events, notify user via DM: "⚠️ Zulip connection degraded — will retry in 30s"
### Rule 3: Long-Poll Timeout (natural)
- **Detect**: Events API response takes > `event_queue_longpoll_timeout_seconds`
- **Not an error**: Server sends heartbeat events when no real events. Simply re-poll.
### Rule 4: Worker Busy Timeout (>5 min)
- **Detect**: Worker `busySince` exceeds 5 minutes
- **Fix**: SIGKILL worker, send error DM, clean up pending replies
### Rule 5: Process Crash (uncaught)
- **Detect**: `uncaughtException` / `unhandledRejection` fires
- **Fix**: Log → clear poll timer → attempt reconnect with backoff → if reconnect fails 3x, exit(1) and let PM2 restart
### Rule 6: Supervisor Detects Router Stall
- **Detect**: External supervisor (`zulip-watchdog`) polls `/health` every 30s. If 3 consecutive failures:
- **Fix**: `pm2 restart abiba-zulip` gracefully (SIGTERM, drain workers, restart)
---
## Implementation Plan
### Phase 1: Rewrite Router Core (circuit-breaker + queue lifecycle)
Replace the poll loop in index.js with a resilience-first event loop:
```js
// Queue lifecycle (Zulip docs pattern)
async function createOrRefreshQueue() {
// POST /register with event_types=["message"]
// Store: queueId, lastEventId, eventQueueLongpollTimeoutSeconds
// NEW: pass idle_queue_timeout parameter (Zulip 12.0+)
}
// Circuit breaker (Opossum pattern, implemented inline to avoid dependency)
class ZulipCircuitBreaker {
constructor({ failureThreshold=0.5, resetTimeout=30000, volumeThreshold=5, windowMs=10000 }) {
this.state = "CLOSED"; // CLOSED | OPEN | HALF_OPEN
this.failures = 0;
this.successes = 0;
this.totalRequests = 0;
this.lastFailureTime = null;
this.openedAt = null;
this.failureThreshold = failureThreshold;
this.resetTimeout = resetTimeout;
this.volumeThreshold = volumeThreshold;
this.windowMs = windowMs;
}
async fire(fn) {
if (this.state === "OPEN") {
if (Date.now() - this.openedAt > this.resetTimeout) {
this.state = "HALF_OPEN";
} else {
throw new CircuitOpenError("Circuit is OPEN");
}
}
try {
const result = await fn();
this.onSuccess();
return result;
} catch (err) {
this.onFailure();
throw err;
}
}
onSuccess() {
this.successes++;
this.totalRequests++;
if (this.state === "HALF_OPEN") {
this.state = "CLOSED";
this.failures = 0;
}
// Reset counters periodically
if (this.totalRequests > this.volumeThreshold * 2) {
this.failures = Math.floor(this.failures / 2);
this.successes = Math.floor(this.successes / 2);
this.totalRequests = Math.floor(this.totalRequests / 2);
}
}
onFailure() {
this.failures++;
this.totalRequests++;
this.lastFailureTime = Date.now();
if (this.totalRequests >= this.volumeThreshold &&
this.failures / this.totalRequests >= this.failureThreshold) {
if (this.state !== "OPEN") {
this.state = "OPEN";
this.openedAt = Date.now();
console.error(`[zulip-ext] CIRCUIT BREAKER OPEN — ${this.failures}/${this.totalRequests} failures`);
}
}
}
}
// Retry with exponential backoff + jitter (from resilience patterns)
async function withRetry(fn, { maxAttempts=3, baseDelay=200, maxDelay=10000, shouldRetry=()=>true }={}) {
let lastError;
for (let attempt = 1; attempt <= maxAttempts; attempt++) {
try {
return await fn();
} catch (err) {
lastError = err;
if (attempt === maxAttempts || !shouldRetry(err)) throw err;
const delay = Math.min(baseDelay * Math.pow(2, attempt - 1), maxDelay);
const jitter = delay * (0.5 + Math.random() * 0.5); // 50-100% of delay
console.warn(`[zulip-ext] Retry ${attempt}/${maxAttempts} after ${Math.round(jitter)}ms: ${err.message.slice(0,80)}`);
await new Promise(r => setTimeout(r, jitter));
}
}
throw lastError;
}
// Resilience-first event loop (Zulip call_on_each_event pattern)
async function resilientPollLoop() {
while (connected) {
try {
const events = await circuitBreaker.fire(() =>
withRetry(() => zulipQueue.poll(), {
maxAttempts: 2,
baseDelay: 1000,
shouldRetry: (err) => {
const msg = err.message || "";
return msg.includes("fetch failed") || msg.includes("ECONN") || msg.includes("network");
}
})
);
lastError = null;
retryCount = 0;
for (const ev of events) {
await processEvent(ev);
}
heartbeat();
} catch (err) {
const msg = err instanceof Error ? err.message : String(err);
if (msg.includes("BAD_EVENT_QUEUE_ID") || msg.includes("deregistered")) {
// Queue expired — re-register (Zulip docs pattern)
console.log(`[zulip-ext] Queue expired, re-registering… (${msg.slice(0,80)})`);
try {
zulipQueue = await createZulipQueue();
console.log(`[zulip-ext] Re-registered, new queue=${zulipQueue.queueId}`);
} catch (reRegErr) {
console.error(`[zulip-ext] Re-registration failed: ${reRegErr.message}`);
connected = false;
retryCount++;
const backoff = Math.min(5000 * Math.pow(2, retryCount), 300000);
console.log(`[zulip-ext] Full reconnect in ${Math.round(backoff/1000)}s`);
await new Promise(r => setTimeout(r, backoff));
await startPolling();
return;
}
} else if (err.name === "CircuitOpenError") {
// Circuit is open — skip this cycle, wait for HALF_OPEN
lastError = "circuit_open";
await new Promise(r => setTimeout(r, POLL_INTERVAL_MS));
} else {
lastError = msg;
retryCount++;
const backoff = Math.min(POLL_INTERVAL_MS * Math.pow(1.5, Math.min(retryCount, 8)), 60000);
console.error(`[zulip-ext] Poll error (retry ${retryCount}, backoff ${backoff}ms): ${msg}`);
await new Promise(r => setTimeout(r, backoff));
}
}
}
}
```
### Phase 2: PM2 Hardening
Create `/root/.pm2/ecosystem.config.cjs`:
```js
module.exports = {
apps: [
{
name: "abiba-zulip",
script: "/bin/pi",
args: "--mode rpc --session-id zulip-service",
env: {
ZULIP_ROLE: "router",
ZULIP_SITE: "https://chat.sysloggh.net",
ZULIP_EMAIL: "abiba-bot@chat.sysloggh.net",
ZULIP_API_KEY: process.env.ZULIP_API_KEY,
AGENT_NAME: "abiba",
AGENT_OWNER_EMAIL: "jerome@sysloggh.com",
},
max_restarts: 100, // Up from default 10 — crash loops won't exhaust
min_uptime: "10s", // Must survive 10s to count as "alive"
max_memory_restart: "500M", // OOM protection
restart_delay: 5000, // 5s between restarts
kill_timeout: 15000, // 15s SIGTERM grace before SIGKILL
listen_timeout: 30000, // 30s to bind health port
log_date_format: "YYYY-MM-DD HH:mm:ss Z",
error_file: "/root/.pm2/logs/abiba-zulip-error.log",
out_file: "/root/.pm2/logs/abiba-zulip-out.log",
merge_logs: true,
autorestart: true,
watch: false,
instances: 1,
exec_mode: "fork",
},
{
name: "zulip-watchdog",
script: "/root/.pi/agent/extensions/zulip/watchdog.js",
max_restarts: 10,
min_uptime: "3s",
restart_delay: 3000,
autorestart: true,
},
],
};
```
### Phase 3: Supervisor Watchdog
Create `/root/.pi/agent/extensions/zulip/watchdog.js`:
```js
// External supervisor — monitors router health and restarts if stalled.
// This is the pattern Hermes uses: an external process that can recover
// the gateway even if the gateway process itself is hung (not just crashed).
const HEALTH_URL = "http://127.0.0.1:9200/health";
const CHECK_INTERVAL_MS = 30_000;
const MAX_FAILURES = 3;
let failures = 0;
async function check() {
try {
const res = await fetch(HEALTH_URL, { signal: AbortSignal.timeout(5000) });
if (res.ok) {
const data = await res.json();
if (data.status === "ok" && data.zulip?.connected) {
if (failures > 0) {
console.log(`[watchdog] Router recovered after ${failures} failures`);
}
failures = 0;
return;
}
}
failures++;
console.warn(`[watchdog] Health check ${failures}/${MAX_FAILURES}: status not ok`);
} catch (err) {
failures++;
console.warn(`[watchdog] Health check ${failures}/${MAX_FAILURES}: ${err.message}`);
}
if (failures >= MAX_FAILURES) {
console.error(`[watchdog] ${MAX_FAILURES} consecutive failures — restarting abiba-zulip`);
const { execSync } = require("child_process");
try {
execSync("pm2 restart abiba-zulip", { timeout: 30000 });
console.log("[watchdog] Restart command sent");
} catch (e) {
console.error(`[watchdog] Restart failed: ${e.message}`);
}
failures = 0;
// Wait for restart to complete before checking again
await new Promise(r => setTimeout(r, 15000));
}
}
console.log("[watchdog] Zulip gateway supervisor started");
setInterval(check, CHECK_INTERVAL_MS);
check(); // Immediate first check
```
### Phase 4: Health Endpoint Enhancement
Add circuit breaker stats to the existing health endpoint:
```js
// In /health response, add:
"circuit_breaker": {
"state": circuitBreaker.state,
"failures": circuitBreaker.failures,
"successes": circuitBreaker.successes,
"total_requests": circuitBreaker.totalRequests,
"failure_rate": circuitBreaker.totalRequests > 0
? (circuitBreaker.failures / circuitBreaker.totalRequests).toFixed(2)
: "0.00"
}
```
---
## Test Plan
### Test 1: Queue Re-registration
1. Manually delete the Zulip event queue via API
2. Next poll should detect BAD_EVENT_QUEUE_ID
3. Router should auto re-register within 1 poll cycle
4. Verify: `/health` shows new queue_id, connected=true
### Test 2: Circuit Breaker Trip
1. Block Zulip server with iptables: `iptables -A OUTPUT -d 192.168.68.19 -j DROP`
2. Router should detect failures, trip circuit after 5 failures
3. `/health` should show circuit_breaker.state = "OPEN"
4. Remove iptables rule
5. Circuit should transition to HALF_OPEN → CLOSED within 60s
6. Verify: messages processed after recovery
### Test 3: Supervisor Recovery
1. Kill the router process: `kill -STOP $(pm2 pid abiba-zulip)` (freeze, don't kill)
2. Watchdog should detect 3 failed health checks in 90s
3. Watchdog should execute `pm2 restart abiba-zulip`
4. Verify: router back online, connected=true
### Test 4: Worker Busy Timeout
1. Send a message that triggers a long-running operation
2. If worker stays busy >5 minutes, should receive SIGKILL
3. User should receive error DM: "Response timed out"
### Test 5: End-to-End Message
1. Send DM "What time is it?" from Jerome
2. Should receive response within 30s
3. `/health` should show messages_processed incremented
---
## Rollback Plan
If v3 causes issues:
1. `pm2 delete abiba-zulip; pm2 delete zulip-watchdog`
2. Restore v2 from git: `cd /root/.pi/agent/extensions/zulip && git checkout index.js`
3. `pm2 resurrect` to reload previous process list
4. Verify: `/health` returns ok
Backup v2 before starting: `cp index.js index.js.v2-backup-$(date +%Y%m%d-%H%M%S)`
---
## Success Metrics
| Metric | Current (v2) | Target (v3) |
|--------|-------------|-------------|
| Uptime between manual interventions | 1-3 days | 30+ days |
| Crash recovery | Manual (PM2 resurrect) | Automatic (circuit breaker + supervisor) |
| Queue expiry handling | Crash | Auto re-register |
| Busy worker deadlock | Router death | Worker SIGKILL + error DM |
| PM2 restart exhaustion | Yes (max_restarts=10) | No (max_restarts=100 + watchdog) |