Compare commits

..
Author SHA1 Message Date
root 2b9b545ca9 no-mistakes(document): Fixed 2 stale model-name references in doc files
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
2026-07-23 09:10:43 +00:00
root e6c52bf071 no-mistakes(review): Fix F1 model rename and F2 128K compaction math in contracts 2026-07-23 09:04:51 +00:00
root 1c44bf1259 contract updates: ornith decommissioned, 256K→128K context, Mumuni Discord disabled
Change 1: Strix Halo — ornith decommissioned
- gpu-fleet: Genesis Hermes V3 APEX → qwen3.6-35B-udq4 throughout
- inference-optimization: ornith→strix-moe/qwen3.6-35B-udq4
- gpu-monitor: ornith status → Strix Halo status
- infrastructure-control: Strix Halo — ornith → qwen3.6-35B-udq4
- infrastructure-update: ornith→strix-moe via router
- proxmox-monitor: Strix Halo LLM (ornith) → (qwen3.6-35B-udq4, strix-moe)

Change 2: GPU context 256K→128K fleet-wide
- hermes-agent-baseline: frontmatter description updated
- litellm-health: GPU Fleet Topology table 256K→128K
- litellm-self-heal: GPU Fleet Topology, engine flags, VRAM alert
- inference-optimization: compress threshold 256K→128K compact at 85K
- gpu-fleet: instability note updated

Change 3: Mumuni Discord platform disabled
- gpu-fleet: Mumuni platforms: removed discord
2026-07-23 09:00:42 +00:00
jerome fc88265e76 Merge pull request 'tune: switch compression model from strix-moe to syslog-auto' (#25) from tune/compression-syslog-auto into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 0s
Reviewed-on: #25
2026-07-19 01:32:44 +00:00
jerome 51da19d92d Merge pull request 'feat(contracts): add infrastructure-maintenance contract' (#24) from fm/infra-maint-contract into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 0s
Reviewed-on: #24
2026-07-19 01:32:29 +00:00
root 6570fd60e7 tune: switch compression model from strix-moe to syslog-auto
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s
Relieve Strix Halo pressure by distributing compression across the
syslog-auto weighted pool (55% RTX 3090, 30% Strix Halo, 15% RTX 5070).

Updates:
- compression.model: strix-moe -> syslog-auto
- auxiliary.compression.model: strix-moe -> syslog-auto
- Rule 7: Updated for syslog-auto compression, removed aux prohibition
- Rule 8: Updated GPU workload distribution
- All docs/comments updated to reflect the change
2026-07-18 23:28:07 +00:00
jerome 5d2ecbace6 Merge branch 'master' into fm/infra-maint-contract
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s
2026-07-18 23:12:20 +00:00
root 99789a00a1 no-mistakes(document): Reframed infra-maintenance contract as deliberate partition
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s
2026-07-18 22:16:09 +00:00
root a68879e904 no-mistakes(review): fix docker rollback command to pin digest then compose up 2026-07-18 22:07:15 +00:00
jerome 06d2bcbc9e Merge pull request 'fix: fleet config issues from 2026-07-18 relay review' (#23) from fix/fleet-config-issues-20260718 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 0s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 0s
Reviewed-on: #23
2026-07-18 22:06:48 +00:00
jerome c4a8c45835 Merge pull request 'zulip-resilience: fleet-wide audit findings and fixes 2026-07-18' (#22) from feat/zulip-resilience-audit-20260718 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 1s
Reviewed-on: #22
2026-07-18 22:06:34 +00:00
root 1d027f71f6 feat(contracts): add infrastructure-maintenance contract
New responsibility contract consolidating host-level system maintenance
and Docker image lifecycle management, filling the gap left by
infrastructure-update (which owns cluster-wide apt waves).

Scope:
- OS package updates on primary host with pre-update snapshot/backup check
- Docker image pulls for LiteLLM, SearXNG, and other running containers
- Container restarts with per-stack health verification
- Post-update verification: LiteLLM, SearXNG, Zulip, Gitea, PM2, Hermes gateways
- Rollback on failure (image/apt/config restore) with circuit breaker

Owner: ops (firstmate secondmate). Trigger: weekly Sunday 2am ET.
Escalation: warning/critical->abiba+mumuni, fatal->abiba+mumuni+kwame.
circuit_breaker: max_retries 2, window 7200, trip_action escalate_to_fatal.
depends_on: infrastructure-monitoring (pre-update health baseline).

Registry:
- Add infrastructure-maintenance to by_category.maintenance, by_domain.infrastructure,
  by_owner.ops (new), by_trigger.scheduled, by_sensitivity.high
- Add 'ops' to owners list
- Move infrastructure-update owner abiba -> ops (in contracts entry + by_owner index)

Also adds '## Maintaining this file' section to AGENTS.md per fm-ensure-agents-md.
2026-07-18 22:02:45 +00:00
jerome a74229ee74 Merge branch 'master' into fix/fleet-config-issues-20260718
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
2026-07-18 18:06:19 +00:00
root aebc98ead6 fix: remove trailing whitespace from litellm-api-keys.prose.md
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 2s
2026-07-18 17:48:52 +00:00
root 17a77e6b3f fix: fleet config issues from 2026-07-18 relay review
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
- hermes-agent-baseline: add gpu-dense and gpu-light to models list
- hermes-config-template: fix pgrep traps (exclude infisical wrapper) + add
  vault empty-key guard documentation in Rule 13
- litellm-api-keys: fix pgrep pattern in auditable check
- scripts/agent-health-check: fix pgrep to exclude infisical bash wrapper

Addresses issues found during relay inbox resolution session:
1. pgrep -f 'hermes_cli.main gateway run' matches both the real python
   gateway and the infisical bash wrapper, causing false health readings
2. infisical vault stores empty key silently — no guard/monitoring
3. gpu-dense/gpu-light stable aliases missing from baseline config
2026-07-18 17:39:46 +00:00
root 23f3f378c5 zulip-resilience: add fleet-wide audit findings and fixes from 2026-07-18
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s
- Added Incident Log section documenting fleet-wide Zulip audit
- Abiba: poll timeout AbortError fix (returns [] instead of error)
- Abiba: credential fallback .env file for Infisical outages
- Tanko: full gateway restart to recover Zulip connection
- Fleet health metrics summary table
- Hermes agent improvement recommendations (circuit breaker, credential fallback, queue re-registration, watchdog)
- Abiba pi extension v2.1 resilience feature matrix
2026-07-18 16:33:12 +00:00
jerome 14d27a09b5 Merge pull request 'gpu-self-heal: refresh to current fleet baseline and topology' (#21) from feat/gpu-self-heal-refresh-20260718 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 1s
Reviewed-on: #21
2026-07-18 08:36:37 +00:00
root bddbb22f03 gpu-self-heal: refresh to current fleet baseline and topology
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
- Synced model assignments to 2026-07-17 swaps (ThinkingCap, HauhauCS QAT, Genesis Hermes V3)
- Added stable role-based aliases from gpu-fleet (gpu-dense, gpu-light, strix-moe)
- Updated benchmark baselines to live values (74.9/169.6/62.9 tok/s)
- Replaced router (port 9000) references with LiteLLM + direct routing
- Replaced Prometheus exporter rule with sidecar health probe
- Updated VRAM thresholds to match operational data (300/300/200 MB/h)
- Added response size limit (1MB) to prevent OOM crashes
- Added Lessons L5 (response size crash) and L6 (stable aliases)
- Removed deprecated Rules 11-12 (router-specific distribution balance)
2026-07-18 08:07:03 +00:00
root 31ec70ae36 gpu-fleet: RTX 3090 swap to ThinkingCap-Qwen3.6-27B Q4_K_M
- Model: Qwopus Q4_K_M (16GB, 63 tok/s) → ThinkingCap Q4_K_M (15.7GB, 68 tok/s)
- RL-finetuned: 50% fewer thinking tokens, 0.85 MMLU-Pro (vs 0.83 base)
- Self-spec MTP (n=4) REQUIRED for stability — segfaults without it
- Added vision via mmproj (0.9GB) — new capability for this GPU
- VRAM: 20.9/24.6GB (85%), tighter but stable
- Outputs reasoning_content (hidden from Hermes agent)
2026-07-17 15:50:50 +00:00
root 5eb6d3bfbd gpu-fleet: RTX 5070 swap to HauhauCS Gemma4-12B QAT Uncensored Balanced
- Model: IQ4_NL (6.3GB, 191 tok/s) → Q4_K_M QAT (6.9GB, 87 tok/s)
- MTP draft: Q8_0 (444MB) → tuned draft (242MB), saves 200MB VRAM
- mmproj: F16 → BF16 (same size, matched to new model)
- Benefits: 0/465 refusals, agent-optimized tuning, QAT quality
- Trade: 54% slower generation (acceptable for gpu-light role)
- Config: --parallel 1, --ctx-size 131072, single-slot full 128K
- VRAM: 10.0/12.2GB (82%), healthy headroom
2026-07-17 15:24:13 +00:00
jerome 33cb88d571 Merge pull request 'feat: GPU context 256K→128K fleet-wide + Genesis Hermes V3 on Strix Halo' (#20) from feat/gpu-128k-genesis-hermes-v3-20260717 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 1s
Reviewed-on: #20
2026-07-17 11:21:32 +00:00
root 4a1f476623 fix: add missing description to inference-optimization frontmatter
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s
2026-07-17 11:11:27 +00:00
root ba2c55c7e6 ci: re-trigger pipeline for PR #20
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Failing after 1s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Has been skipped
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Has been skipped
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Has been skipped
2026-07-17 11:10:29 +00:00
root b9149bce47 feat: GPU context 256K→128K fleet-wide + Genesis Hermes V3 on Strix Halo
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Failing after 1s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Has been skipped
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Has been skipped
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Has been skipped
GPU Changes:
- All 3 GPUs reduced from 256K (-c 262144) to 128K (-c 131072) for stability
- Observed instability near 100K at 256K — 128K is the stable ceiling
- VRAM improved: RTX 3090 ~70% (was 90%), RTX 5070 ~65% (was 88%)
- Strix Halo swapped to LuffyTheFox/Genesis Hermes V3 APEX
  - Hermes agent fine-tune, tensor repair (3 SSM layers, 76% W1 improvement)
  - Uncensored (0/465 refusals), multimodal (mmproj F16)
  - Speed: 65 tok/s gen, 140 tok/s prompt
  - Alias strix-moe maintained

Agent Updates:
- Mumuni: max_context_window 262144→131072, already aligned on strix-moe/0.65
- Tanko: max_context_window 262144→131072
- Koonimo: max_context_window + context_length 262144→131072
- CT114 SSH access confirmed (was 'Zulip only')

LiteLLM (CT116):
- Updated backend model references qwen3.6-35B-udq4→strix-moe
- Removed stale ornith-1.0-35b from model_cost
- Fallback chains updated

Contracts Updated:
- gpu-fleet.prose.md: topology, VRAM, benchmarks, config lines, model assignments
- gpu-self-heal.prose.md: Rule 9/10 context targets
- hermes-config-template.prose.md: template values, Rules 7-9, compression thresholds
- inference-optimization.prose.md: added to repo, 128K recommendation

Compression: 0.65 fires at ~85K (~43K headroom before 128K ceiling)
For >128K workloads: route to external providers (deepseek)
2026-07-17 10:59:45 +00:00
root 65c99dab50 contract: update litellm-api-keys to v2026-07-17 — fleet standardization
- Document canonical systemd drop-in pattern (ExecStart= reset + wrapper)
- Hardcode venv paths — never use variables in single-quoted bash -c
- Update migration status: all 4 agents on while-true wrapper + st.8e848433
- Infisical CLI update procedure (0.38.0 → 0.43.109 via artifacts-cli)
- Service token inventory + .env fallback inventory
- Key rotation log: fleet standardize + tanko fix-zulip entries
- Remove git merge conflict artifacts
- Add fleet-wide standardization lessons section
2026-07-17 10:17:23 +00:00
jerome 20cbb96e2d Merge pull request 'Vault cleanup + contract sync (WAL #1316)' (#19) from fix/vault-cleanup-contract-sync into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 0s
Reviewed-on: #19
2026-07-16 21:24:59 +00:00
root 8215b84f88 merge: resolve conflicts with master (delegation delete + gpu-fleet)
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 2s
2026-07-16 21:22:34 +00:00
jerome 24cd7f72ac Merge branch 'master' into feat/zulip-v3-resilience 2026-07-16 21:20:24 +00:00
jerome d6e85e68e9 Merge pull request 'UPDATE 2026-07-15: Full GPU fleet rebuild + stable aliases' (#18) from feat/gpu-fleet-rebuild-20260715 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 0s
Reviewed-on: #18
2026-07-16 21:17:12 +00:00
root b7e23e2592 vault: Infisical cleanup + contract sync (WAL #1316)
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s
- litellm-api-keys: vault audit, tanko/koby/koonimo migration, key rotation log
- gpu-fleet: ornith-1.0-35b→strix-moe (remaining refs)
- hermes-agent-baseline: 256K all GPUs, CT IPs updated, Shumba retired
- delegation-prose-contract: removed (renamed to mumuni-delegation)
- contract-registry.yaml + cron-prompts-review.md: from feat/contract-registry
2026-07-16 21:17:09 +00:00
root 369c312ccb vault: BAGGY_LITELLM_API_KEY deleted — one secret per agent
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s
2026-07-16 19:54:51 +00:00
root 7f39a444d6 vault: BAGGY+LITELLM synced, deletion needs Infisical UI
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
2026-07-16 19:28:28 +00:00
Abiba 17f4c79fe4 vault: sync BAGGY_LITELLM_API_KEY = KOONIMO_LITELLM_API_KEY (same agent CT113)
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s
Baggy = Koonimo (CT 113). The vault had two divergent keys:
- BAGGY_LITELLM_API_KEY = sk-QT-kDt8Szo... (stale, 401)
- KOONIMO_LITELLM_API_KEY = sk-OEK7z26n6... (valid, 200, session-13 rotation)

Synced BAGGY to match KOONIMO (the valid key).
Added contract note: both secrets MUST mirror each other.
This was the root cause of Koonimo's 401 after migration —
the ghost process had BAGGY in its env instead of KOONIMO.
2026-07-16 19:22:48 +00:00
Abiba a9caf5b216 vault: fix koby platform connectivity + document migration lessons
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s
Koby was broken after migration — Zulip credentials lost (only in process memory),
Telegram token overwritten (only existed in old .env.bak-20260603).

Fix: Koby shares Tanko's Zulip bot (tanko-bot@, TANKO_ZULIP_API_KEY in vault).
Telegram token recovered from .env.bak-20260603. Both platforms now connected.

Added Koby migration lessons to contract (back up .env before migration,
inject ALL platform env vars, cat /proc/pid/environ before killing old gateway).
Added KOBY_ZULIP_API_KEY to vault (shares Tanko's key).
2026-07-16 18:19:27 +00:00
Abiba 15a826457e vault: canonical Production Vault Access Process + sync vault + migrate koby/koonimo
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s
- Vault SYNCED: mumuni/koby/koonimo rotated keys written to Infisical (all validate 200).
  Vault is source of truth again (was stale since session-13 rotations).
- Koby (.129) + Koonimo (.114) migrated from hardcoded systemd drop-ins to
  infisical-gateway.sh wrapper (live vault injection, .env fallback safety net).
  4/5 agents now vault-backed (abiba/mumuni/koby/koonimo).
- New § Production Vault Access Process: canonical 7-step pattern, migration status
  table, key rotation procedure, why-it's-non-fail rationale.
- Fixed stale notes: 'vault sync PENDING' (now synced), 'Abiba key IS master key'
  (now a proper agent key sk-sxbphLvk1OU).
- Tanko (user jerome, not systemd) documented as pending migration.
- hermes-key-enforcement: cross-ref to canonical process, drop-ins deprecated.
2026-07-16 17:51:53 +00:00
Abiba cd52deda92 contracts: stable alias chain + no-master-key-for-inference rule
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s
Agent configs aligned to stable aliases (gpu-light/gpu-dense/strix-moe/syslog-auto):
- Tanko fixed: base_url :4000->/v1 (nginx), max_context_window 131072->262144,
  aux.compression gemma->strix-moe, aux gemma->gpu-light
- Mumuni aligned: aux gemma->gpu-light, delegation/x_search qwen->gpu-dense
- Koby/Koonimo already used gpu-light (validated)
- Zero raw model names remain in any agent config

No-master-key rule: litellm_proxy_master_key (sk-litellm-...) is admin-only
(/key/list, /key/generate). All inference uses agent keys. Health-check model
tests use dedicated monitor key (/etc/litellm-monitor.env on CT 116).
Fixed: daily-infra-report.py, gpu-self-heal.py, litellm-health-check.sh.

Fixed wrong key-scope note: qwen3.6-35B-udq4 IS in scope for baggy/koby/mumuni/
abiba-pi keys (not 'NOT in allowlist' as previously documented).
2026-07-16 17:39:05 +00:00
Abiba 97f1cd77e4 litellm-self-heal: document 2026-07-16 ops sync (script fixes, systemd gpu-monitor, live-key agent check)
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
Health-check script: critical-only gpu-fleet alerting, strix-moe model test.
gpu-monitor: systemd unit (was bare &), Strix counted in gpu_count, VRAM
thresholds raised (93/97 — 256K ctx steady-state ~96% on RTX 3090).
agent-health-check.py: reads live LITELLM_API_KEY from gateway env (no hardcoded keys).
daily-infra-report.py: stale SYNTHETIC key → env-based master key.
2026-07-16 17:23:37 +00:00
Abiba dc572889f8 contracts: sync to ground truth — ornith-1.0-35b→strix-moe, 256K all GPUs, real LiteLLM timeouts
Verified on ground 2026-07-16 against CT 116 litellm_config.yaml + GPU hosts:
- AMD host serves qwen3.6-35B-udq4 (LiteLLM alias strix-moe); ornith-1.0-35b does NOT exist
- All 3 GPUs at 256K ctx, parallel 2 (RTX 3090 was listed 128K/parallel 1)
- LiteLLM timeouts: qwen 300s, gemma 120s, strix 300s (were stale 90s/120s)
- Added LiteLLM model surface + key scoping to litellm-self-heal
- Patched health-check script path ref

Files: litellm-self-heal, litellm-health, gpu-fleet, gpu-self-heal,
zulip-adapter-lessons, abiba-zulip-restore, hermes-agent-baseline,
delegation-prose-contract, mumuni-delegation-prose-contract
2026-07-16 17:02:44 +00:00
Abiba d0feb7881e Contracts: Rule 13 two key-injection patterns + koby/baggy rotation log (WAL #1300) 2026-07-16 15:10:50 +00:00
Abiba 9fd8c68bd2 Contracts: replace remaining ornith-1.0-35b refs with strix-moe 2026-07-16 13:57:57 +00:00
Abiba 39e2fa0cfa Contracts: Mumuni context-fix (WAL #1300) — strix-moe, 256K all GPUs, Rule 12/13, machine-identity vault procedure 2026-07-16 13:57:31 +00:00
root 622cf7b176 UPDATE 2026-07-15: Full fleet rebuild
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Failing after 12m4s
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Has been skipped
- Stable role-based aliases: strix-moe, gpu-dense, gpu-light
- Strix Halo: ornith-1.0-35b -> unsloth/Qwen3.6-35B-A3B-MTP (UD-Q4_K_M, 256K)
- RTX 5070: Q4_K_M -> IQ4_NL+MTP (256K, 125-213 tok/s, 2x faster)
- RTX 3090: bumped to 256K context
- Provider rename: harness -> litellm (all agents)
- syslog-auto: weighted pool restored with direct GPU routing
- Mumuni agent profile documented
- Self-heal thresholds updated for 256K context
- All model contexts corrected (RTX 3090 was 131K, not 256K)
2026-07-15 16:04:21 +00:00
root 30bf42b841 feat: Zulip v3 resilience contract + restore playbook
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
- zulip-resilience-v3.prose.md: Production resilience rewrite contract covering
  circuit breaker, retry with jitter, queue lifecycle management, supervisor
  watchdog, and PM2 hardening. Research-backed from Zulip event system docs.
- abiba-zulip-restore.prose.md: Quick-restore playbook for recovery scenarios.
2026-07-13 21:29:35 +00:00
abiba-bot 190ceb9be4 Merge pull request 'GPU workload redistribution + lessons learned July 2026' (#15) from feat/gpu-workload-compression-jul2026 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 2s
2026-07-12 23:40:12 +00:00
root 0c298eb9d9 feat: add routing configuration with RPM caps (ornith syslog-auto: 40→60)
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 2s
- Strix Halo ornith RPM increased from 40 to 60 inside syslog-auto pool
- Prevented unbounded ornith traffic when multiple agents use syslog-auto
- Direct ornith remains at RPM=40
- All models verified healthy (ornith 80ms, gemma 110ms, qwen 840ms)
- Added full routing table to gpu-fleet contract
2026-07-12 23:38:21 +00:00
root af41f8f57a fix: remaining template YAML and topology table corrections
- hermes-config-template: Template YAML compression model gemma→ornith.
  Context guidance updated (qwen=256K, only gemma limited to 131K).
- litellm-self-heal: GPU topology table fixed (RTX 3090: 128K→256K, parallel 2→1)
2026-07-12 22:55:24 +00:00
root be02b0e843 fix: stale dates, context values, and architecture references
- gpu-fleet: Updated all context references (128K→256K for RTX 3090).
  Fixed parallel count (2→1 for RTX 3090). Updated architecture diagram
  (router labeled deprecated/not in path). Corrected VRAM values.
  Updated dates: June→July 2026.
- hermes-config-template: Fixed description (removed incorrect '128K on
  NVIDIA' claim). Updated to reflect verified 256K on RTX 3090.
- litellm-self-heal: Updated last-verified date.
- litellm-api-keys: Updated last-verified date.
2026-07-12 22:54:30 +00:00
root 79d4a73895 docs: lessons learned from 2026-07-12 session
- gpu-fleet: Corrected architecture (direct GPU, no router in path).
  Updated context values (RTX 3090=256K, not 128K). Added api-key
  standardization requirement.
- gpu-self-heal: Added Lessons Learned section with 5 critical findings:
  L1: API key standardization (RTX 5070 sk-loc...5678 vs not-needed)
  L2: Fallback chain cascading failure loop detection
  L3: Verify running state, not documentation
  L4: Infisical fallback requirement (.env must have uncommented key)
  L5: Zulip event queue can silently die after ~40 reconnects
- litellm-self-heal: Updated status manual-only→deployed, cron schedule
- litellm-api-keys: Added Infisical token expiry warning + .env fallback
- hermes-config-template: Rule 3 updated with .env fallback requirement
2026-07-12 22:49:39 +00:00
root 19b6db9891 feat: GPU workload redistribution — compression → Strix Halo
- Move compression model from gemma-4-12b (RTX 5070) to ornith-1.0-35b (Strix Halo)
- Add Rule 8: GPU Workload Distribution — per-GPU role assignment
- Add Rule 9: Compression Threshold for 256K models
- Update Rule 7: Auxiliary Model Consistency with new compression routing
- Add gpu-self-heal.prose.md contract with 10 remediation rules
- Strix Halo (64GB, 256K, 72.4 tok/s) → compression specialist
- RTX 5070 (12GB) → vision/web search specialist
- RTX 3090 (24GB, 256K) → heavy reasoning specialist
- All rules grilled and confirmed with Kwame 2026-07-12
2026-07-12 22:27:15 +00:00
jerome 22eaaf4254 Merge pull request 'fix: add data source integrity rule to delegation contract' (#14) from feat/data-source-integrity-rule into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 0s
Reviewed-on: #14
2026-07-12 18:09:10 +00:00
jerome fdb22948d9 fix: add data source integrity rule to delegation contract
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
Workers MUST use provided input data, not fetch external sources.
Adds: Data Source Integrity section, dispatch guidance example,
anti-pattern entry. Root cause of e2e test discrepancies where
writer worker queried Proxmox API instead of using raw data.
2026-07-11 18:51:30 -04:00
25 changed files with 2459 additions and 252 deletions
+7
View File
@@ -128,3 +128,10 @@ safe-mutate --verify "CMD" [--expect "PATTERN"] --mutate "CMD" [--reason "WHY"]
Read the [Authoring Guide](docs/AUTHORING-GUIDE.md) before writing any new contract.
It covers the full process: verify → draft → lint → review → ship, with templates
and style rules.
## Maintaining this file
Keep this file for knowledge useful to almost every future agent session in this project.
Do not repeat what the codebase already shows; point to the authoritative file or command instead.
Prefer rewriting or pruning existing entries over appending new ones.
When updating this file, preserve this bar for all agents and keep entries concise.
Symlink
+1
View File
@@ -0,0 +1 @@
AGENTS.md
+309
View File
@@ -0,0 +1,309 @@
---
kind: function
name: abiba-zulip-restore
description: >
Restores Zulip connectivity for Abiba (pi agent). Verifies the v2 router-worker
extension code, starts a PM2 process as the Zulip gateway with correct env vars,
validates health endpoint, and confirms DM delivery. Run this whenever Abiba
stops responding on Zulip or after system restart.
agent: abiba
version: 1.0.0
status: active
runtime_contract: 2
---
# Abiba Zulip Restore — Resume pi Zulip Communication
Single-shot function that restores full Zulip connectivity for the Abiba pi agent.
Covers extension code validation, PM2 process management, health endpoint
verification, and DM loopback testing.
## Live-State Fields
| Field | Value | Trust |
|-------|-------|-------|
| Agent name | abiba | ✅ |
| Bot email | abiba-bot@chat.sysloggh.net | ✅ |
| Zulip server | https://chat.sysloggh.net | ✅ |
| Extension path | /root/.pi/agent/extensions/zulip/index.js | ✅ |
| Config path | /root/.pi/agent/extensions/zulip/config.yaml | ✅ |
| Health port | 9200 | ✅ |
| @all-bots user ID | 20 | ✅ (config, verified by API at runtime) |
| PM2 process name | abiba-zulip | ✅ |
| Provider | syslog-harness (http://192.168.68.116/v1) | ✅ |
| Default model | deepseek-v4-pro | ✅ (settings.json) |
## Architecture
The pi Zulip extension uses a **router-worker architecture** (v2):
- **Router** (PM2 `abiba-zulip`, `ZULIP_ROLE=router`): Polls Zulip for events,
maintains a per-sender pool of pi RPC worker processes. Each sender gets
their own `pi --mode rpc --session-dir` process with persistent sessions.
Handles streaming edits back to Zulip.
- **Worker** (child `pi --mode rpc`): Runs the agent with per-sender persistent
sessions. No Zulip logic in the worker — the router handles all Zulip I/O.
The extension loads in ALL pi sessions (via `settings.json` extensions array)
but is a **NO-OP** unless `ZULIP_ROLE=router` is set. Only the PM2 router process
carries the env var.
## Maintains
- extension_valid: bool — Extension code imports without errors
- pm2_running: bool — PM2 process `abiba-zulip` is online
- health_responding: bool — GET :9200/health returns "ok"
- zulip_connected: bool — Queue registered, bot identity resolved
- loopback_delivered: bool — Test DM sent and received
- model_valid: bool — Configured models match LiteLLM authorized models
### Postconditions
- PM2 process `abiba-zulip` online and stable (uptime > 30s)
- Health endpoint returns `{ status: "ok", connected: true }`
- Worker pool creates sessions on demand
- Echo prevention active (BOT_EMAILS includes all known bots)
- PM2 saved for auto-restart on boot
## Requires
- Node.js with `yaml` module available
- PM2 installed globally
- Zulip API key in `config.yaml`
- `pi` CLI available on PATH
- Zulip server accessible at https://chat.sysloggh.net
- LiteLLM provider accessible at http://192.168.68.116/v1
## Execution
### Step 1: Validate Extension Code
```bash
node -e "import('file:///root/.pi/agent/extensions/zulip/index.js').then(() => console.log('OK')).catch(e => {console.error('FAIL:', e.message); process.exit(1)})"
```
Expected: `OK`. If FAIL → check for missing dependencies, syntax errors.
### Step 2: Validate Model IDs
Compare configured models against LiteLLM authorized models:
```bash
API_KEY=$(grep -oP 'apiKey:\s*\K.*' /root/.pi/agent/models.json | head -1)
curl -s -H "Authorization: Bearer $API_KEY" http://192.168.68.116/v1/models | \
python3 -c "import json,sys; d=json.load(sys.stdin); [print(m['id']) for m in d.get('data',[])]" 2>/dev/null
```
Check: Every model ID in `models.json` must appear in the authorized list.
If not → fix `models.json` to only include authorized models (prefer `syslog-auto`).
### Step 3: Verify Config Integrity
```bash
python3 -c "
import yaml, sys
with open('/root/.pi/agent/extensions/zulip/config.yaml') as f:
cfg = yaml.safe_load(f)
required = ['zulip.site', 'zulip.email', 'zulip.api_key', 'agent.name']
for k in required:
keys = k.split('.')
v = cfg
for kk in keys:
v = v.get(kk)
if v is None:
print(f'MISSING: {k}')
sys.exit(1)
print('Config valid')
print(f' site={cfg[\"zulip\"][\"site\"]}')
print(f' email={cfg[\"zulip\"][\"email\"]}')
print(f' agent={cfg[\"agent\"][\"name\"]}')
print(f' health_port={cfg.get(\"health_port\", 9200)}')
"
```
Expected: Config valid with all fields non-empty.
### Step 4: Remove Stale Systemd Service
The old `abiba-zulip.service` points to `/opt/abiba-zulip/dist/index.js` (compiled
TypeScript, not the v2 extension). It's disabled and stale. Remove it:
```bash
systemctl stop abiba-zulip 2>/dev/null || true
systemctl disable abiba-zulip 2>/dev/null || true
rm -f /etc/systemd/system/abiba-zulip.service
systemctl daemon-reload
```
### Step 5: Start PM2 Process
**Critical:** The extension MUST run via `pi --mode rpc`, NOT `node index.js` directly.
The extension exports a function that requires pi's session lifecycle. Running
`node index.js` loads the module but never calls the export, so nothing happens.
`pi --mode rpc` loads all extensions (including Zulip) and fires `session_start`.
```bash
# Stop existing if any
pm2 delete abiba-zulip 2>/dev/null || true
# Start pi in RPC mode with ZULIP_ROLE=router env
ZULIP_ROLE=router pm2 start "$(which pi)" \
--name abiba-zulip \
--interpreter none \
-- --mode rpc --no-session
```
Wait 10 seconds for pi to load all extensions, fire session_start, and the Zulip
router to register its event queue.
### Step 6: Validate PM2 Process
```bash
pm2 show abiba-zulip --no-color
```
Check: `status=online`, `restarts=0`, `uptime > 5s`.
### Step 7: Check Logs for Connection
```bash
sleep 3
tail -20 /root/.pm2/logs/abiba-zulip-out.log
```
Look for:
- `[zulip-ext] Connecting to https://chat.sysloggh.net as abiba-bot@chat.sysloggh.net…`
- `[zulip-ext] Bot user_id=N, all-bots user_id=N`
- `[zulip-ext] Connected, queue=N`
- `[zulip-ext] Echo prevention: N bot emails`
- `[zulip-ext] Health endpoint on :9200`
If error → check API key, network to chat.sysloggh.net.
### Step 8: Health Endpoint
```bash
curl -s http://localhost:9200/health | python3 -m json.tool
```
Check:
- `status: "ok"` (not "down")
- `zulip.connected: true`
- `zulip.queue_id` is non-null string
- `zulip.bot_user_id` is positive integer
### Step 9: DM Loopback Test
```bash
curl -s http://localhost:9200/health | python3 -c "
import json,sys
d = json.load(sys.stdin)
if d.get('zulip',{}).get('connected'):
print(f'✅ Zulip connected. Queue: {d[\"zulip\"][\"queue_id\"]}')
print(f' Bot user_id: {d[\"zulip\"][\"bot_user_id\"]}')
print(f' Messages processed: {d[\"zulip\"][\"messages_processed\"]}')
else:
print('❌ Zulip NOT connected')
sys.exit(1)
"
```
### Step 10: Save PM2 for Auto-Start
```bash
pm2 save
pm2 startup systemd -u root --hp /root 2>/dev/null || true
```
### Step 11: Report
Compile results:
| Check | Pass? |
|-------|-------|
| Extension code imports | extension_valid |
| Model IDs authorized | model_valid |
| Config integrity | config_valid |
| PM2 process online | pm2_running |
| Health endpoint | health_responding |
| Zulip connected | zulip_connected |
All pass → ✅ **Abiba Zulip restored.** Relay success to user.
Partial failure → see recovery matrix below.
## Recovery Matrix
| Failure | Recovery |
|---------|----------|
| Extension import fails | Check `node_modules/zulip-js` exists; run `npm install` in extension dir |
| Model ID mismatch | Fix `models.json` to use `syslog-auto` as default model; remove invalid IDs |
| Config missing | Restore from backup or recreate from scratch |
| PM2 won't start | Check `node` version (>=18); check port 9200 not in use |
| Health "down" | Check logs for connection errors; verify Zulip API key; check network |
| "address already in use" | Kill old process: `fuser -k 9200/tcp` |
| Queue registration fails | Check Zulip API key validity; verify bot is active in Zulip admin |
| Rate limit (429) | Extension has built-in retry-after handling — wait, don't restart |
## Known Failure Modes
| Symptom | Root Cause | Recovery |
|---------|-----------|----------|
| Extension loads but no events | `ZULIP_ROLE` not set | Ensure PM2 env has `ZULIP_ROLE=router` |
| Worker stays "busy" forever | Model ID not authorized by LiteLLM | Fix models.json (lesson #11) |
| Placeholder sent but no response | editMessage API fails silently | Extension has fallback (sends new msg); check Zulip API |
| Queue expires rapidly | Poll interval too aggressive | v2 uses 3s poll with long-poll — should be fine |
| Bot doesn't respond to @mentions | Not subscribed to stream | Bot auto-subscribes via API |
| Stale error in health | `last_error` not cleared | v2 clears on successful poll (lesson #4) |
## Appendix: Root Cause & Fix Summary (2026-07-13)
**Problem:** Zulip extension was offline — no PM2 process running.
**Root cause:** The PM2 command ran `node index.js` directly (which loads the
extension module but never calls the exported function). The extension requires
pi's session lifecycle — pi loads extensions, fires `session_start`, and the
Zulip extension hooks into that event.
**Fix:** Run `pi --mode rpc` (not `node index.js`). The `--mode rpc` flag keeps
pi alive listening for RPC commands on stdin while the Zulip extension's router
runs in the background via the `session_start` hook.
```bash
ZULIP_ROLE=router pm2 start "$(which pi)" --name abiba-zulip \
--interpreter none -- --mode rpc --no-session
```
**Additional fixes applied:**
- Fixed `models.json`: replaced `qwen3.6-35B-A3B` (not authorized by LiteLLM)
with `syslog-auto` + `strix-moe` (prevents silent worker failure per
Lesson #11)
- Removed stale systemd unit `abiba-zulip.service` (pointed to old TS code)
- PM2 saved for auto-restart on boot
## Appendix: PM2 Ecosystem Config (Optional)
If preferred over manual `pm2 start`, create `/root/ecosystem.config.js` entry:
```js
module.exports = {
apps: [{
name: 'abiba-zulip',
script: '/bin/pi',
interpreter: 'none',
args: '--mode rpc --no-session',
cwd: '/root',
env: {
ZULIP_ROLE: 'router',
},
log_file: '/root/.pm2/logs/abiba-zulip-out.log',
error_file: '/root/.pm2/logs/abiba-zulip-error.log',
max_restarts: 20,
restart_delay: 5000,
}]
};
```
---
**Last verified good state**: 2026-07-13 — Extension v2 running via `pi --mode rpc`, health endpoint :9200 returning `{status:"ok",connected:true}`, queue a669f21e.
+104 -2
View File
@@ -22,6 +22,7 @@ owners:
- abiba
- mumuni
- kwame
- ops
trigger_types:
- scheduled
- event_driven
@@ -58,6 +59,7 @@ index:
- memory-audit-maintenance
- gpu-fleet
- infrastructure-update
- infrastructure-maintenance
reference:
- infrastructure-control
- ra-h-os-custodianship-contract
@@ -90,6 +92,7 @@ index:
- infrastructure-control
- infrastructure-monitoring
- infrastructure-update
- infrastructure-maintenance
- pm2-self-heal
- disk-gc-threat-response
gpu:
@@ -130,7 +133,6 @@ index:
- build-zulip-plugin
- stirling-pdf-agent-access
- gpu-fleet
- infrastructure-update
- infrastructure-control
- zulip-adapter-lessons
- pi-approval-architecture
@@ -144,6 +146,9 @@ index:
- mumuni-delegation
kwame:
- hello-world
ops:
- infrastructure-maintenance
- infrastructure-update
by_trigger:
scheduled:
- hermes-key-enforcement
@@ -156,6 +161,7 @@ index:
- litellm-health
- memory-audit-maintenance
- infrastructure-update
- infrastructure-maintenance
event_driven:
- litellm-self-heal
- pm2-self-heal
@@ -197,6 +203,7 @@ index:
- hermes-zulip-plugin
- build-zulip-plugin
- infrastructure-update
- infrastructure-maintenance
- ra-h-os-custodianship-contract
- mumuni-delegation
normal:
@@ -1256,13 +1263,108 @@ contracts:
last_run: null
last_status: null
drift_alerts: []
- name: infrastructure-maintenance
file: infrastructure-maintenance.prose.md
kind: responsibility
category: maintenance
sensitivity: high
status: active
owner: ops
version: 1.0.0
trigger:
type: scheduled
cadence: 0 2 * * 0
description: Weekly host-level maintenance Sunday at 2am ET (replaces infrastructure-update
build-phase role; infra-update moves to ops)
cron_job_id: null
execution:
agent: ops
timeout: 3600
requires:
- infrastructure-monitoring run within last 30 minutes (pre-update health baseline)
- Proxmox snapshot of primary host OR /tmp backup dir created this run
- LiteLLM master key from Infisical vault for health verification
protocol:
- Load contract from prose-contracts/main
- Phase 0 preflight — capture health baseline, backup check, record image baseline
- Phase 1 apt update && apt upgrade -y on primary host
- Phase 2 docker compose pull for LiteLLM, SearXNG, and other running containers
- Phase 3 restart stacks one at a time with per-stack health verification
- Phase 4 verify every critical service (LiteLLM, SearXNG, Zulip, Gitea, PM2, Hermes gateways)
- On failure — rollback per protocol, escalate, do not loop beyond circuit breaker
- Log actions to ~/.hermes/runs/infrastructure-maintenance/
verification:
postconditions:
- check: all critical services running after update
verify: 'curl -sf http://192.168.68.116/litellm/v1/models && curl -sf https://chat.sysloggh.net/api/v1/server_settings && curl -sf https://git.sysloggh.net/api/v1/version && curl -sf http://192.168.68.7:8888 && pm2 jlist'
expect: all probes 200 OK / processes online
- check: no regressions from pre-update health baseline
verify: diff Phase 0 health-baseline against Phase 4 results
expect: no GREEN service turned RED
- check: docker containers on latest stable tags
verify: docker inspect --format '{{.Config.Image}}' <container> per service matches image-baseline.pulled_tag
expect: all containers running pulled tags
- check: APT packages up to date with no held broken packages
verify: apt list --upgradable 2>/dev/null | wc -l and apt-get -s upgrade | grep -ci broken
expect: upgradable == 0, broken == 0
artifact: maintenance run report with phase results and any rollback/escalation
verify_commands:
- curl -sf http://192.168.68.116/litellm/v1/models
- curl -sf http://192.168.68.7:8888
- curl -sf https://chat.sysloggh.net/api/v1/server_settings
- curl -sf https://git.sysloggh.net/api/v1/version
- pm2 jlist
receipt:
format: json
storage: ~/.hermes/runs/infrastructure-maintenance/
graph_node: true
schema:
contract: string
run_id: string
timestamp: ISO 8601
agent: string
status: pass|fail|escalated
phase: preflight|apt|images|restarts|verify|rollback|done|failed
actions_taken: array
postconditions: array
drift_alerts: array
evidence_path: string
escalation:
info:
action: log_to_receipt
notify: []
warning:
action: relay_alert
notify:
- abiba
- mumuni
critical:
action: relay_alert
notify:
- abiba
- mumuni
fatal:
action: relay_alert + pause + human_required
notify:
- abiba
- mumuni
- kwame
circuit_breaker:
max_retries: 2
window: 7200
trip_action: escalate_to_fatal
depends_on:
- infrastructure-monitoring
last_run: null
last_status: null
drift_alerts: []
- name: infrastructure-update
file: infrastructure-update.prose.md
kind: responsibility
category: maintenance
sensitivity: high
status: active
owner: abiba
owner: ops
version: 1.0.0
trigger:
type: scheduled
+161 -50
View File
@@ -5,9 +5,15 @@ description: >
Manages the GPU inference fleet across all hosts. Handles model deployment,
registration, health checks, LiteLLM sync, agent key management, GPU
saturation watchdog, Prometheus/Grafana monitoring, and self-healing.
Current as of 2026-07-08: context reduced to 128K on NVIDIA GPUs, parallel 2
on all GPUs, LiteLLM timeouts tuned (gemma 25→120s, qwen 40→90s), router fully
deprecated — nginx routes /v1 → LiteLLM directly.
UPDATED 2026-07-15: Stable role-based aliases introduced: strix-moe,
gpu-dense, gpu-light. These never change — only the underlying model does.
Strix Halo: strix-moe → unsloth/Qwen3.6-35B-A3B-MTP (UD-Q4_K_M, 22GB).
RTX 5070: gemma-4-12b Q4_K_M → IQ4_NL + MTP draft (122 tok/s, 2x faster).
UPDATED 2026-07-17: Context reduced fleet-wide from 256K to 128K for stability.
Strix Halo model swapped to qwen3.6-35B-udq4 (22GB, strix-moe alias).
Instability observed near 100K at 256K (now all GPUs at 128K). 128K is the stable ceiling.
For larger context needs → fall back to external providers (deepseek).
VRAM headroom improved: RTX 3090 ~70%, RTX 5070 ~65%.
agent: abiba
triggers:
- on model add/remove
@@ -30,7 +36,7 @@ triggers:
- prometheus: { status: "running", targets: 5 } — Scrapes GPU :9400 exporters + LiteLLM
- port_conflict_detection: { status: "active" } — All 3 GPU wrappers detect ghost processes before binding
## Fleet Topology (Current — June 2026)
## Fleet Topology (Current — July 2026)
```
┌──────────────────────────────────────────────────────────────────┐
@@ -47,9 +53,9 @@ triggers:
│ Containers: │
│ ┌──────────┐ ┌──────────┐ ┌──────────┐ ┌──────────┐ │
│ │ LiteLLM │ │ Router │ │Dashboard │ │ Grafana │ │
│ │ :4000 │─▶│ :9000 │ │ :3000 │ │ :3000 │ │
│ │ keys+sync│ │internal │ │ harness │ │ Prometheus│ │
│ │ fallback │ │only! │ │ UI │ │ data src │ │
│ │ :4000 │ │ :9000 │ │ :3000 │ │ :3000 │ │
│ │ keys+sync│ │deprecated│ │ harness │ │ Prometheus│ │
│ │ fallback │ │not in │ │ UI │ │ data src │ │
│ └──────────┘ └───┬──────┘ └──────────┘ └──────────┘ │
│ │ │
│ ┌──────────┐ ┌──────────┐ ┌──────────┐ │
@@ -64,21 +70,74 @@ triggers:
│ CT 8 │ │ CT 110 │ │ CT 15 │ │ pi (.24) │
│ RTX 3090 │ │ RTX 5070 │ │ Strix Halo│ │ GPU Monitor │
│ 24GB │ │ 12GB │ │ 64GB UMA │ │ :9100 │
│ 128K ctx │ │ 128K ctx │ │ 256K ctx │ │ Watchdog │
│ qwen3.6 │ │ gemma-4-12b │ │ ornith35B │ │ Prometheus │
│ 27B-code │ │ :8080 │ │ :8080 │ │ exporter │
│ 128K ctx │ │ 128K ctx │ │ 128K ctx │ │ Watchdog │
│ qwen3.6 │ │ gemma-4-12b │ │ qwen3.6 │ │ Prometheus │
│ 27B-code │ │ :8080 │ │ -35B-udq4 │ │ exporter │
│ :8080 │ │ :9400 (exp) │ │ :9400(exp)│ │ :9401 │
│ :9400 │ └─────────────┘ └───────────┘ └──────────────┘
└──────────┘
```
## Current Model Assignments (2026-07-08)
## Stable Role-Based Aliases (Introduced 2026-07-15)
Agent configs, cron jobs, and workflows MUST use these aliases, never model-specific names.
When a model is swapped on a GPU, ONLY the infrastructure layer changes — agent configs are untouched.
| Alias | GPU | Current Model | Will Route To |
|-------|-----|---------------|---------------|
| `strix-moe` | Strix Halo (.15) | qwen3.6-35B-udq4 | Whatever runs on Strix Halo |
| `gpu-dense` | RTX 3090 (.8) | qwen3.6-27B-code | Whatever runs on RTX 3090 |
| `gpu-light` | RTX 5070 (.110) | gemma-4-12b | Whatever runs on RTX 5070 |
**Backward compatibility**: Old model-specific names (qwen3.6-27B-code, gemma-4-12b, qwen3.6-35B-udq4) still work
but are deprecated for agent configs. Only the stable aliases survive model swaps.
## Current Model Assignments (2026-07-15)
| Model | GPU | Host | VRAM | Ctx | KV Cache | Parallel | Batch/Ubatch | Status |
|-------|-----|------|------|-----|----------|----------|-------------|--------|
| qwen3.6-27B-code | RTX 3090 | .8 (llm-gpu) | 20.3/24GB (83%) | 128K | turbo4 | 2 | 512/512 | ✅ healthy |
| gemma-4-12b | RTX 5070 | .110 (ocu-llm) | 9.4/12.2GB (77%) | 128K | q4_0 | 2 | 2048/512 | ✅ healthy |
| ornith-1.0-35b | Strix Halo Vulkan | .15 (amdpve) | 24.4/64GB (35%) | 256K | q8_0 | 2 | 2048/512 | ✅ healthy |
| qwen3.6-27B-code (MTP) | RTX 3090 | .8 (llm-gpu) | ~17/24.6GB (70%) | **128K** | turbo4 | 2 | default | ✅ 63 tok/s |
| gemma-4-12b | RTX 5070 | .110 (ocu-llm) | ~7.8/12.2GB (65%) | 128K | q4_0 | 2 | 2048/1024 | ✅ healthy |
| qwen3.6-35B-udq4 | Strix Halo Vulkan | .15 (amdpve) | ~22GB/64GB | 128K | q4_0 | 1 | 4096/1024 | ✅ 65 tok/s |
## Routing Configuration (LiteLLM — July 2026)
### syslog-auto Weighted Pool (Direct GPU — bypasses router)
| Model | GPU | Weight | RPM Cap | Timeout |
|-------|-----|--------|---------|---------|
| qwen3.6-27B-code | RTX 3090 (.8:8080) | **0.55** | 500 | **300s** |
| qwen3.6-35B-udq4 | Strix Halo (.15:8080) | **0.30** | 60 | **300s** |
| gemma-4-12b | RTX 5070 (.110:8080) | **0.15** | 200 | **120s** |
Note: All syslog-auto entries route directly to GPUs with `api_key: not-needed`. The router (port 9000) is NOT in the inference path.
### Direct Model Endpoints
| Model | RPM Cap | Notes |
|-------|---------|-------|
| strix-moe (qwen3.6-35B-udq4) | 40 | Tight cap — prevents Strix overload |
| qwen3.6-27B-code | 500 | High cap — primary workhorse |
| gemma-4-12b | 500 | High cap — IQ4_NL+MTP, 122 tok/s |
### Stable Aliases (for agent configs — never change)
| Alias | RPM Cap | Routes To | Purpose |
|-------|---------|-----------|---------|
| `strix-moe` | 40 | Strix Halo | Compression tasks (MoE models) |
| `gpu-dense` | 500 | RTX 3090 | Heavy reasoning |
| `gpu-light` | 500 | RTX 5070 | Vision, web extract, light tasks |
### Fallback Chains
- gemma → qwen
- qwen → gemma
- strix-moe → qwen → gemma
- syslog-auto → qwen → gemma → qwen3.6-35B-udq4
### Why Strix Halo RPM Is Capped
- Direct (strix-moe): 40 RPM (tight) — Strix Halo is shared with compression tasks
- Via syslog-auto: 60 RPM (moderate) — prevents flooding when multiple agents use syslog-auto simultaneously
- Combined max: ~100 RPM across both paths — Strix Halo can sustain this at 80°C
## Operations
@@ -114,10 +173,10 @@ triggers:
### sync-keys
1. List all agent keys in LiteLLM DB via `GET /key/list`
2. Compare against expected agent list: [tanko, mumuni, abiba, tdunna, baggy, kagenz0]
2. Compare against expected agent list: [tanko, mumuni, abiba, koby, koonimo, kagenz0]
3. Generate missing keys via `POST /key/generate` with unlimited budget
4. Update agent configs — `/etc/environment` LITELLM_API_KEY
5. Send Zulip DM to agents that can't be reached via SSH
4. Update Infisical vault: `infisical secrets set LITELLM_API_KEY=<key> --project=agents --env=production`
5. Send Zulip DM to agents that can't be reached via SSH (provide vault login instructions)
6. Verify each key with test request through full chain
7. Document keys in knowledge graph
@@ -130,25 +189,31 @@ Show full fleet status: GPUs, models, VRAM, context windows, parallel slots, act
3. Check LiteLLM: `curl http://192.168.68.116/health` (expect "I'm alive!")
4. Check LiteLLM models: `curl -H "Authorization: Bearer $MASTER_KEY" http://192.168.68.116/v1/models`
5. Check LiteLLM timeouts: `grep -n 'timeout:' /opt/inference-harness/litellm_config.yaml`
- gemma-4-12b: 120s, qwen3.6-27B: 90s, ornith-1.0-35b: 120s
- gemma-4-12b: 120s, qwen3.6-27B-code: 300s, qwen3.6-35B-udq4/strix-moe: 300s (strix-moe does NOT exist — legacy name, do not use)
- global request_timeout: 300s, nginx proxy_read_timeout: 600s
6. Check AMD metrics: `curl http://192.168.68.15:9400/metrics` (Radeon 8060S, util%, VRAM, temp, power)
7. Check port conflicts: verify only one llama-server on :8080 per host
8. Verify agent keys: 9 keys in LiteLLM DB (`GET /key/list`)
## Agent Keys (LiteLLM DB — Current 2026-06-30)
## Agent Keys (LiteLLM DB — Current 2026-07-11)
| Agent | CT | IP | Key | Access |
|-------|-----|-----|-----|--------|
| Tanko | 112 | .122 | `sk-CggiHWlamQyShxWC3Hx6uw` | SSH jerome |
| Mumuni | 114 | .123 | `sk-VrqCNlwUgzoNGOpikJ7nwQ` | SSH root |
| Abiba | 100 | .24 | `sk-Qvzi4uYQBhlSK_XstEhcyQ` | local (pi agent) |
| Tdunna | 111 | ? | `sk-6sbCNjz2T6lTVDBdlNHXsA` | Zulip DM |
| Baggy | 113 | ? | `sk-krnw_zGBwvvL5b7l2t-s-A` | no SSH |
| Kagenz0 | 105 | ? | `sk-Dh4CDkaHebMLEp8qqq20qA` | no SSH |
Keys stored in Infisical vault (project=agents, env=production, secret=LITELLM_API_KEY).
Agent gateways inject keys at runtime via `infisical run --` wrapper.
Plaintext keys removed from this contract post-vault-migration.
**Key update procedure**: Update `/etc/environment``LITELLM_API_KEY=sk-...` → restart Hermes.
If no SSH access, send Zulip DM via abiba-bot.
| Agent | CT | IP | LiteLLM Alias | Key Source | Access |
|-------|-----|-----|---------------|------------|--------|
| Tanko | 112 | .122 | `tanko` | Infisical vault | SSH jerome |
| Mumuni | 114 | .123 | `mumuni` | Infisical vault | SSH root |
| Abiba | 100 | .24 | `abiba-pi` | Infisical vault | local (pi agent) |
| Koby | 111 | ? | `koby` | Infisical vault | Zulip DM |
| Koonimo | 113 | ? | `koonimo` | Infisical vault (migrated 2026-07-11) | no SSH |
| Kagenz0 | 105 | ? | `kagenz0` | Infisical vault | no SSH |
> **Note**: CT hostnames differ from agent identities. CT111=tdunna runs koby; CT113=baggy runs koonimo.
**Key update procedure**: Update Infisical vault → `infisical secrets set LITELLM_API_KEY=sk-... --project=agents --env=production` → restart agent gateway. Agent picks up new key via `infisical run --` wrapper at startup.
If no SSH access, send Zulip DM via abiba-bot with vault update instructions.
## Configuration Files
@@ -164,13 +229,13 @@ If no SSH access, send Zulip DM via abiba-bot.
| `/root/scripts/gpu-saturation-watchdog.py` | pi (.24) | Auto-restart stuck llama-server |
| `/root/dashboard/gpu-fleet.html` | pi (.24) | Live HTML dashboard |
| `/etc/systemd/system/llama-server.service` | .8, .110 | llama-server daemons (Nvidia GPUs) |
| `/etc/systemd/system/ornith-server.service` | .15 (amdpve) | llama-server daemon (Vulkan, Strix Halo). Note: `llama-server.service` and `llama-server@.service` are **masked** on .15 to prevent port 8080 collisions. |
| `/etc/systemd/system/strix-server.service` | .15 (amdpve) | llama-server daemon (Vulkan, Strix Halo) running unsloth/Qwen3.6-35B-A3B-MTP-GGUF. Note: `llama-server.service` and `llama-server@.service` are **masked** on .15 to prevent port 8080 collisions. |
## Prometheus & Grafana
| Component | URL | Details |
|-----------|-----|---------|
| Grafana | `http://192.168.68.116:3001/` | admin / syslog-grafana-2026 |
| Grafana | `http://192.168.68.116:3001/` | admin / vault (`GRAFANA_ADMIN_PASSWORD`) |
| GPU Dashboard | `http://192.168.68.116:3001/d/gpu-fleet` | Gauges + time series |
| Prometheus | `http://192.168.68.116:9090/` (internal) | 5 scrape targets |
| GPU Exporters | `:9400/metrics` on .8, .110, .15 | NVIDIA/AMD GPU metrics |
@@ -181,14 +246,14 @@ If no SSH access, send Zulip DM via abiba-bot.
- **Router startup race**: Compose router.py doesn't call load_roster(). Reload thread sleeps 30s first.
Fix: trigger roster reload via SSH after restart, or rebuild image with startup load_roster().
- **LiteLLM /metrics**: Requires auth. Prometheus uses `/health/liveliness` as workaround.
- **VRAM (2026-07-08)**: RTX 3090 at 20.3/24GB (83%), RTX 5070 at 9.4/12.2GB (77%), Strix Halo at 24.4/64GB (35%). Context reduced from 256K→128K on NVIDIA GPUs freed ~3.3GB (.8) and ~1.5GB (.110).
- **All GPUs at `--parallel 2` (2026-07-08)**: Fleet serves 6 concurrent requests (was 3). 2× throughput.
- **RTX 3090 config**: `-c 131072 -ctk turbo4 -ctv turbo4 --parallel 2`. No explicit batch flags (512/512 default). Service: `/home/llmuser/llama-wrapper.sh`.
- **RTX 5070 config**: `--ctx-size 131072 --cache-type-k q4_0 --cache-type-v q4_0 --batch-size 2048 --ubatch-size 512 --parallel 2`. Ubatch fixed 4096→512 (was inverted — ubatch > batch killed prompt throughput). Service: `/home/llmuser/llama-wrapper.sh`.
- **LiteLLM timeout tuning (2026-07-08)**: gemma-4-12b 25→120s, qwen3.6-27B-code 40→90s, syslog-auto (qwen route) 40→90s. Nginx proxy_read_timeout: 600s. Global request_timeout: 300s. Config at `/opt/inference-harness/litellm_config.yaml`.
- **Strix Halo GPU**: Vulkan is the working backend (ROCm/HIP path abandoned — HSA runtime blocked on Debian 13). Build at `/root/llama.cpp/build-vk/`, commit `4fc4ec5` (2026-07-01), ggml 0.15.3 shared-lib arch. Mesa RADV 25.0.7, KHR_coopmat fast path active. ~70 tok/s gen, 532 tok/s prompt. Service: `ornith-server.service` on port 8080, 256K context, flash-attn + q8 KV.
- **VRAM (2026-07-15)**: RTX 3090 at ~17/24.6GB (~70%) with **128K context** (reduced from 256K 2026-07-17). RTX 5070 at ~7.8/12.2GB (~65%) with 128K context + MTP. Strix Halo at ~7GB/64GB.
- **RTX 3090 runs `--parallel 2`** with MTP draft (spec-type draft-mtp, spec-draft-n-max 2).
- **RTX 3090 config**: `-c 131072 -ctk turbo4 -ctv turbo4 --parallel 2 --flash-attn on --cont-batching --spec-type draft-mtp`. Context reduced to 128K (2026-07-17, was 256K). VRAM: ~70%. Service: `/home/llmuser/llama-wrapper.sh`.
- **RTX 5070 config (2026-07-15)**: Switched to IQ4_NL + MTP draft (Q8_0) at 128K context. Gen speed: 122 tok/s. VRAM: ~7.8/12.2GB (~65%). Service: `/home/llmuser/llama-wrapper.sh`. Config: `--model gemma-4-12b-it-IQ4_NL.gguf --spec-draft-model gemma-4-12b-it-Q8_0-MTP.gguf --spec-type draft-mtp --spec-draft-n-max 4 --ctx-size 131072`.
- **LiteLLM timeout tuning (verified 2026-07-16 against `/opt/inference-harness/litellm_config.yaml` on CT 116)**: gemma-4-12b 120s, qwen3.6-27B-code 300s, qwen3.6-35B-udq4 300s, strix-moe 300s, syslog-auto routes all 300s. Nginx proxy_read_timeout: 600s. Global request_timeout: 300s.
- **Strix Halo GPU**: Vulkan is the working backend (ROCm/HIP path abandoned — HSA runtime blocked on Debian 13). Build at `/root/llama.cpp/build-vk/`, commit `4fc4ec5` (2026-07-01), ggml 0.15.3 shared-lib arch. Mesa RADV 25.0.7, KHR_coopmat fast path active. ~70 tok/s gen, 532 tok/s prompt. Service: `strix-server.service` on port 8080, model: `qwen3.6-35B-udq4`, alias `strix-moe`, 128K context, flash-attn + q4 KV, multimodal (mmproj loaded).
- **Port conflict detection (2026-07-05)**: All 3 GPU wrappers now detect ghost processes squatting port 8080 before starting. `.8` and `.110` use inline pre-start check in `llama-wrapper.sh`; `.15` uses `/usr/local/bin/port-cleanup.sh` ExecStartPre. Replaces the blanket `pkill -9 -x llama-server` on .15 which would kill ALL llama-server instances regardless of port. Ghost detection was the root cause of .8 crash-looping for 27+ restarts (stale pid 25836 squatting 8080 after OOM kill).
- **Strix Halo thermal safeguard (2026-07-02)**: `ornith-server.service` has `-n 8192` (hard generation cap per request). Without it, `--predict` defaults to -1 (infinity) — a runaway request from .123 (Mumuni) decoded 39,868 tokens over 24 min, pushing Tctl to 98°C (crit 89.8°C) and throttling 70→29 t/s. The cap bounds worst-case generation to ~5 min. Do NOT remove `-n` without a replacement ceiling. Sustained load hits ~84°C even at 92s; the APU is fanless/low-flow. Clients MUST also set `max_tokens`.
- **Strix Halo thermal safeguard (2026-07-02)**: `strix-server.service` has `-n 8192` (hard generation cap per request). Without it, `--predict` defaults to -1 (infinity) — a runaway request from .123 (Mumuni) decoded 39,868 tokens over 24 min, pushing Tctl to 98°C (crit 89.8°C) and throttling 70→29 t/s. The cap bounds worst-case generation to ~5 min. Do NOT remove `-n` without a replacement ceiling. Sustained load hits ~84°C even at 92s; the APU is fanless/low-flow. Clients MUST also set `max_tokens`.
- **Port 8080 firewall**: amdpve iptables restricts 8080 to 192.168.68.116 (LiteLLM/router host) only. All inbound connections are from .116 (LiteLLM proxied via nginx). Localhost curls hang (SYN dropped). Always test from .116.
- **Router sidecar fallback**: `router.py` `check_gpu_health()` now probes GPU `/health` directly when sidecar at :8090 is absent. Sidecar JSON exporters not deployed on any GPU host — router relies on GPU-direct fallback.
- **Router GPU_MOE_URL bug (fixed 2026-07-01)**: docker-compose had `GPU_MOE_URL=.110:8080` (gemma host) instead of `.15:8080` (amdpve). Corrected.
@@ -198,11 +263,14 @@ If no SSH access, send Zulip DM via abiba-bot.
## GPU Inference Benchmarks (Current)
| GPU | Model | Gen tok/s | Prompt tok/s | Baseline | Samples |
| GPU | Model | Gen tok/s | Prompt tok/s | Baseline | Context |
|-----|-------|-----------|--------------|----------|---------|
| RTX 3090 (.8) | qwen3.6-27B-code | 75 | 305 | 74 | 6 |
| RTX 5070 (.110) | gemma-4-12b | 75 | 323 | 75 | 6 |
| Strix Halo (.15) | ornith-1.0-35b | 70 | 532 | 70 | 6 |
| RTX 3090 (.8) | qwen3.6-27B-code (MTP) | **63** | — | | **128K** |
| RTX 5070 (.110) | gemma-4-12b (IQ4_NL+MTP) | **191** | — | | **128K** |
| Strix Halo (.15) | qwen3.6-35B-udq4 | **65** | 140 | | **128K** |
Benchmarks from 2026-07-17. Strix Halo model: qwen3.6-35B-udq4. RTX 5070 MTP provides 2.7x speedup over pre-upgrade 70 tok/s.
All 3 GPUs now at 128K context (2026-07-17, reduced from 256K for stability).
Benchmarks run through LiteLLM proxy (192.168.68.116:4001) every 5 minutes.
Degradation alerts fire at 30% (warning) and 50% (critical) below baseline.
@@ -210,12 +278,55 @@ History stored at `/root/data/toks-history.json` with 7-day rolling window.
**Note (2026-07-01)**: Strix Halo prompt tok/s jumped 209→532 after Vulkan rebuild (cooperative-matrix fast path now active on GFX1151). Baseline may need re-calibration.
## Agent Config Implications (2026-07-08)
## Agent Config Implications (2026-07-15)
With NVIDIA GPUs at 128K context:
- Agents using `syslog-auto` (50/50 qwen+ornith): keep `context_length: 262144` — ornith supports it, Litellm fallbacks handle qwen overflow
- Agents using `qwen3.6-27B-code` directly: set `context_length: 131072` and `max_tokens: 4096` per thermal safety rule
- Agents using `gemma-4-12b` directly (auxiliary tasks): set `context_length: 131072`
- Compression threshold at 0.65: fires at ~170K for syslog-auto (262K ctx), ~85K for direct qwen/gemma (128K ctx)
- All Hermes clients MUST set `max_tokens: 4096` — first line of defense before server-side `-n 8192` cap
- Port 8080 is used on all 3 GPU hosts (not 8090 as previously documented)
### Stable Aliases — CRITICAL
All agent configs MUST use stable role-based aliases, never model-specific names:
- `compression.model: strix-moe` (NOT `qwen3.6-35B-udq4`)
- `auxiliary.vision.model: gpu-light` (NOT `gemma-4-12b`)
- `delegation.model: gpu-dense` (NOT `qwen3.6-27B-code`)
- `auxiliary.web_extract.model: gpu-light`
When the underlying model is swapped, only the LiteLLM config changes — agent configs are untouched.
### Context Windows
- RTX 3090: **128K** (reduced from 256K 2026-07-17) | RTX 5070: **128K** (reduced from 256K) | Strix Halo: **128K**
- **All agents**: 128K ceiling — stable margin. For >128K workloads, use external providers (deepseek)
- Compression threshold 0.65: fires at ~85K (~43K headroom before 128K ceiling)
- Mumuni compression model alias: `strix-moe` with 300s timeout
### Mumuni Agent Profile
Mumuni (CT114, 192.168.68.123) is the primary business assistant. This profile is the reference for all agent configs:
| Setting | Value | Notes |
|---------|-------|-------|
| `model.default` | `syslog-auto` | Weighted pool (55% qwen, 30% strix, 15% gemma) |
| `model.provider` | `custom:litellm` | LiteLLM on CT116 |
| `compression.model` | `strix-moe` | Stable alias — survives model swaps |
| `aux.compression.model` | `strix-moe` | Compression auxiliary model |
| `aux.vision.model` | `gpu-light` | Vision tasks (RTX 5070) |
| `aux.web_extract.model` | `gpu-light` | Web extraction |
| `delegation.model` | `gpu-dense` | Sub-agent reasoning (RTX 3090) |
| `context.max_context_window` | 131072 (128K) | Reduced from 256K 2026-07-17 — stable 128K ceiling |
| `compression.threshold` | 0.65 | Triggers at ~85K |
| `compression.target_ratio` | 0.3 | Compresses to ~38K |
| `compression.protect_last_n` | 40 | Preserves last 40 messages |
| `memory.memory_char_limit` | 800 | Brief memory entries |
| `personalities` | `creative` | Creative assistant personality |
| Platforms | cli, homeassistant, signal, telegram, zulip | All Hermes platforms |
| Main model timeout | 300s | LiteLLM global timeout |
| Compression model timeout | 300s | strix-moe timeout increased from 120s |
### Agent Update Status (2026-07-15)
| Agent | Host | Status |
|-------|------|--------|
| **Mumuni** | CT114 (.123) | ✅ Updated to stable aliases |
| **Tanko** | CT112 (.122) | ✅ Updated to stable aliases |
| **Koby** | CT111 (.129) | ❌ SSH unreachable — needs Zulip DM |
| **Koonimo** | CT113 | ❌ SSH unreachable — needs Zulip DM |
| **Kagenz0** | CT105 | ❌ SSH unreachable — needs Zulip DM |
All Hermes clients MUST set `max_tokens: 4096` — first line of defense before server-side `-n 8192` cap.
+1 -1
View File
@@ -47,7 +47,7 @@ poll .15:8080 directly; must go through router on .116.
| Router (unified) | `http://192.168.68.116/health/unified` | 15s | models, CB, scores, GPU status |
| Router (basic) | `http://192.168.68.116/health` | 15s | basic aliveness |
| LiteLLM | `http://192.168.68.116/litellm/health` | 15s | proxy health, model count |
| Strix Halo | `http://192.168.68.116/health/unified` (router) | 15s | ornith status via router — cannot poll .15:8080 directly (firewalled to .116 only) |
| Strix Halo | `http://192.168.68.116/health/unified` (router) | 15s | Strix Halo status via router — cannot poll .15:8080 directly (firewalled to .116 only) |
| Dashboard | `http://192.168.68.116/dashboard/` | 15s | harness-dashboard aliveness |
### Alert Delivery
+311
View File
@@ -0,0 +1,311 @@
---
kind: responsibility
name: gpu-self-heal
description: >
GPU fleet self-healing — detects anomalies, applies remediation, tracks
benchmarks, and predicts failures before they happen. Extends gpu-monitor
(v2.1.0) with active remediation rules, Prometheus metrics consumption,
VRAM trend analysis, and predictive alerting.
UPDATED 2026-07-18: Model assignments synced to 2026-07-17 swaps.
Router (port 9000) references replaced with direct GPU routing.
Benchmark baselines refreshed to live values.
Prometheus exporters removed — not deployed; fall back to direct sidecar probes.
Stable role-based aliases (strix-moe, gpu-dense, gpu-light) from gpu-fleet.
agent: abiba
depends_on:
- gpu-monitor.prose.md (live data source on .24:9100)
- gpu-fleet.prose.md (source of truth for topology, aliases, model assignments)
---
## Maintains
- gpu-health: { status: "healthy"|"degraded"|"down", issues: array, actions: array }
- gpu-self-heal-log: array of { timestamp, gpu, issue, action, result } — audit trail
- benchmark-regression: { gpu, baseline_tok_sec, current_tok_sec, trend, alerts }
- vram-trend: { gpu, current_mb, rate_mb_per_hour, projected_full_in_hours }
- circuit-breaker-status: { gpu, open, auto_reset_attempted, last_reset }
## Requires
- gpu-monitor:function — Live fleet data from localhost:9100/gpu-data
- Direct sidecar probe access to all GPU hosts (:8080/health)
- SSH access to GPU hosts for restart operations
## Continuity
- Self-driven: check every 60 seconds against GPU monitor data
- Also wakes on gpu-fleet health degradation
- On fix: verify with benchmark inference test before declaring resolved
- Escalate: after 3 failed remediation attempts → Zulip #agent-hub alert
---
## Current Fleet Baseline (2026-07-18)
| Alias | GPU | Host | Model | VRAM | Ctx | tok/s | Role |
|-------|-----|------|-------|------|-----|-------|------|
| `gpu-dense` | RTX 3090 24GB | ct8 (.8:8080) | ThinkingCap Qwen3.6-27B Q4_K_M + MTP + vision | 21.6/24.6GB (88%) | 128K | 74.9 | Heavy reasoning, code gen |
| `gpu-light` | RTX 5070 12GB | ct110 (.110:8080) | HauhauCS Gemma4-12B QAT Q4_K_M + MTP draft | 10.1/12.2GB (83%) | 128K | 169.6 | Vision, web extract, light tasks |
| `strix-moe` | Strix Halo 64GB | ct15 (.15:8080) | qwen3.6-35B-udq4 | ~10/64GB (16%) | 128K | 62.9 | Compression, summarization, long docs |
Key notes:
- All models use direct GPU routing via LiteLLM (`api_key: not-needed`). Router (port 9000) is deprecated and NOT in the inference path.
- Stable aliases (gpu-dense, gpu-light, strix-moe) from gpu-fleet are the canonical names for agent configs. Model-specific names still work but are deprecated.
- RTX 5070 tok/s is 2.3x faster than RTX 3090 for its model — gpu-light is the fastest endpoint. Route vision/web/light work there first.
- Strix Halo is 62.9 tok/s (89% of 70.5 baseline) — below optimal but stable. Check for competing workloads.
- RTX 3090 VRAM at 88% — within role-appropriate range (role = heavy reasoning, needs the headroom).
- RTX 5070 VRAM at 83% — role-appropriate for vision/web (smaller batch sizes).
## Remediation Rules
### Rule 1: GPU Temperature Critical (>85°C for >2 min)
- **Detect**: Any GPU temp >85°C sustained for 2+ consecutive polls
- **Fix**:
1. Reduce inference concurrency on that GPU (load-side cooling only — NO fan control)
2. Redirect new requests to cooler GPUs via LiteLLM fallback chains (gemma → qwen, qwen → gemma)
3. If all GPUs hot, alert about cooling infrastructure
- **Verify**: Temp drops below 80°C within 5 minutes
- **Escalate after**: 3 verification failures → Zulip alert
### Rule 2: VRAM Leak Detection (tiered by GPU capacity)
- **Detect**: VRAM growing at sustained rate over 6+ hour window
- RTX 3090 (24GB): ≥300MB/hour
- RTX 5070 (12GB): ≥300MB/hour
- Strix Halo (64GB UMA): ≥200MB/hour
- **Fix**:
1. Log VRAM snapshot with process list (nvidia-smi/rocm-smi + ps aux)
2. If llama-server is the growth source → restart with memory cap flag
3. If unknown process → kill and alert
- **Verify**: VRAM growth rate drops below threshold
- **Escalate after**: persistent leak after restart → hardware investigation
### Rule 3: Model Inference Timeout / GPU Stuck
- **Detect**: >50% failure rate over 60s window + 30s grace period (not just single stuck request)
- **Fix**:
1. Restart llama-server on affected GPU host
2. Wait 15s for model to reload
3. Run benchmark inference test
- **Verify**: Model returns 200 with <30s response, failure rate drops to 0%
- **Escalate after**: 3 restarts in 1 hour → GPU hardware check
### Rule 4: Benchmark Regression (>20% drop)
- **Detect**: gen_tok_per_sec drops >20% below baseline over 3+ benchmarks
- RTX 3090 baseline: 74.8 tok/s → alert at <59.8 tok/s
- RTX 5070 baseline: 165.2 tok/s → alert at <132.2 tok/s
- Strix Halo baseline: 70.5 tok/s → alert at <56.4 tok/s
- **Fix**:
1. Check GPU utilization — if >90%, other process is competing
2. Check power limit — if throttled, restore to max
3. Check thermal — if hot, apply Rule 1
- **Verify**: Benchmark returns to within 10% of baseline
- **Escalate after**: persistent regression → possible hardware degradation
### Rule 5: Circuit Breaker Stuck Open
- **Detect**: Circuit breaker open >10 minutes with GPU reporting healthy
- **Note**: Router (port 9000) is deprecated. If circuit breakers are reported by gpu-monitor, they come from LiteLLM's internal tracking, not the old router.
- **Fix**:
1. Verify GPU /health returns 200 on direct port (:8080)
2. If GPU healthy, alert but do NOT reset via router API (deprecated)
3. Check LiteLLM health directly: http://192.168.68.116/litellm/health/liveliness
4. Restart LiteLLM container on CT 116 if circuit breakers are stuck
- **Verify**: LiteLLM returns healthy, circuit breaker clears within 60s
- **Escalate after**: LiteLLM restart doesn't clear → human investigation
### Rule 6: Strix Halo Unreachable
- **Detect**: Strix not responding — probe .15:8080 directly (firewall opened .24→.15)
- **Fix**:
1. SSH to .15 → check llama-server process
2. Restart llama-server if not running
3. Verify through both direct probe AND LiteLLM health
- **Verify**: Direct health probe returns 200, LiteLLM reports model healthy
- **Escalate**: If host .15 itself is unreachable → infrastructure alert
### Rule 7: GPU Data Source Unreachable (replaces old Prometheus rule)
- **Detect**: gpu-monitor endpoint (localhost:9100/gpu-data) or sidecar port (:8080) on any GPU unreachable for >2 polls
- **Fix**:
1. If gpu-monitor is down: restart systemd service `gpu-monitor.service` on this host
2. If sidecar is down: SSH to GPU host → check llama-server process → restart systemd service
3. Fall back to direct nvidia-smi/rocm-smi probe via SSH if all API paths fail
- **Verify**: gpu-monitor returns healthy + all sidecars reachable
- **Escalate after**: 3 failed restarts → networking issue
### Rule 8: Predictive Thermal Warning (two-tier)
- **Detect**:
- Tier 1 (warning): temp >70°C AND rising >2°C/min → reduce concurrency, no alert
- Tier 2 (critical): temp >80°C AND still rising → full alert + aggressive load shedding
- **Fix**:
- Tier 1: silently reduce parallel requests to that GPU by 50%
- Tier 2: redirect all new requests away, alert #agent-hub, apply Rule 1 logic
- **Verify**: Temp rise rate drops below 1°C/min (Tier 1) or temp drops below 80°C (Tier 2)
- **Escalate**: If Tier 2 triggers and temp still rising after 5 min → possible hardware failure
### Rule 9: Context Window Optimization
- **Detect**: Benchmark tok/s vs baseline for each GPU at current context (all 128K)
- RTX 3090 (128K ctx, ThinkingCap): baseline 74.8 tok/s — currently at 74.9 (100%)
- RTX 5070 (128K ctx, HauhauCS QAT): baseline 165.2 tok/s — currently at 169.6 (103%)
- Strix Halo (128K ctx, qwen3.6-35B-udq4): baseline 70.5 tok/s — currently at 62.9 (89%)
- **Fix**:
- If tok/s > baseline → context has headroom, consider increasing
- If tok/s < 90% baseline → reduce context by 25% and retest
- If tok/s within 10% of baseline → optimal, no change
- Strix Halo at 89% of baseline → MONITOR but do not reduce yet (recent model swap may still be settling)
- **Verify**: Re-benchmark after context change, confirm within 10% of target
- **Escalate**: If context can't be adjusted without significant perf loss
### Rule 10: Workload Distribution Optimization (updated 2026-07-18)
- **Detect**: GPU roles misaligned with hardware capabilities
- **Target distribution**:
- RTX 3090 (gpu-dense, 24GB, 74.9 tok/s) → Heavy reasoning, code gen, long conversations (slowest per-token but largest context capacity). Weight: 0.55 (LiteLLM).
- RTX 5070 (gpu-light, 12GB, 169.6 tok/s) → Vision/image, web search, lightweight tasks (2.3x faster than 3090 per token). Weight: 0.15 (LiteLLM).
- Strix Halo (strix-moe, 64GB, 62.9 tok/s) → Context compression, summarization, long docs (MoE model). Weight: 0.30 (LiteLLM).
- **Note**: RTX 5070 is the fastest endpoint per token. Route high-volume, low-complexity work there first.
- **Fix**:
- Alert if any GPU is handling workload outside its designated role
- Recommend agent alias updates to match workload to GPU role (use stable aliases: gpu-dense, gpu-light, strix-moe)
- Track per-GPU request distribution via LiteLLM spend logs
- **Verify**: Each GPU's request pattern matches its designated role within 24h
- **Escalate**: If role mismatch persists >48h → agent alias audit needed
---
## Execution
```prose
-- Phase 1: Fetch live GPU data
let fleet = call gpu-monitor
endpoint: "http://localhost:9100/gpu-data"
-- Phase 2: Evaluate each GPU against remediation rules
let actions = []
for gpu in fleet.gpus:
-- Rule 1: Thermal critical
if gpu.temp_c > 85 and sustained_for(gpu, 120):
push actions apply-thermal-fix(gpu)
-- Rule 2: VRAM leak
let vram_rate = calculate-vram-trend(gpu, hours=6)
if vram_rate > 50:
push actions apply-vram-fix(gpu, vram_rate)
-- Rule 4: Benchmark regression
let bench = fleet.benchmarks[gpu.hostname]
if bench.current_tok_s < bench.baseline_tok_s * 0.8:
push actions apply-benchmark-fix(gpu, bench)
-- Rule 3: Model stuck
for model in fleet.summary.available_models:
if model.consecutive_timeouts >= 3:
push actions apply-model-restart(model)
-- Rule 5: Circuit breaker check via LiteLLM (router deprecated)
if fleet.summary.circuit_breakers_open > 0:
push actions check-litellm-circuit-breakers()
-- Rule 6: Strix Halo
if not fleet.strix.running and pingable("192.168.68.15"):
push actions apply-strix-restart()
-- Rule 7: GPU data source
if not fleet.gpus or len(fleet.gpus) < 2:
push actions check-gpu-monitor-service()
-- Rule 8: Predictive thermal
for gpu in fleet.gpus:
let rise_rate = calculate-temp-rise(gpu, minutes=5)
if rise_rate > 2.0 and gpu.temp_c < 80:
push actions apply-proactive-cooling(gpu)
-- Phase 3: Execute actions, verify, log
for action in actions:
let result = execute-with-verify(action)
log-to-kg(action, result)
if result.failed:
escalate-if-needed(action)
-- Phase 4: Update health state
call update-gpu-health
gpus: fleet.gpus
actions: actions
status: derive-overall-status(fleet, actions)
-- Wait 60s and repeat
```
## Audit Trail Format
```json
{
"run_id": "gpu-self-heal-20260718-001",
"timestamp": "2026-07-18T08:00:00Z",
"gpu": "ct8-rtx3090",
"issue": "thermal-critical",
"detected": { "temp_c": 87, "duration_s": 180 },
"action": "load-shedding",
"result": "resolved",
"verification": { "temp_c": 76, "after_s": 300 },
"escalated": false
}
```
---
## Reporting
### 1. Knowledge Graph
Every action logged as `[GPU-SELF-HEAL] <run_id>` node with full audit trail.
### 2. Zulip Alerts (#agent-hub → alerts-gpu)
- `issues_fixed > 0` → "🛠 GPU Self-Heal — <gpu> <issue> resolved"
- `issues_escalated > 0` → "⚠ GPU Self-Heal — <gpu> needs attention"
- Every 100th clean cycle → "✅ GPU Fleet: All Clear"
### 3. Weekly Benchmark Report
- Per-GPU tok/s trend over 7 days
- Regression alerts if any GPU degrades >10% week-over-week
---
## Design Decisions (Verified 2026-07-12, Reaffirmed 2026-07-18)
1. **Fan control**: ❌ NO auto fan control. Load-side cooling only (reduce concurrency, redirect).
2. **Model restart**: ✅ Only if >50% failure rate over 60s + 30s grace period. Not on single stuck request.
3. **Strix direct access**: ✅ Open firewall .15:8080 → .24 for direct health probe + restart.
4. **VRAM thresholds**: Tiered — **300MB/h** (RTX 3090), **300MB/h** (RTX 5070), 200MB/h (Strix). Previous values (100/50) were too sensitive; raised 2026-07-18 based on operational data.
5. **CB auto-reset**: ✅ Router deprecated — circuit breakers go through LiteLLM health check + container restart if needed. No per-GPU auto-reset.
6. **Benchmark baseline**: Rolling 30-day average, recalculated weekly. Current baselines live in gpu-monitor.
7. **Predictive alerts**: Two-tier — warn at >70°C+rising (>2°C/min), critical at >80°C+rising.
8. **Prometheus**: ❌ Not deployed. Use direct sidecar probes (:8080/health) and gpu-monitor API. Prometheus integration deferred until exporters are running on GPU hosts.
## Lessons Learned (2026-07-12, Updated 2026-07-18)
### L1: API Key Standardization Is Critical
- All GPU llama-servers MUST use the same api-key as the LiteLLM config.
- RTX 5070 had `--api-key sk-loc...5678` while LiteLLM sent `not-needed`.
This caused cascading 401 → fallback → timeout → 401 loops.
- **Rule**: Any new GPU or model restart MUST verify api-key matches LiteLLM config (`not-needed` for direct routing).
### L2: Fallback Chain Cascading Failures
- When one model returns 401 (auth) and another is slow (timeout), the fallback
chain creates an infinite loop.
- **Rule**: If a model returns 401 (auth error), do NOT fall back to it again.
Mark it as permanently failed for this request.
### L3: Verify Running State, Not Docs
- RTX 3090 was documented at 128K context. Running at 128K (verified 2026-07-18).
- Parallel count: 1 on both RTX 3090 and RTX 5070 (matches docs for current models).
- **Rule**: Before making decisions, check `/proc/PID/cmdline` on GPU hosts.
### L4: Infisical Is Not Always Available
- Keep a local `.env` fallback for `LITELLM_API_KEY`.
- **Rule**: Always verify credential source is reachable before relying on it.
### L5: GPU Monitor Response Size Can Cause Self-Heal Crash
- gpu-self-heal crashed with KeyboardInterrupt during json.loads() of 20MB response.
- Root cause: router poll returns accumulated data → cache balloons.
- **Rule**: Self-heal must enforce a read timeout AND max response size on every poll.
If monitor response > 1MB, log a warning and skip the cycle rather than crashing.
### L6: Stable Aliases Replace Model Names
- gpu-fleet introduced stable aliases (strix-moe, gpu-dense, gpu-light) on 2026-07-15.
- Self-heal must use aliases for reporting and alerting, not model-specific names.
- **Rule**: All alert messages and KG nodes use the stable alias as the GPU identifier.
+50 -27
View File
@@ -5,7 +5,7 @@ version: 1.0.0
description: >
Canonical known-good baseline for all Syslog Hermes agents. Captures the exact
configuration state, keys, workarounds, and audit procedure. When an agent's
configuration goes sideways, restore from this baseline. Last verified 2026-07-08. GPU context reduced to 128K on .8/.110, parallel 2 fleet-wide.
configuration goes sideways, restore from this baseline. Last verified 2026-07-16. All GPUs 128K context (reduced from 256K for stability Jul 2026) (RTX 3090 .8, RTX 5070 .110, Strix Halo .15). Parallel 1 fleet-wide (Strix Halo handles compression solo).
author: Abiba (pi agent)
---
@@ -22,30 +22,38 @@ done
## Agent Map
| Agent | CT | Node | IP | LiteLLM Key | LiteLLM Alias | Platform |
|-------|-----|------|-----|-------------|---------------|----------|
| Tanko | 112 | amdpve | .122 | `sk-CggiHWlamQyShxWC3Hx6uw` | `tanko` | Hermes |
| Mumuni | 114 | minipve | .123 | `sk-VrqCNlwUgzoNGOpikJ7nwQ` | `mumuni` | Hermes |
| Tdunna | 111 | amdpve | srv1079750 | `sk-Qvzi4uYQBhlSK_XstEhcyQ` | `tdunna` | **pi** |
| Baggy | 113 | amdpve | ? | `sk-krnw_zGBwvvL5b7l2t-s-A` | `baggy` | Hermes |
| Agent | CT | Node | IP | LiteLLM Alias | Key Source | Platform |
|-------|-----|------|-----|---------------|------------|----------|
| Tanko | 112 | amdpve | .122 | `tanko` | Infisical vault | Hermes |
| Mumuni | 114 | minipve | .123 | `mumuni` | Infisical vault | Hermes |
| Koby | 129 | amdpve | srv1079750 | `koby` | Infisical vault | **Hermes** |
| Koonimo | 114 | amdpve | ? | `koonimo` | Infisical vault | Hermes |
| Shumba | — | 192.168.68.119 | N/A | N/A (DeepSeek) | Hermes (RETIRED — CT119 now Infisical vault) |
> **Note**: CT hostnames (tdunna→CT129, baggy→CT114) differ from agent identities (koby, koonimo).
Access: `pct-run <CT_ID> <command>` — no IPs needed. GPU hosts (.8, .110, .15) use SSH.
Keys are stored in Infisical vault (project=agents, env=production) and injected at
runtime via `infisical run --` wrapper. Plaintext keys removed from this baseline.
## Key Architecture
```
Infisical vault → infisical run -- hermes gateway → LITELLM_API_KEY (runtime)
Agent (systemd) → LITELLM_API_KEY → LiteLLM (:116/v1) → Router (:9000) → GPU (llama-server)
└── Key DB (Postgres)
└── Key DB (Postgres)
```
- **Master key**: `sk-litellm-7f96080dd99b15c36bd4b333b58a6796` — ADMIN ONLY, never in agent configs
- **Master key**: stored in Infisical vault (project=infrastructure, secret=LITELLM_MASTER_KEY) — ADMIN ONLY
- **Agent keys**: Each agent has a dedicated key in LiteLLM's database with alias matching the agent name
- **Key source**: `/etc/environment``LITELLM_API_KEY=sk-...` (systemd service sources this)
- **Override**: `/home/jerome/.config/systemd/user/hermes-gateway.service.d/env.conf` (if present, must match)
- **Key injection**: `infisical run --project=agents --env=production -- hermes gateway run` injects `LITELLM_API_KEY` at runtime
- **Key source**: Infisical vault → runtime env var. /etc/environment is CLEAN (stripped, tagged `# [INFISICAL]`)
- **Legacy override** (pre-migration): `/home/jerome/.config/systemd/user/hermes-gateway.service.d/env.conf` — should be REMOVED
## Config Pattern — Mandatory Fields
### For Hermes Agents (Tanko, Mumuni, Baggy)
### For Hermes Agents (Tanko, Mumuni, Koonimo)
Every agent's `/root/.hermes/config.yaml` (or `/home/jerome/.hermes/config.yaml`) MUST have:
@@ -76,7 +84,7 @@ auxiliary:
model: gemma-4-12b # or syslog-auto
base_url: http://192.168.68.116/v1
api_key_env: LITELLM_API_KEY
api_key: <ACTUAL_KEY_FROM_/etc/environment> # ← MANDATORY workaround
api_key: <value from: infisical secrets get LITELLM_API_KEY --project=agents --env=production> # ← MANDATORY workaround
timeout: 60
download_timeout: 30
```
@@ -91,7 +99,7 @@ auxiliary:
model: syslog-auto # or gemma-4-12b
base_url: http://192.168.68.116/v1
api_key_env: LITELLM_API_KEY
api_key: <ACTUAL_KEY_FROM_/etc/environment> # ← MANDATORY workaround
api_key: <value from: infisical secrets get LITELLM_API_KEY --project=agents --env=production> # ← MANDATORY workaround
timeout: 120
```
@@ -110,8 +118,7 @@ LiteLLM/harness will fail with:
401: LiteLLM Virtual Key expected. Received=no-k****ired, expected to start with 'sk-'
```
**Workaround**: Set `api_key` directly (copy the value from `/etc/environment`) alongside
`api_key_env` in every auxiliary task config that uses the harness provider.
**Workaround**: Set `api_key` directly (copy the value from Infisical vault: `infisical secrets get LITELLM_API_KEY --project=agents --env=production`) alongside `api_key_env` in every auxiliary task config that uses the harness provider.
**Permanent fix**: Patch `_resolve_task_provider_model()` to resolve `api_key_env` when
`api_key` is empty:
@@ -129,7 +136,11 @@ if not cfg_api_key:
```bash
for ct in 112 114 111 113; do
echo "=== CT $ct ==="
pct-run $ct grep LITELLM_API_KEY /etc/environment
# Verify /etc/environment is CLEAN (no LITELLM_API_KEY)
pct-run $ct "grep -c LITELLM_API_KEY /etc/environment 2>/dev/null || echo '0 (clean)'"
# Verify gateway uses infisical run wrapper
pct-run $ct "ps aux | grep 'infisical run' | grep -v grep"
# Check for hardcoded harness keys
pct-run $ct grep "api_key: sk-" /root/.hermes/config.yaml | grep -v api_key_env
echo ""
done
@@ -137,15 +148,17 @@ done
### Master Key Leak Check
```bash
# On every agent:
pct-run <CT> grep -rl "sk-litellm-7f96080dd" /root/ /etc/ 2>/dev/null
# Must return empty
# On every agent — must return empty:
pct-run <CT> grep -rl "sk-litellm" /root/ /etc/ 2>/dev/null
# Vault is the only place the master key should exist
```
### Verify Key Works
```bash
# Retrieve key from vault and test:
KEY=$(infisical secrets get LITELLM_API_KEY --project=agents --env=production --plain)
curl -s http://192.168.68.116:80/v1/models \
-H "Authorization: Bearer <AGENT_KEY>" | grep syslog-auto
-H "Authorization: Bearer $KEY" | grep syslog-auto
# Must return model list
```
@@ -156,22 +169,32 @@ pct-run <CT> grep -A8 "vision:" /root/.hermes/config.yaml | grep api_key
# Must show both api_key: sk-... and api_key_env: LITELLM_API_KEY
```
### For pi Agents (Tdunna)
### For Koby (CT 129 / tdunna)
Tdunna (CT111) runs pi 0.80.3 via PM2 with the Zulip extension (router-worker architecture).
Koby runs Hermes on CT 129 (tdunna). Config files at `/root/.hermes/config.yaml`.
Same Hermes pattern as Tanko/Mumuni/Koonimo — see config sections above.
**LiteLLM key**: alias `koby` in LiteLLM DB, injected via `infisical run --` wrapper.
### For pi Agents (Abiba)
Abiba (CT100) runs pi via PM2 with the Zulip extension.
Config files: `~/.pi/agent/models.json`, `~/.pi/agent/settings.json`.
**models.json** — Must only list models authorized for the agent's LiteLLM key:
**models.json** — Must only list models authorized for the agent's LiteLLM key.
Key is injected via `infisical run --` wrapper at PM2 startup:
```json
{
"providers": {
"syslog-harness": {
"baseUrl": "http://192.168.68.116/v1",
"api": "openai-completions",
"apiKey": "sk-...",
"apiKey": "${LITELLM_API_KEY}",
"models": [
{ "id": "syslog-auto" },
{ "id": "ornith-1.0-35b" },
{ "id": "strix-moe" },
{ "id": "gpu-dense" },
{ "id": "gpu-light" },
{ "id": "qwen3.6-27B-code" },
{ "id": "gemma-4-12b" }
]
@@ -257,6 +280,6 @@ If they differ → ghost detected → kill ghost → start fresh.
| Date | Change |
|------|--------|
| 2026-07-08 | Tdunna: fixed model mismatch (qwen3.6-35B-A3B→syslog-auto), added pi-specific config section. Key updated to sk-Qvzi4uYQBhlSK_XstEhcyQ. Added Failure Mode #11 to zulip-adapter-lessons. |
| 2026-07-08 | Koby: fixed model mismatch (qwen3.6-35B-A3B→syslog-auto), added config section. Key rotated and stored in vault. Added Failure Mode #11 to zulip-adapter-lessons. |
| 2026-07-06 | Port conflict detection added to all 3 GPU wrappers. Consolidated health check script deployed. Zulip streaming edit_message enabled for Tanko/Mumuni. |
| 2026-07-05 | Baseline created. All 4 agents audited, master key removed, api_key workaround applied |
+165 -48
View File
@@ -5,35 +5,44 @@ description: >
Standard Hermes configuration template for Syslog Solution LLC agents.
Enforces shared infrastructure setup (Firecrawl, SearXNG, local models,
RA-H OS MCP) while keeping agent-specific API keys and model choices.
Updated 2026-07-08: GPU context reduced to 128K on NVIDIA (.8, .110),
parallel 2 on all GPUs, LiteLLM timeouts tuned, context_length guidance added.
UPDATED 2026-07-18: Compression model switched to `syslog-auto` (was `strix-moe`)
to relieve Strix Halo pressure. syslog-auto distributes compression across the
weighted pool (55% RTX 3090, 30% Strix Halo, 15% RTX 5070).
UPDATED 2026-07-16: Compression model was the stable alias `strix-moe` (NOT `ornith-1.0-35b`,
which LiteLLM does not serve). All 3 GPUs verified at 128K (reduced from 256K 2026-07-17 for stability).
Added Rule 12 (Context-Issue Diagnostic) + Rule 13 (.env fallback enforcement) from the
2026-07-16 Mumuni root-cause investigation (WAL #1300).
UPDATED 2026-07-12: GPU workload redistributed. Compression → Strix Halo (later switched to syslog-auto 2026-07-18). RTX 3090 context verified at 128K. Infisical .env fallback required (Rule 3/13).
---
## Maintains
- template_version: "2.1.0"
- last_applied: timestamp
- agents_configured: ["tanko", "mumuni", "abiba", "tdunna", "baggy", "kagenz0"]
- agents_configured: ["tanko", "mumuni", "abiba", "koby", "koonimo", "kagenz0"]
- agent_keys: map (see Agent Keys section)
- infra_endpoints_verified: array
## Agent Keys (LiteLLM — Current 2026-07-04)
## Agent Keys (LiteLLM — Current 2026-07-11)
Each agent has a unique LiteLLM API key (virtual key) generated against the LiteLLM
PostgreSQL DB via `POST /key/generate` on CT 116. Keys are stored in the DB, not in
config files. The env var `LITELLM_API_KEY` is set in `/etc/environment` on each agent
host AND in `~/.hermes/.env` for gateway env propagation.
PostgreSQL DB via `POST /key/generate` on CT 116. Keys are stored in the DB.
The env var `LITELLM_API_KEY` is injected at runtime via `infisical run --` wrapper
(project=agents, env=production). /etc/environment and ~/.hermes/.env are NO LONGER
used for agent keys — stripped and tagged `# [INFISICAL]` post-migration.
Sub-agent profiles inherit auth from the main config — no separate keys needed.
| Agent | Key Alias | Host | SSH | Sub-Agents |
|-------|-----------|------|-----|-----------|
| Tanko | `tanko-*` | 192.168.68.122 | jerome@.122 | — |
| Mumuni | `mumuni-jul2026` | 192.168.68.123 | root@.123 | 6 profiles ✱ |
| Abiba | `abiba-*` | 192.168.68.24 | local | — |
| Tdunna | `tdunna-*` | ? | Zulip | — |
| Baggy | `baggy-*` | ? | Zulip | — |
| Mumuni | `mumuni` | 192.168.68.123 | root@.123 | 6 profiles ✱ |
| Abiba | `abiba-pi` | 192.168.68.24 | local | — |
| Koby | `koby` | CT 111 (tdunna) | Zulip | — |
| Koonimo | `koonimo` | CT 114 (baggy) | SSH root | — |
| Kagenz0 | `kagenz0-*` | ? | Zulip | — |
> CT hostnames (tdunna, baggy) differ from agent identities (koby, koonimo).
✱ Mumuni sub-agents: syslog-code, syslog-devops, syslog-email, syslog-research,
syslog-review, syslog-writer — all at `/root/.hermes/profiles/<name>/config.yaml`
@@ -51,11 +60,12 @@ Sub-agent profiles inherit auth from the main config — no separate keys needed
## API Key Rules
- `api_key_env: LITELLM_API_KEY` — Use env var for main model auth (preferred)
Key is injected at runtime via `infisical run --` wrapper — never in /etc/environment
- `api_key: ''` — Sub-agents leave empty to inherit from main config's custom_provider
- `api_key: sk-...` — Hardcoded key only as fallback when env var not possible
- Set `LITELLM_API_KEY` in `/etc/environment` on each host
- Store `LITELLM_API_KEY` in Infisical vault (project=agents, env=production)
- Sub-agents NEVER get their own key — they share the host agent's key
- Restart Hermes after updating `/etc/environment`
- Restart Hermes gateway after updating vault secret (key auto-injected via wrapper)
### Sub-Agent Profiles (Mumuni pattern)
@@ -79,7 +89,7 @@ Sub-agent profile rules:
5. **Auxiliary tasks** (vision, compression, etc.) also leave `api_key` empty
6. **Never hardcode a key** in sub-agent profiles
This ensures all 6 sub-agents use the same LiteLLM key set in `/etc/environment`.
This ensures all 6 sub-agents use the same LiteLLM key injected via `infisical run --` wrapper.
When the key is rotated, only the env var needs updating — all 7 configs (main + 6 subs)
work immediately after restart.
@@ -88,13 +98,13 @@ work immediately after restart.
```yaml
# ─── Model Selection ───
model:
default: <agent_model> # e.g., ornith-1.0-35b, qwen3.6-27B-code
default: <agent_model> # e.g., strix-moe, qwen3.6-27B-code, syslog-auto
provider: harness
base_url: http://192.168.68.116/v1
api_key_env: LITELLM_API_KEY # Set in /etc/environment AND ~/.hermes/.env
api_key_env: LITELLM_API_KEY # Injected via infisical run -- wrapper
max_tokens: 4096 # ⚠️ CRITICAL: Prevents unbounded generation
context_length: 262144 # For syslog-auto (ornith route supports 256K).
# Set 131072 if using qwen3.6-27B-code or gemma-4-12b directly.
context_length: 131072 # For syslog-auto (all GPUs at 128K for stability).
# Set 65536 if using gemma-4-12b directly (tight VRAM).
fallback_providers:
provider: deepseek
@@ -123,10 +133,11 @@ mcp_servers:
# ─── Compression ───
compression:
enabled: true
model: gemma-4-12b # ⚠️ Must match auxiliary.compression.model
model: syslog-auto # ⚠️ Switched from strix-moe 2026-07-18 to relieve Strix Halo.
# syslog-auto distributes across weighted pool (55% RTX 3090,
# 30% Strix Halo, 15% RTX 5070). All GPUs at 128K.
provider: harness
max_context_window: 262144 # For syslog-auto (ornith supports 256K).
# Set 131072 if using qwen or gemma directly.
max_context_window: 131072 # MUST match actual GPU capacity. All 3 GPUs are 128K (Jul 17).
threshold: 0.65 # Fires at ~170K for 262K window, ~85K for 128K
target_ratio: 0.30
protect_last_n: 40
@@ -136,37 +147,49 @@ compression:
# ─── Auxiliary Tasks (CONSISTENCY RULE) ───
# All auxiliary services MUST use identical model, base_url, and api_key_env:
# model: gemma-4-12b
# model: gpu-light # stable alias (NOT raw "gemma-4-12b")
# base_url: http://192.168.68.116/v1
# api_key_env: LITELLM_API_KEY
# Do NOT use syslog-auto for auxiliary tasks — it routes to the primary GPU.
# gemma-4-12b is a lightweight 12B model on the RTX 5070, freeing the Strix Halo
# for agent reasoning.
# Compression uses syslog-auto (switched from strix-moe 2026-07-18) to distribute
# load across the weighted pool and relieve Strix Halo pressure.
# Vision and web_extract use gpu-light = RTX 5070 (12B).
# Heavy aux (delegation, x_search) use gpu-dense (RTX 3090) instead.
# NEVER use raw model names (gemma-4-12b, qwen3.6-27B-code, qwen3.6-35B-udq4)
# in agent configs — use the stable aliases so model swaps don't break agents.
auxiliary:
vision:
provider: harness
model: gemma-4-12b
model: gpu-light # stable alias for RTX 5070 (was raw gemma-4-12b)
base_url: http://192.168.68.116/v1
api_key_env: LITELLM_API_KEY
timeout: 60
download_timeout: 30
web_extract:
provider: harness
model: gemma-4-12b
model: gpu-light # stable alias for RTX 5070
base_url: http://192.168.68.116/v1
api_key_env: LITELLM_API_KEY
timeout: 30
compression:
provider: harness
model: gemma-4-12b
base_url: http://192.168.68.116/v1
model: syslog-auto # Switched from strix-moe 2026-07-18. Relieves Strix Halo pressure.
base_url: http://192.168.68.116/v1 # Rule 5: /v1 NOT /litellm/v1
api_key_env: LITELLM_API_KEY
timeout: 60
timeout: 300 # gpu-fleet: 300s for large-history summarization (was 60)
# ─── Delegation / Heavy Aux (use gpu-dense = RTX 3090) ───
# delegation.model and x_search.model use gpu-dense (NOT raw qwen3.6-27B-code).
delegation:
model: gpu-dense # stable alias for RTX 3090 (was raw qwen3.6-27B-code)
provider: harness
base_url: http://192.168.68.116/v1
api_key_env: LITELLM_API_KEY
# ─── Custom Provider ───
custom_providers:
- name: harness
model: <agent_model>
model: syslog-auto # weighted pool (default)
base_url: http://192.168.68.116/v1
api_key_env: LITELLM_API_KEY
api_mode: chat_completions
@@ -176,7 +199,7 @@ custom_providers:
When LiteLLM keys are regenerated (e.g., after infrastructure changes):
1. **If SSH available**: `ssh <host> "sudo sed -i 's/LITELLM_API_KEY=.*/LITELLM_API_KEY=sk-<NEW>/' /etc/environment"`
1. **If SSH available**: Update Infisical vault: `infisical secrets set LITELLM_API_KEY=sk-<NEW> --project=agents --env=production`, then `ssh <host> "systemctl restart hermes-gateway"`
2. **If SSH unavailable**: Send Zulip DM via abiba-bot with update command
3. **After update**: Restart Hermes on the agent host
4. **Verify**: `curl -H "Authorization: Bearer sk-<KEY>" http://192.168.68.116/v1/models`
@@ -198,7 +221,14 @@ The following MUST be identical across ALL profiles:
### Rule 3: API Keys via Environment
- Prefer `api_key_env: LITELLM_API_KEY` over hardcoded keys
- Hardcoded keys in config.yaml become stale after key rotation
- `/etc/environment` persists across config updates
- Infisical vault secrets persist across config updates / reinstalls
- **NEW (July 2026): Always keep a local `.env` fallback.** Infisical service tokens
can expire/404 (tanko incident: token not found, gateway ran without key for hours).
The `.env` file should have the key uncommented as a fallback:
```
LITELLM_API_KEY=sk-...
# [INFISICAL] Also sourced from vault.sysloggh.net
```
- Restart Hermes after env var updates
### Rule 4: Sub-Agent Profiles Inherit Auth
@@ -208,7 +238,7 @@ The following MUST be identical across ALL profiles:
- Auxiliary tasks: `api_key: ''`, `provider: harness`
- Never hardcode a key in sub-agent profiles
- When main config uses `api_key_env`, sub-agents automatically use it
- This means key rotation only touches ONE file (`/etc/environment`)
- This means key rotation only touches ONE vault secret (`LITELLM_API_KEY`)
### Rule 5: Main Config Base URL
|- Use direct IP: `http://192.168.68.116/v1`
@@ -223,26 +253,48 @@ The following MUST be identical across ALL profiles:
- Apply to BOTH main config AND all sub-agent profiles
- For agents needing longer outputs: raise to 8192, but never omit
### Rule 7: Auxiliary Model Consistency
- All auxiliary services (vision, web_extract, compression) MUST use the same model:
- `model: gemma-4-12b`
- `base_url: http://192.168.68.116/v1`
### Rule 7: Auxiliary Model Consistency (UPDATED 2026-07-18)
- Vision and web_extract use `gpu-light` (stable alias, RTX 5070 — 12GB, vision-optimized)
- Compression now uses `syslog-auto` (switched from `strix-moe` 2026-07-18) to distribute
compression load across the weighted pool (55% RTX 3090, 30% Strix Halo, 15% RTX 5070).
This relieves Strix Halo pressure while keeping compression functional on all GPUs.
- **`syslog-auto` is the valid compression model** — LiteLLM serves it as the weighted pool.
Old configs with `strix-moe` for compression should be updated to `syslog-auto`.
- All auxiliary services MUST use identical routing:
- `base_url: http://192.168.68.116/v1` (Rule 5: `/v1`, NOT `/litellm/v1`)
- `api_key_env: LITELLM_API_KEY`
- **Do NOT use `syslog-auto`** for auxiliary tasks — it routes to the primary 35B reasoning GPU
- gemma-4-12b is a lightweight 12B model on the RTX 5070, keeping the Strix Halo free for reasoning
- The `compression:` block's `model` MUST match `auxiliary: compression: model` — they are two different configs for the same service
- **Compression via syslog-auto**: Routes through the weighted pool. Strix Halo still handles
~30% of compression calls (at 60 RPM via pool vs 40 RPM direct), but the bulk (55%)
goes to RTX 3090 which has ample spare capacity.
- The `compression:` block's `model` MUST match `auxiliary: compression: model`
- The `compression: max_context_window: 131072` MUST match actual GPU capacity (128K)
### Rule 8: Compression Threshold for 256K Models
- For 262K context window: `threshold: 0.65` (fires at ~170K tokens)
### Rule 8: GPU Workload Distribution (UPDATED 2026-07-18)
- **RTX 3090 (24GB, 128K ctx, qwen3.6-27B-code)**: Heavy reasoning, code gen, long conversations — also handles ~55% of compression via syslog-auto pool
- **RTX 5070 (12GB, 128K ctx, gemma-4-12b)**: Vision, web search, quick tasks, web_extract — handles ~15% of compression via syslog-auto pool
- **Strix Halo (64GB, 128K ctx, qwen3.6-35B-udq4)**: Agent reasoning, compression (~30% via syslog-auto pool), fallback for other GPUs
- Agent profiles MUST route auxiliary tasks to the correct GPU:
- `auxiliary.vision.model: gpu-light` (RTX 5070)
- `auxiliary.web_extract.model: gpu-light` (RTX 5070)
- `auxiliary.compression.model: syslog-auto` (distributed pool, switched from strix-moe 2026-07-18)
- Default model (`model.default`) and custom_provider remain `syslog-auto` for auto-routing
- For 128K context window: `threshold: 0.65` (fires at ~85K tokens)
- Do NOT use `threshold: 0.25` — this fires at 65K, causing premature context loss
- Do NOT use `threshold: 0.80` — this delays until 209K, risking the gateway hygiene layer
- `max_context_window: 262144` MUST match the model's actual capacity
- Do NOT use `threshold: 0.80` — this delays until ~105K, leaving only 23K margin
- `max_context_window: 131072` MUST match the model's actual capacity (128K)
- See `devops-hermes-compression` skill for full reference
### Rule 9: Default Model Must Be `syslog-auto` (All Agents)
### Rule 9: Compression Threshold for 128K Models
- For 128K context window: `threshold: 0.65` (fires at ~85K tokens)
- Do NOT use `threshold: 0.25` — this fires at 65K, causing premature context loss
- Do NOT use `threshold: 0.80` — this delays until ~105K, leaving only 23K margin
- `max_context_window: 131072` MUST match the model's actual capacity (all GPUs = 128K)
- See `devops-hermes-compression` skill for full reference
### Rule 10: Default Model Must Be `syslog-auto` (All Agents)
- **Hermes agents**: `model.default: syslog-auto`, `custom_providers[0].model: syslog-auto`
- **pi agents**: `defaultModel: syslog-auto` in `settings.json`, first model in `models.json`
- `syslog-auto` is the LiteLLM routing model — it load-balances between ornith-1.0-35b
- `syslog-auto` is the LiteLLM routing model — it load-balances between strix-moe
and qwen3.6-27B-code, with gemma-4-12b as fallback. Using it protects against:
- Model name typos that cause 403 errors and silent worker failures
- Single GPU downtime (routing falls back automatically)
@@ -250,7 +302,7 @@ The following MUST be identical across ALL profiles:
- **Exception**: Sub-agent profiles (Mumuni's 6 profiles) may specify explicit models
for specialized tasks, but MUST validate those models exist in the key's authorized list
### Rule 10: Validate Model IDs Before Deployment (pi Agents)
### Rule 11: Validate Model IDs Before Deployment (pi Agents)
- After configuring a pi agent's `models.json`, verify every model ID:
```bash
curl -s http://192.168.68.116:4000/v1/models \
@@ -261,6 +313,71 @@ The following MUST be identical across ALL profiles:
- A non-existent model ID causes 403 errors that silently break the pi RPC worker
(no `agent_end` emitted, worker stays "busy", Zulip messages pile up unprocessed)
### Rule 12: Context-Issue Diagnostic Checklist (ADDED 2026-07-16, WAL #1300)
When an agent shows "context issues" (premature compression, 401s, 504s, DeepSeek fallback),
verify ALL FOUR of these against the live config. They are the only root causes found in production:
1. **max_context_window correct?** — BOTH `compression.max_context_window` AND `context.max_context_window`
MUST be `131072` (all GPUs are 128K). A value of `262144` causes instability near 100K and must NOT be used.
~83K instead of ~170K. Check: `grep -n max_context_window ~/.hermes/config.yaml`
2. **base_url uses /v1 NOT /litellm/v1?** — `custom_providers[0].base_url`, `delegation.base_url`,
and ALL `auxiliary.*.base_url` MUST be `http://192.168.68.116/v1` (Rule 5). nginx `/litellm/`
has a 60s default timeout → 504 on any inference >60s; `/v1/` has 600s.
Check: `grep -n 'litellm/v1' ~/.hermes/config.yaml` (must return NOTHING)
3. **LITELLM_API_KEY valid?** — The key must be a real LiteLLM key (`sk-` + 64 hex, 67 chars).
Malformed values (e.g. `sk-_SWAl_Vu_…`, 47 chars) return 401 → DeepSeek fallback.
Verify: `curl -s -o /dev/null -w '%{http_code}' -H "Authorization: Bearer $LITELLM_API_KEY" http://192.168.68.116/v1/models` (must be 200)
4. **custom_providers aligned?** — `model: syslog-auto` (Rule 10), `api_mode: chat_completions`
(NOT `responses`). A wrong api_mode causes silent request failures.
One-line agent health check (run on the agent host):
```bash
# Use grep -v infisical to avoid matching the bash wrapper that contains the same string
PID=$(pgrep -f "python -m hermes_cli.main gateway run" | grep -v infisical | head -1)
cat /proc/$PID/environ | tr '\0' '\n' | grep ^LITELLM_API_KEY= | sed 's/=.*/<set>/'
curl -s -o /dev/null -w 'key_health: %{http_code}\n' -H "Authorization: Bearer $(cat /proc/$PID/environ | tr '\0' '\n' | grep ^LITELLM_API_KEY= | cut -d= -f2)" http://192.168.68.116/v1/models
```
### Rule 13: API Key Injection — Two Patterns (UPDATED 2026-07-16, WAL #1300)
Agents inject `LITELLM_API_KEY` via ONE of two mechanisms. Both are valid; the contract
requirement is that the key is a **valid LiteLLM virtual key** (HTTP 200 on /v1/models).
**Pattern A — systemd drop-in (Koby, Koonimo, and any agent without infisical wrapper):**
A systemd drop-in `/etc/systemd/system/hermes-gateway.service.d/litellm-key.conf` sets the key:
```ini
[Service]
Environment="LITELLM_API_KEY=sk-<VALID_KEY>"
```
The service unit `hermes-gateway.service` runs `python -m hermes_cli.main gateway run --replace`
directly (no infisical). Apply with `systemctl daemon-reload && systemctl restart hermes-gateway`.
- Koonimo (CT113/.114): service = `hermes-gateway.service`, drop-in has the key.
- Koby (CT111/.129): service = `hermes-gateway.service` (created 2026-07-16), ExecStart uses `--replace`
to win the lock against stray `hermes gateway restart` invocations. Key also in `/etc/environment`.
**Pattern B — infisical-gateway.sh wrapper (Mumuni):**
The wrapper sources `~/.hermes/.env` then exports `LITELLM_API_KEY="$<AGENT>_LITELLM_API_KEY"`.
See litellm-api-keys.prose.md § Machine Identity for Vault Writes for vault sync.
**⚠️ Vault empty-key guard:** If the vault stores the secret as an empty string,
the wrapper will inject an empty key and the gateway will silently get 401 errors
on all LiteLLM requests (triggering silent DeepSeek fallback). The `.env` fallback
is present but the vault takes precedence when the secret key exists (even if empty).
**Fix:** The wrapper MUST validate the key length after injection. If LITELLM_API_KEY
is empty or shorter than 20 chars, log a warning and either fail with a clear error
message or fall back to the `.env` value before starting the gateway.
**Verification (all agents):**
```bash
# Use grep -v infisical to avoid matching the bash wrapper that contains the same string
GP=$(pgrep -f "python -m hermes_cli.main gateway run" | grep -v infisical | head -1)
K=$(cat /proc/$GP/environ | tr '\0' '\n' | grep '^LITELLM_API_KEY=' | cut -d= -f2)
curl -s -o /dev/null -w '%{http_code}' -H "Authorization: Bearer $K" http://192.168.68.116/v1/models # must be 200
```
- `/etc/environment` is NO LONGER the canonical key source (stale values there caused 401s).
- Do NOT leave a hardcoded stale key in `/etc/environment` — it shadows the drop-in/wrapper.
## Execution
1. **Check current config** — Read the target agent's config.yaml
+124 -48
View File
@@ -5,8 +5,10 @@ version: 1.0.0
description: >
Enforces standardized API key configuration across all Hermes agents. Harness/LiteLLM
providers MUST use api_key_env indirection. External providers (DeepSeek, OpenAI,
Anthropic) may use hardcoded keys. Single source of truth: /etc/environment on each
agent host. Designed to make key rotation a one-step operation.
Anthropic) may use hardcoded keys. Single source of truth: Infisical vault
(project=agents, env=production) — injected at runtime via `infisical run --` wrapper.
/etc/environment is DEPRECATED for agent keys post-migration. Designed to make key
rotation a one-step vault operation.
author: Abiba (pi agent)
---
@@ -14,7 +16,7 @@ author: Abiba (pi agent)
## Rule (One Sentence)
**Any `api_key` pointing to a Syslog-hosted LiteLLM/harness provider MUST be replaced with `api_key_env: LITELLM_API_KEY` — hardcoded harness keys are forbidden.**
**All harness/litellm providers MUST use `api_key_env: LITELLM_API_KEY` with authenticated path `http://192.168.68.116/litellm/v1/responses` — hardcoded keys AND unauthenticated `/v1` direct access are both forbidden.**
## Scope
@@ -26,6 +28,39 @@ Applies to all Hermes agent configs across all hosts. Covers these config sectio
- `compression.api_key` (when `provider` is `harness` or contains `litellm`)
- `fallback_providers[].api_key` (when provider is harness)
## Architecture (2026-07-10)
Syslog is migrating away from **unauthenticated direct access** to the shared inference harness.
| Path | Auth | Status |
|------|------|--------|
| `http://192.168.68.116/v1` | None (direct) | ❌ **DEPRECATED** — being phased out |
| `http://192.168.68.116/litellm/v1/responses` | Bearer `sk-*` key | ✅ **CURRENT** — authenticated LiteLLM proxy |
All harness/litellm providers MUST use the authenticated `/litellm/v1/responses` path.
Any `base_url` pointing to bare `/v1` on 192.168.68.116 is a **migration violation**.
### 🔥 CRITICAL: Double-Path Bug (2026-07-10)
When `api_mode: responses` is set, Hermes **appends `/v1/responses`** to `base_url`.
If `base_url` already includes `/litellm/v1/responses`, the result is:
```
http://192.168.68.116/litellm/v1/responses/v1/responses → 404
```
**The `base_url` must end at `/v1` — never include `/responses`:**
```yaml
# ✅ CORRECT — Hermes appends /v1/responses for api_mode: responses
base_url: http://192.168.68.116/litellm/v1
# ❌ WRONG — produces double path
base_url: http://192.168.68.116/litellm/v1/responses
```
This applies to ALL sections using the harness provider: `custom_providers`, `delegation`, `auxiliary.*`.
## Exemptions
External providers are **explicitly exempt** and may use hardcoded keys:
@@ -37,36 +72,48 @@ External providers are **explicitly exempt** and may use hardcoded keys:
## Standard Pattern
> **Canonical vault process (2026-07-16):** see `litellm-api-keys` § Production Vault Access Process.
> All agents MUST use the `infisical-gateway.sh` wrapper (live vault injection). Hardcoded systemd
> drop-ins / config.yaml keys are DEPRECATED — they rot on rotation (root cause of the 2026-07-16 401 storm).
> 4/5 agents migrated; tanko (user jerome) pending.
```yaml
# ✅ CORRECT — all harness/litellm providers
# ✅ CORRECT — all harness/litellm providers (authenticated path, NO /responses suffix)
model:
provider: harness # or custom:litellm.sysloggh.net
base_url: http://192.168.68.116/v1
api_key_env: LITELLM_API_KEY # ← indirection
provider: harness
base_url: http://192.168.68.116/litellm/v1 # ← Hermes appends /v1/responses
api_key_env: LITELLM_API_KEY
custom_providers:
- name: harness
base_url: http://192.168.68.116/v1
api_key_env: LITELLM_API_KEY # ← indirection
api_mode: responses
base_url: http://192.168.68.116/litellm/v1 # ← NO /responses suffix!
api_key_env: LITELLM_API_KEY
auxiliary:
compression:
provider: harness
api_key_env: LITELLM_API_KEY # ← indirection
base_url: http://192.168.68.116/litellm/v1 # ← NO /responses suffix!
api_key_env: LITELLM_API_KEY
# ✅ ALSO CORRECT — external providers
fallback_providers:
- provider: deepseek
base_url: https://api.deepseek.com
api_key: sk-b7d9... # ← hardcoded OK (external)
api_key_env: DEEPSEEK_API_KEY # ← also OK if set in /etc/environment
api_key_env: DEEPSEEK_API_KEY # ← also OK if set in environment (vault or /etc/environment)
```
```yaml
# ❌ FORBIDDEN — hardcoded harness/litellm key
# ❌ FORBIDDEN — hardcoded key (top) OR unauthenticated path (bottom)
model:
provider: harness
api_key: sk-Flc62smlegyMEaSo1ka8JA # ← RULE VIOLATION
api_key: sk-Flc62smlegyMEaSo1ka8JA # ← RULE VIOLATION: hardcoded key
model:
provider: harness
base_url: http://192.168.68.116/v1 # ← RULE VIOLATION: unauthenticated path
api_key_env: LITELLM_API_KEY
```
## Detection Query
@@ -75,13 +122,18 @@ Run on any Hermes host to detect violations:
```bash
# 1. Check config.yaml for hardcoded harness keys
grep -rn 'api_key: sk-' /home/jerome/.hermes/ \
grep -rn 'api_key: sk-' /root/.hermes/ \
--include='config.yaml' \
| grep -v 'deepseek\|openai\|anthropic\|DEEPSEEK'
# 1b. Check for double-path bug: base_url ending with /responses
# (Hermes appends /v1/responses when api_mode=responses, so base_url must end at /v1)
grep -rn 'litellm/v1/responses' /root/.hermes/config.yaml
# ANY output here = WRONG. Must be 'litellm/v1' without /responses suffix.
# 2. Check systemd drop-ins for master key leaks (2026-07-05: Tanko had this)
grep -rn 'LITELLM_API_KEY' /home/jerome/.config/systemd/user/ 2>/dev/null
grep -rn 'LITELLM_API_KEY=sk-litellm-7f96080d' /home/jerome/.config/systemd/ 2>/dev/null
grep -rn 'LITELLM_API_KEY' /root/.config/systemd/user/ 2>/dev/null
grep -rn 'LITELLM_API_KEY=sk-litellm-7f96080d' /root/.config/systemd/ 2>/dev/null
# 3. Verify running process env matches dedicated key
cat /proc/$(cat /home/jerome/.hermes/gateway.pid | python3 -c "import sys,json; print(json.load(sys.stdin)['pid'])")/environ \
@@ -92,26 +144,32 @@ If any output from step 2 — **critical violation** (master key leaked). Fix im
## Rotation Procedure
With this standard enforced, key rotation is one step:
With this standard enforced, key rotation is one vault update:
```bash
# 1. Generate new key in LiteLLM
# 2. Update /etc/environment on agent host
ssh root@<host> "sed -i 's/LITELLM_API_KEY=.*/LITELLM_API_KEY=sk-NEW_KEY/' /etc/environment"
# 3. Restart agent gateway
ssh root@<host> "pkill -f 'hermes_cli.main gateway run'; sleep 2; nohup ... &"
# 1. Generate new key in LiteLLM: POST /key/generate with agent alias
# 2. Update Infisical vault secret
infisical secrets set LITELLM_API_KEY=sk-NEW_KEY \
--project=agents --env=production
# 3. Restart agent gateway (key auto-injected via infisical run -- wrapper)
ssh root@<host> "systemctl restart hermes-gateway"
# 4. Verify
curl -s -H "Authorization: Bearer sk-NEW_KEY" http://192.168.68.116/v1/models
curl -s -H "Authorization: Bearer sk-NEW_KEY" http://192.168.68.116/litellm/v1/models
```
**Done.** No config file changes needed. The agent picks up the new key on restart.
**Done.** No config file changes needed. No /etc/environment edits needed.
The agent picks up the new key via `infisical run --` at gateway startup.
> **Post-migration note**: /etc/environment is NO LONGER the key source.
> Strip all `LITELLM_API_KEY` lines from /etc/environment (comment out with `# [INFISICAL]`)
> and let the `infisical run --` wrapper inject the key at runtime.
## Key Longevity Policy (2026-07-04)
**Keys are permanent and use bare agent name aliases.**
- **Duration**: `null` — keys never expire. This is enforced by `default_key_generate_params` in `litellm_config.yaml`.
- **Alias convention**: bare agent name only (e.g., `tanko`, `mumuni`, `tdunna`, `baggy`). No dates, no versions. The alias IS the identity.
- **Alias convention**: bare agent name only (e.g., `tanko`, `mumuni`, `koby`, `koonimo`). No dates, no versions. The alias IS the identity.
- **Rotation triggers**: compromise, personnel departure, or quarterly security hygiene. NOT calendar-driven.
- **Max budget**: $100 per key (config default).
@@ -128,36 +186,54 @@ litellm_settings:
## Verified Agents (2026-07-05 update)
| Agent | CT | IP | LiteLLM Alias | Key | Status | Systemd Source | Last Verified |
|-------|-----|-----|---------------|-----|--------|----------------|---------------|
| Tanko | 112 | .122 | `tanko` | `sk-CggiHWlamQyShxWC3Hx6uw` | ✅ Fixed | User drop-in `env.conf` | 20:17 UTC Jul 5 |
| Mumuni | 114 | .123 | `mumuni` | `sk-XY2aUfvy2BIs6kp1ZPh6VA` | ⚠️ Unverified | `/etc/environment` | 23:00 EDT Jul 4 |
| Tdunna | 111 | ? | `tdunna` | `sk-6sbCNjz2T6lTVDBdlNHXsA` | ✅ Fixed | `/etc/environment` + drop-in | 23:30 UTC Jul 5 |
| Baggy | 113 | ? | `baggy` | `sk-krnw_zGBwvvL5b7l2t-s-A` | ✅ Fixed | `/etc/environment` | 23:30 UTC Jul 5 |
| Abiba | 100 | .65 | — | — | ✅ N/A (pi native) | — | 19:44 UTC Jul 5 |
| Agent | CT | IP | LiteLLM Alias | Key Source | Status | Gateway Wrapper | Last Verified |
|-------|-----|-----|---------------|------------|--------|-----------------|---------------|
| Tanko | 112 | .122 | `tanko` | Infisical vault | ✅ Fixed | `infisical run` | 20:17 UTC Jul 5 |
| Mumuni | 114 | .123 | `mumuni` | Infisical vault | ✅ Fixed | `infisical run` | 01:46 EDT Jul 10 |
| Koby | 111 | ? | `koby` | Infisical vault | ✅ Fixed | `infisical run` | 23:30 UTC Jul 5 |
| Koonimo | 113 | ? | `koonimo` | Infisical vault | ✅ Fixed | `infisical run` (migrated 2026-07-11) | 2026-07-11 |
| Abiba | 100 | .65 | `abiba-pi` | Infisical vault | ✅ N/A (pi native) | — | 19:44 UTC Jul 5 |
| Kagenz0 | 105 | ? | — | — | ❌ DOWN | — | 19:14 EDT Jul 4 |
### Systemd Service Pattern (2026-07-04 fix)
> **Note**: CT hostnames (tdunna, baggy) differ from agent identities (koby, koonimo).
> LiteLLM key aliases use agent identity, not CT hostname.
All Hermes agents use systemd to manage their gateway. Two issues were fixed:
### Migration Status: Authenticated Path
1. **Drop-in override**`/etc/systemd/system/hermes-gateway.service.d/litellm-key.conf` (or user equivalent) had hardcoded `LITELLM_API_KEY` that bypassed `/etc/environment`.
2. **Missing EnvironmentFile** — Services did not source `/etc/environment`.
| Agent | `/litellm/v1/responses` | Deprecated `/v1` | Status |
|-------|--------------------------|--------------------|--------|
| Mumuni | ✅ 5 sections | 0 | ✅ Authenticated |
| Tanko | ⚠️ No SSH access | — | Needs check |
| Koby | ⚠️ No route to host | — | Needs check |
| Koonimo | ⚠️ Connection timed out | — | Needs check |
**Correct pattern:**
### Systemd Service Pattern (2026-07-11 — vault migration)
All Hermes agents use systemd to manage their gateway. The gateway service is wrapped
with `infisical run --` to inject secrets at runtime.
**Correct pattern (post-migration):**
```ini
# In service file:
EnvironmentFile=/etc/environment
# Service file wraps gateway with Infisical:
[Service]
ExecStart=/usr/bin/infisical run --project=agents --env=production -- \
/usr/bin/hermes gateway run
# Drop-in only for overrides, NOT primary key storage.
# If a drop-in exists, it must match /etc/environment.
# /etc/environment is CLEAN — no LITELLM_API_KEY present
# (strip it and tag with # [INFISICAL] if present)
```
**Rotation procedure** (one step with this standard):
**Legacy pattern (deprecated — pre-migration only):**
```ini
# DO NOT USE post-migration:
EnvironmentFile=/etc/environment
# This pattern was replaced by infisical run -- wrapper
```
**Rotation procedure** (one vault operation with this standard):
1. Generate new key in LiteLLM: `curl /key/generate` with agent alias
2. Update `/etc/environment`: `sed -i 's/LITELLM_API_KEY=.*/LITELLM_API_KEY=sk-NEW/' /etc/environment`
3. Update drop-in (if exists): same sed on `litellm-key.conf`
4. Restart: `systemctl [--user] restart hermes-gateway`
2. Update Infisical vault: `infisical secrets set LITELLM_API_KEY=sk-NEW --project=agents --env=production`
3. Restart: `systemctl restart hermes-gateway` (key auto-injected via wrapper)
## Violation Response
@@ -167,7 +243,7 @@ EnvironmentFile=/etc/environment
4. **Restart** — gateway must restart to pick up env var
5. **Confirm** — test key against LiteLLM: `curl -H "Authorization: Bearer $KEY" .../v1/models` → 200
6. **Update** — bump the verified table above
7. **Use safe-mutate** — if the fix requires changing `/etc/environment` or restarting the gateway on a remote host, use `safe-mutate` to verify current state before mutating.
7. **Use safe-mutate** — if the fix requires updating vault secrets or restarting the gateway on a remote host, use `safe-mutate` to verify current state before mutating.
## Related Contracts
@@ -206,13 +282,13 @@ task config:
```yaml
auxiliary:
vision:
api_key: sk-CggiHWlamQyShxWC3Hx6uw # ← workaround
api_key: sk-<agent-key-from-vault> # ← workaround (get via: infisical secrets get LITELLM_API_KEY --project=agents --env=production --plain)
api_key_env: LITELLM_API_KEY
base_url: http://192.168.68.116/v1
model: gemma-4-12b
provider: harness
compression:
api_key: sk-CggiHWlamQyShxWC3Hx6uw # ← workaround
api_key: sk-<agent-key-from-vault> # ← workaround (same as above)
api_key_env: LITELLM_API_KEY
base_url: http://192.168.68.116/v1
model: gemma-4-12b
+2 -1
View File
@@ -28,7 +28,7 @@ connectivity recovery including end-to-end DM validation.
| Param | Type | Required | Default | Description |
|-------|------|----------|---------|-------------|
| `target` | string | yes | — | Agent name: `mumuni`, `tanko`, or `koby` |
| `target` | string | yes | — | Agent name: `mumuni`, `tanko`, `koby`, or `shumba` |
| `branch` | string | no | `master` | Git branch to pull (overridable for pinning) |
## Maintains
@@ -58,6 +58,7 @@ connectivity recovery including end-to-end DM validation.
| Mumuni | CT114 | — | 192.168.68.123 | /root/.hermes | root |
| Tanko | CT112 | amdpve | 192.168.68.122 | /home/jerome/.hermes | jerome |
| Koby | CT111 | amdpve | 192.168.68.129 | /root/.hermes | root |
| Shumba | — | — | 192.168.68.119 | /home/lucky/.hermes | lucky |
| Field | Value | Trust |
|-------|-------|-------|
+3 -2
View File
@@ -3,7 +3,7 @@ kind: function
name: hermes-zulip-restore
description: >
Restores Zulip connectivity for any Hermes agent (Mumuni CT114, Tanko CT112,
Koby CT111). Deploys the zulip-platform adapter to the correct bundled plugin
Koby CT111, Shumba on Lucky's mini PC). Deploys the zulip-platform adapter to the correct bundled plugin
path, verifies env credentials, restarts the gateway, and confirms Zulip
connects. Run this whenever a Hermes agent stops responding on Zulip or after
a fresh agent deployment.
@@ -23,7 +23,7 @@ gateway restart, and connection validation.
| Param | Type | Required | Default | Description |
|-------|------|----------|---------|-------------|
| `target` | string | yes | — | Agent name: `mumuni`, `tanko`, or `koby` |
| `target` | string | yes | — | Agent name: `mumuni`, `tanko`, `koby`, or `shumba` |
## Maintains
@@ -54,6 +54,7 @@ gateway restart, and connection validation.
| Mumuni | CT114 | — | 192.168.68.123 | /root/.hermes | root |
| Tanko | CT112 | amdpve | 192.168.68.122 | /home/jerome/.hermes | jerome |
| Koby | CT111 | amdpve | 192.168.68.129 | /root/.hermes | root |
| Shumba | — | — | 192.168.68.119 | /home/lucky/.hermes | lucky |
| Field | Value | Trust |
|-------|-------|-------|
+103
View File
@@ -0,0 +1,103 @@
---
name: inference-optimization
kind: responsibility
description: >
Optimizes the full Syslog inference stack — LiteLLM routing weights, GPU model
assignments, agent context management, and prompt caching — to reduce response
times to sub-15s average. All GPUs now at 128K context (stable ceiling).
id: 067NC6KP02RG60S50M40E30928
---
### Goal
Syslog inference response times reduced to sub-15s average by optimizing the full
stack: LiteLLM routing weights, GPU model assignments, Hermes agent context
management, and prompt caching — without sacrificing agent capability.
### Requires
- `inference-metrics`: current SpendLogs from CT116 LiteLLM Postgres — avg
request_duration_ms, prompt_tokens, completion_tokens, model_group breakdown,
cache_hit rate over the last 3 hours
- `agent-configs`: current config.yaml from each active Hermes agent (Mumuni
.123, any others on .129/.122) including compression, model, context_window,
prompt_caching, memory settings
- `gpu-health`: health check response from all 3 GPU backends (strix-moe .15:8080,
qwen .8:8080, gemma .110:8080)
### Maintains
The optimized inference stack configuration — every change is applied and
verified end-to-end. Postcondition: avg request_duration_ms ≤ 15000 for 90% of
inference calls.
#### liteLLM-routing
The syslog-auto routing weights, model-specific timeouts, RPM limits, and
model_list entries on CT116 `/opt/inference-harness/litellm_config.yaml`.
#### agent-compression
Each Hermes agent's `~/.hermes/config.yaml` compression, context_window,
prompt_caching, and model sections.
#### prompt-caching
LiteLLM cache configuration and llama.cpp `--cache-prompt` flag on GPU hosts.
#### verification
End-to-end latency measurements after changes applied — at least 3 test
inference calls per model path measuring ttft (time-to-first-token) and total
duration.
### Continuity
- input-driven
### Strategies
**Context is the root cause.** Every ~46K prompt token costs ~87s of
prefill time at 532 tok/s. Fix context first, routing second.
- **Route by task**: qwen for code/standard queries; gemma for
compression/auxiliary; strix-moe for compression tasks.
- **Compress aggressively**: threshold at 40% (not 65%) — a 128K window should
compact at 51K, not 85K. Target 15% tail (not 30%).
- **Cache everything repeated**: system prompts, skill docs, AGENTS.md — these
never change between turns. Single-digit cache hit rate is unacceptable.
- **Lower context ceiling**: 128K window is the stable ceiling for agent conversations.
GPUs reduced from 256K to 128K (2026-07-17). 128K window should compact at 85K (0.65 threshold). For larger contexts, route to external providers.
### Shape
- `self`: analyze metrics, compute optimal configs, apply changes, verify
- `delegates`:
- `apply-liteLLM`: update litellm_config.yaml and reload
- `apply-agent-config`: update hermes config.yaml per agent
- `verify-latency`: run test inference calls and measure response
### Execution
```prose
-- Phase 1: Analyze current state (already complete)
-- Phase 2: Apply LiteLLM routing optimization
call apply-liteLLM-routing
config_path: /opt/inference-harness/litellm_config.yaml
host: 192.168.68.116
-- Phase 3: Apply agent context compression optimization
call apply-agent-compression
agent: mumuni
host: 192.168.68.123
config_path: /root/.hermes/config.yaml
-- Phase 4: Enable llama.cpp prompt caching on GPU hosts
call enable-prompt-caching
hosts: [192.168.68.15, 192.168.68.8, 192.168.68.110]
-- Phase 5: Verify end-to-end latency
call verify-latency
host: 192.168.68.116
models: [syslog-auto, qwen3.6-27B-code, gemma-4-12b]
```
+10 -10
View File
@@ -62,19 +62,19 @@ description: >
| Resource | Auth Method | Credential Source | Status |
|----------|------------|-------------------|--------|
| Proxmox Cluster | PVE API Token | `monitoring@pve!mumuni=...` | ✅ |
| Proxmox Root | Password via API ticket | `root@pam:kakashi19` | ✅ |
| Proxmox Cluster | PVE API Token | Infisical vault (`PROXMOX_API_TOKEN`) | ✅ |
| Proxmox Root | Password via API ticket | Infisical vault (`PROXMOX_ROOT_PASSWORD`) | ✅ |
| docker-vm (.7) | SSH root | SSH key | ✅ |
| CT 116 (syslog-api) | SSH root | SSH key | ✅ |
| Tanko CT (.122) | SSH jerome | id_ed25519 | ✅ |
| Mumuni CT (.123) | SSH root | id_ed25519 | ✅ |
| Baggy CT (113) | SSH jerome | ❌ no key access |
| Netbird (.17) | SSH root | SSH key | ✅ |
| Gitea | API token | abiba-bot token | ✅ |
| Zulip | Bot API key | abiba-bot@chat.sysloggh.net | ✅ |
| Gitea | API token | Infisical vault (`GITEA_BOT_TOKEN`) | ✅ |
| Zulip | Bot API key | Infisical vault (`ZULIP_BOT_KEY`) | ✅ |
| RA-H OS | MCP bridge | port 3100 | ✅ |
| LiteLLM Admin | master key | `sk-litellm-7f96...` | ✅ VERIFY-BEFORE-USE |
| Grafana | admin password | `syslog-grafana-2026` | ✅ VERIFY-BEFORE-USE |
| LiteLLM Admin | master key | Infisical vault (`LITELLM_MASTER_KEY`) | ✅ VERIFY-BEFORE-USE |
| Grafana | admin password | Infisical vault (`GRAFANA_ADMIN_PASSWORD`) | ✅ VERIFY-BEFORE-USE |
> **VERIFY-BEFORE-USE**: Credentials, IPs, ports, and hostnames in this
> contract are live-state fields. Test them against the live system before
@@ -199,7 +199,7 @@ description: >
**Prometheus targets**:
- 192.168.68.8:9400 (RTX 3090 — qwen)
- 192.168.68.110:9400 (RTX 5070 — gemma)
- 192.168.68.15:9400 (Strix Halo — ornith)
- 192.168.68.15:9400 (Strix Halo — qwen3.6-35B-udq4)
- 192.168.68.24:9401 (Router metrics exporter)
- harness-litellm:4000 (LiteLLM health)
@@ -550,13 +550,13 @@ curl -s http://192.168.68.116/health/unified | jq .status
curl -s http://192.168.68.24:9100/gpu-data | jq .summary
# Grafana status
curl -s http://admin:syslog-grafana-2026@192.168.68.116:3001/api/health
curl -s http://admin:$(infisical secrets get GRAFANA_ADMIN_PASSWORD --project=infrastructure --env=production --plain)@192.168.68.116:3001/api/health
# Prometheus targets
curl -s http://192.168.68.116:9090/api/v1/targets | jq '.data.activeTargets[] | {job: .labels.job, health: .health}'
# LiteLLM key check
curl -s -H "Authorization: Bearer sk-litellm-7f96080dd99b15c36bd4b333b58a6796" \
curl -s -H "Authorization: Bearer $(infisical secrets get LITELLM_MASTER_KEY --project=infrastructure --env=production --plain)" \
http://192.168.68.116/litellm/key/list | jq '.keys[] | {alias: .key_alias, models: .models}'
# Storage check
@@ -586,7 +586,7 @@ ssh root@192.168.68.110 "systemctl restart llama-server"
| 108 | media | storepve | — | Media | ❌ |
| 109 | docker-vm | storepve | .7 | Docker host | ❌ |
| 110 | gitea | minipve | — | Git | ❌ |
| 111 | tdunna | amdpve | — | ? | |
| 111 | tdunna | amdpve | .129 | Hermes agent | |
| 112 | tanko | amdpve | .122 | Hermes agent | ✅ |
| 113 | baggy | amdpve | ? | Hermes agent | ✅ |
| 114 | mumuni | minipve | .123 | Hermes agent | ✅ |
+200
View File
@@ -0,0 +1,200 @@
---
kind: responsibility
name: infrastructure-maintenance
description: >
Weekly system-level maintenance for the Syslog inference fleet: OS package
updates on the primary host, Docker image pulls for LiteLLM/SearXNG and other
running containers, container restarts with health verification, post-update
verification of every critical service (LiteLLM proxy, SearXNG, Zulip, Gitea,
PM2 processes, Hermes gateways), and rollback on failure. Consolidates the
raw shell scripts that previously did this piecemeal. This contract owns the
HOST-LEVEL weekly maintenance loop on the primary host plus Docker image
pulls ONLY for .116 and .7, while infrastructure-update owns the FULL-FLEET
cluster-wide wave (apt across the full PVE cluster + CTs/VMs AND its Docker
image Wave 3 across all stacks). Runs Sunday 2am ET. Owner:
ops (firstmate secondmate). Blast radius: an unverified image pull can break
LiteLLM (all agents lose inference) or SearXNG (search-stack down); a bad apt
upgrade can leave the host in a half-upgraded state. Pre-update backup check
and rollback are mandatory for this reason.
agent: ops
triggers:
- weekly (Sunday 02:00 ET) via cron
- on demand when ops/abiba triggers "infra maintenance"
version: 1.0.0
---
## Maintains
- maintenance-status: { phase: idle|preflight|apt|images|restarts|verify|rollback|done|failed, host, step, result, timestamp }
- image-baseline: { service, current_tag, pulled_tag, digest, updated_at } — last known-good image per container
- apt-state: { upgradable_before, upgradable_after, held_broken, kernel_reboot_required }
- health-baseline: snapshot of critical-service health captured pre-update (used for regression check post-update)
- rollback-snapshot: { backup_path, configs, image_digests, timestamp } — restore point created in preflight
- maintenance-history: array of past runs with phase results and any escalations
## Scope
Primary host is the maintenance host where apt updates apply. Docker image pulls
span the two Docker ecosystems that run critical services. infrastructure-update
runs the full-fleet cluster-wide wave (including its Wave 3 Docker pulls across
all stacks/hosts); this contract runs a narrower host-level weekly pull limited
to .116 and .7. Topology, CT IDs, and IPs are live-state fields — verify against
`infrastructure-control.prose.md` (the source of truth) and the live system
before mutating.
| Host | IP | Role | Trust |
|------|----|------|-------|
| CT 116 (syslog-api) | 192.168.68.116 | LiteLLM proxy + Grafana + Prometheus (inference harness) | ⚠️ VERIFY-BEFORE-USE |
| VM 109 (docker-vm) | 192.168.68.7 | SearXNG + Firecrawl + home stack (Docker host) | ⚠️ VERIFY-BEFORE-USE |
| CT 117 (zulip) | 192.168.68.19 | Zulip (storepve bridge IP .19) | ⚠️ VERIFY-BEFORE-USE |
| Gitea | https://git.sysloggh.net | Prose-contracts + agent configs source control | ⚠️ VERIFY-BEFORE-USE |
| CT 100 (abiba/pi) | 192.168.68.24 | PM2 processes (pi agent harness) | ⚠️ VERIFY-BEFORE-USE |
> "Primary host" for the apt phase is the host the ops agent runs maintenance
> from. Confirm which host that is against infrastructure-control before
> running; do not assume. If the ops agent is containerized/CT-based, apt runs
> inside that CT.
## Requires
- SSH/exec access to CT 116 (.116) and VM 109 (.7) for Docker operations
- `apt`, `docker`, `docker compose` available on target hosts
- LiteLLM master key available (Infisical vault, `LITELLM_API_KEY`) for health verification
- `infrastructure-monitoring` run completed within the last 30 minutes — provides the pre-update health baseline used by the regression check
- Writable backup directory `/tmp/infra-maintenance-backup-<date>/` on each mutated host
- Proxmox snapshot of the primary host available (or confirmed not required) before apt phase
## Continuity
- Self-driven: weekly cron `0 2 * * 0` (Sunday 02:00 ET)
- Also wakes on: explicit "infra maintenance" trigger from ops/abiba
- Depends on `infrastructure-monitoring` for the pre-update health baseline — do not run if the last monitoring run is stale (>30 min) or RED; abort and escalate instead
## Execution
### Phase 0 — Preflight (snapshot/backup check + health baseline)
1. **Capture health baseline** — run the `infrastructure-monitoring` postcondition checks (LiteLLM, Zulip, Gitea, SearXNG, Proxmox API) and record results as `health-baseline`. If any critical service is already down, **abort**: maintenance must not run on a degraded fleet.
2. **Backup check** — confirm a Proxmox snapshot of the primary host exists OR `/tmp/infra-maintenance-backup-<date>/` was created this run. Snapshot critical config files into the backup dir:
- `/opt/inference-harness/docker-compose.yml`, `/opt/inference-harness/litellm_config.yaml` (CT 116)
- `/opt/search-stack/searxng/docker-compose.yml`, `/opt/search-stack/firecrawl-source/docker-compose.yaml` (VM 109)
3. **Record image baseline**`docker inspect --format '{{.Image}} {{.Config.Image}}' <container>` for every running container on .116 and .7; store digests in `image-baseline` so rollback can restore them.
4. **Disk check**`df -h` on each mutated host; abort if free space <20% (apt upgrade + image pulls need headroom).
### Phase 1 — OS package updates (primary host)
```bash
# On the primary host only (VERIFY host against infrastructure-control first)
apt update
apt upgrade -y
```
- Capture `apt list --upgradable` before and after → store in `apt-state`.
- If apt reports held/broken packages (`apt-get -s upgrade | grep -i broken`, or non-zero exit), **stop** — do not force. Record `held_broken` and go to rollback/escalate.
- If `/var/run/reboot-required` exists after upgrade, flag `kernel_reboot_required: true` in `apt-state` but **do not reboot automatically** — that's a separate coordinated action (see infra-update Wave 4). Note it in the report.
### Phase 2 — Docker image pulls
Pull latest stable tags for every running container. Do NOT pin to `:main`/`:nightly` — use stable tags where the compose file specifies them; otherwise `latest`.
```bash
# CT 116 (.116) — inference harness
cd /opt/inference-harness && docker compose pull
# VM 109 (.7) — search + home stacks
cd /opt/search-stack/searxng && docker compose pull
cd /opt/search-stack/firecrawl-source && docker compose pull
# any other running stacks on .7 (home stack, audiobookshelf) — pull per their compose files
```
- LiteLLM and SearXNG are the two explicitly required pulls; "any other running containers" means every stack with a compose file on .116 and .7.
- Record pulled tag + digest per service in `image-baseline`.
### Phase 3 — Container restarts with health verification
Restart one stack at a time, verify health before moving to the next. Do not restart everything at once — a failure mid-wave must leave the rest running.
```bash
# CT 116
cd /opt/inference-harness && docker compose up -d
# VM 109
cd /opt/search-stack/searxng && docker compose up -d
cd /opt/search-stack/firecrawl-source && docker compose up -d
```
After each stack comes up, wait for health (max 120s):
- `docker ps` shows the container `Up` (and `healthy` if a healthcheck is defined)
- Service-specific probe passes (see Phase 4 probes)
If a stack fails to come up within 120s, **stop the wave** and go to rollback for that stack only; do not proceed to the next.
### Phase 4 — Post-update service verification
After ALL updates (apt + images + restarts), verify every critical service is back up and matches the pre-update baseline. This is the regression gate.
| Service | Probe | Expect |
|---------|-------|--------|
| LiteLLM proxy | `curl -sf http://192.168.68.116/litellm/v1/models` | 200 OK, models returned |
| LiteLLM MCP gateway | `curl -sf http://192.168.68.116:4000/mcp-rest/tools/list -H "Authorization: Bearer $MASTER_KEY"` | 90 tools (23 RA-H OS + 67 GitHub) |
| SearXNG | `curl -sf http://192.168.68.7:8888` | 200 OK |
| Zulip | `curl -sf https://chat.sysloggh.net/api/v1/server_settings` | 200 OK |
| Gitea | `curl -sf https://git.sysloggh.net/api/v1/version` | 200 OK |
| PM2 processes | `pm2 jlist` (CT 100) | all pi-agent processes `online` |
| Hermes gateways | SSH to Mumuni CT 114, Tanko CT 112; `systemctl is-active hermes-gateway` | `active` for each |
Regression check: every service that was GREEN in `health-baseline` must still be GREEN. A service that was already RED (and caused a preflight abort) is excluded — but Phase 0 should have aborted before we got here.
## Rollback Protocol
If ANY service in Phase 4 fails to come back up (or regresses vs baseline):
1. **Image rollback** — for the failing stack, restore the previous image:
```bash
# Restore from recorded image-baseline digest
docker compose down
# Pin the service image to the recorded digest in compose, then recreate
# image: <name>@sha256:<previous_digest>
docker compose pull && docker compose up -d
```
2. **APT rollback** — restore the primary host from the Proxmox snapshot taken/confirmed in Phase 0. If no snapshot, `apt install <pkg>=<old_version>` per package using apt history (`/var/log/apt/history.log`).
3. **Config rollback** — restore configs from `/tmp/infra-maintenance-backup-<date>/`.
4. **Re-verify** — re-run the Phase 4 probes on the rolled-back service. If still failing, escalate (do not loop — circuit breaker below).
5. **Escalate** — send a Zulip DM to abiba + mumuni with: failing service, phase, baseline vs current, rollback actions taken, backup path.
## Circuit Breaker
- `max_retries: 2` per failing phase — after 2 rollback attempts on the same service, stop and escalate.
- `window: 7200` seconds — no more than 2 retries within a 2-hour window.
- `trip_action: escalate_to_fatal` — when tripped, escalate to fatal (abiba + mumuni + kwame) and pause; a human must clear before the next scheduled run.
## Report
After completion (or on abort), emit a receipt (JSON) to `~/.hermes/runs/infrastructure-maintenance/` and send a Zulip DM summary:
```
🛠 Infrastructure Maintenance — YYYY-MM-DD
Phase: apt | images | restarts | verify | rollback
Primary host: <host>
APT: <N> packages upgraded, <M> held/broken, kernel_reboot_required=<bool>
Images pulled: LiteLLM <tag>, SearXNG <tag>, <others>
Services: all GREEN | <service> FAILED (rolled back)
Baseline regression: none | <details>
Backup: /tmp/infra-maintenance-backup-YYYYMMDD/
Escalation: none | warning | critical | fatal
```
## Verification Postconditions
- All critical services running after update (Phase 4 all GREEN)
- No regressions from pre-update health baseline (Phase 0 baseline)
- Docker containers on latest stable tags (`image-baseline.pulled_tag` recorded)
- APT packages up to date with no held broken packages (`apt-state.held_broken == 0`)
## Related Contracts
- `infrastructure-update.prose.md` — owns the full-fleet cluster-wide wave INCLUDING its Wave 3 Docker image updates across all stacks (SearXNG, Firecrawl, Inference Harness on .116, home stack, audiobookshelf); infrastructure-maintenance is a deliberately narrower host-level weekly pull scoped to .116 and .7.
- `infrastructure-monitoring.prose.md` — provides the pre-update health baseline (depends_on).
- `infrastructure-control.prose.md` — topology source of truth (CT IDs, IPs, hostnames).
- `litellm-health.prose.md` — LiteLLM probe details.
- `proxmox-monitor.prose.md` — Docker stats + monitoring stack health.
+57 -11
View File
@@ -11,7 +11,7 @@ triggers:
- on "infra update" command
- weekly (Sunday 03:00 EDT) via cron
- on security advisory relay from Mumuni
version: 1.1.0
version: 1.2.0
---
## Maintains
@@ -27,11 +27,12 @@ Before ANY update wave:
2. ✅ All critical VMs/CTs running (VM 109 docker-vm, CT 116 syslog-api, CT 117 zulip, CT 106 ra-h-os)
3. ✅ GPU bare-metal hosts reachable: .8 (RTX 3090), .110 (RTX 5070), .15 (Strix Halo)
4. ✅ Docker healthy on VM 109 (.7), CT 116 (.116)
5. ✅ LiteLLM health check passing
6. ✅ Zulip server reachable
6. ✅ GPU fleet healthy (all 3 GPUs: RTX 3090, RTX 5070, RX 7600)
7. ✅ Disk >20% free on all nodes
8. 📋 Snapshot critical configs (LiteLLM, nginx, docker-compose files)
5. ✅ LiteLLM health check passing (port 4000, /mcp-rest/tools/list with master key)
6. ✅ LiteLLM MCP gateway serving RA-H OS tools (90 tools)
7. ✅ Zulip server reachable
8. ✅ GPU fleet healthy (all 3 GPUs: RTX 3090, RTX 5070, RX 7600)
9. ✅ Disk >20% free on all nodes
10. 📋 Snapshot critical configs (LiteLLM, nginx, docker-compose files)
## Wave 1: Storage & Infra Nodes (lowest impact)
@@ -65,7 +66,8 @@ Before ANY update wave:
**Verify after Wave 2:**
- All VMs/CTs running: check via Proxmox API
- LiteLLM healthy: `curl localhost:4000/health/liveliness` (via CT 116)
- GPU servers responding: check :8080 on VM 101, VM 103; check ornith via router (http://192.168.68.116/health/unified — .15:8080 is firewalled to .116 only)
- LiteLLM MCP tools: `curl localhost:4000/mcp-rest/tools/list -H "Authorization: Bearer $MASTER_KEY"` → 90 tools
- GPU servers responding: check :8080 on VM 101, VM 103; check strix-moe via router (http://192.168.68.116/health/unified — .15:8080 is firewalled to .116 only)
- Zulip agents connected: check Mumuni/Tanko gateway state
- Abiba PM2 processes online: `pm2 status`
@@ -83,6 +85,7 @@ Before ANY update wave:
**Verify after Wave 3:**
- All containers healthy: `docker ps` on each host
- End-to-end inference test: `curl localhost:4000/v1/chat/completions` (via CT 116) with syslog-auto
- MCP integration test: `curl localhost:4000/mcp-rest/tools/list -H "Authorization: Bearer $MASTER_KEY"` → 90 tools (23 RA-H OS + 67 GitHub)
- Zulip test: send test message to #agent-hub
- Dashboard loading: `curl localhost:3001/` (via CT 116)
- Firecrawl test: `curl :3002/`
@@ -112,8 +115,8 @@ If ANY verification fails:
Before Wave 1, snapshot these files:
```
/opt/inference-harness/docker-compose.yml (CT 116 .116)
/opt/inference-harness/litellm_config.yaml (CT 116 .116)
/opt/inference-harness/docker-compose.yml (CT 116 .116) ⚡ contains MCP_SERVER env vars
/opt/inference-harness/litellm_config.yaml (CT 116 .116) ⚡ contains mcp_servers.ra_h_os
/opt/monitoring/prometheus.yml (CT 116 .116)
/etc/nginx/nginx.conf (harness-nginx on CT 116)
/opt/search-stack/firecrawl-source/docker-compose.yaml (VM 109 .7)
@@ -121,11 +124,53 @@ Before Wave 1, snapshot these files:
/opt/home_stack/docker-compose.yml (VM 109 .7)
/opt/audiobookshelf/docker-compose.yml (VM 109 .7)
/root/.pi/agent/extensions/config.yaml (CT 100 .24)
/etc/systemd/system/ornith-server.service (amdpve .15)
/etc/systemd/system/ornith-server.service (amdpve .15 — strix-moe)
/etc/systemd/system/llama-server.service (VM 101 .8, VM 103 .110)
# Hermes agent configs (key enforcement — 2026-07-10)
/root/.hermes/config.yaml (Mumuni CT 114, Tanko CT 112, etc.)
/root/.config/systemd/user/hermes-gateway.service (Mumuni CT 114 — EnvironmentFile fixed)
/etc/environment (Mumuni CT 114 — LITELLM_API_KEY)
```
Run: `mkdir -p /tmp/infra-update-backup-$(date +%Y%m%d) && rsync -av ...`
## MCP Gateway (2026-07-10)
LiteLLM CT 116 now serves as an authenticated MCP gateway for RA-H OS tools.
### Configuration
**litellm_config.yaml** (`/opt/inference-harness/litellm_config.yaml`):
```yaml
mcp_servers:
ra_h_os:
url: "http://192.168.68.65:3100/mcp"
transport: "http"
auth_type: "none"
```
**docker-compose.yml** env vars:
```yaml
- MCP_SERVER_RAHOS_URL=http://192.168.68.65:3100/mcp
- MCP_SERVER_RAHOS_TRANSPORT=http
```
### Access
| Key | MCP Access |
|-----|-----------|
| Master key | ✅ Full — 90 tools (vault-injected) |
| Agent keys (mumuni, tanko, etc.) | ❌ Per-key grants not supported in v1.90.0-rc.1 |
### Known Limitations
- Per-key MCP server grants not functional — only master key has access
- Responses API (`/v1/responses`) with MCP tools broken on llama.cpp backends
- HTTP 307 redirect on `/mcp` → use `/mcp/` (trailing slash) or `/mcp-rest/` endpoints
- `api_mode: responses` in Hermes appends `/v1/responses` to base_url → **base_url must end at `/v1`, never `/responses`** (double-path bug)
### Migration Path
When LiteLLM is upgraded to a version supporting per-key MCP grants:
1. Grant agent keys `mcp_servers: ["ra_h_os"]`
2. Update Hermes `mcp_servers.ra-h-os.url` from `http://192.168.68.65:3100/mcp``http://192.168.68.116:4000/mcp/`
3. Add `headers: {x-litellm-api-key: "Bearer $LITELLM_API_KEY"}` to MCP config
## Security-Specific Updates
@@ -144,6 +189,7 @@ Run: `mkdir -p /tmp/infra-update-backup-$(date +%Y%m%d) && rsync -av ...`
- [ ] LiteLLM inference passing (syslog-auto test)
- [ ] Zulip server + all 3 agents connected
- [ ] GPU fleet at full capacity (3/3)
- [ ] LiteLLM MCP gateway healthy (90 tools via master key)
- [ ] Zero security CVEs remaining
- [ ] <10 min total downtime per service
+238 -10
View File
@@ -8,10 +8,32 @@ description: >
Ensures agents never use the master key directly. Rotation is event-driven,
not calendar-driven — rotate only on compromise, personnel change, or
periodic security hygiene (quarterly/annually).
UPDATED 2026-07-12: Keys are stored in Infisical vault (project=agents, env=production)
BUT each agent host MUST keep a local .env fallback. Infisical service tokens can
expire/404. The .env fallback prevents agents from running without keys.
Tanko incident: token 404 → gateway had no LITELLM_API_KEY for hours.
UPDATED 2026-07-16: Vault is SYNCED (session-13 keys written to vault via abiba service
token, all validate 200). Koby/Koonimo migrated from hardcoded drop-ins to the
infisical-gateway.sh wrapper (live vault injection). 4/5 agents now vault-backed.
Canonical process: see § Production Vault Access Process. Tanko (user jerome) pending.
Abiba's key is now a proper agent key (NOT the master key — stale note removed).
UPDATED 2026-07-17: FLEET-WIDE STANDARDIZATION. All 4 agents (Mumuni, Tanko, Koby, Koonimo)
standardized on a single pattern: systemd drop-in (ExecStart= reset + wrapper path) →
infisical-gateway.sh while-true loop → /usr/bin/infisical run --token → bash -c key
injection → .env fallback → exec python. Systemd drop-ins are IMMUNE to hermes gateway
install which overwrites the unit file ExecStart. Infisical CLI updated to 0.43.109 on
all agents (was 0.38.0). Service token st.8e848433 shared across fleet (st.353699cd
for tanko was deleted). .env fallback on every agent protects against token loss.
Critical lessons: (1) NEVER use shell variables inside single-quoted bash -c in wrappers
— hardcode absolute paths. (2) Drop-ins override unit file ExecStart permanently.
(3) Capture /proc/<pid>/environ before gateway restarts to preserve running env set.
Current key inventory and agent list: see gpu-fleet.prose.md § Agent Keys.
Source of truth for LiteLLM config: /opt/inference-harness/litellm_config.yaml
on CT 116. Last verified: 2026-07-09.
on CT 116. Last verified: 2026-07-17.
---
## Parameters
@@ -19,7 +41,10 @@ description: >
- agent_name: string — The agent to manage keys for (e.g., "tanko", "mumuni")
- action: "create" | "rotate" | "verify" | "list" — What to do (default: "create")
- litellm_host: string — LiteLLM admin endpoint (default: "192.168.68.116:4000")
- master_key: string — LiteLLM master key (default from environment)
- master_key: string — LiteLLM master key (default from Infisical vault: project=infrastructure, env=production, secret=LITELLM_MASTER_KEY)
- vault_url: string — Infisical vault URL (default: "https://vault.sysloggh.net")
- vault_project: string — Infisical project slug (default: "infrastructure")
- vault_env: string — Infisical environment (default: "production")
- agent_host: string — Agent's IP for SSH (default: resolved from infra)
- agent_user: string — SSH user (default: "jerome")
@@ -30,28 +55,231 @@ description: >
- key_prefix: string — First 10 chars of the new key (for identification)
- previous_key_alias: string | null — Previous key alias if rotating
- litellm_response: object — Raw response from LiteLLM /key/generate
- agent_config_updated: boolean — Whether /etc/environment was updated
- vault_updated: boolean — Whether Infisical vault secret was updated
- agent_config_updated: boolean — Legacy: whether /etc/environment was updated (deprecated, always false post-migration)
- verification: { status: string, detail: string } — Final health check
## Execution
1. **Authenticate**Verify master_key works against LiteLLM /key/list
1. **Authenticate**Retrieve master key from Infisical vault via `infisical export --project=<vault_project> --env=<vault_env>`, verify against LiteLLM /key/list
2. **Check existing keys** — List all keys, find any with agent_name alias
3. **If action == "list"**: Return all keys with their aliases and spend
4. **If action == "create"**:
- Generate new key with key_alias: "{agent_name}" (e.g., "tanko" — bare name, no date)
- Set metadata: { "agent": "{agent_name}", "purpose": "agent-inference" }
- Duration is null (permanent) — inherited from litellm default_key_generate_params
- Set models: ["syslog-auto", "qwen3.6-27B-code", "gemma-4-12b", "ornith-1.0-35b"]
- Note: qwen3.6-35B-A3B removed from fleet (was never deployed on any GPU)
- Set models: ["syslog-auto", "qwen3.6-27B-code", "gemma-4-12b", "strix-moe", "gpu-dense", "gpu-light", "qwen3.6-35B-udq4"]
- Note: `ornith-1.0-35b` is NOT a valid LiteLLM model name (use `strix-moe`, the stable alias). qwen3.6-35B-A3B removed from fleet (was never deployed).
- Return the new key
5. **If action == "rotate"**:
- Generate new key with same alias (LiteLLM replaces the old key)
- SSH to agent_host, update /etc/environment LITELLM_API_KEY
- Restart agent gateway (hermes gateway restart for Hermes agents)
- Update secret in Infisical vault: `infisical secrets set LITELLM_API_KEY=<new_key> --project=<vault_project> --env=<vault_env>`
- Restart agent gateway (Hermes: `systemctl restart hermes-gateway`; pi: restart PM2 process)
The gateway automatically picks up the new key via `infisical run --` wrapper
- Verify: curl test against /v1/models with new key
- Rotation policy: on-demand only (compromise, departure, quarterly hygiene)
- Note: /etc/environment is NO LONGER used for LiteLLM keys. Agents inject keys at runtime via vault wrapper.
6. **If action == "verify"**:
- SSH to agent, read /etc/environment
- Retrieve key from Infisical vault: `infisical secrets get LITELLM_API_KEY --project=<vault_project> --env=<vault_env>`
- Test the key against LiteLLM /v1/models
- Confirm key alias matches agent_name in LiteLLM key list
- Verify agent gateway uses vault wrapper: `cat /proc/<pid>/cmdline` shows `infisical run`
## Production Vault Access Process (canonical, 2026-07-17)
The non-fail approach to agentic vault access. Deployed on all 4 Hermes agents
(Mumuni, Tanko, Koby, Koonimo) as of 2026-07-17. Abiba (pi) uses a similar pattern
through its agent wrapper.
### The canonical pattern
1. **infisical CLI** installed on the host at `/usr/bin/infisical` (v0.43.109+, from
artifacts-cli.infisical.com apt repo). Update procedure:
```bash
curl -1sLf 'https://artifacts-cli.infisical.com/setup.deb.sh' | sudo -E bash
sudo apt-get update && sudo apt-get install -y infisical
# Remove stale old binary if present
rm -f /usr/local/bin/infisical /bin/infisical
```
Wrappers use absolute path `/usr/bin/infisical run`. Never rely on PATH resolution.
2. **Service token** (Infisical Machine Identity, `st.…`) stored at `~/.infisical-token`
(`chmod 600`). Current: shared `st.8e848433…` (abiba, READ+WRITE on agents project).
Tanko's `st.353699cd…` (tanko-agent) was deleted — reverted to shared token.
Proper: one machine identity per agent (create in Infisical UI → Project Settings →
Machine Identities).
3. **`infisical-gateway.sh` wrapper** at `~/.hermes/infisical-gateway.sh` (`chmod 700`):
```bash
#!/bin/bash
export INFISICAL_API_URL="https://vault.sysloggh.net"
TOKEN=$(cat $HOME/.infisical-token)
LOG=$HOME/.hermes/logs/gateway.log; mkdir -p $HOME/.hermes/logs
while true; do
echo "[$(date -Iseconds)] Starting gateway with Infisical injection..." >> $LOG
/usr/bin/infisical run --token="$TOKEN" \
--projectId=322fceab-39da-4854-a55a-568e76c0f13f \
--env=prod --domain=https://vault.sysloggh.net -- bash -c '
. $HOME/.hermes/.env 2>/dev/null # [FALLBACK Rule 3]
export LITELLM_API_KEY="${<AGENT>_LITELLM_API_KEY}"
export ZULIP_API_KEY="${<AGENT>_ZULIP_API_KEY}"
export ZULIP_SITE="https://chat.sysloggh.net"
export ZULIP_EMAIL="<agent>-bot@chat.sysloggh.net"
export SEARXNG_URL="http://192.168.68.7:8888"
# ⚠️ HARDCODE the full venv path. NEVER use $VENV inside single quotes.
exec /root/.hermes/hermes-agent/venv/bin/python -m hermes_cli.main gateway run
' >> $LOG 2>&1
EXIT_CODE=$?
echo "[$(date -Iseconds)] Gateway exited with code $EXIT_CODE — restarting in 5s..." >> $LOG
sleep 5
done
```
**CRITICAL: VENV PATH.** The inner `bash -c '...'` uses single quotes. Shell
variables set in the outer wrapper are NOT expanded inside single quotes.
`$VENV/bin/python` resolves to `/bin/python` (file not found). Always hardcode
the absolute path to the venv python binary.
4. **Agent key in vault** as `<AGENT>_LITELLM_API_KEY` and `<AGENT>_ZULIP_API_KEY`.
Vault = source of truth for ALL platform credentials.
5. **`.env` fallback** at `~/.hermes/.env` (`chmod 600`) with agent-specific keys —
safety net for vault outage or token revocation. Must be kept in sync on rotation.
Example:
```bash
MUMUNI_LITELLM_API_KEY=sk-OzuWsoX22Hmb3Ps3JY01gw
MUMUNI_ZULIP_API_KEY=H8dY6V7aHmWNcfgNtJaDBPZ1dGWn0Ttt
```
6. **systemd drop-in** at `~/.config/systemd/user/hermes-gateway.service.d/50-vault-wrapper.conf`:
```ini
[Service]
ExecStart=
ExecStart=/root/.hermes/infisical-gateway.sh
```
The `ExecStart=` (empty reset) clears any ExecStart from the main unit file,
then the second `ExecStart=` sets the wrapper. This drop-in **survives unit file
regeneration** by `hermes gateway install` — the drop-in always wins.
**Why a drop-in instead of editing the unit file:** `hermes gateway install`
(called during Hermes updates and some self-heal operations) regenerates the
systemd unit file with `ExecStart=/path/to/python -m hermes_cli.main gateway run`.
Editing the unit file directly is futile — it will be overwritten. The drop-in
approach explicitly resets ExecStart and sets the wrapper regardless of what the
main unit file says.
7. **NEVER hardcode** API keys in systemd drop-ins, config.yaml, or /etc/environment.
The wrapper injects live from vault at every start.
### Why this is non-fail
- **No rot**: keys pulled live from vault at every gateway start. Rotation = one `infisical secrets set` + `systemctl restart`. No per-host file edits.
- **Survives vault outage**: the `.env` fallback (Rule 3) keeps the gateway running if Infisical is unreachable or the service token is revoked.
- **Survives gateway crash**: the wrapper's `while true` + systemd `Restart=always` revive the gateway. Two-layer defense.
- **Survives Hermes updates**: systemd drop-in overrides unit file ExecStart — `hermes gateway install` cannot break the vault injection.
- **Survives reboot**: systemd user service + `loginctl enable-linger` ensures gateway starts at boot without a login session.
- **Auditable**: `cat /proc/$(pgrep -f 'python.*hermes_cli.main.gateway.run' | grep -v infisical | head -1)/environ` shows all injected keys (note: pipe through grep -v infisical to avoid matching the bash wrapper); `infisical secrets` shows the vault source.
### Migration status (2026-07-17)
| Agent | Host | Pattern | Keys | Status |
|-------|------|---------|------|--------|
| abiba | .24 | pi agent wrapper | ABIBA_LITELLM_API_KEY + ABIBA_ZULIP_API_KEY | ✅ vault-backed |
| mumuni | .123 | systemd drop-in + while-true wrapper + st.8e848433 | MUMUNI_LITELLM_API_KEY + MUMUNI_ZULIP_API_KEY | ✅ vault-backed + .env fallback |
| tanko | .122 | systemd drop-in + while-true wrapper + st.8e848433 (user jerome) | TANKO_LITELLM_API_KEY + TANKO_ZULIP_API_KEY | ✅ vault-backed + .env fallback |
| koby | .129 | systemd drop-in + while-true wrapper + st.8e848433 | KOBY_LITELLM_API_KEY, shares TANKO_ZULIP_API_KEY (tanko-bot) | ✅ vault-backed |
| koonimo | .114 | systemd drop-in + while-true wrapper + st.8e848433 | KOONIMO_LITELLM_API_KEY + KOONIMO_ZULIP_API_KEY | ✅ vault-backed |
> Tanko runs as user `jerome` — wrapper/token at `~/.hermes/infisical-gateway.sh` and
> `~/.infisical-token`. Linger enabled (`loginctl enable-linger jerome`) for boot startup.
### Tanko migration (COMPLETED 2026-07-17)
Tanko was the last agent migrated from hardcoded keys to vault wrapper.
Previously: key hardcoded in `/home/jerome/.hermes/config.yaml` (`api_key: sk-CggiHWlamQy…`)
and `zulip-env.conf` systemd drop-in. Now: user-scope systemd service with drop-in
`50-vault-wrapper.conf`, `infisical-gateway.sh` wrapper with while-true loop, token at
`~/.infisical-token`, `.env` fallback at `~/.hermes/.env`. Keys injected live from vault.
### Koby migration lessons (2026-07-16, updated 2026-07-17)
Migrated Koby from hardcoded systemd drop-in → `infisical-gateway.sh` wrapper.
**Three mistakes made:**
1. **Overwrote `/root/.hermes/.env`** without backing it up. The Zulip API key only existed
in the running process memory — the old .env was minimal (just LiteLLM key). Zulip creds were
inherited from the pre-migration gateway env, not stored in any file. Lost on restart.
2. **Only injected `LITELLM_API_KEY`** in the wrapper — forgot Zulip + Telegram credentials.
Agents need ALL their platform env vars. Missing vars cause silent adapter failures.
3. (2026-07-17 fix) **VENV variable in single-quoted bash -c**: `exec "$VENV/bin/python"`
inside single quotes resolved to `exec "/bin/python"` (file not found). Hardcoded full path.
**How Koby actually connects:**
- Zulip: shares **Tanko's bot** (`tanko-bot@chat.sysloggh.net`, `TANKO_ZULIP_API_KEY=5PeD6f3zo…`).
- Telegram: token from `.env` fallback. Allowed users: 6679773481.
- Both platforms now connect through the wrapper's env injection.
**Golden rules for gateway restarts:**
1. Always `cat /proc/<pid>/environ` before killing the old process — captures the live env set.
2. Hardcode venv python path in wrapper — never use variables inside single-quoted bash -c.
3. Use systemd drop-ins (not unit file edits) to override ExecStart — survives Hermes updates.
### Fleet-wide standardization lessons (2026-07-17)
After auditing all 4 agents, five systemic patterns caused repeated failures:
1. **Three incompatible startup patterns** coexisted (systemd drop-in, direct python, orphaned wrapper)
2. **Systemd unit files reverted** by `hermes gateway install` during updates
3. **VENV variable scoping** broke wrappers on Koby and Mumuni (single-quote bash -c)
4. **Service token expiry** — Tanko's `st.353699cd` was deleted from Infisical
5. **No ZULIP_API_KEY** in env on Tanko — wrapper bypassed by systemd direct python
All resolved by the canonical drop-in + while-true wrapper pattern documented above.
### Key rotation procedure (one vault operation with this standard)
1. Generate new key: `POST /key/generate` (master key, admin).
2. Update vault: `infisical secrets set <AGENT>_LITELLM_API_KEY=sk-NEW --token=$TOKEN --projectId=322fceab… --env=prod --domain=https://vault.sysloggh.net`.
3. Update `.env` fallback: `echo '<AGENT>_LITELLM_API_KEY=sk-NEW' > /root/.hermes/.env && chmod 600 /root/.hermes/.env`.
4. Restart: `systemctl restart hermes-gateway`. The wrapper pulls the new key live.
5. Verify: `curl -H "Authorization: Bearer sk-NEW" http://192.168.68.116/v1/models` → 200.
## Machine Identity for Vault Writes (UPDATED 2026-07-17)
**Current state:** Infisical CLI updated to v0.43.109 on all agents (from v0.38.0).
The v0.38.0 bug (user-session auth fails for `secrets set`/`export`) is resolved.
Service token `st.8e848433…` (abiba, READ+WRITE) can write to vault from CLI.
**Proper fix — per-agent Machine Identities:**
Create machine identities in Infisical UI → Project Settings → Machine Identities
for each agent with READ-only scope on the `agents` project. Store client_id +
client_secret per agent. Then vault writes use the shared abiba identity, and
reads use per-agent identities. This eliminates the single shared token risk.
**Service Token Inventory (2026-07-17):**
| Token ID | Name | Permissions | Used By | Status |
|----------|------|-------------|---------|--------|
| `st.8e848433…` | tanko-gateway | READ+WRITE | Mumuni, Tanko, Koby, Koonimo, Abiba | ✅ Active |
| `st.353699cd…` | tanko-agent | READ-only | — | ❌ Deleted from Infisical |
**Per-agent .env fallback inventory (2026-07-17):**
| Agent | .env Keys |
|-------|-----------|
| Mumuni | MUMUNI_LITELLM_API_KEY, MUMUNI_ZULIP_API_KEY |
| Tanko | TANKO_LITELLM_API_KEY, TANKO_ZULIP_API_KEY |
| Koby | (wrapper injects from vault — .env has Telegram token) |
| Koonimo | KOONIMO_LITELLM_API_KEY, KOONIMO_ZULIP_API_KEY |
## Key Rotation Log
| Date | Agent | Action | Notes |
|------|-------|--------|-------|
| 2026-07-17 | fleet | standardize | All 4 agents standardized on systemd drop-in + while-true wrapper + infisical v0.43.109. Removed conflicting zulip-env.conf + litellm-key.conf drop-ins. Added .env fallbacks with ZULIP keys. WAL #1322. |
| 2026-07-17 | tanko | fix-zulip | Added ZULIP_API_KEY to env (was missing — systemd bypassed vault). Updated wrapper from exec to while-true. Created .env fallback. Removed hardcoded zulip-env.conf drop-in. WAL #1321. |
| 2026-07-16 | vault | cleanup | 4 stale secrets deprecated. 5 personal creds flagged. |
| 2026-07-16 | koonimo | add-zulip | Added KOONIMO_ZULIP_API_KEY to vault. Wrapper injects ZULIP_API_KEY + ZULIP_EMAIL. 3 platforms. |
| 2026-07-16 | tanko | migrate | Migrated from hardcoded config.yaml to infisical-gateway.sh + st.353699cd. NOTE: st.353699cd later deleted — reverted to st.8e848433 on 2026-07-17. |
| 2026-07-16 | mumuni | rotate | Old key malformed (sk-_SWAl_Vu_, 47 chars, not LiteLLM format) → 401. Deleted old `mumuni` key (token 15cbca18…), generated fresh (alias `mumuni`, 7 models: syslog-auto, qwen3.6-27B-code, gemma-4-12b, strix-moe, gpu-dense, gpu-light, qwen3.6-35B-udq4). New key sk-OzuWsoX2… written to /root/.hermes/.env (Rule 3/13 fallback). Vault sync PENDING (needs machine identity). WAL #1300. |
| 2026-07-16 | koby | rotate | Old key sk-6sbCNjz (401, stale in /etc/environment). Deleted old `koby` key, generated fresh (alias `koby`). New key sk-BqRRMboTI… in systemd drop-in `hermes-gateway.service.d/litellm-key.conf` + /etc/environment. Created `hermes-gateway.service` unit (was missing — gateway wasn't persistent) with `--replace`. Verified HTTP 200, Telegram connected. |
| 2026-07-16 | baggy (koonimo) | rotate | Old key sk-krnw_zGB (401, hardcoded in systemd drop-in). Deleted old `baggy` key, generated fresh (alias `baggy`, metadata agent=koonimo). New key sk-OEK7z26n6E… in drop-in `hermes-gateway.service.d/litellm-key.conf`. CT113 IP changed .113→.114. Verified HTTP 200, Zulip connected. |
## LiteLLM Master Key (use sparingly — agents should NOT use it directly)
- Master key: `sk-litellm-7f96080dd99b15c36bd4b333b58a6796` (in /opt/inference-harness/.env on CT116, Infisical project=infrastructure env=production secret=LITELLM_MASTER_KEY)
- Used for /key/generate, /key/delete, /key/list (GET), DB queries
- **Known violation (RESOLVED 2026-07-16):** Abiba's LITELLM_API_KEY was previously the master key.
It is now a dedicated agent key `sk-sxbphLvk1OU…` (vault secret `ABIBA_LITELLM_API_KEY`, alias `abiba-pi`).
The master key is admin-only (/key/generate, /key/delete, /key/list). NEVER use it for inference —
see `litellm-self-heal` § "NEVER use litellm_proxy_master_key for inference".
- LiteLLM key DB: `harness-postgres` container on CT116, table `"LiteLLM_VerificationToken"` (columns: token, key_alias, key_name, created_at, expires). Query: `docker exec harness-postgres psql -U litellm -d litellm -t -c "SELECT key_alias, substr(token,1,16) FROM \"LiteLLM_VerificationToken\" ORDER BY created_at;"`
+10 -10
View File
@@ -42,8 +42,8 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2)
**What changed (v3.2.0 → v4.0.0 — 2026-07-08)**:
- Router REMOVED from request path — LiteLLM proxies directly to GPU
- All GPUs at parallel 2 (was parallel 1)
- NVIDIA context reduced 256K→128K to free VRAM
- LiteLLM timeouts tuned: gemma 25→120s, qwen 40→90s
- NVIDIA context reduced 256K→128K to free VRAM — now the stable ceiling across all GPUs (2026-07-17)
- LiteLLM timeouts tuned: gemma 25→120s, qwen 40→90s (SUPERSEDED 2026-07-16: qwen 300s, gemma 120s, strix 300s — see litellm-self-heal)
- nginx proxy_read_timeout: 600s, LiteLLM request_timeout: 300s
## Parameters
@@ -72,18 +72,18 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2)
| Host | IP | Hardware | Models Served | Engine | Context | Parallel |
|------|-----|----------|---------------|--------|---------|----------|
| llm-gpu | 192.168.68.8 | NVIDIA RTX 3090 (24 GB) | qwen3.6-27B-code | llama-server systemd | 128K | 2 |
| ocu-llm | 192.168.68.110 | NVIDIA RTX 5070 (12 GB) | gemma-4-12b | llama-server systemd | 128K | 2 |
| amdpve | 192.168.68.15 | AMD Strix Halo 64GB UMA | ornith-1.0-35b | llama-server systemd (Vulkan) | 256K | 2 |
| llm-gpu | 192.168.68.8 | NVIDIA RTX 3090 (24 GB) | qwen3.6-27B-code | llama-server systemd | **128K** | 2 |
| ocu-llm | 192.168.68.110 | NVIDIA RTX 5070 (12 GB) | gemma-4-12b | llama-server systemd | **128K** | 2 |
| amdpve | 192.168.68.15 | AMD Strix Halo 64GB UMA | qwen3.6-35B-udq4 (LiteLLM alias: strix-moe) | llama-server systemd (Vulkan) | 128K | 2 |
## Model Fallback Chains (LiteLLM)
| Primary | Timeout | Fallback | Timeout |
|---------|---------|----------|---------|
| qwen3.6-27B-code | 90s | gemma-4-12b | 120s |
| gemma-4-12b | 120s | qwen3.6-27B-code | 90s |
| ornith-1.0-35b | 120s | qwen → gemma | — |
| syslog-auto (balanced) | 90s | qwen → gemma | — |
| qwen3.6-27B-code | 300s | gemma-4-12b | 120s |
| gemma-4-12b | 120s | qwen3.6-27B-code | 300s |
| qwen3.6-35B-udq4 / strix-moe | 300s | qwen → gemma | — |
| syslog-auto (balanced) | 300s | qwen → gemma | — |
> Global: request_timeout=300s, nginx proxy_read_timeout=600s
@@ -129,7 +129,7 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2)
7. **Check model inference via LiteLLM** — Test each model:
- POST /v1/chat/completions model=gemma-4-12b → expect 200
- POST /v1/chat/completions model=qwen3.6-27B-code → expect 200
- POST /v1/chat/completions model=ornith-1.0-35b → expect 200
- POST /v1/chat/completions model=strix-moe → expect 200
- Use master key for auth
8. **Check agent keys**:
+34 -13
View File
@@ -1,18 +1,21 @@
---
kind: responsibility
name: litellm-self-heal
status: manual-only
status: deployed
note: >
Auto-remediation code was removed from the pi Zulip extension (retired 2026-07-04).
This contract is now manual-only — triggers require explicit user request.
Consider reimplementing as a standalone cron job or prose contract.
DEPLOYED 2026-07-12 on CT 116 cron: 0 */6 * * *
Auto-remediation code was removed from the pi Zulip extension (retired 2026-07-04),
now reimplemented as `litellm-health-check.sh` on CT 116.
Script: `/opt/inference-harness/scripts/litellm-health-check.sh` on CT 116 (cron `0 */6 * * *`).
Reports to /var/log/litellm/health-*.json and RA-H OS knowledge graph.
GPU monitoring integrated from gpu-monitor on .24:9100.
Consolidated from litellm-health + litellm-self-heal on 2026-07-09 to eliminate
duplication of architecture diagrams, GPU topology, timeout tables, and container
lists. Health check is now § Health Check within this contract.
Source of truth for GPU topology and keys: gpu-fleet.prose.md
Last verified: 2026-07-09
Last verified: 2026-07-12
description: >
LiteLLM inference stack health monitoring + self-healing. Verifies the full
nginx → LiteLLM → GPU chain, 8 containers on CT 116, 3 GPU hosts, model
@@ -54,18 +57,29 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2)
| Host | IP | Hardware | Models Served | Engine | Context | Parallel |
|------|-----|----------|---------------|--------|---------|----------|
| llm-gpu | 192.168.68.8 | NVIDIA RTX 3090 (24 GB) | qwen3.6-27B-code | llama-server systemd | 128K | 2 |
| ocu-llm | 192.168.68.110 | NVIDIA RTX 5070 (12 GB) | gemma-4-12b | llama-server systemd | 128K | 2 |
| amdpve | 192.168.68.15 | AMD Strix Halo 64GB UMA | ornith-1.0-35b | llama-server systemd (Vulkan) | 256K | 2 |
| llm-gpu | 192.168.68.8 | NVIDIA RTX 3090 (24 GB) | qwen3.6-27B-code | llama-server systemd (`/home/llmuser/llama-wrapper.sh`, `-c 131072 --parallel 2 --ngl 99`) | **128K** | 2 |
| ocu-llm | 192.168.68.110 | NVIDIA RTX 5070 (12 GB) | gemma-4-12b | llama-server systemd (`/home/llmuser/llama-wrapper.sh`, `--ctx-size 131072 --parallel 2`, IQ4_NL + MTP draft) | **128K** | 2 |
| amdpve | 192.168.68.15 | AMD Strix Halo 64GB UMA | qwen3.6-35B-udq4 (LiteLLM alias: `strix-moe`) | llama-server systemd (Vulkan) | 128K | 2 |
> Verified on ground 2026-07-16 via `curl /v1/models` on each host + `llama-wrapper.sh`. The AMD host's underlying model is `qwen3.6-35B-udq4`; LiteLLM exposes it under two `model_name`s: `qwen3.6-35B-udq4` and `strix-moe` (rpm 40). The legacy name `ornith-1.0-35b` does NOT exist in LiteLLM and must not be referenced.
## LiteLLM Model Surface (ground truth — `/opt/inference-harness/litellm_config.yaml` on CT 116)
`model_name`s served: `qwen3.6-27B-code`, `gemma-4-12b`, `qwen3.6-35B-udq4`, `strix-moe`, `gpu-dense`, `gpu-light`, `syslog-auto`.
- `syslog-auto` is a weighted router model: qwen3.6-27B-code (0.55, rpm 500) + qwen3.6-35B-udq4 (0.30, rpm 60) + gemma-4-12b (0.15, rpm 200).
- `gpu-dense` / `gpu-light` are high-rpm aliases (rpm 500) onto qwen3.6-27B-code / gemma-4-12b respectively.
- Key scoping: agent keys are restricted to `['syslog-auto','qwen3.6-27B-code','gemma-4-12b','strix-moe','gpu-dense','gpu-light']`. As of 2026-07-16 the `baggy`/`koby`/`mumuni`/`abiba-pi` keys ALSO include `qwen3.6-35B-udq4`; `abiba-pi` additionally includes `deepseek-v4-pro` (cloud fallback). `kagenz0`/`koonimo`/`pi-agents-unified` have the standard 6 only. Agents should still use the stable alias `strix-moe` (not the raw `qwen3.6-35B-udq4`) so model swaps don't break them.
- **NEVER use `litellm_proxy_master_key` (the `sk-litellm-...` master key) for inference.** It is for admin endpoints only (`/key/list`, `/key/generate`, `/key/info`). All inference — agent traffic, health-check model tests, monitor scripts — uses agent-specific keys. The health-check script's model tests use a dedicated `monitor` agent key stored at `/etc/litellm-monitor.env` on CT 116 (root-only, `chmod 600`); `/key/list` is the only call that legitimately uses `$MASTER_KEY`.
## Model Fallback Chains (LiteLLM)
| Primary | Timeout | Fallback | Timeout |
|---------|---------|----------|---------|
| qwen3.6-27B-code | 90s | gemma-4-12b | 120s |
| gemma-4-12b | 120s | qwen3.6-27B-code | 90s |
| ornith-1.0-35b | 120s | qwen → gemma | — |
| syslog-auto (balanced) | 90s | qwen → gemma | — |
| qwen3.6-27B-code | 300s | gemma-4-12b | 120s |
| gemma-4-12b | 120s | qwen3.6-27B-code | 300s |
| qwen3.6-35B-udq4 / strix-moe | 300s | qwen → gemma | — |
| syslog-auto (balanced) | 300s | qwen → gemma | — |
> Global: request_timeout=300s, nginx proxy_read_timeout=600s
@@ -84,6 +98,13 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2)
| harness-docker-stats | python:3.12-alpine | — | container stats exporter |
| harness-pve-exporter | prompve/prometheus-pve-exporter | — | Proxmox metrics → Prometheus |
## Script Operations (synced 2026-07-16)
- **Health-check script** (`/opt/inference-harness/scripts/litellm-health-check.sh` on CT 116): `gpu-fleet` check fails only on **critical** alerts (warnings are informational). Tests `strix-moe` (not `ornith-1.0-35b`).
- **GPU monitor** (`/root/scripts/gpu-monitor-server.py` on pi .24): runs as **systemd unit `gpu-monitor.service`** (was bare `&` process). `gpu_count` includes Strix Halo (was 2, now 3). VRAM alert thresholds: warning 93%, critical 97% (raised from 90/95 — 128K context steady-state is ~70% on RTX 3090, not a fault).
- **Agent key monitor** (`/root/scripts/agent-health-check.py` on pi .24, cron `*/10`): reads each agent's **live** `LITELLM_API_KEY` from its gateway process env via SSH — never hardcodes keys (hardcoded keys rot on rotation and caused 9×401/30min). Fleet roster: abiba, tanko, mumuni, koby, koonimo (legacy `tdunna`/`baggy` removed — never existed).
- **Stale keys cleaned**: `daily-infra-report.py` SYNTHETIC_API_KEY was stale (`sk-U_ydi3B` → 401); now reads `LITELLM_MASTER_KEY` from env. Deprecated scripts (`router-original.py`, `router-phase0-backup.py`, `apply-fixes.py`) still reference `sk-syslog-local-master-key` but do not actively poll LiteLLM.
## Maintains
- litellm-admin-ui: { status: "healthy", last_check: timestamp }
@@ -126,7 +147,7 @@ Run this first on every cycle. Results feed into remediation rules below.
### 5. Check model inference via LiteLLM — test each model
- POST /v1/chat/completions model=gemma-4-12b → expect 200
- POST /v1/chat/completions model=qwen3.6-27B-code → expect 200
- POST /v1/chat/completions model=ornith-1.0-35b → expect 200
- POST /v1/chat/completions model=strix-moe → expect 200
- Use master key for auth
### 6. Check agent keys
+4 -4
View File
@@ -1,7 +1,7 @@
---
name: memory-audit-maintenance
kind: responsibility
description: Shared memory audit and maintenance contract for all Hermes agents (Mumuni, Tanko, Tdunna, Baggy). Each agent runs it against its own isolated memory files — no cross-agent access, no shared state. Detects staleness, enforces writer registry, and rotates canary tokens.
description: Shared memory audit and maintenance contract for all Hermes agents (Mumuni, Tanko, Koby, Koonimo). Each agent runs it against its own isolated memory files — no cross-agent access, no shared state. Detects staleness, enforces writer registry, and rotates canary tokens.
id: 067NC4KG01RG50R40M30E20918
---
@@ -15,13 +15,13 @@ Autonomously audit and reorganize an agent's native memory (MEMORY.md, USER.md,
### Scope
This contract is the **Hermes Agent standard** for memory maintenance. It is shared across all Hermes agents (Mumuni, Tanko, Tdunna, Baggy). Each agent runs it against its own memory files only no cross-agent access, no shared state, no shared ledger, no shared canary. The contract is the standard; each agent enforces it independently with fully isolated data.
This contract is the **Hermes Agent standard** for memory maintenance. It is shared across all Hermes agents (Mumuni, Tanko, Koby, Koonimo). Each agent runs it against its own memory files only no cross-agent access, no shared state, no shared ledger, no shared canary. The contract is the standard; each agent enforces it independently with fully isolated data.
**Agent Roster:**
- Mumuni
- Tanko
- Tdunna
- Baggy
- Koby (CT 111 / tdunna)
- Koonimo (CT 113 / baggy)
**Isolation Principle:** Each agent has its own:
- `MEMORY.md` and `USER.md`
+38 -4
View File
@@ -66,16 +66,36 @@ Delegation is **mandatory** when any of these apply:
`curl`, `hermes tools list` — these are decision-making tools. The manager
reads them directly.
## Data Source Integrity (CRITICAL)
**Workers MUST use the data provided in their task context. They MUST NOT
fetch their own data from external sources unless explicitly told to.**
When a task says "Read file X and format it", the worker reads file X. It does
not query a separate API, run its own diagnostics, or pull data from a different
system. This is the #1 source of cross-worker inconsistency: one worker gathers
SSH data, another queries the Proxmox API, and the report merges two incompatible
datasets.
**Rule:** If a worker needs additional data beyond what's in its task description,
it asks the manager (via relay) — it doesn't go find it on its own.
**This is a hard rule, not a recommendation.** Violating it produces the exact
type of discrepancy the kanban pipeline exists to prevent: a review worker finds
"5 nodes present" in the raw data but "5/5 online" in the report — even though
one of those nodes was unreachable. The report lied because it used data the
raw data never provided.
## Worker Selection Matrix
| Worker | Model | Toolsets | Role | Use When |
|--------|-------|----------|------|----------|
| `syslog-code` | qwen3.6-27B-code | terminal, file, web, memory, skills | Code patches, automation, scripts | Writing/modifying code, creating scripts, debugging, reading/writing files |
| `syslog-devops` | qwen3.6-27B-code | terminal, file, web, memory, skills | Infrastructure, DB, bridge, Proxmox | Server ops, SSH, Docker, Proxmox, DB queries, hardware checks |
| `syslog-email` | ornith-1.0-35b | terminal, file, web, memory, skills | Email automation, mail operations | Sending/receiving email, inbox management, SMTP operations |
| `syslog-research` | ornith-1.0-35b | terminal, file, web, memory, skills, **browser** | Analysis, classification, data processing | Web research, browser tasks, data analysis, classification, reading docs |
| `syslog-review` | ornith-1.0-35b | terminal, file, web, memory, skills | Verification, QA, audit validation | **ALWAYS** verify worker output before delivery — especially for infra changes, code builds, and research findings |
| `syslog-writer` | ornith-1.0-35b | terminal, file, web, memory, skills | Docs, content, branding, reports | Writing docs, reports, proposals, content, markdown formatting |
| `syslog-email` | strix-moe | terminal, file, web, memory, skills | Email automation, mail operations | Sending/receiving email, inbox management, SMTP operations |
| `syslog-research` | strix-moe | terminal, file, web, memory, skills, **browser** | Analysis, classification, data processing | Web research, browser tasks, data analysis, classification, reading docs |
| `syslog-review` | strix-moe | terminal, file, web, memory, skills | Verification, QA, audit validation | **ALWAYS** verify worker output before delivery — especially for infra changes, code builds, and research findings |
| `syslog-writer` | strix-moe | terminal, file, web, memory, skills | Docs, content, branding, reports | Writing docs, reports, proposals, content, markdown formatting |
### Selection Rules
@@ -100,6 +120,19 @@ dispatch sequentially.
Fire workers via `delegate_task`:
**Critical: Pass the data, not just the goal.** When dispatching a worker that
processes output from another worker, include the file path AND explicit
instructions to use ONLY that source. Example:
```
delegate_task(
goal="Format the cluster check into a clean report",
context="Source data is at /tmp/proxmox-check-raw.md. Format ONLY the data
in that file. Do NOT query the Proxmox API or any other data source. Use the
file as your sole source of truth."
)
```
**Parallel (independent lanes):**
```
delegate_task(
@@ -203,6 +236,7 @@ Only verified results reach Kwame. Format per channel:
- ❌ Skipping verification → raw worker output never reaches the user
- ❌ Delegating single tool calls → keep quick reads/writes at manager level
- ❌ Firing more than 3 workers in parallel → hard limit
- ❌ **Workers fetching their own data sources** → a writer worker that queries the Proxmox API when told to "format the raw file" is fabricating data. Use the input given, not external sources
## Emergency Exception
+3 -3
View File
@@ -41,7 +41,7 @@ agent: abiba
- User: `monitoring@pve` (cluster-replicated)
- Role: `PVEAuditor` on `/` (read-only, whole cluster)
- Token: `monitoring@pve!prometheus` = `2c74ceb6-f905-444a-94f9-1c4f7889b68c`
- Token: `monitoring@pve!prometheus` — stored in Infisical vault (`PROXMOX_MONITOR_TOKEN`)
- `verify_ssl: false` (proxmoxer uses `verify_ssl`, NOT `verify_tls`)
## Grafana Dashboards (file-provisioned, folder "Syslog Fleet")
@@ -62,7 +62,7 @@ agent: abiba
- **URL**: `http://192.168.68.116:3001/` (LAN, direct — Grafana bound to `0.0.0.0:3001`)
- **Dashboards**: `http://192.168.68.116:3001/d/gpu-fleet`, `.../d/proxmox-cluster`, `.../d/proxmox-node`, `.../d/docker-containers`
- **Credentials**: admin / syslog-grafana-2026
- **Credentials**: admin / password stored in Infisical vault (`GRAFANA_ADMIN_PASSWORD`)
- Grafana is NOT behind nginx — access port 3001 directly. The `harness-nginx` `/grafana/` sub-path route was tried and reverted (broke the existing `:3001` URL and gpu-fleet path). Do not re-add `GF_SERVER_SERVE_FROM_SUB_PATH` or an nginx `/grafana/` route.
- grafana compose port mapping: `"3001:3000"` (0.0.0.0, not 127.0.0.1)
@@ -87,7 +87,7 @@ agent: abiba
| storepve | 192.168.68.6 | PVE |
| acerpve | 192.168.68.9 | PVE (hosts llm-gpu qemu/101) |
| minipve | 192.168.68.12 | PVE |
| amdpve | 192.168.68.15 | PVE + Strix Halo LLM (ornith) |
| amdpve | 192.168.68.15 | PVE + Strix Halo LLM (qwen3.6-35B-udq4, strix-moe) |
## Operations
+46 -6
View File
@@ -19,13 +19,53 @@ from datetime import datetime
LITELLM = "http://192.168.68.116:80"
def _get_agent_key(agent_name):
"""Retrieve agent key from Infisical vault."""
try:
result = subprocess.run(
["infisical", "secrets", "get", "LITELLM_API_KEY",
"--project=agents", "--env=production", "--plain"],
capture_output=True, text=True, timeout=10
)
if result.returncode == 0:
return result.stdout.strip()
except Exception:
pass
# Fallback: try exporting all secrets
try:
result = subprocess.run(
["infisical", "export", "--project=agents", "--env=production",
"--format=dotenv"],
capture_output=True, text=True, timeout=10
)
if result.returncode == 0:
for line in result.stdout.splitlines():
if line.startswith(f"LITELLM_API_KEY_{agent_name.upper()}") or \
(line.startswith("LITELLM_API_KEY=") and agent_name == os.uname().nodename):
return line.split("=", 1)[1].strip().strip('"').strip("'")
except Exception:
pass
return None
# Agent keys are pulled from Infisical vault at runtime.
# The 'key' field is populated dynamically below.
AGENTS = {
"tanko": {"ct": 112, "host": "192.168.68.122", "key": "sk-CggiHWlamQyShxWC3Hx6uw", "user": "jerome"},
"mumuni": {"ct": 114, "host": "192.168.68.123", "key": "sk-VrqCNlwUgzoNGOpikJ7nwQ", "user": "root"},
"tdunna": {"ct": 111, "host": None, "key": "sk-6sbCNjz2T6lTVDBdlNHXsA", "user": None},
"baggy": {"ct": 113, "host": None, "key": "sk-krnw_zGBwvvL5b7l2t-s-A", "user": None},
"tanko": {"ct": 112, "host": "192.168.68.122", "user": "jerome"},
"mumuni": {"ct": 114, "host": "192.168.68.123", "user": "root"},
"koby": {"ct": 111, "host": None, "user": None},
"koonimo": {"ct": 113, "host": None, "user": None},
}
# Inject keys from vault
for agent_name in AGENTS:
key = _get_agent_key(agent_name)
if key:
AGENTS[agent_name]["key"] = key
else:
AGENTS[agent_name]["key"] = None
GPU_HOSTS = {
"gpu-rtx3090 (.8)": {"host": "192.168.68.8", "port": 8080, "service": "llama-server"},
"gpu-rtx5070 (.110)": {"host": "192.168.68.110", "port": 8080, "service": "llama-server"},
@@ -144,8 +184,8 @@ def check_agents():
print(f"{name} (CT {ct}): cannot SSH — skip liveness check")
continue
# Gateway process
pid = ssh(host, "pgrep -f 'hermes_cli.main gateway run' | head -1", user=user)
# Gateway process (exclude the infisical bash wrapper that contains the same string)
pid = ssh(host, "pgrep -f 'hermes_cli.main gateway run' | grep -v infisical | head -1", user=user)
if not pid:
print(f"{name}: GATEWAY NOT RUNNING")
FAIL.append(f"gateway-down:{name}")
+2 -2
View File
@@ -79,7 +79,7 @@ description: >
`messages_processed` stalls.
- **Root cause**: Agent's model config (`models.json` or `settings.json`) references a
model ID that doesn't exist in LiteLLM's authorized model list. Example: `qwen3.6-35B-A3B`
configured but LiteLLM only exposes `ornith-1.0-35b` under that key. pi's session
configured but LiteLLM only exposes `strix-moe` (alias for qwen3.6-35B-udq4) under that key. pi's session
workers emit 403 on first prompt, then never recover because the error doesn't trigger
`agent_end` — worker stays `busy` and all subsequent messages pile up in the steer queue.
- **Detection**: Compare `~/.pi/agent/models.json` model IDs against `curl -H "Authorization: Bearer <KEY>" http://192.168.68.116/v1/models` output. A stuck worker shows
@@ -88,7 +88,7 @@ description: >
(2) Set `defaultModel` to `syslog-auto` (safe routing model). (3) Delete stale session
JSONL files from `~/.pi/agent/sessions/zulip/`. (4) Restart PM2 process.
- **Prevention**: Use `syslog-auto` as default model for all agents — it handles model
routing and fallback automatically. Direct model IDs (`ornith-1.0-35b`, etc.) should
routing and fallback automatically. Direct model IDs (`strix-moe`, etc.) should
only be used when explicitly requested. Validate model IDs at agent setup time.
- **Applies to**: pi extension (Tdunna CT111, fixed 2026-07-08), any agent using `syslog-harness` provider
+476
View File
@@ -0,0 +1,476 @@
---
kind: responsibility
name: zulip-resilience-v3
description: >
Rewrite the pi Zulip gateway with production-grade resilience patterns drawn from
Zulip's own event system docs (queue lifecycle, heartbeat monitoring, BAD_EVENT_QUEUE_ID
handling, idle_queue_timeout) and battle-tested Node.js resilience patterns
(circuit breaker, exponential backoff with jitter, bulkhead isolation, supervisor watchdog).
replaces: zulip-self-heal (retired)
agent: abiba
triggers:
- "/zulip self-heal v3"
- "zulip stopped responding"
- "PM2 abiba-zulip crashed"
---
# Zulip Gateway v3 — Production Resilience
## Architecture Overview
The current v2 gateway (`/root/.pi/agent/extensions/zulip/index.js`) has three structural
weaknesses that cause repeated deaths:
1. **No crash recovery** — uncaught errors kill the Node process, PM2 exhausts max_restarts
2. **No circuit breaker** — 502/fetch-failed errors escalate to process death with no fallback
3. **No queue lifecycle management** — doesn't use Zulip's documented heartbeat protocol or
idle_queue_timeout, so BAD_EVENT_QUEUE_ID errors cascade into crashes
The v3 rewrite addresses all three, following patterns from:
- [Zulip Events System docs](https://zulip.readthedocs.io/en/11.6/subsystems/events-system.html) —
queue registration, heartbeat, BAD_EVENT_QUEUE_ID recovery, call_on_each_event loop
- [Zulip API: Get Events](https://zulip.com/api/get-events) — long-poll timeout, dont_block, event ack
- [Circuit Breaker & Retry Patterns in Node.js 2026](https://1xapi.com/blog/resilient-api-circuit-breaker-bulkhead-retry-nodejs-2026) —
Opossum-based circuit breaker with fallback, retry with jitter, bulkhead isolation
---
## Maintains
- `zulip-gateway`: { status: "healthy" | "degraded" | "down" }
- `circuit-breaker`: { state: "CLOSED" | "OPEN" | "HALF_OPEN", failures, successes }
- `queue-lifecycle`: { queue_id, last_event_id, idle_timeout, heartbeat_age }
- `workers`: { count, busy, idle, stuck }
- `supervisor`: { pid, last_check, health_failures }
---
## Detection Rules
### Rule 1: Queue Expired (BAD_EVENT_QUEUE_ID)
- **Detect**: Events API returns error with BAD_EVENT_QUEUE_ID in body
- **Fix**: Call `POST /register` to create new queue, update queue_id and last_event_id
- **Debounce**: If 3 re-registrations fail within 60s, escalate (server may be down)
- **Ref**: Zulip docs: "Your software will need to handle that error condition by re-initializing itself"
### Rule 2: Network Degradation (502/ECONNREFUSED/fetch failed)
- **Detect**: Events API returns 502 or network error
- **Circuit breaker**: Track failure rate over 10s rolling window
- CLOSED → OPEN: 50% failure rate with ≥5 requests
- OPEN → HALF_OPEN: After 30s reset timeout
- HALF_OPEN → CLOSED: Probe succeeds
- HALF_OPEN → OPEN: Probe fails
- **While OPEN**: Log errors, skip events, notify user via DM: "⚠️ Zulip connection degraded — will retry in 30s"
### Rule 3: Long-Poll Timeout (natural)
- **Detect**: Events API response takes > `event_queue_longpoll_timeout_seconds`
- **Not an error**: Server sends heartbeat events when no real events. Simply re-poll.
### Rule 4: Worker Busy Timeout (>5 min)
- **Detect**: Worker `busySince` exceeds 5 minutes
- **Fix**: SIGKILL worker, send error DM, clean up pending replies
### Rule 5: Process Crash (uncaught)
- **Detect**: `uncaughtException` / `unhandledRejection` fires
- **Fix**: Log → clear poll timer → attempt reconnect with backoff → if reconnect fails 3x, exit(1) and let PM2 restart
### Rule 6: Supervisor Detects Router Stall
- **Detect**: External supervisor (`zulip-watchdog`) polls `/health` every 30s. If 3 consecutive failures:
- **Fix**: `pm2 restart abiba-zulip` gracefully (SIGTERM, drain workers, restart)
---
## Implementation Plan
### Phase 1: Rewrite Router Core (circuit-breaker + queue lifecycle)
Replace the poll loop in index.js with a resilience-first event loop:
```js
// Queue lifecycle (Zulip docs pattern)
async function createOrRefreshQueue() {
// POST /register with event_types=["message"]
// Store: queueId, lastEventId, eventQueueLongpollTimeoutSeconds
// NEW: pass idle_queue_timeout parameter (Zulip 12.0+)
}
// Circuit breaker (Opossum pattern, implemented inline to avoid dependency)
class ZulipCircuitBreaker {
constructor({ failureThreshold=0.5, resetTimeout=30000, volumeThreshold=5, windowMs=10000 }) {
this.state = "CLOSED"; // CLOSED | OPEN | HALF_OPEN
this.failures = 0;
this.successes = 0;
this.totalRequests = 0;
this.lastFailureTime = null;
this.openedAt = null;
this.failureThreshold = failureThreshold;
this.resetTimeout = resetTimeout;
this.volumeThreshold = volumeThreshold;
this.windowMs = windowMs;
}
async fire(fn) {
if (this.state === "OPEN") {
if (Date.now() - this.openedAt > this.resetTimeout) {
this.state = "HALF_OPEN";
} else {
throw new CircuitOpenError("Circuit is OPEN");
}
}
try {
const result = await fn();
this.onSuccess();
return result;
} catch (err) {
this.onFailure();
throw err;
}
}
onSuccess() {
this.successes++;
this.totalRequests++;
if (this.state === "HALF_OPEN") {
this.state = "CLOSED";
this.failures = 0;
}
// Reset counters periodically
if (this.totalRequests > this.volumeThreshold * 2) {
this.failures = Math.floor(this.failures / 2);
this.successes = Math.floor(this.successes / 2);
this.totalRequests = Math.floor(this.totalRequests / 2);
}
}
onFailure() {
this.failures++;
this.totalRequests++;
this.lastFailureTime = Date.now();
if (this.totalRequests >= this.volumeThreshold &&
this.failures / this.totalRequests >= this.failureThreshold) {
if (this.state !== "OPEN") {
this.state = "OPEN";
this.openedAt = Date.now();
console.error(`[zulip-ext] CIRCUIT BREAKER OPEN — ${this.failures}/${this.totalRequests} failures`);
}
}
}
}
// Retry with exponential backoff + jitter (from resilience patterns)
async function withRetry(fn, { maxAttempts=3, baseDelay=200, maxDelay=10000, shouldRetry=()=>true }={}) {
let lastError;
for (let attempt = 1; attempt <= maxAttempts; attempt++) {
try {
return await fn();
} catch (err) {
lastError = err;
if (attempt === maxAttempts || !shouldRetry(err)) throw err;
const delay = Math.min(baseDelay * Math.pow(2, attempt - 1), maxDelay);
const jitter = delay * (0.5 + Math.random() * 0.5); // 50-100% of delay
console.warn(`[zulip-ext] Retry ${attempt}/${maxAttempts} after ${Math.round(jitter)}ms: ${err.message.slice(0,80)}`);
await new Promise(r => setTimeout(r, jitter));
}
}
throw lastError;
}
// Resilience-first event loop (Zulip call_on_each_event pattern)
async function resilientPollLoop() {
while (connected) {
try {
const events = await circuitBreaker.fire(() =>
withRetry(() => zulipQueue.poll(), {
maxAttempts: 2,
baseDelay: 1000,
shouldRetry: (err) => {
const msg = err.message || "";
return msg.includes("fetch failed") || msg.includes("ECONN") || msg.includes("network");
}
})
);
lastError = null;
retryCount = 0;
for (const ev of events) {
await processEvent(ev);
}
heartbeat();
} catch (err) {
const msg = err instanceof Error ? err.message : String(err);
if (msg.includes("BAD_EVENT_QUEUE_ID") || msg.includes("deregistered")) {
// Queue expired — re-register (Zulip docs pattern)
console.log(`[zulip-ext] Queue expired, re-registering… (${msg.slice(0,80)})`);
try {
zulipQueue = await createZulipQueue();
console.log(`[zulip-ext] Re-registered, new queue=${zulipQueue.queueId}`);
} catch (reRegErr) {
console.error(`[zulip-ext] Re-registration failed: ${reRegErr.message}`);
connected = false;
retryCount++;
const backoff = Math.min(5000 * Math.pow(2, retryCount), 300000);
console.log(`[zulip-ext] Full reconnect in ${Math.round(backoff/1000)}s`);
await new Promise(r => setTimeout(r, backoff));
await startPolling();
return;
}
} else if (err.name === "CircuitOpenError") {
// Circuit is open — skip this cycle, wait for HALF_OPEN
lastError = "circuit_open";
await new Promise(r => setTimeout(r, POLL_INTERVAL_MS));
} else {
lastError = msg;
retryCount++;
const backoff = Math.min(POLL_INTERVAL_MS * Math.pow(1.5, Math.min(retryCount, 8)), 60000);
console.error(`[zulip-ext] Poll error (retry ${retryCount}, backoff ${backoff}ms): ${msg}`);
await new Promise(r => setTimeout(r, backoff));
}
}
}
}
```
### Phase 2: PM2 Hardening
Create `/root/.pm2/ecosystem.config.cjs`:
```js
module.exports = {
apps: [
{
name: "abiba-zulip",
script: "/bin/pi",
args: "--mode rpc --session-id zulip-service",
env: {
ZULIP_ROLE: "router",
ZULIP_SITE: "https://chat.sysloggh.net",
ZULIP_EMAIL: "abiba-bot@chat.sysloggh.net",
ZULIP_API_KEY: process.env.ZULIP_API_KEY,
AGENT_NAME: "abiba",
AGENT_OWNER_EMAIL: "jerome@sysloggh.com",
},
max_restarts: 100, // Up from default 10 — crash loops won't exhaust
min_uptime: "10s", // Must survive 10s to count as "alive"
max_memory_restart: "500M", // OOM protection
restart_delay: 5000, // 5s between restarts
kill_timeout: 15000, // 15s SIGTERM grace before SIGKILL
listen_timeout: 30000, // 30s to bind health port
log_date_format: "YYYY-MM-DD HH:mm:ss Z",
error_file: "/root/.pm2/logs/abiba-zulip-error.log",
out_file: "/root/.pm2/logs/abiba-zulip-out.log",
merge_logs: true,
autorestart: true,
watch: false,
instances: 1,
exec_mode: "fork",
},
{
name: "zulip-watchdog",
script: "/root/.pi/agent/extensions/zulip/watchdog.js",
max_restarts: 10,
min_uptime: "3s",
restart_delay: 3000,
autorestart: true,
},
],
};
```
### Phase 3: Supervisor Watchdog
Create `/root/.pi/agent/extensions/zulip/watchdog.js`:
```js
// External supervisor — monitors router health and restarts if stalled.
// This is the pattern Hermes uses: an external process that can recover
// the gateway even if the gateway process itself is hung (not just crashed).
const HEALTH_URL = "http://127.0.0.1:9200/health";
const CHECK_INTERVAL_MS = 30_000;
const MAX_FAILURES = 3;
let failures = 0;
async function check() {
try {
const res = await fetch(HEALTH_URL, { signal: AbortSignal.timeout(5000) });
if (res.ok) {
const data = await res.json();
if (data.status === "ok" && data.zulip?.connected) {
if (failures > 0) {
console.log(`[watchdog] Router recovered after ${failures} failures`);
}
failures = 0;
return;
}
}
failures++;
console.warn(`[watchdog] Health check ${failures}/${MAX_FAILURES}: status not ok`);
} catch (err) {
failures++;
console.warn(`[watchdog] Health check ${failures}/${MAX_FAILURES}: ${err.message}`);
}
if (failures >= MAX_FAILURES) {
console.error(`[watchdog] ${MAX_FAILURES} consecutive failures — restarting abiba-zulip`);
const { execSync } = require("child_process");
try {
execSync("pm2 restart abiba-zulip", { timeout: 30000 });
console.log("[watchdog] Restart command sent");
} catch (e) {
console.error(`[watchdog] Restart failed: ${e.message}`);
}
failures = 0;
// Wait for restart to complete before checking again
await new Promise(r => setTimeout(r, 15000));
}
}
console.log("[watchdog] Zulip gateway supervisor started");
setInterval(check, CHECK_INTERVAL_MS);
check(); // Immediate first check
```
### Phase 4: Health Endpoint Enhancement
Add circuit breaker stats to the existing health endpoint:
```js
// In /health response, add:
"circuit_breaker": {
"state": circuitBreaker.state,
"failures": circuitBreaker.failures,
"successes": circuitBreaker.successes,
"total_requests": circuitBreaker.totalRequests,
"failure_rate": circuitBreaker.totalRequests > 0
? (circuitBreaker.failures / circuitBreaker.totalRequests).toFixed(2)
: "0.00"
}
```
---
## Test Plan
### Test 1: Queue Re-registration
1. Manually delete the Zulip event queue via API
2. Next poll should detect BAD_EVENT_QUEUE_ID
3. Router should auto re-register within 1 poll cycle
4. Verify: `/health` shows new queue_id, connected=true
### Test 2: Circuit Breaker Trip
1. Block Zulip server with iptables: `iptables -A OUTPUT -d 192.168.68.19 -j DROP`
2. Router should detect failures, trip circuit after 5 failures
3. `/health` should show circuit_breaker.state = "OPEN"
4. Remove iptables rule
5. Circuit should transition to HALF_OPEN → CLOSED within 60s
6. Verify: messages processed after recovery
### Test 3: Supervisor Recovery
1. Kill the router process: `kill -STOP $(pm2 pid abiba-zulip)` (freeze, don't kill)
2. Watchdog should detect 3 failed health checks in 90s
3. Watchdog should execute `pm2 restart abiba-zulip`
4. Verify: router back online, connected=true
### Test 4: Worker Busy Timeout
1. Send a message that triggers a long-running operation
2. If worker stays busy >5 minutes, should receive SIGKILL
3. User should receive error DM: "Response timed out"
### Test 5: End-to-End Message
1. Send DM "What time is it?" from Jerome
2. Should receive response within 30s
3. `/health` should show messages_processed incremented
---
## Rollback Plan
If v3 causes issues:
1. `pm2 delete abiba-zulip; pm2 delete zulip-watchdog`
2. Restore v2 from git: `cd /root/.pi/agent/extensions/zulip && git checkout index.js`
3. `pm2 resurrect` to reload previous process list
4. Verify: `/health` returns ok
Backup v2 before starting: `cp index.js index.js.v2-backup-$(date +%Y%m%d-%H%M%S)`
---
## Success Metrics
| Metric | Current (v2) | Target (v3) |
|--------|-------------|-------------|
| Uptime between manual interventions | 1-3 days | 30+ days |
| Crash recovery | Manual (PM2 resurrect) | Automatic (circuit breaker + supervisor) |
| Queue expiry handling | Crash | Auto re-register |
| Busy worker deadlock | Router death | Worker SIGKILL + error DM |
| PM2 restart exhaustion | Yes (max_restarts=10) | No (max_restarts=100 + watchdog) |
---
## Incident Log — 2026-07-18 Fleet-Wide Audit
### Fleet State After Audit
| Agent | Platform | Zulip State | Issues Found | Fix Applied |
|-------|----------|-------------|--------------|-------------|
| **Abiba** | pi (CT 100) | ✅ Connected | API key missing from Infisical injection; poll timeout noise | Added .env fallback; AbortError treated as empty poll (no retry); poll timeout 65s→90s |
| **Tanko** | Hermes (CT 112) | ✅ Connected | Gateway disconnected since Jul 11; watchdog restart didn't re-establish Zulip | Full gateway restart (kill wrapper, let infisical-gateway.sh respawn) |
| **Mumuni** | Hermes (CT 114) | ✅ Connected | No issues found | None needed |
### Key Fixes Applied
**1. Abiba — Credential Fallback (L4 Pattern)**
- Root cause: `zulip.api_key` in config.yaml is `""` (expected from Infisical). Infisical vault `ABIBA_ZULIP_API_KEY` wasn't being injected into the process environment.
- Fix: Added `.env` file fallback at `/root/.pi/agent/extensions/zulip/.env` with known-working key, sourced before the Infisical `exec`.
- Lesson: Per L4 from gpu-self-heal, Infisical is not always available — always keep a local `.env` fallback.
**2. Abiba — Poll Timeout Handling**
- Root cause: Zulip long-poll uses `AbortSignal.timeout(65000)`. Zulip's default `event_queue_longpoll_timeout_seconds` can exceed 65s. When the signal fires, an `AbortError` is thrown and caught by the circuit breaker as a failure.
- Fix: Caught `AbortError` inside `poll()` and return empty array (no events) instead of throwing. Extended timeout to 90s to match Zulip server default.
- Reference: [Zulip Events System — long-poll timeout](https://zulip.readthedocs.io/en/11.6/subsystems/events-system.html)
**3. Tanko — Gateway Restart**
- Root cause: Gateway process was running but Zulip platform stayed in "disconnected" state since Jul 11, 2026. The wrapper script (`infisical-gateway.sh`) restarts on crash but the gateway wasn't re-establishing Zulip on restart.
- Fix: Killed gateway PID to trigger wrapper restart. New gateway (PID 331991) established Zulip connection successfully.
### Fleet-Wide Zulip Health Metrics (as of 2026-07-18)
| Metric | Value |
|--------|-------|
| Zulip server | ✅ HTTP 200 |
| Agents connected | 3/3 (Abiba, Tanko, Mumuni) |
| Abiba circuit breaker | CLOSED (0 failures) |
| Abiba uptime | 2D (post-restart) |
| Tanko gateway uptime | Ongoing |
| Mumuni gateway uptime | Ongoing |
| Watchdog status | ✅ Online (2D uptime) |
### Hermes Agent Zulip Plugin Improvements
Based on the audit, improvements that should be ported to all Hermes Zulip adapters:
1. **Circuit breaker pattern** — Already in Abiba's pi extension. Hermes adapters should add the same CLOSED→OPEN→HALF_OPEN state machine with exponential backoff.
2. **Credential fallback** — All Hermes agents use Infisical for credentials. Add `.env` local fallback per L4 pattern for `ZULIP_API_KEY`.
3. **Queue re-registration** — Handle `BAD_EVENT_QUEUE_ID` with automatic re-registration instead of gateway restart.
4. **Supervisor watchdog** — Hermes uses PM2 which auto-restarts on crash, but has no health-check watchdog. Add lightweight external health checks.
5. **Streaming** — All agents have `streaming: true` in their zulip config. Verify `edit_message()` is implemented in each adapter.
### Abiba pi Zulip Extension v2 — Implemented Resilience Summary
| Feature | Status | Notes |
|---------|--------|-------|
| Circuit breaker | ✅ | CLOSED→OPEN→HALF_OPEN; 50% failure threshold; 30s reset timeout |
| Retry with jitter | ✅ | 2 attempts, 200ms base, 50-100% jitter |
| Queue lifecycle | ✅ | 10min idle_queue_timeout; BAD_EVENT_QUEUE_ID handling |
| Crash prevention | ✅ | uncaughtException + unhandledRejection recovery |
| Worker timeout | ✅ | 5min busy timeout → SIGKILL + error DM |
| Health endpoint | ✅ | :9200 with circuit breaker metrics |
| Echo prevention | ✅ | Dynamic bot user resolution |
| Poll timeout (AbortError) | ✅ v2.1 | Normal timeout returns [] instead of error |
| Credential fallback | ✅ v2.1 | .env file before Infisical exec |
| Provider auto-fix | ✅ | Detects reasoning_content models, switches to compatible |