Compare commits

..
Author SHA1 Message Date
Abiba Bot 6d975360d8 no-mistakes(document): Fix stale compression-model docs across 4 contracts 2026-07-19 00:31:54 +00:00
Abiba Bot 3e246835d5 tune: switch compression model from strix-moe to syslog-auto (v2)
Relieves Strix Halo thermal pressure by routing compression through
syslog-auto weighted pool (55% RTX 3090, 30% Strix Halo, 15% RTX 5070).

- compression.model: strix-moe -> syslog-auto
- auxiliary.compression.model: strix-moe -> syslog-auto
- Rule 7: Updated for syslog-auto compression distribution
- Rule 8: Updated GPU workload distribution with syslog-auto pool
- Description/version banners: Added 2026-07-18 change record
2026-07-18 23:57:07 +00:00
10 changed files with 27 additions and 526 deletions
-7
View File
@@ -128,10 +128,3 @@ safe-mutate --verify "CMD" [--expect "PATTERN"] --mutate "CMD" [--reason "WHY"]
Read the [Authoring Guide](docs/AUTHORING-GUIDE.md) before writing any new contract.
It covers the full process: verify → draft → lint → review → ship, with templates
and style rules.
## Maintaining this file
Keep this file for knowledge useful to almost every future agent session in this project.
Do not repeat what the codebase already shows; point to the authoritative file or command instead.
Prefer rewriting or pruning existing entries over appending new ones.
When updating this file, preserve this bar for all agents and keep entries concise.
-1
View File
@@ -1 +0,0 @@
AGENTS.md
+4 -106
View File
@@ -1,5 +1,5 @@
registry_version: 0.1.0
last_updated: '2026-07-23T00:00:00Z'
last_updated: '2026-07-13T00:00:00Z'
updated_by: mumuni
categories:
- compliance
@@ -22,7 +22,6 @@ owners:
- abiba
- mumuni
- kwame
- ops
trigger_types:
- scheduled
- event_driven
@@ -59,7 +58,6 @@ index:
- memory-audit-maintenance
- gpu-fleet
- infrastructure-update
- infrastructure-maintenance
reference:
- infrastructure-control
- ra-h-os-custodianship-contract
@@ -92,7 +90,6 @@ index:
- infrastructure-control
- infrastructure-monitoring
- infrastructure-update
- infrastructure-maintenance
- pm2-self-heal
- disk-gc-threat-response
gpu:
@@ -133,6 +130,7 @@ index:
- build-zulip-plugin
- stirling-pdf-agent-access
- gpu-fleet
- infrastructure-update
- infrastructure-control
- zulip-adapter-lessons
- pi-approval-architecture
@@ -146,9 +144,6 @@ index:
- mumuni-delegation
kwame:
- hello-world
ops:
- infrastructure-maintenance
- infrastructure-update
by_trigger:
scheduled:
- hermes-key-enforcement
@@ -161,7 +156,6 @@ index:
- litellm-health
- memory-audit-maintenance
- infrastructure-update
- infrastructure-maintenance
event_driven:
- litellm-self-heal
- pm2-self-heal
@@ -203,7 +197,6 @@ index:
- hermes-zulip-plugin
- build-zulip-plugin
- infrastructure-update
- infrastructure-maintenance
- ra-h-os-custodianship-contract
- mumuni-delegation
normal:
@@ -628,7 +621,7 @@ contracts:
sensitivity: high
status: active
owner: abiba
version: 3.1.0
version: 3.0.0
trigger:
type: scheduled
cadence: '*/15 * * * *'
@@ -1263,108 +1256,13 @@ contracts:
last_run: null
last_status: null
drift_alerts: []
- name: infrastructure-maintenance
file: infrastructure-maintenance.prose.md
kind: responsibility
category: maintenance
sensitivity: high
status: active
owner: ops
version: 1.0.0
trigger:
type: scheduled
cadence: 0 2 * * 0
description: Weekly host-level maintenance Sunday at 2am ET (replaces infrastructure-update
build-phase role; infra-update moves to ops)
cron_job_id: null
execution:
agent: ops
timeout: 3600
requires:
- infrastructure-monitoring run within last 30 minutes (pre-update health baseline)
- Proxmox snapshot of primary host OR /tmp backup dir created this run
- LiteLLM master key from Infisical vault for health verification
protocol:
- Load contract from prose-contracts/main
- Phase 0 preflight — capture health baseline, backup check, record image baseline
- Phase 1 apt update && apt upgrade -y on primary host
- Phase 2 docker compose pull for LiteLLM, SearXNG, and other running containers
- Phase 3 restart stacks one at a time with per-stack health verification
- Phase 4 verify every critical service (LiteLLM, SearXNG, Zulip, Gitea, PM2, Hermes gateways)
- On failure — rollback per protocol, escalate, do not loop beyond circuit breaker
- Log actions to ~/.hermes/runs/infrastructure-maintenance/
verification:
postconditions:
- check: all critical services running after update
verify: 'curl -sf http://192.168.68.116/litellm/v1/models && curl -sf https://chat.sysloggh.net/api/v1/server_settings && curl -sf https://git.sysloggh.net/api/v1/version && curl -sf http://192.168.68.7:8888 && pm2 jlist'
expect: all probes 200 OK / processes online
- check: no regressions from pre-update health baseline
verify: diff Phase 0 health-baseline against Phase 4 results
expect: no GREEN service turned RED
- check: docker containers on latest stable tags
verify: docker inspect --format '{{.Config.Image}}' <container> per service matches image-baseline.pulled_tag
expect: all containers running pulled tags
- check: APT packages up to date with no held broken packages
verify: apt list --upgradable 2>/dev/null | wc -l and apt-get -s upgrade | grep -ci broken
expect: upgradable == 0, broken == 0
artifact: maintenance run report with phase results and any rollback/escalation
verify_commands:
- curl -sf http://192.168.68.116/litellm/v1/models
- curl -sf http://192.168.68.7:8888
- curl -sf https://chat.sysloggh.net/api/v1/server_settings
- curl -sf https://git.sysloggh.net/api/v1/version
- pm2 jlist
receipt:
format: json
storage: ~/.hermes/runs/infrastructure-maintenance/
graph_node: true
schema:
contract: string
run_id: string
timestamp: ISO 8601
agent: string
status: pass|fail|escalated
phase: preflight|apt|images|restarts|verify|rollback|done|failed
actions_taken: array
postconditions: array
drift_alerts: array
evidence_path: string
escalation:
info:
action: log_to_receipt
notify: []
warning:
action: relay_alert
notify:
- abiba
- mumuni
critical:
action: relay_alert
notify:
- abiba
- mumuni
fatal:
action: relay_alert + pause + human_required
notify:
- abiba
- mumuni
- kwame
circuit_breaker:
max_retries: 2
window: 7200
trip_action: escalate_to_fatal
depends_on:
- infrastructure-monitoring
last_run: null
last_status: null
drift_alerts: []
- name: infrastructure-update
file: infrastructure-update.prose.md
kind: responsibility
category: maintenance
sensitivity: high
status: active
owner: ops
owner: abiba
version: 1.0.0
trigger:
type: scheduled
+7 -7
View File
@@ -125,7 +125,7 @@ Note: All syslog-auto entries route directly to GPUs with `api_key: not-needed`.
| Alias | RPM Cap | Routes To | Purpose |
|-------|---------|-----------|---------|
| `strix-moe` | 40 | Strix Halo | Compression tasks (MoE models) |
| `strix-moe` | 40 | Strix Halo | Agent reasoning, compression (~30% via syslog-auto pool) (MoE models) |
| `gpu-dense` | 500 | RTX 3090 | Heavy reasoning |
| `gpu-light` | 500 | RTX 5070 | Vision, web extract, light tasks |
@@ -136,7 +136,7 @@ Note: All syslog-auto entries route directly to GPUs with `api_key: not-needed`.
- syslog-auto → qwen → gemma → qwen3.6-35B-udq4
### Why Strix Halo RPM Is Capped
- Direct (strix-moe): 40 RPM (tight) — Strix Halo is shared with compression tasks
- Direct (strix-moe): 40 RPM (tight) — Strix Halo handles agent reasoning + compression via syslog-auto pool
- Via syslog-auto: 60 RPM (moderate) — prevents flooding when multiple agents use syslog-auto simultaneously
- Combined max: ~100 RPM across both paths — Strix Halo can sustain this at 80°C
@@ -284,7 +284,7 @@ History stored at `/root/data/toks-history.json` with 7-day rolling window.
### Stable Aliases — CRITICAL
All agent configs MUST use stable role-based aliases, never model-specific names:
- `compression.model: strix-moe` (NOT `qwen3.6-35B-udq4`)
- `compression.model: syslog-auto` — switched from strix-moe 2026-07-18 to distribute across weighted pool (55% RTX 3090, 30% Strix Halo, 15% RTX 5070); relieves Strix Halo thermal pressure
- `auxiliary.vision.model: gpu-light` (NOT `gemma-4-12b`)
- `delegation.model: gpu-dense` (NOT `qwen3.6-27B-code`)
- `auxiliary.web_extract.model: gpu-light`
@@ -295,7 +295,7 @@ When the underlying model is swapped, only the LiteLLM config changes — agent
- RTX 3090: **128K** (reduced from 256K 2026-07-17) | RTX 5070: **128K** (reduced from 256K) | Strix Halo: **128K**
- **All agents**: 128K ceiling — stable margin. For >128K workloads, use external providers (deepseek)
- Compression threshold 0.65: fires at ~85K (~43K headroom before 128K ceiling)
- Mumuni compression model alias: `strix-moe` with 300s timeout
- Mumuni compression model alias: `syslog-auto` (switched from strix-moe 2026-07-18) with 300s timeout
### Mumuni Agent Profile
@@ -305,8 +305,8 @@ Mumuni (CT114, 192.168.68.123) is the primary business assistant. This profile i
|---------|-------|-------|
| `model.default` | `syslog-auto` | Weighted pool (55% qwen, 30% strix, 15% gemma) |
| `model.provider` | `custom:litellm` | LiteLLM on CT116 |
| `compression.model` | `strix-moe` | Stable alias — survives model swaps |
| `aux.compression.model` | `strix-moe` | Compression auxiliary model |
| `compression.model` | `syslog-auto` | Switched from strix-moe 2026-07-18 — distributes across weighted pool |
| `aux.compression.model` | `syslog-auto` | Compression auxiliary — switched from strix-moe 2026-07-18 |
| `aux.vision.model` | `gpu-light` | Vision tasks (RTX 5070) |
| `aux.web_extract.model` | `gpu-light` | Web extraction |
| `delegation.model` | `gpu-dense` | Sub-agent reasoning (RTX 3090) |
@@ -318,7 +318,7 @@ Mumuni (CT114, 192.168.68.123) is the primary business assistant. This profile i
| `personalities` | `creative` | Creative assistant personality |
| Platforms | cli, discord, homeassistant, signal, telegram, zulip | All Hermes platforms |
| Main model timeout | 300s | LiteLLM global timeout |
| Compression model timeout | 300s | strix-moe timeout increased from 120s |
| Compression model timeout | 300s | syslog-auto timeout (switched from strix-moe 2026-07-18) |
### Agent Update Status (2026-07-15)
+2 -2
View File
@@ -46,7 +46,7 @@ depends_on:
|-------|-----|------|-------|------|-----|-------|------|
| `gpu-dense` | RTX 3090 24GB | ct8 (.8:8080) | ThinkingCap Qwen3.6-27B Q4_K_M + MTP + vision | 21.6/24.6GB (88%) | 128K | 74.9 | Heavy reasoning, code gen |
| `gpu-light` | RTX 5070 12GB | ct110 (.110:8080) | HauhauCS Gemma4-12B QAT Q4_K_M + MTP draft | 10.1/12.2GB (83%) | 128K | 169.6 | Vision, web extract, light tasks |
| `strix-moe` | Strix Halo 64GB | ct15 (.15:8080) | Genesis Hermes V3 APEX (LuffyTheFox, 24GB) | ~10/64GB (16%) | 128K | 62.9 | Compression, summarization, long docs |
| `strix-moe` | Strix Halo 64GB | ct15 (.15:8080) | Genesis Hermes V3 APEX (LuffyTheFox, 24GB) | ~10/64GB (16%) | 128K | 62.9 | Agent reasoning, compression (~30% via syslog-auto pool), summarization, long docs |
Key notes:
- All models use direct GPU routing via LiteLLM (`api_key: not-needed`). Router (port 9000) is deprecated and NOT in the inference path.
@@ -157,7 +157,7 @@ Key notes:
- **Target distribution**:
- RTX 3090 (gpu-dense, 24GB, 74.9 tok/s) → Heavy reasoning, code gen, long conversations (slowest per-token but largest context capacity). Weight: 0.55 (LiteLLM).
- RTX 5070 (gpu-light, 12GB, 169.6 tok/s) → Vision/image, web search, lightweight tasks (2.3x faster than 3090 per token). Weight: 0.15 (LiteLLM).
- Strix Halo (strix-moe, 64GB, 62.9 tok/s) → Context compression, summarization, long docs (MoE model). Weight: 0.30 (LiteLLM).
- Strix Halo (strix-moe, 64GB, 62.9 tok/s) → Agent reasoning, compression (~30% via syslog-auto pool), summarization, long docs (MoE model). Weight: 0.30 (LiteLLM).
- **Note**: RTX 5070 is the fastest endpoint per token. Route high-volume, low-complexity work there first.
- **Fix**:
- Alert if any GPU is handling workload outside its designated role
+1 -1
View File
@@ -5,7 +5,7 @@ version: 1.0.0
description: >
Canonical known-good baseline for all Syslog Hermes agents. Captures the exact
configuration state, keys, workarounds, and audit procedure. When an agent's
configuration goes sideways, restore from this baseline. Last verified 2026-07-16. All GPUs 256K context (RTX 3090 .8, RTX 5070 .110, Strix Halo .15). Parallel 1 fleet-wide (Strix Halo handles compression solo).
configuration goes sideways, restore from this baseline. Last verified 2026-07-16. All GPUs 128K context (RTX 3090 .8, RTX 5070 .110, Strix Halo .15). Parallel 1 fleet-wide (compression via syslog-auto pool — switched from strix-moe 2026-07-18).
author: Abiba (pi agent)
---
+1 -1
View File
@@ -272,7 +272,7 @@ The following MUST be identical across ALL profiles:
### Rule 8: GPU Workload Distribution (UPDATED 2026-07-18)
- **RTX 3090 (24GB, 128K ctx, qwen3.6-27B-code)**: Heavy reasoning, code gen, long conversations — also handles ~55% of compression via syslog-auto pool
- **RTX 5070 (12GB, 128K ctx, gemma-4-12b)**: Vision, web search, quick tasks, web_extract — handles ~15% of compression via syslog-auto pool
- **Strix Halo (64GB, 128K ctx, Geneis Hermes V3 APEX)**: Agent reasoning, compression (~30% via syslog-auto pool), fallback for other GPUs
- **Strix Halo (64GB, 128K ctx, Genesis Hermes V3 APEX)**: Agent reasoning, compression (~30% via syslog-auto pool), fallback for other GPUs
- Agent profiles MUST route auxiliary tasks to the correct GPU:
- `auxiliary.vision.model: gpu-light` (RTX 5070)
- `auxiliary.web_extract.model: gpu-light` (RTX 5070)
+1 -1
View File
@@ -57,7 +57,7 @@ duration.
prefill time at 532 tok/s. Fix context first, routing second.
- **Route by task**: ornith for multi-step reasoning only; qwen for code/standard
queries; gemma for compression/auxiliary. Never send simple completion to a
queries; gemma for vision/web extraction; syslog-auto for compression (via weighted pool). Never send simple completion to a
35B MoE.
- **Compress aggressively**: threshold at 40% (not 65%) — a 256K window should
compact at 102K, not 166K. Target 15% tail (not 30%).
-200
View File
@@ -1,200 +0,0 @@
---
kind: responsibility
name: infrastructure-maintenance
description: >
Weekly system-level maintenance for the Syslog inference fleet: OS package
updates on the primary host, Docker image pulls for LiteLLM/SearXNG and other
running containers, container restarts with health verification, post-update
verification of every critical service (LiteLLM proxy, SearXNG, Zulip, Gitea,
PM2 processes, Hermes gateways), and rollback on failure. Consolidates the
raw shell scripts that previously did this piecemeal. This contract owns the
HOST-LEVEL weekly maintenance loop on the primary host plus Docker image
pulls ONLY for .116 and .7, while infrastructure-update owns the FULL-FLEET
cluster-wide wave (apt across the full PVE cluster + CTs/VMs AND its Docker
image Wave 3 across all stacks). Runs Sunday 2am ET. Owner:
ops (firstmate secondmate). Blast radius: an unverified image pull can break
LiteLLM (all agents lose inference) or SearXNG (search-stack down); a bad apt
upgrade can leave the host in a half-upgraded state. Pre-update backup check
and rollback are mandatory for this reason.
agent: ops
triggers:
- weekly (Sunday 02:00 ET) via cron
- on demand when ops/abiba triggers "infra maintenance"
version: 1.0.0
---
## Maintains
- maintenance-status: { phase: idle|preflight|apt|images|restarts|verify|rollback|done|failed, host, step, result, timestamp }
- image-baseline: { service, current_tag, pulled_tag, digest, updated_at } — last known-good image per container
- apt-state: { upgradable_before, upgradable_after, held_broken, kernel_reboot_required }
- health-baseline: snapshot of critical-service health captured pre-update (used for regression check post-update)
- rollback-snapshot: { backup_path, configs, image_digests, timestamp } — restore point created in preflight
- maintenance-history: array of past runs with phase results and any escalations
## Scope
Primary host is the maintenance host where apt updates apply. Docker image pulls
span the two Docker ecosystems that run critical services. infrastructure-update
runs the full-fleet cluster-wide wave (including its Wave 3 Docker pulls across
all stacks/hosts); this contract runs a narrower host-level weekly pull limited
to .116 and .7. Topology, CT IDs, and IPs are live-state fields — verify against
`infrastructure-control.prose.md` (the source of truth) and the live system
before mutating.
| Host | IP | Role | Trust |
|------|----|------|-------|
| CT 116 (syslog-api) | 192.168.68.116 | LiteLLM proxy + Grafana + Prometheus (inference harness) | ⚠️ VERIFY-BEFORE-USE |
| VM 109 (docker-vm) | 192.168.68.7 | SearXNG + Firecrawl + home stack (Docker host) | ⚠️ VERIFY-BEFORE-USE |
| CT 117 (zulip) | 192.168.68.19 | Zulip (storepve bridge IP .19) | ⚠️ VERIFY-BEFORE-USE |
| Gitea | https://git.sysloggh.net | Prose-contracts + agent configs source control | ⚠️ VERIFY-BEFORE-USE |
| CT 100 (abiba/pi) | 192.168.68.24 | PM2 processes (pi agent harness) | ⚠️ VERIFY-BEFORE-USE |
> "Primary host" for the apt phase is the host the ops agent runs maintenance
> from. Confirm which host that is against infrastructure-control before
> running; do not assume. If the ops agent is containerized/CT-based, apt runs
> inside that CT.
## Requires
- SSH/exec access to CT 116 (.116) and VM 109 (.7) for Docker operations
- `apt`, `docker`, `docker compose` available on target hosts
- LiteLLM master key available (Infisical vault, `LITELLM_API_KEY`) for health verification
- `infrastructure-monitoring` run completed within the last 30 minutes — provides the pre-update health baseline used by the regression check
- Writable backup directory `/tmp/infra-maintenance-backup-<date>/` on each mutated host
- Proxmox snapshot of the primary host available (or confirmed not required) before apt phase
## Continuity
- Self-driven: weekly cron `0 2 * * 0` (Sunday 02:00 ET)
- Also wakes on: explicit "infra maintenance" trigger from ops/abiba
- Depends on `infrastructure-monitoring` for the pre-update health baseline — do not run if the last monitoring run is stale (>30 min) or RED; abort and escalate instead
## Execution
### Phase 0 — Preflight (snapshot/backup check + health baseline)
1. **Capture health baseline** — run the `infrastructure-monitoring` postcondition checks (LiteLLM, Zulip, Gitea, SearXNG, Proxmox API) and record results as `health-baseline`. If any critical service is already down, **abort**: maintenance must not run on a degraded fleet.
2. **Backup check** — confirm a Proxmox snapshot of the primary host exists OR `/tmp/infra-maintenance-backup-<date>/` was created this run. Snapshot critical config files into the backup dir:
- `/opt/inference-harness/docker-compose.yml`, `/opt/inference-harness/litellm_config.yaml` (CT 116)
- `/opt/search-stack/searxng/docker-compose.yml`, `/opt/search-stack/firecrawl-source/docker-compose.yaml` (VM 109)
3. **Record image baseline**`docker inspect --format '{{.Image}} {{.Config.Image}}' <container>` for every running container on .116 and .7; store digests in `image-baseline` so rollback can restore them.
4. **Disk check**`df -h` on each mutated host; abort if free space <20% (apt upgrade + image pulls need headroom).
### Phase 1 — OS package updates (primary host)
```bash
# On the primary host only (VERIFY host against infrastructure-control first)
apt update
apt upgrade -y
```
- Capture `apt list --upgradable` before and after → store in `apt-state`.
- If apt reports held/broken packages (`apt-get -s upgrade | grep -i broken`, or non-zero exit), **stop** — do not force. Record `held_broken` and go to rollback/escalate.
- If `/var/run/reboot-required` exists after upgrade, flag `kernel_reboot_required: true` in `apt-state` but **do not reboot automatically** — that's a separate coordinated action (see infra-update Wave 4). Note it in the report.
### Phase 2 — Docker image pulls
Pull latest stable tags for every running container. Do NOT pin to `:main`/`:nightly` — use stable tags where the compose file specifies them; otherwise `latest`.
```bash
# CT 116 (.116) — inference harness
cd /opt/inference-harness && docker compose pull
# VM 109 (.7) — search + home stacks
cd /opt/search-stack/searxng && docker compose pull
cd /opt/search-stack/firecrawl-source && docker compose pull
# any other running stacks on .7 (home stack, audiobookshelf) — pull per their compose files
```
- LiteLLM and SearXNG are the two explicitly required pulls; "any other running containers" means every stack with a compose file on .116 and .7.
- Record pulled tag + digest per service in `image-baseline`.
### Phase 3 — Container restarts with health verification
Restart one stack at a time, verify health before moving to the next. Do not restart everything at once — a failure mid-wave must leave the rest running.
```bash
# CT 116
cd /opt/inference-harness && docker compose up -d
# VM 109
cd /opt/search-stack/searxng && docker compose up -d
cd /opt/search-stack/firecrawl-source && docker compose up -d
```
After each stack comes up, wait for health (max 120s):
- `docker ps` shows the container `Up` (and `healthy` if a healthcheck is defined)
- Service-specific probe passes (see Phase 4 probes)
If a stack fails to come up within 120s, **stop the wave** and go to rollback for that stack only; do not proceed to the next.
### Phase 4 — Post-update service verification
After ALL updates (apt + images + restarts), verify every critical service is back up and matches the pre-update baseline. This is the regression gate.
| Service | Probe | Expect |
|---------|-------|--------|
| LiteLLM proxy | `curl -sf http://192.168.68.116/litellm/v1/models` | 200 OK, models returned |
| LiteLLM MCP gateway | `curl -sf http://192.168.68.116:4000/mcp-rest/tools/list -H "Authorization: Bearer $MASTER_KEY"` | 90 tools (23 RA-H OS + 67 GitHub) |
| SearXNG | `curl -sf http://192.168.68.7:8888` | 200 OK |
| Zulip | `curl -sf https://chat.sysloggh.net/api/v1/server_settings` | 200 OK |
| Gitea | `curl -sf https://git.sysloggh.net/api/v1/version` | 200 OK |
| PM2 processes | `pm2 jlist` (CT 100) | all pi-agent processes `online` |
| Hermes gateways | SSH to Mumuni CT 114, Tanko CT 112; `systemctl is-active hermes-gateway` | `active` for each |
Regression check: every service that was GREEN in `health-baseline` must still be GREEN. A service that was already RED (and caused a preflight abort) is excluded — but Phase 0 should have aborted before we got here.
## Rollback Protocol
If ANY service in Phase 4 fails to come back up (or regresses vs baseline):
1. **Image rollback** — for the failing stack, restore the previous image:
```bash
# Restore from recorded image-baseline digest
docker compose down
# Pin the service image to the recorded digest in compose, then recreate
# image: <name>@sha256:<previous_digest>
docker compose pull && docker compose up -d
```
2. **APT rollback** — restore the primary host from the Proxmox snapshot taken/confirmed in Phase 0. If no snapshot, `apt install <pkg>=<old_version>` per package using apt history (`/var/log/apt/history.log`).
3. **Config rollback** — restore configs from `/tmp/infra-maintenance-backup-<date>/`.
4. **Re-verify** — re-run the Phase 4 probes on the rolled-back service. If still failing, escalate (do not loop — circuit breaker below).
5. **Escalate** — send a Zulip DM to abiba + mumuni with: failing service, phase, baseline vs current, rollback actions taken, backup path.
## Circuit Breaker
- `max_retries: 2` per failing phase — after 2 rollback attempts on the same service, stop and escalate.
- `window: 7200` seconds — no more than 2 retries within a 2-hour window.
- `trip_action: escalate_to_fatal` — when tripped, escalate to fatal (abiba + mumuni + kwame) and pause; a human must clear before the next scheduled run.
## Report
After completion (or on abort), emit a receipt (JSON) to `~/.hermes/runs/infrastructure-maintenance/` and send a Zulip DM summary:
```
🛠 Infrastructure Maintenance — YYYY-MM-DD
Phase: apt | images | restarts | verify | rollback
Primary host: <host>
APT: <N> packages upgraded, <M> held/broken, kernel_reboot_required=<bool>
Images pulled: LiteLLM <tag>, SearXNG <tag>, <others>
Services: all GREEN | <service> FAILED (rolled back)
Baseline regression: none | <details>
Backup: /tmp/infra-maintenance-backup-YYYYMMDD/
Escalation: none | warning | critical | fatal
```
## Verification Postconditions
- All critical services running after update (Phase 4 all GREEN)
- No regressions from pre-update health baseline (Phase 0 baseline)
- Docker containers on latest stable tags (`image-baseline.pulled_tag` recorded)
- APT packages up to date with no held broken packages (`apt-state.held_broken == 0`)
## Related Contracts
- `infrastructure-update.prose.md` — owns the full-fleet cluster-wide wave INCLUDING its Wave 3 Docker image updates across all stacks (SearXNG, Firecrawl, Inference Harness on .116, home stack, audiobookshelf); infrastructure-maintenance is a deliberately narrower host-level weekly pull scoped to .116 and .7.
- `infrastructure-monitoring.prose.md` — provides the pre-update health baseline (depends_on).
- `infrastructure-control.prose.md` — topology source of truth (CT IDs, IPs, hostnames).
- `litellm-health.prose.md` — LiteLLM probe details.
- `proxmox-monitor.prose.md` — Docker stats + monitoring stack health.
+11 -200
View File
@@ -1,9 +1,9 @@
---
kind: responsibility
name: zulip-health
description: Multi-platform health monitor for the Zulip messaging mesh spanning Platform A (Abiba pi), Platform B (Hermes agents Tanko/Mumuni/Koonimo/Koby), and Platform C (Agent Zero). Verifies bot registration, DM delivery, cross-platform connectivity, secret injection, and YAML config integrity.
description: Multi-platform health monitor for the Zulip messaging mesh spanning Platform A (Agent Zero Docker), Platform B (Hermes agents Tanko/Mumuni), and the Zulip bridge. Verifies bot registration, DM delivery, and cross-platform connectivity.
title: Zulip Mesh Health Monitor — Multi-Platform
version: 3.1.0
version: 3.0.0
runtime_contract: 2
agent: abiba
---
@@ -12,15 +12,11 @@ agent: abiba
Monitors ALL Zulip-connected agents across three platforms (pi, Hermes, Agent Zero).
Runs every 15 minutes in the background. Also triggers on session start.
v3.1.0 adds Koonimo+Koby to Platform B, Infisical dependency checks, config YAML
validation, stale PID/lock detection, Telegram adapter health, and the
cli_agent_setup_mixin patch verification.
## Requires
- **Zulip API key** for `abiba-bot@chat.sysloggh.net` in `$ZULIP_API_KEY`
- **SSH access** to Tanko (192.168.68.122), Mumuni (192.168.68.123), and Agent Zero Docker host (192.168.68.14)
- **SSH access to amdpve (192.168.68.15)** for `pct exec` access to Koonimo (CT 113) and Koby (CT 111)
- **PM2** on localhost for pi process management
- **Network access** to `chat.sysloggh.net`, `localhost:9200`
- **Write access** to `/root/zulip-health-monitor.log` and `/tmp/zulip-monitor-debounce`
@@ -53,22 +49,12 @@ cli_agent_setup_mixin patch verification.
"zulip_state": "connected",
"heartbeat_age_seconds": 45,
"gateway_pid": 1234,
"infisical_present": true,
"config_valid": true,
"telegram_state": "connected",
"no_key_required_count": 0,
"edit_fail_rate_pct": 0,
"severity": "healthy"
}
}
```
New fields in v3.1.0:
- `infisical_present` — /usr/local/bin/infisical exists on the agent CT
- `config_valid` — /root/.hermes/config.yaml passes YAML validation
- `telegram_state` — Telegram adapter status from gateway_state.json
- `no_key_required_count` — count of `no-key-required` in gateway logs
### Postconditions
- Every platform is independently checked; one failure doesn't block others
@@ -95,13 +81,13 @@ Log as "unreachable" — don't treat as critical unless it persists for 3+ conse
## Streaming Support (2026-07-05)
Zulip agents now support progressive message editing during agent generation.
When a Hermes agent (Tanko, Mumuni, Koonimo, Koby) processes a message, the
response is streamed in real-time via Zulip's `PATCH /api/v1/messages/{id}` API:
When a Hermes agent (Tanko, Mumuni) processes a message, the response is
streamed in real-time via Zulip's `PATCH /api/v1/messages/{id}` API:
- Adapter implements `edit_message()` using `_api_patch()` helper
- Gateway stream consumer progressively edits the Zulip message
- User sees real-time agent thinking instead of waiting for full response
- Verified: Tanko (CT 112), Mumuni (CT 114), Koonimo (CT 113), Koby (CT 111)
- Verified: Tanko (CT 112) and Mumuni (CT 114) both have streaming active
### Verification
```bash
@@ -196,217 +182,50 @@ grep -a "Finalized\|Failed to finalize" /root/.pm2/logs/abiba-zulip-out.log | ta
| `last_error` set | Log and monitor |
| Crash loop >10/h | Alert user |
### Step 3: Platform B — Hermes (Tanko .122, Mumuni .123, Koonimo .114, Koby .129)
Platform B now monitors four Hermes agents:
- Tanko (CT 112, 192.168.68.122) — Zulip + Telegram
- Mumuni (CT 114, 192.168.68.123) — Zulip + Telegram + Email
- Koonimo (CT 113, 192.168.68.114, hostname "baggy") — Zulip + Telegram
- Koby (CT 111, 192.168.68.129, hostname "tdunna") — Zulip + Telegram
SSH access: Koonimo is reachable at .114; Koby has no direct SSH. Use `pct exec`
from amdpve as the primary access method for both:
```bash
ssh root@192.168.68.15 "pct exec 113 -- <command>" # Koonimo (or ssh .114)
ssh root@192.168.68.15 "pct exec 111 -- <command>" # Koby (pct exec only)
```
### Step 3: Platform B — Hermes (Tanko .122, Mumuni .123)
**B1: Gateway State**
```bash
ssh root@192.168.68.122 "cat ~/.hermes/gateway_state.json"
ssh root@192.168.68.123 "cat ~/.hermes/gateway_state.json"
ssh root@192.168.68.15 "pct exec 113 -- cat /root/.hermes/gateway_state.json"
ssh root@192.168.68.15 "pct exec 111 -- cat /root/.hermes/gateway_state.json"
```
Check `platforms.zulip.state`: `connected` ✅ | `disconnected` ❌ | `error` ❌.
**B1.5: Infisical Dependency Check**
```bash
ssh root@<CT> "test -f /usr/local/bin/infisical && echo OK || echo MISSING"
# pct exec variant for Koonimo/Koby:
ssh root@192.168.68.15 "pct exec 113 -- test -f /usr/local/bin/infisical && echo OK || echo MISSING"
ssh root@192.168.68.15 "pct exec 111 -- test -f /usr/local/bin/infisical && echo OK || echo MISSING"
```
If MISSING → flag `infisical_present: false`, note as degraded — gateway cannot
auto-start on reboot without the infisical binary.
Check `platforms.zulip.state`: `connected` ✅ | `disconnected` ❌ | `error` | missing → not installed.
**B2: Agent Process**
```bash
ssh root@<CT> "ps aux | grep 'gateway run' | grep -v grep"
# pct exec variant:
ssh root@192.168.68.15 "pct exec 113 -- ps aux | grep 'gateway run' | grep -v grep"
ssh root@192.168.68.15 "pct exec 111 -- ps aux | grep 'gateway run' | grep -v grep"
```
Gateway PID should exist with uptime > 60s. Check for stale PIDs:
- `gateway.pid` and `gateway.lock` files that reference a dead process
- Multiple gateway processes (duplicate PIDs)
**B2.5: Config YAML Validation**
```bash
ssh root@<CT> "python3 -c 'import yaml; yaml.safe_load(open(\"/root/.hermes/config.yaml\"))' 2>&1"
# pct exec variant:
ssh root@192.168.68.15 "pct exec 113 -- python3 -c 'import yaml; yaml.safe_load(open(\"/root/.hermes/config.yaml\"))' 2>&1"
ssh root@192.168.68.15 "pct exec 111 -- python3 -c 'import yaml; yaml.safe_load(open(\"/root/.hermes/config.yaml\"))' 2>&1"
```
Expected: no output (clean parse). If parse fails → flag `config_valid: false`,
degraded — gateway is running on stale in-memory config.
Check specifically for:
- Stray `api_key: sk-...` lines indented under `api_key_env` entries in
`custom_providers` section (hardcoded keys violate hermes-key-enforcement)
- Indentation errors in `custom_providers`, `auxiliary`, or `compression` blocks
Gateway PID should exist with uptime > 60s.
**B3: Heartbeat Verification**
```bash
ssh root@<CT> "grep Heartbeat ~/.hermes/logs/agent.log | tail -3"
# pct exec variant:
ssh root@192.168.68.15 "pct exec 113 -- grep Heartbeat /root/.hermes/logs/agent.log | tail -3"
ssh root@192.168.68.15 "pct exec 111 -- grep Heartbeat /root/.hermes/logs/agent.log | tail -3"
```
Expected: recent heartbeat (within 5 min), `polls=N` incrementing.
Silence > 300s → warning. Silence > 600s → critical.
**B3.5: Stale PID/Lock Detection**
Before any restart action, check for stale pid/lock files:
```bash
ssh root@<CT> "ls -la /root/.hermes/gateway.pid /root/.hermes/gateway.lock 2>/dev/null"
# pct exec variant:
ssh root@192.168.68.15 "pct exec 113 -- ls -la /root/.hermes/gateway.pid /root/.hermes/gateway.lock 2>/dev/null"
ssh root@192.168.68.15 "pct exec 111 -- ls -la /root/.hermes/gateway.pid /root/.hermes/gateway.lock 2>/dev/null"
```
If gateway process is dead (no PID) but pid/lock files exist:
```bash
ssh root@<CT> "rm -f /root/.hermes/gateway.pid /root/.hermes/gateway.lock"
# pct exec variant:
ssh root@192.168.68.15 "pct exec 113 -- rm -f /root/.hermes/gateway.pid /root/.hermes/gateway.lock"
ssh root@192.168.68.15 "pct exec 111 -- rm -f /root/.hermes/gateway.pid /root/.hermes/gateway.lock"
```
Pid/lock files blocking restart → clear them before restart attempt.
**B4: Response Delivery**
```bash
ssh root@<CT> "grep -E 'Finalized|Failed to finalize|Replied to' ~/.hermes/logs/agent.log | tail -10"
# pct exec variant for Koonimo/Koby:
ssh root@192.168.68.15 "pct exec 113 -- grep -E 'Finalized|Failed to finalize|Replied to' /root/.hermes/logs/agent.log | tail -10"
ssh root@192.168.68.15 "pct exec 111 -- grep -E 'Finalized|Failed to finalize|Replied to' /root/.hermes/logs/agent.log | tail -10"
```
> 50% fail rate → critical.
**B4.5: LiteLLM Key Injection Verification**
Check gateway logs for `no-key-required` failure pattern (indicates the
cli_agent_setup_mixin.py patch is missing):
```bash
ssh root@<CT> "grep -c 'no-key-required' /root/.hermes/logs/gateway.log 2>/dev/null || echo 0"
# pct exec variant for Koonimo/Koby:
ssh root@192.168.68.15 "pct exec 113 -- grep -c 'no-key-required' /root/.hermes/logs/gateway.log 2>/dev/null || echo 0"
ssh root@192.168.68.15 "pct exec 111 -- grep -c 'no-key-required' /root/.hermes/logs/gateway.log 2>/dev/null || echo 0"
```
If > 0 → flag `no_key_required_count: <count>`, note as degraded — provider
requests silently fall back to `no-key-required` when LITELLM_API_KEY env var
resolves empty.
**B5: Telegram Adapter Health**
Check Telegram connectivity in gateway state or logs:
```bash
# From gateway_state.json (all agents):
ssh root@<CT> "cat ~/.hermes/gateway_state.json | python3 -c 'import json,sys;d=json.load(sys.stdin);print(d[\"platforms\"].get(\"telegram\",{}).get(\"state\",\"missing\"))'"
# pct exec variant for Koonimo/Koby (cat the file; read platforms.telegram.state):
ssh root@192.168.68.15 "pct exec 113 -- cat /root/.hermes/gateway_state.json"
ssh root@192.168.68.15 "pct exec 111 -- cat /root/.hermes/gateway_state.json"
# From logs (check for stuck DNS resolution):
ssh root@<CT> "grep -E 'Telegram.*Connecting|Telegram.*Connected|attempt 1/8' /root/.hermes/logs/gateway.log | tail -5"
ssh root@192.168.68.15 "pct exec 113 -- grep -E 'Telegram.*Connecting|Telegram.*Connected|attempt 1/8' /root/.hermes/logs/gateway.log | tail -5"
ssh root@192.168.68.15 "pct exec 111 -- grep -E 'Telegram.*Connecting|Telegram.*Connected|attempt 1/8' /root/.hermes/logs/gateway.log | tail -5"
```
Telegram states: `connected` ✅ | `disconnected` ❌ | `retrying` ⚠️ | `fatal` ❌ | `paused` ⚠️
If stuck on "attempt 1/8" for > 60s → flag Telegram as degraded (Zulip may
still be fine — do NOT treat as Zulip outage).
**B6: cli_agent_setup_mixin.py Patch Verification**
Check whether the `no-key-required` fallback string exists without the
LiteLLM-specific guard (only needed when LiteLLM key injection failures are
suspected). A bare count of the fallback string cannot distinguish a guarded
occurrence from an unguarded one, so count both the fallback string and the
LiteLLM guard token:
```bash
ssh root@<CT> "f=\$(grep -c 'no-key-required' /usr/local/lib/hermes-agent/hermes_cli/cli_agent_setup_mixin.py 2>/dev/null || echo 0); g=\$(grep -c 'LITELLM_API_KEY' /usr/local/lib/hermes-agent/hermes_cli/cli_agent_setup_mixin.py 2>/dev/null || echo 0); echo fallback=\$f guard=\$g"
# pct exec variant:
ssh root@192.168.68.15 "pct exec 113 -- bash -c 'f=\$(grep -c no-key-required /usr/local/lib/hermes-agent/hermes_cli/cli_agent_setup_mixin.py 2>/dev/null || echo 0); g=\$(grep -c LITELLM_API_KEY /usr/local/lib/hermes-agent/hermes_cli/cli_agent_setup_mixin.py 2>/dev/null || echo 0); echo fallback=\$f guard=\$g'"
ssh root@192.168.68.15 "pct exec 111 -- bash -c 'f=\$(grep -c no-key-required /usr/local/lib/hermes-agent/hermes_cli/cli_agent_setup_mixin.py 2>/dev/null || echo 0); g=\$(grep -c LITELLM_API_KEY /usr/local/lib/hermes-agent/hermes_cli/cli_agent_setup_mixin.py 2>/dev/null || echo 0); echo fallback=\$f guard=\$g'"
```
If `fallback > 0` and `guard == 0` → the fallback string is present without
the LiteLLM-specific guard, so the patch is missing. If unsure, inspect each
occurrence with `grep -n -B2 -A2 'no-key-required'` to confirm the guard
wraps it.
**Platform B Actions**
| Condition | Action |
|-----------|--------|
| `zulip.state != "connected"` | Restart gateway (see B2 restart commands below) |
| No heartbeat in 10min | Restart gateway |
| `zulip.state != "connected"` | `ssh root@<CT> "pkill -f 'gateway run'; sleep 2; hermes gateway restart"` |
| No heartbeat in 10min | Same as above |
| `Failed to finalize` > 50% | Check PATCH API, Zulip server |
| Response empty/short | Check A2A endpoint / LiteLLM model |
| `infisical_present: false` | Flag as degraded — log and alert; do NOT auto-restart via the standard command (it cannot inject the key without infisical). Operator may use the Infisical-missing fallback restart below. |
| `config_valid: false` | Flag as degraded — alert user, gateway running on stale config |
| Stale pid/lock files detected | Clean files before restart |
| `no_key_required_count > 0` | Flag as degraded — check LITELLM_API_KEY injection |
| Telegram stuck on attempt 1/8 | Flag Telegram as degraded, no Zulip action needed |
**Restart Commands**
Standard restart (Infisical present):
```bash
ssh root@<CT> "pkill -f 'gateway run'; sleep 2; hermes gateway restart"
# pct exec variant:
ssh root@192.168.68.15 "pct exec 113 -- bash -c 'pkill -f \"gateway run\"; sleep 2; systemctl restart hermes-gateway'"
ssh root@192.168.68.15 "pct exec 111 -- bash -c 'pkill -f \"gateway run\"; sleep 2; systemctl restart hermes-gateway'"
```
The mechanisms differ by design: Tanko/Mumuni launch the gateway through a
wrapper loop, so the `hermes gateway restart` CLI is the correct entry point
(it re-arms the wrapper). Koonimo/Koby run the gateway as a systemd unit
(`hermes-gateway.service`), so `systemctl restart hermes-gateway` is the
correct entry point under `pct exec`. Do not swap the two — using the CLI on
Koonimo/Koby would bypass the unit, and using `systemctl` on Tanko/Mumuni
would miss the wrapper loop.
Fallback: Infisical missing → start gateway directly from venv with env vars:
```bash
ssh root@<CT> "source /usr/local/lib/hermes-agent/venv/bin/activate && \
export LITELLM_API_KEY=\$(grep -E '^LITELLM_API_KEY=' /root/.hermes/.env | cut -d= -f2-) && \
export ZULIP_API_KEY=\$(grep -E '^ZULIP_API_KEY=' /root/.hermes/.env | cut -d= -f2-) && \
cd /root/.hermes && nohup hermes gateway run > logs/gateway-manual-start.log 2>&1 &"
# pct exec variant:
ssh root@192.168.68.15 "pct exec 113 -- bash -c 'source /usr/local/lib/hermes-agent/venv/bin/activate; export LITELLM_API_KEY=\$(grep -E ^LITELLM_API_KEY= /root/.hermes/.env | cut -d= -f2-); export ZULIP_API_KEY=\$(grep -E ^ZULIP_API_KEY= /root/.hermes/.env | cut -d= -f2-); cd /root/.hermes; nohup hermes gateway run > logs/gateway-manual-start.log 2>&1 &'"
ssh root@192.168.68.15 "pct exec 111 -- bash -c 'source /usr/local/lib/hermes-agent/venv/bin/activate; export LITELLM_API_KEY=\$(grep -E ^LITELLM_API_KEY= /root/.hermes/.env | cut -d= -f2-); export ZULIP_API_KEY=\$(grep -E ^ZULIP_API_KEY= /root/.hermes/.env | cut -d= -f2-); cd /root/.hermes; nohup hermes gateway run > logs/gateway-manual-start.log 2>&1 &'"
```
### Step 4: Platform C — Agent Zero (kagentz, CT 105 via Docker host .14)
@@ -459,7 +278,7 @@ Expected: task ID with "working" status. Poll for completion with `tasks/get`.
Check each agent's log for excessive bot-to-bot chatter:
- Abiba: `Skipped.*bot msgs` count
- Tanko/Mumuni/Koonimo/Koby: Repeated DM exchanges between bots
- Tanko/Mumuni: Repeated DM exchanges between bots
- kagentz: Adapter log for bot DMs being processed
If any bot processes >50 bot-originated messages in 15min → warning.
@@ -482,14 +301,6 @@ Track via `/tmp/zulip-monitor-debounce` (unix timestamp of last restart).
## History
### v3.1.0 (2026-07-23) — Koonimo Outage Lessons
Added Koonimo (CT 113) and Koby (CT 111) to Platform B monitoring.
Added Infisical dependency check, config YAML validation, stale PID/lock
detection, Telegram adapter health, LiteLLM key injection verification, and
cli_agent_setup_mixin.py patch verification. Updated restart commands with
Infisical-missing fallback path.
### Gen 5 (2026-07-02) — Rate Limit Death Spiral Fix
**Root Cause**: Proactive Queue Rotation at 25 min triggered queue re-registration every cycle. Each re-registration + retry loop (3 attempts) + monitor restart = 8-12 API calls per cycle. Combined with monitor's own API calls (server check, stream alerts), `abiba-bot` hit Zulip's rate limit (429 RATE_LIMIT_HIT). Each restart reset the cycle, creating a death spiral: 111 restarts in 24 hours.