--- kind: responsibility name: zulip-health description: Multi-platform health monitor for the Zulip messaging mesh spanning Platform A (pi/Abiba Zulip bridge), Platform B (Tanko on DSH), and Platform C (Agent Zero Docker). Verifies bot registration, DM delivery, and cross-platform connectivity. The kagentz Zulip adapter leg is retired (its code no longer exists) — Agent Zero is probed for A2A liveness, A2A response, and public-path access only. Mumuni is no longer monitored from this host — she runs on her own container (kagentz CT 105 on minipve, .14) and is monitored on her side. title: Zulip Mesh Health Monitor — Multi-Platform version: 3.4.0 runtime_contract: 2 agent: abiba report_only_agents: - koby # ⛔ KOBY IS NEVER REPAIRED (Rule 17, 2026-08-17) — detect + report, never fix on .129 --- # Zulip Mesh Health Monitor Monitors the Zulip-connected agents under this host's operational control (pi, DSH, Agent Zero). Runs every 15 minutes in the background. Also triggers on session start. > **Mumuni is NOT monitored from this host (captain ruling 2026-09-10).** She > moved off this host onto her own container — kagentz CT 105 on minipve > (192.168.68.14), running a dedicated `hermes` user — and is monitored on her > side. No step in this contract, and no leg of `scripts/zulip-monitor.sh`, may > ssh to her old deployment, read her `~/.hermes/gateway_state.json`, or alert on > her state. The former Platform-B-for-Mumuni steps (gateway process, heartbeat, > response delivery) are retired: they always read "unknown" against the > decommissioned deployment and produced a false 🔴 alert on every run. ## Requires - **Zulip API key** for `abiba-bot@chat.sysloggh.net` in `$ZULIP_API_KEY` - **SSH access** to minipve (192.168.68.12) for Tanko — CT 112 reached via `pct exec` (direct SSH to .122 is not a dependency of this contract: per-worker key availability varies); and the Agent Zero Docker host (192.168.68.14) - **PM2** on localhost for pi process management - **Network access** to `chat.sysloggh.net`, `kagentz.sysloggh.net` (C3 public path), `localhost:9200` - **Write access** to `/root/zulip-health-monitor.log` and `/tmp/zulip-monitor-debounce` - **Relay access** via RA-H OS MCP for alert delivery ## Maintains - zulip_server_status: "healthy" | "down" - agents: map of per-platform agent snapshots (see schema below) - last_check: timestamp — When the last full diagnostic ran - overall_severity: "healthy" | "degraded" | "critical" - restart_debounce: timestamp — Last restart action (enforces 300s minimum) ### Per-Agent Snapshot Schema ```json { "abiba": { "platform": "pi", "connected": true, "response_pipeline": "healthy", "queue_healthy": true, "pm2_status": "online", "echo_loop_bot_msgs_15min": 12, "edit_fail_rate_pct": 3, "severity": "healthy" }, "tanko": { "platform": "dsh", "service_state": "active", "http_status": 200, "severity": "healthy" } } ``` ### Postconditions - Every platform is independently checked; one failure doesn't block others - Restart actions respect 300s debounce window - Critical conditions generate relay alerts to user - All checks logged to `/root/zulip-health-monitor.log` with timestamps ## Strategies ### When Zulip server is unreachable Skip all per-platform checks — they'll all fail downstream. Report "Zulip server down" and alert. ### When all agents are silent Differentiate: if Zulip server returns 200, likely a shared infrastructure issue (Netbird, DNS). If server is down, it's upstream — wait. ### When restart is indicated but debounce window hasn't passed Log the condition as "pending restart" with the timestamp. If the condition persists after debounce window, apply restart. Never bypass debounce for non-critical conditions. ### When a platform agent is unreachable via SSH Log as "unreachable" — don't treat as critical unless it persists for 3+ consecutive checks. ### Severity escalation ## Streaming Support (2026-07-05) Zulip agents now support progressive message editing during agent generation. When a Zulip agent under this monitor's scope (Tanko on DSH) processes a message, the response is streamed in real-time via Zulip's `PATCH /api/v1/messages/{id}` API: - Adapter implements `edit_message()` using `_api_patch()` helper - Gateway stream consumer progressively edits the Zulip message - User sees real-time agent thinking instead of waiting for full response - Verified: Tanko (CT 112) has streaming active; Mumuni's (kagentz CT 105) is verified on her own host, not from here ### Verification ```bash # Check if agent has streaming: grep -c "async def edit_message" ~/.hermes/plugins/*/zulip*/adapter.py # Must return 1 — streaming is active ``` - Single agent warning → log only - Single agent critical → relay message to user - Two or more agents critical → immediate relay + attempt auto-recovery - All agents critical + server up → Netbird/DNS likely down ## Invariants - **Never restart more than once per 300s** per agent - **Never restart if PM2 crashes > 10/h** — alert user instead - **Never send duplicate alerts** — check last alert timestamp before relaying - **Health checks are read-only** — diagnostics don't mutate state except for logged restarts - **Respect Zulip API rate limits** — no more than 200 requests in rapid succession ## Continuity - **On session start**: Run full diagnostic pass - **Every 15 minutes**: Scheduled background check while Abiba is running - **On `zulip-status` command**: Run on-demand and report to user - **On critical alert**: Escalate to relay message immediately, don't wait for schedule ## Execution **Execution model**: This contract is executed by a host-scheduled cron job (see `scripts/contract-run.sh`). The cron job runs the monitoring script directly on the target host and appends the result to `/var/log/contract-runs/.log`. An agent-session acknowledgement (a `done:` line in the ops status log) is NOT execution — it only proves the agent read the result and reported it. The actual monitoring work happens in the host cron job. ### Liveness rule (scoped) Any HTTP response proves the service is ALIVE. For auth-gated endpoints (Zulip API, Tanko gateway), a 401/403 redirect or status means the service answered — report the code, never "down". Only a failed CONNECTION (curl status 000, timeout, refused) is a failed probe. **Proxy-fronted exception:** a reverse proxy's `502`/`504` (upstream refused/unreachable) is not the backend answering — for proxy-fronted public endpoints such as `https://kagentz.sysloggh.net/` (C3), treat it as DOWN. **STANDING PROBE RULES (2026-09-14, from defect report 1150.msg):** 1. **Any HTTP status means ALIVE.** 200, 301, 302, 401, 403, 404 all prove the service answered — report the code, never "down". A redirect is not a failure. Only a failed CONNECTION (curl status 000, timeout, refused) is a failed probe (proxy `502`/`504` excepted, above). 2. **A failed probe is never a service verdict.** Print `probe-failed: ` naming the exact URL/host/port and the failure kind (timeout, refused, no-route, dns), retry once at a longer timeout, and only then report. 3. **Say which probe produced each number.** "API: 000" is unusable; "API https://chat.sysloggh.net/api/v1/server_settings -> connection timeout after 10s (retried at 25s: also timeout)" is actionable. ### Step 1: Zulip Server Liveness ```bash # Probe the Zulip API (authenticated, any HTTP status = ALIVE) code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 https://chat.sysloggh.net/api/v1/server_settings -u 'abiba-bot@chat.sysloggh.net:$ZULIP_API_KEY') if [ "$code" == "000" ]; then code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 25 https://chat.sysloggh.net/api/v1/server_settings -u 'abiba-bot@chat.sysloggh.net:$ZULIP_API_KEY') echo "Zulip API https://chat.sysloggh.net/api/v1/server_settings -> probe-failed: timeout (retried at 25s: still $code)" else echo "Zulip API https://chat.sysloggh.net/api/v1/server_settings -> $code" fi ``` Expected: `200` (authenticated). Any HTTP status = ALIVE; only 000/timeout = probe-failed. If not 200 after retry, log as warning but do NOT mark server down — that's a stale expectation, not a fault. ### Step 2: Platform A — pi (Abiba, localhost) **A1: Health Endpoint** ```bash # Probe the Abiba extension health endpoint (loopback, any HTTP status = ALIVE) code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 http://127.0.0.1:9200/health) if [ "$code" == "000" ]; then code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 25 http://127.0.0.1:9200/health) echo "Abiba extension http://127.0.0.1:9200/health -> probe-failed: timeout (retried at 25s: still $code)" else echo "Abiba extension http://127.0.0.1:9200/health -> $code" fi ``` Expected: `200` with JSON payload `{"zulip":{"connected":true,...}}`. Any HTTP status = ALIVE; only 000/timeout = probe-failed. **NOTE: This probe MUST run on the Abiba host (CT 100) where 127.0.0.1:9200 is the extension. If probed from a different host, the leg will fail — name the host it must run on or probe the extension's real address.** Check the JSON payload: | Field | Healthy | Critical | |-------|---------|----------| | `connected` | `true` | `false` | | `response_pipeline` | `healthy` | `blocked` | | `queue_healthy` | `true` | `false` | | `stuck` | `false` | `true` | | `idle_seconds` | < 1800 | ≥ 1800 | | `pending_count` | 0 | > 0 with `response_pipeline: blocked` | | `agent_busy_duration_seconds` | < 300 | ≥ 600 | | `last_error` | `null` | non-null string | | `retry_count` | 0–2 | 3+ | | `queue_id` | non-null string | `null` | **A2: PM2 Process** ```bash pm2 show abiba-zulip --no-color 2>/dev/null ``` Check: `status=online`, `restarts < 10/h`, `uptime > 60s`. **A3: Echo Loop Detection** ```bash grep -a "Skipped.*bot msgs" /root/.pm2/logs/abiba-zulip-out.log | tail -5 ``` > 100 skipped in 15min → info only (echo loop prevention working). **A4: Response Delivery** ```bash grep -a "Finalized\|Failed to finalize" /root/.pm2/logs/abiba-zulip-out.log | tail -20 ``` > 50% fail rate → critical — check editMessage API. **Platform A Actions** | Condition | Action | |-----------|--------| | `response_pipeline: blocked` | `pm2 restart abiba-zulip` | | `connected: false` | `pm2 restart abiba-zulip` | | `queue_healthy: false` | `pm2 restart abiba-zulip` | | `stuck: true` | `pm2 restart abiba-zulip` | | `retry_count >= 3` | `pm2 restart abiba-zulip` | | `response_pipeline: degraded` | Monitor — no action, watchdog handles | | `last_error` set | Log and monitor | | Crash loop >10/h | Alert user | ### Step 3: Platform B — Tanko (DSH on minipve CT 112) Mumuni is out of scope for this host (see the note above): she runs on her own container and is monitored on her side. Tanko runs on DSH (DeepSeek Harness) — it no longer runs a Hermes gateway, so there is no `~/.hermes/gateway_state.json` on CT 112. Tanko's Zulip gateway runs as the `dsh-web` systemd unit inside **CT 112**, which resides on the **minipve** PVE host (**192.168.68.12**). Direct SSH to 192.168.68.122 is not a dependency of this contract — per-worker key availability varies — so CT 112 probes run from the minipve vantage via `pct exec`: ```bash ssh root@192.168.68.12 "pct exec 112 -- " ``` > **By design (verified 2026-09-08):** the `dsh-web` gateway binds > `127.0.0.1:3080` **loopback-only**. A remote probe against > `192.168.68.122:3080` gets connection-refused — that is EXPECTED, NOT a fault, > and must never be raised as Tanko down. Only loopback probes from inside > CT 112 (or the public-URL fallback below) are valid health signals. **B1: Gateway Service State (Tanko)** ```bash ssh root@192.168.68.12 "pct exec 112 -- systemctl is-active dsh-web" ``` Expected: `active`. Anything else → gateway service down → apply the Tanko heal (restart via DSH service, Platform B Actions table below). **B2: Gateway HTTP Liveness (Tanko — loopback-only :3080)** ```bash ssh root@192.168.68.12 "pct exec 112 -- curl -s --connect-timeout 5 --max-time 10 -o /dev/null -w '%{http_code}' http://127.0.0.1:3080/" ``` Alive = **ANY** HTTP status response from the endpoint — the expected set is `200`/`301`/`302`/`307`/`308`/`401`/`403` (the gateway UI is token-gated and legitimately answers with redirects/auth-challenges, so never require a bare `200`), and any other status, including `404`/`5xx`, also counts alive: a process answering `503` is running and self-heal must NOT restart-loop it. Down = connection refused (`000`) or timeout only. Statuses outside the expected set are logged/reported as a warning — reported, never healed on. **B3: Public-URL Fallback Probe (Tanko — for nodes without pct/ssh access to minipve)** ```bash curl -s --connect-timeout 10 --max-time 15 -o /dev/null -w '%{http_code}' https://tankodhs.sysloggh.net/ ``` Fallback only — used when the monitoring node has no pct/SSH path to minipve. Alive = **ANY** HTTP status response from the endpoint — healthy signals are `302` (authentik proxy-auth redirect) and `401` (auth-gated), and any other status, including `404`/`5xx`, also counts alive: the endpoint is up and answering and must NOT be restart-looped. Down = connection refused (`000`) or timeout only. Never expect a bare `200` — the public URL terminates in the token-gated authentik chain. Statuses outside the healthy set are logged/reported as a warning — reported, never healed on. **Platform B Actions** | Condition | Action | |-----------|--------| | `dsh-web` service not `active` | Restart Tanko via DSH service | | HTTP `:3080` connection refused/timeout (`000`) | Same as above | | HTTP status outside the expected set | Log/report as a warning — reported, never healed on | **B4: dsh-web Authentication (Tanko — restart-persistent login)** The dsh-web UI is token-gated. On every start the process prints a random launch token to the journal: ``` dsh web: http://127.0.0.1:3080/?token= ``` The token only bootstraps an authority-bound, HMAC-signed browser cookie with a 30-day lifetime. The signing secret is durable in `/root/.dsh/.credentials.yaml` (key `client-connection/browser-session`), so a cookie minted once keeps working across `dsh-web` restarts; the launch token itself rotates on every restart. **Login endpoint (public, Authentik-gated):** `https://tankodhs.sysloggh.net/dsh-web-login` It lives inside the Authentik-gated `:80` server block (`/etc/nginx/sites-available/dsh`, symlinked from `/etc/nginx/sites-enabled/dsh`) as `location = /dsh-web-login`, guarded by `auth_request /outpost.goauthentik.io/auth/nginx`. It proxies to dsh-web with `Host: tankodhs.sysloggh.net`, so the minted cookie is bound to the public authority — never to `127.0.0.1:3080`. The token-dependent line is isolated in the generated include `/etc/dsh-web/nginx-login.conf`: ``` proxy_pass http://127.0.0.1:3080/?token=; ``` **Token refresh (non-disruptive):** `/opt/deepseek-harness/capture-dsh-token.sh` (source: `scripts/capture-dsh-token.sh`) reads candidate launch tokens from the journal **scoped to the service's current systemd invocation** (`systemctl show -p InvocationID` + `_SYSTEMD_INVOCATION_ID=`), re-sampling the invocation on every pass so a restart that lands during the wait switches to the new invocation; a restarted process's stale token is never considered while its new startup banner is still pending and there is no whole-journal or cross-invocation fallback. Each candidate is then functionally verified against dsh-web with `Host: tankodhs.sysloggh.net`, using the first the running process accepts with `303`. It waits up to 120s for a restarted process to accept a token and re-probes every current-invocation candidate on each pass, so a token that briefly returns `000` while the service is still starting is not disqualified. If none is accepted it leaves the include untouched and exits so the timer retries (exiting non-zero when a pending reload is still outstanding). It writes `/etc/dsh-web/launch-token` and regenerates `/etc/dsh-web/nginx-login.conf`, reloading nginx only when the on-disk include differs from the generated one or the applied-state stamp does not match the token (`nginx -t` guards the reload, and the stamp is written only after a successful `nginx -s reload`, so a failed or interrupted reload is retried on the next run). Any failed reload records a pending-reload marker under `/etc/dsh-web/`; the next run attempts the reload before the token wait, independent of token state, and clears the marker only once the reload succeeds, so a disabled legacy `:8081` file can never leave the running nginx unreloaded. The generated include is recreated before any `nginx -t` if it is missing, so a failed run cannot wedge recovery. Runs are serialized with `flock` on `/run/capture-dsh-token.lock`. It **never stops or starts `dsh-web`**. It is triggered by the `dsh-web.service` drop-in `/etc/systemd/system/dsh-web.service.d/20-token-refresh.conf` (`ExecStartPost=/bin/systemctl --no-block start dsh-web-token.service`) and by `dsh-web-token.timer` every 2 minutes for reconciliation.
Installed systemd wiring (CT 112) ```ini # /etc/systemd/system/dsh-web-token.service [Unit] Description=Refresh the dsh-web launch token for the nginx login endpoint After=dsh-web.service [Service] Type=oneshot TimeoutStartSec=180 ExecStart=/opt/deepseek-harness/capture-dsh-token.sh # /etc/systemd/system/dsh-web-token.timer [Unit] Description=Periodically refresh the dsh-web login token [Timer] OnBootSec=90s OnUnitActiveSec=120s AccuracySec=10s Persistent=true [Install] WantedBy=timers.target # /etc/systemd/system/dsh-web.service.d/20-token-refresh.conf [Service] ExecStartPost=/bin/systemctl --no-block start dsh-web-token.service ```
> **Do NOT reintroduce the `:8081` endpoint.** It listened on `0.0.0.0:8081` > with no `auth_request` and was a full Authentik bypass for anyone on the LAN. > The script now removes `/etc/nginx/sites-enabled/dsh.token` automatically if > it ever reappears. **Authentication flow:** 1. `GET https://tankodhs.sysloggh.net/dsh-web-login` 2. Unauthenticated → Authentik sign-in; once authenticated the request reaches dsh-web with `Host: tankodhs.sysloggh.net`. 3. dsh-web accepts the launch token on `GET /`, writes the `dsh-auth-` cookie (30 days, `HttpOnly`, `SameSite=Strict`) and returns `303` to `/`. 4. Every later request through `/` presents that cookie; the token is not needed again until the cookie expires or a new browser is used. **Verification** (minipve vantage): ```bash # 1. Login endpoint is Authentik-gated: unauthenticated -> 302 (not 200/303). ssh root@192.168.68.12 "pct exec 112 -- curl -s -o /dev/null -w '%{http_code}\n' \ -H 'Host: tankodhs.sysloggh.net' http://127.0.0.1/dsh-web-login" # Expected: 302 # 2. Legacy :8081 endpoint is gone (connection refused -> 000). ssh root@192.168.68.12 "pct exec 112 -- curl -s --max-time 3 -o /dev/null \ -w '%{http_code}\n' http://192.168.68.122:8081/" # Expected: 000 # 3. Backend cookie mint + reuse (exactly what /dsh-web-login proxies to). TOKEN=$(ssh root@192.168.68.12 "pct exec 112 -- cat /etc/dsh-web/launch-token") ssh root@192.168.68.12 "pct exec 112 -- curl -s -c /tmp/dsh.jar -o /dev/null \ -H 'Host: tankodhs.sysloggh.net' 'http://127.0.0.1:3080/?token=$TOKEN'" ssh root@192.168.68.12 "pct exec 112 -- curl -s -b /tmp/dsh.jar -o /dev/null \ -w '%{http_code}\n' -H 'Host: tankodhs.sysloggh.net' http://127.0.0.1:3080/" # Expected: 200 — the minted dsh-auth-... cookie (authority # tankodhs.sysloggh.net) is replayed on the next request and accepted. # 4. Token refresh is non-disruptive and idempotent. ssh root@192.168.68.12 "pct exec 112 -- /opt/deepseek-harness/capture-dsh-token.sh" # Expected: "token unchanged; nginx not reloaded" when nothing changed ``` **Restart durability (acceptance):** after `systemctl restart dsh-web`, (a) the cookie minted before the restart still returns `200` on `/`, and (b) the refreshed `/etc/dsh-web/nginx-login.conf` carries the new token and mints a fresh cookie. Both verified live 2026-09-11. ```bash # 5. Cookie survives a dsh-web restart, and the new token mints a new cookie. ssh root@192.168.68.12 "pct exec 112 -- systemctl restart dsh-web" # dsh-web is Type=simple: restart returns before :3080 is listening. Bounded-poll # until the socket answers (any status but 000) before asserting the cookie. for i in $(seq 1 60); do UP=$(ssh root@192.168.68.12 "pct exec 112 -- curl -s -o /dev/null -w '%{http_code}' \ -H 'Host: tankodhs.sysloggh.net' http://127.0.0.1:3080/") [ "$UP" != "000" ] && break sleep 2 done ssh root@192.168.68.12 "pct exec 112 -- curl -s -b /tmp/dsh.jar -o /dev/null \ -w '%{http_code}\n' -H 'Host: tankodhs.sysloggh.net' http://127.0.0.1:3080/" # Expected: 200 — the pre-restart cookie is still accepted. # The restart's ExecStartPost (or the 2-minute timer) refreshes the include. A # manual run may no-op on the flock, so poll until the include carries a token # the running process accepts (bounded wait) before the mint+reuse check. for i in $(seq 1 60); do TOKEN=$(ssh root@192.168.68.12 "pct exec 112 -- sed -n 's/.*token=//p' /etc/dsh-web/nginx-login.conf | tr -d ';\n'") CODE=$(ssh root@192.168.68.12 "pct exec 112 -- curl -s -o /dev/null -w '%{http_code}' \ -H 'Host: tankodhs.sysloggh.net' 'http://127.0.0.1:3080/?token=$TOKEN'") [ "$CODE" = "303" ] && break sleep 2 done # Expected: 303 — the include now holds the token the running process accepts. ssh root@192.168.68.12 "pct exec 112 -- curl -s -c /tmp/dsh-new.jar -o /dev/null \ -H 'Host: tankodhs.sysloggh.net' 'http://127.0.0.1:3080/?token=$TOKEN'" ssh root@192.168.68.12 "pct exec 112 -- curl -s -b /tmp/dsh-new.jar -o /dev/null \ -w '%{http_code}\n' -H 'Host: tankodhs.sysloggh.net' http://127.0.0.1:3080/" # Expected: 200 — the refreshed token minted a fresh cookie. ``` ### Step 4: Platform C — Agent Zero (kagentz, CT 105 via Docker host .14) > **The kagentz Zulip adapter leg is retired (2026-09-12).** Its code > (`/a0/usr/kagentz-zulip/`) no longer exists in the agent-zero container, so > the former adapter-process and heartbeat/queue checks always failed and the > monitor issued a restart for something that could not start, posting a false > kagentz-adapter-down alert on every run. Do NOT re-add an adapter-process, > heartbeat/queue, or adapter-restart step. Agent Zero is probed for A2A > liveness/response and public-path access only, and a probe must never restart > a platform. **C1: A2A Server Health (no credential needed)** ```bash # A2A listens on :80 inside the agent-zero container (host-mapped to :50080) and # is auth-gated: an unauthenticated probe gets 401, which means the server is up. ssh root@192.168.68.14 "docker exec agent-zero curl -s --connect-timeout 5 -o /dev/null -w '%{http_code}' http://127.0.0.1:80/a2a/" ``` Expected: `401` (auth-gated, A2A server is up and responding) or `200` (if no auth required). Connection refused (`000`) → A2A server down (INCIDENT). Any other status → running but unexpected: log/report it, never restart. **C2: A2A Response Verification (requires LITELLM_KEY)** ```bash # A2A listens on :80 inside the container and is auth-gated (401 expected unauthenticated). ssh root@192.168.68.14 "docker exec agent-zero curl -s -X POST http://127.0.0.1:80/a2a \ -H 'Content-Type: application/json' \ -H 'Authorization: Bearer $LITELLM_KEY' \ -d '{\"jsonrpc\":\"2.0\",\"method\":\"tasks/send\",\"params\":{\"message\":{\"role\":\"user\",\"parts\":[{\"text\":\"ping\"}]}},\"id\":1}'" ``` Expected: task ID with "working" status. Poll for completion with `tasks/get`. If 401, check LITELLM_KEY is set (this is a credential issue, NOT a server-down incident). **C3: Public Access Path (no credential needed)** ```bash # Probes the public URL that NetBird proxies to the agent-zero container. # This is the captain's point of view: if the captain can't reach it, it's down. # 200/302/401 = alive, 502 = proxy's "upstream refused" page (INCIDENT), # connection failed (000) = INCIDENT. Never restarts anything. curl -s -o /dev/null --connect-timeout 10 --max-time 15 -w '%{http_code}' https://kagentz.sysloggh.net/ ``` Expected: `200` (Agent Zero login page), `302` (redirect), or `401` (auth-gated) = alive. `502` = NetBird proxy's "upstream refused" page (the container's port 80 is not listening) = INCIDENT. `000` (connection failed) = INCIDENT. Any other status → running but unexpected: log/report it, never restart. **Platform C Actions** | Condition | Action | |-----------|--------| | C1 A2A returns `000` (connection refused/timeout) | Alert only — never restart the platform; investigate the agent-zero container | | C1 A2A returns a status other than `200`/`401` | Log/report as a warning — reported, never healed on | | C2 A2A returns 401 | Check `LITELLM_KEY` is set (credential issue, NOT server-down) | | C3 public URL returns `502` (upstream refused) | Alert only — the container's port 80 is not listening; investigate the agent-zero container | | C3 public URL returns `000` (connection failed) | Alert only — the public path is down; investigate the NetBird proxy or the container | | C3 public URL returns a status other than `200`/`302`/`401` | Log/report as a warning — reported, never healed on | ### Step 5: Global Checks **Cross-Agent Echo Loop Detection** Check each agent's log for excessive bot-to-bot chatter: - Abiba: `Skipped.*bot msgs` count - Tanko: Repeated DM exchanges between bots If any bot processes >50 bot-originated messages in 15min → warning. ### Step 6: Compile and Report 1. Run `scripts/zulip-monitor.sh` and take its final `Result:` line as the authoritative run verdict. The verdict line is either `Result: ✅ 0 issues (all healthy)` or `Result: 🔴 INCIDENT — N issue(s) found`. 2. Quote that `Result:` line verbatim in the status report. When it says `INCIDENT`, the run MUST be reported as an incident — never summarised as OK/healthy and never annotated as "expected". 3. Compile all platform checks and severity 4. Determine `overall_severity` from worst per-agent severity 5. If restart action needed, check `/tmp/zulip-monitor-debounce` — apply only if >300s since last restart 6. Log full diagnostic to `/root/zulip-health-monitor.log` with timestamp 7. If any agent critical or >2 degraded: send relay message to user 8. Update `last_check` timestamp in `### Maintains` snapshot ### Restart Debounce All restart actions MUST debounce: minimum 300s between restarts. Track via `/tmp/zulip-monitor-debounce` (unix timestamp of last restart). --- ## History ### Gen 5 (2026-07-02) — Rate Limit Death Spiral Fix **Root Cause**: Proactive Queue Rotation at 25 min triggered queue re-registration every cycle. Each re-registration + retry loop (3 attempts) + monitor restart = 8-12 API calls per cycle. Combined with monitor's own API calls (server check, stream alerts), `abiba-bot` hit Zulip's rate limit (429 RATE_LIMIT_HIT). Each restart reset the cycle, creating a death spiral: 111 restarts in 24 hours. **Fixes — Extension (`index.js`)**: 1. **Rotation extended to 55 min** (from 25) with **±90s jitter** — avoids aligning with cron/monitor cycles 2. **Rate-limit-aware retry** — if connection fails with 429, skip the retry loop entirely, wait 120s, try once **Fixes — Monitor (`zulip-monitor.sh`)**: 3. **Smart Triage** instead of instant restart: - **Rate-limit detection**: If error log shows recent 429s, wait — don't add more API load - **Self-healing detection**: If `retry_count` is 1-2, extension is already retrying — don't interrupt - **Rotation window awareness**: If `queue_age` is 25-35min, disconnection is likely transient rotation — wait - **Persistent failure threshold**: Only restart after 3 consecutive failed checks (45 min) AND no self-healing in progress - **Debounce gate**: Skip all checks entirely if recently restarted ### Gen 4 (2026-06-29) — Response Pipeline Fix **Root Cause**: LLM hangs due to GPU saturation (503 QUEUE_TIMEOUT) → pi's `agent_end` never fires → pending Zulip replies accumulate with no timeout. Health endpoint showed `connected: true, stuck: false` while user experienced complete silence. **Four-Layer Defense**: 1. **Response Watchdog** — 30s deadline timer; 5min placeholder edit; 10min error message + dequeue 2. **Queue Health Ping** — Every 2min `/api/v1/events?dont_block=true` to detect silent expiry 3. **Proactive Queue Rotation** — New queue every 25min + 600 empty poll reconnection threshold 4. **Health Endpoint v2** — Added `response_pipeline`, `pending_count`, `queue_healthy`, `agent_busy_duration_seconds` ### Gen 3 (2026-06-15) — Stuck Detection Added `stuck: bool` and `idle_seconds` to health endpoint. Monitor restarts on `stuck: true`. Added 300s restart debounce. Queue re-registers after 30min of no events.