Files
prose-contracts/zulip-health.prose.md
T

23 KiB
Raw Blame History

kind, name, description, title, version, runtime_contract, agent, report_only_agents
kind name description title version runtime_contract agent report_only_agents
responsibility zulip-health Multi-platform health monitor for the Zulip messaging mesh spanning Platform A (pi/Abiba Zulip bridge), Platform B (Tanko on DSH), and Platform C (Agent Zero Docker). Verifies bot registration, DM delivery, and cross-platform connectivity. Mumuni is no longer monitored from this host — she runs on her own container (kagentz CT 105 on minipve, .14) and is monitored on her side. Zulip Mesh Health Monitor — Multi-Platform 3.2.0 2 abiba
koby

Zulip Mesh Health Monitor

Monitors the Zulip-connected agents under this host's operational control (pi, DSH, Agent Zero). Runs every 15 minutes in the background. Also triggers on session start.

Mumuni is NOT monitored from this host (captain ruling 2026-09-10). She moved off this host onto her own container — kagentz CT 105 on minipve (192.168.68.14), running a dedicated hermes user — and is monitored on her side. No step in this contract, and no leg of scripts/zulip-monitor.sh, may ssh to her old deployment, read her ~/.hermes/gateway_state.json, or alert on her state. The former Platform-B-for-Mumuni steps (gateway process, heartbeat, response delivery) are retired: they always read "unknown" against the decommissioned deployment and produced a false 🔴 alert on every run.

Requires

  • Zulip API key for abiba-bot@chat.sysloggh.net in $ZULIP_API_KEY
  • SSH access to amdpve (192.168.68.15) for Tanko — CT 112 reached via pct exec (direct SSH to .122 is not a dependency of this contract: per-worker key availability varies); and the Agent Zero Docker host (192.168.68.14)
  • PM2 on localhost for pi process management
  • Network access to chat.sysloggh.net, localhost:9200
  • Write access to /root/zulip-health-monitor.log and /tmp/zulip-monitor-debounce
  • Relay access via RA-H OS MCP for alert delivery

Maintains

  • zulip_server_status: "healthy" | "down"
  • agents: map of per-platform agent snapshots (see schema below)
  • last_check: timestamp — When the last full diagnostic ran
  • overall_severity: "healthy" | "degraded" | "critical"
  • restart_debounce: timestamp — Last restart action (enforces 300s minimum)

Per-Agent Snapshot Schema

{
  "abiba": {
    "platform": "pi",
    "connected": true,
    "response_pipeline": "healthy",
    "queue_healthy": true,
    "pm2_status": "online",
    "echo_loop_bot_msgs_15min": 12,
    "edit_fail_rate_pct": 3,
    "severity": "healthy"
  },
  "tanko": {
    "platform": "dsh",
    "service_state": "active",
    "http_status": 200,
    "severity": "healthy"
  }
}

Postconditions

  • Every platform is independently checked; one failure doesn't block others
  • Restart actions respect 300s debounce window
  • Critical conditions generate relay alerts to user
  • All checks logged to /root/zulip-health-monitor.log with timestamps

Strategies

When Zulip server is unreachable

Skip all per-platform checks — they'll all fail downstream. Report "Zulip server down" and alert.

When all agents are silent

Differentiate: if Zulip server returns 200, likely a shared infrastructure issue (Netbird, DNS). If server is down, it's upstream — wait.

When restart is indicated but debounce window hasn't passed

Log the condition as "pending restart" with the timestamp. If the condition persists after debounce window, apply restart. Never bypass debounce for non-critical conditions.

When a platform agent is unreachable via SSH

Log as "unreachable" — don't treat as critical unless it persists for 3+ consecutive checks.

Severity escalation

Streaming Support (2026-07-05)

Zulip agents now support progressive message editing during agent generation. When a Zulip agent under this monitor's scope (Tanko on DSH) processes a message, the response is streamed in real-time via Zulip's PATCH /api/v1/messages/{id} API:

  • Adapter implements edit_message() using _api_patch() helper
  • Gateway stream consumer progressively edits the Zulip message
  • User sees real-time agent thinking instead of waiting for full response
  • Verified: Tanko (CT 112) has streaming active; Mumuni's (kagentz CT 105) is verified on her own host, not from here

Verification

# Check if agent has streaming:
grep -c "async def edit_message" ~/.hermes/plugins/*/zulip*/adapter.py
# Must return 1 — streaming is active
  • Single agent warning → log only
  • Single agent critical → relay message to user
  • Two or more agents critical → immediate relay + attempt auto-recovery
  • All agents critical + server up → Netbird/DNS likely down

Invariants

  • Never restart more than once per 300s per agent
  • Never restart if PM2 crashes > 10/h — alert user instead
  • Never send duplicate alerts — check last alert timestamp before relaying
  • Health checks are read-only — diagnostics don't mutate state except for logged restarts
  • Respect Zulip API rate limits — no more than 200 requests in rapid succession

Continuity

  • On session start: Run full diagnostic pass
  • Every 15 minutes: Scheduled background check while Abiba is running
  • On zulip-status command: Run on-demand and report to user
  • On critical alert: Escalate to relay message immediately, don't wait for schedule

Execution

Step 1: Zulip Server Liveness

curl -s -o /dev/null -w "%{http_code}" https://chat.sysloggh.net/api/v1/server_settings \
  -u 'abiba-bot@chat.sysloggh.net:$ZULIP_API_KEY'

Expected: 200. If not → mark zulip_server_status: "down", skip per-platform checks, alert.

Step 2: Platform A — pi (Abiba, localhost)

A1: Health Endpoint

Fetch http://localhost:9200/health as JSON. Check:

Field Healthy Critical
connected true false
response_pipeline healthy blocked
queue_healthy true false
stuck false true
idle_seconds < 1800 ≥ 1800
pending_count 0 > 0 with response_pipeline: blocked
agent_busy_duration_seconds < 300 ≥ 600
last_error null non-null string
retry_count 0–2 3+
queue_id non-null string null

A2: PM2 Process

pm2 show abiba-zulip --no-color 2>/dev/null

Check: status=online, restarts < 10/h, uptime > 60s.

A3: Echo Loop Detection

grep -a "Skipped.*bot msgs" /root/.pm2/logs/abiba-zulip-out.log | tail -5

100 skipped in 15min → info only (echo loop prevention working).

A4: Response Delivery

grep -a "Finalized\|Failed to finalize" /root/.pm2/logs/abiba-zulip-out.log | tail -20

50% fail rate → critical — check editMessage API.

Platform A Actions

Condition Action
response_pipeline: blocked pm2 restart abiba-zulip
connected: false pm2 restart abiba-zulip
queue_healthy: false pm2 restart abiba-zulip
stuck: true pm2 restart abiba-zulip
retry_count >= 3 pm2 restart abiba-zulip
response_pipeline: degraded Monitor — no action, watchdog handles
last_error set Log and monitor
Crash loop >10/h Alert user

Step 3: Platform B — Tanko (DSH on amdpve CT 112)

Mumuni is out of scope for this host (see the note above): she runs on her own container and is monitored on her side.

Tanko runs on DSH (DeepSeek Harness) — it no longer runs a Hermes gateway, so there is no ~/.hermes/gateway_state.json on CT 112. Tanko's Zulip gateway runs as the dsh-web systemd unit inside CT 112, which resides on the amdpve PVE host (192.168.68.15). Direct SSH to 192.168.68.122 is not a dependency of this contract — per-worker key availability varies — so CT 112 probes run from the amdpve vantage via pct exec:

ssh root@192.168.68.15 "pct exec 112 -- <command>"

By design (verified 2026-09-08): the dsh-web gateway binds 127.0.0.1:3080 loopback-only. A remote probe against 192.168.68.122:3080 gets connection-refused — that is EXPECTED, NOT a fault, and must never be raised as Tanko down. Only loopback probes from inside CT 112 (or the public-URL fallback below) are valid health signals.

B1: Gateway Service State (Tanko)

ssh root@192.168.68.15 "pct exec 112 -- systemctl is-active dsh-web"

Expected: active. Anything else → gateway service down → apply the Tanko heal (restart via DSH service, Platform B Actions table below).

B2: Gateway HTTP Liveness (Tanko — loopback-only :3080)

ssh root@192.168.68.15 "pct exec 112 -- curl -s --connect-timeout 5 --max-time 10 -o /dev/null -w '%{http_code}' http://127.0.0.1:3080/"

Alive = ANY HTTP status response from the endpoint — the expected set is 200/301/302/307/308/401/403 (the gateway UI is token-gated and legitimately answers with redirects/auth-challenges, so never require a bare 200), and any other status, including 404/5xx, also counts alive: a process answering 503 is running and self-heal must NOT restart-loop it. Down = connection refused (000) or timeout only. Statuses outside the expected set are logged/reported as a warning — reported, never healed on.

B3: Public-URL Fallback Probe (Tanko — for nodes without pct/ssh access to amdpve)

curl -s --connect-timeout 10 --max-time 15 -o /dev/null -w '%{http_code}' https://tankodhs.sysloggh.net/

Fallback only — used when the monitoring node has no pct/SSH path to amdpve. Alive = ANY HTTP status response from the endpoint — healthy signals are 302 (authentik proxy-auth redirect) and 401 (auth-gated), and any other status, including 404/5xx, also counts alive: the endpoint is up and answering and must NOT be restart-looped. Down = connection refused (000) or timeout only. Never expect a bare 200 — the public URL terminates in the token-gated authentik chain. Statuses outside the healthy set are logged/reported as a warning — reported, never healed on.

Platform B Actions

Condition Action
dsh-web service not active Restart Tanko via DSH service
HTTP :3080 connection refused/timeout (000) Same as above
HTTP status outside the expected set Log/report as a warning — reported, never healed on
B4: dsh-web Authentication (Tanko — restart-persistent login)

The dsh-web UI is token-gated. On every start the process prints a random launch token to the journal:

dsh web: http://127.0.0.1:3080/?token=<TOKEN>

The token only bootstraps an authority-bound, HMAC-signed browser cookie with a 30-day lifetime. The signing secret is durable in /root/.dsh/.credentials.yaml (key client-connection/browser-session), so a cookie minted once keeps working across dsh-web restarts; the launch token itself rotates on every restart.

Login endpoint (public, Authentik-gated): https://tankodhs.sysloggh.net/dsh-web-login

It lives inside the Authentik-gated :80 server block (/etc/nginx/sites-available/dsh, symlinked from /etc/nginx/sites-enabled/dsh) as location = /dsh-web-login, guarded by auth_request /outpost.goauthentik.io/auth/nginx. It proxies to dsh-web with Host: tankodhs.sysloggh.net, so the minted cookie is bound to the public authority — never to 127.0.0.1:3080. The token-dependent line is isolated in the generated include /etc/dsh-web/nginx-login.conf:

proxy_pass http://127.0.0.1:3080/?token=<TOKEN>;

Token refresh (non-disruptive): /opt/deepseek-harness/capture-dsh-token.sh (source: scripts/capture-dsh-token.sh) reads candidate launch tokens from the journal scoped to the service's current systemd invocation (systemctl show -p InvocationID + _SYSTEMD_INVOCATION_ID=), so a restarted process's stale token is never considered while its new startup banner is still pending; there is no whole-journal or cross-invocation fallback. Each candidate is then functionally verified against dsh-web with Host: tankodhs.sysloggh.net, using the first the running process accepts with 303. It waits up to 120s for a restarted process to accept a token and re-probes every current-invocation candidate on each pass, so a token that briefly returns 000 while the service is still starting is not disqualified. If none is accepted it leaves the include untouched and exits so the timer retries (exiting non-zero when a pending reload is still outstanding). It writes /etc/dsh-web/launch-token and regenerates /etc/dsh-web/nginx-login.conf, reloading nginx only when the on-disk include differs from the generated one or the applied-state stamp does not match the token (nginx -t guards the reload, and the stamp is written only after a successful nginx -s reload, so a failed or interrupted reload is retried on the next run). Any failed reload records a pending-reload marker under /etc/dsh-web/; the next run attempts the reload before the token wait, independent of token state, and clears the marker only once the reload succeeds, so a disabled legacy :8081 file can never leave the running nginx unreloaded. The generated include is recreated before any nginx -t if it is missing, so a failed run cannot wedge recovery. Runs are serialized with flock on /run/capture-dsh-token.lock. It never stops or starts dsh-web. It is triggered by the dsh-web.service drop-in /etc/systemd/system/dsh-web.service.d/20-token-refresh.conf (ExecStartPost=/bin/systemctl --no-block start dsh-web-token.service) and by dsh-web-token.timer every 2 minutes for reconciliation.

Installed systemd wiring (CT 112)
# /etc/systemd/system/dsh-web-token.service
[Unit]
Description=Refresh the dsh-web launch token for the nginx login endpoint
After=dsh-web.service
[Service]
Type=oneshot
TimeoutStartSec=180
ExecStart=/opt/deepseek-harness/capture-dsh-token.sh

# /etc/systemd/system/dsh-web-token.timer
[Unit]
Description=Periodically refresh the dsh-web login token
[Timer]
OnBootSec=90s
OnUnitActiveSec=120s
AccuracySec=10s
Persistent=true
[Install]
WantedBy=timers.target

# /etc/systemd/system/dsh-web.service.d/20-token-refresh.conf
[Service]
ExecStartPost=/bin/systemctl --no-block start dsh-web-token.service

Do NOT reintroduce the :8081 endpoint. It listened on 0.0.0.0:8081 with no auth_request and was a full Authentik bypass for anyone on the LAN. The script now removes /etc/nginx/sites-enabled/dsh.token automatically if it ever reappears.

Authentication flow:

  1. GET https://tankodhs.sysloggh.net/dsh-web-login
  2. Unauthenticated → Authentik sign-in; once authenticated the request reaches dsh-web with Host: tankodhs.sysloggh.net.
  3. dsh-web accepts the launch token on GET /, writes the dsh-auth-<authority-hash> cookie (30 days, HttpOnly, SameSite=Strict) and returns 303 to /.
  4. Every later request through / presents that cookie; the token is not needed again until the cookie expires or a new browser is used.

Verification (amdpve vantage):

# 1. Login endpoint is Authentik-gated: unauthenticated -> 302 (not 200/303).
ssh root@192.168.68.15 "pct exec 112 -- curl -s -o /dev/null -w '%{http_code}\n' \
  -H 'Host: tankodhs.sysloggh.net' http://127.0.0.1/dsh-web-login"
# Expected: 302

# 2. Legacy :8081 endpoint is gone (connection refused -> 000).
ssh root@192.168.68.15 "pct exec 112 -- curl -s --max-time 3 -o /dev/null \
  -w '%{http_code}\n' http://192.168.68.122:8081/"
# Expected: 000

# 3. Backend cookie mint + reuse (exactly what /dsh-web-login proxies to).
TOKEN=$(ssh root@192.168.68.15 "pct exec 112 -- cat /etc/dsh-web/launch-token")
ssh root@192.168.68.15 "pct exec 112 -- curl -s -c /tmp/dsh.jar -o /dev/null \
  -H 'Host: tankodhs.sysloggh.net' 'http://127.0.0.1:3080/?token=$TOKEN'"
ssh root@192.168.68.15 "pct exec 112 -- curl -s -b /tmp/dsh.jar -o /dev/null \
  -w '%{http_code}\n' -H 'Host: tankodhs.sysloggh.net' http://127.0.0.1:3080/"
# Expected: 200 — the minted dsh-auth-... cookie (authority
# tankodhs.sysloggh.net) is replayed on the next request and accepted.

# 4. Token refresh is non-disruptive and idempotent.
ssh root@192.168.68.15 "pct exec 112 -- /opt/deepseek-harness/capture-dsh-token.sh"
# Expected: "token unchanged; nginx not reloaded" when nothing changed

Restart durability (acceptance): after systemctl restart dsh-web, (a) the cookie minted before the restart still returns 200 on /, and (b) the refreshed /etc/dsh-web/nginx-login.conf carries the new token and mints a fresh cookie. Both verified live 2026-09-11.

# 5. Cookie survives a dsh-web restart, and the new token mints a new cookie.
ssh root@192.168.68.15 "pct exec 112 -- systemctl restart dsh-web"
# dsh-web is Type=simple: restart returns before :3080 is listening. Bounded-poll
# until the socket answers (any status but 000) before asserting the cookie.
for i in $(seq 1 60); do
  UP=$(ssh root@192.168.68.15 "pct exec 112 -- curl -s -o /dev/null -w '%{http_code}' \
    -H 'Host: tankodhs.sysloggh.net' http://127.0.0.1:3080/")
  [ "$UP" != "000" ] && break
  sleep 2
done
ssh root@192.168.68.15 "pct exec 112 -- curl -s -b /tmp/dsh.jar -o /dev/null \
  -w '%{http_code}\n' -H 'Host: tankodhs.sysloggh.net' http://127.0.0.1:3080/"
# Expected: 200 — the pre-restart cookie is still accepted.
# The restart's ExecStartPost (or the 2-minute timer) refreshes the include. A
# manual run may no-op on the flock, so poll until the include carries a token
# the running process accepts (bounded wait) before the mint+reuse check.
for i in $(seq 1 60); do
  TOKEN=$(ssh root@192.168.68.15 "pct exec 112 -- sed -n 's/.*token=//p' /etc/dsh-web/nginx-login.conf | tr -d ';\n'")
  CODE=$(ssh root@192.168.68.15 "pct exec 112 -- curl -s -o /dev/null -w '%{http_code}' \
    -H 'Host: tankodhs.sysloggh.net' 'http://127.0.0.1:3080/?token=$TOKEN'")
  [ "$CODE" = "303" ] && break
  sleep 2
done
# Expected: 303 — the include now holds the token the running process accepts.
ssh root@192.168.68.15 "pct exec 112 -- curl -s -c /tmp/dsh-new.jar -o /dev/null \
  -H 'Host: tankodhs.sysloggh.net' 'http://127.0.0.1:3080/?token=$TOKEN'"
ssh root@192.168.68.15 "pct exec 112 -- curl -s -b /tmp/dsh-new.jar -o /dev/null \
  -w '%{http_code}\n' -H 'Host: tankodhs.sysloggh.net' http://127.0.0.1:3080/"
# Expected: 200 — the refreshed token minted a fresh cookie.

Step 4: Platform C — Agent Zero (kagentz, CT 105 via Docker host .14)

C1: A2A Server Health

# A2A server is on :50080 (not :8001) and is auth-gated (401 expected for unauthenticated)
ssh root@192.168.68.14 "curl -s --connect-timeout 5 -o /dev/null -w '%{http_code}' http://127.0.0.1:50080/a2a/"

Expected: 401 (auth-gated, A2A server is up and responding) or 200 (if no auth required). Connection refused (000) → A2A server down.

C2: Adapter Process

ssh root@192.168.68.14 "docker exec agent-zero ps aux | grep adapter | grep -v grep"

Adapter should be running. Missing → restart inside container.

C3: Heartbeat & Queue

ssh root@192.168.68.14 "docker exec agent-zero grep Heartbeat /tmp/zulip-adapter.log | tail -3"

Check: processed=N incrementing, silence < 600s, reconnects ≈ 0.

C4: A2A Response Verification

# A2A server is on :50080 (not :8001) and is auth-gated (401 expected for unauthenticated)
ssh root@192.168.68.14 "curl -s -X POST http://127.0.0.1:50080/a2a \
  -H 'Content-Type: application/json' \
  -H 'Authorization: Bearer $LITELLM_KEY' \
  -d '{\"jsonrpc\":\"2.0\",\"method\":\"tasks/send\",\"params\":{\"message\":{\"role\":\"user\",\"parts\":[{\"text\":\"ping\"}]}},\"id\":1}'"

Expected: task ID with "working" status. Poll for completion with tasks/get. If 401, check LITELLM_KEY is set.

Platform C Actions

Condition Action
A2A .well-known/agent.json fails docker exec agent-zero bash -c "pkill -9 -f a2a_agent; cd /a0 && /opt/venv-a0/bin/python3 -u /a0/usr/a2a_agent.py > /tmp/a2a.log 2>&1 &"
Adapter process missing Restart adapter inside container with env vars
Silence > 600s Restart adapter (auto-reconnect handles BAD_EVENT_QUEUE_ID)
LiteLLM 401 Check API key in a2a_agent.py LITELLM_KEY

Step 5: Global Checks

Cross-Agent Echo Loop Detection

Check each agent's log for excessive bot-to-bot chatter:

  • Abiba: Skipped.*bot msgs count
  • Tanko: Repeated DM exchanges between bots
  • kagentz: Adapter log for bot DMs being processed

If any bot processes >50 bot-originated messages in 15min → warning.

Step 6: Compile and Report

  1. Compile all platform checks and severity
  2. Determine overall_severity from worst per-agent severity
  3. If restart action needed, check /tmp/zulip-monitor-debounce — apply only if >300s since last restart
  4. Log full diagnostic to /root/zulip-health-monitor.log with timestamp
  5. If any agent critical or >2 degraded: send relay message to user
  6. Update last_check timestamp in ### Maintains snapshot

Restart Debounce

All restart actions MUST debounce: minimum 300s between restarts. Track via /tmp/zulip-monitor-debounce (unix timestamp of last restart).


History

Gen 5 (2026-07-02) — Rate Limit Death Spiral Fix

Root Cause: Proactive Queue Rotation at 25 min triggered queue re-registration every cycle. Each re-registration + retry loop (3 attempts) + monitor restart = 8-12 API calls per cycle. Combined with monitor's own API calls (server check, stream alerts), abiba-bot hit Zulip's rate limit (429 RATE_LIMIT_HIT). Each restart reset the cycle, creating a death spiral: 111 restarts in 24 hours.

Fixes — Extension (index.js):

  1. Rotation extended to 55 min (from 25) with ±90s jitter — avoids aligning with cron/monitor cycles
  2. Rate-limit-aware retry — if connection fails with 429, skip the retry loop entirely, wait 120s, try once

Fixes — Monitor (zulip-monitor.sh): 3. Smart Triage instead of instant restart:

  • Rate-limit detection: If error log shows recent 429s, wait — don't add more API load
  • Self-healing detection: If retry_count is 1-2, extension is already retrying — don't interrupt
  • Rotation window awareness: If queue_age is 25-35min, disconnection is likely transient rotation — wait
  • Persistent failure threshold: Only restart after 3 consecutive failed checks (45 min) AND no self-healing in progress
  • Debounce gate: Skip all checks entirely if recently restarted

Gen 4 (2026-06-29) — Response Pipeline Fix

Root Cause: LLM hangs due to GPU saturation (503 QUEUE_TIMEOUT) → pi's agent_end never fires → pending Zulip replies accumulate with no timeout. Health endpoint showed connected: true, stuck: false while user experienced complete silence.

Four-Layer Defense:

  1. Response Watchdog — 30s deadline timer; 5min placeholder edit; 10min error message + dequeue
  2. Queue Health Ping — Every 2min /api/v1/events?dont_block=true to detect silent expiry
  3. Proactive Queue Rotation — New queue every 25min + 600 empty poll reconnection threshold
  4. Health Endpoint v2 — Added response_pipeline, pending_count, queue_healthy, agent_busy_duration_seconds

Gen 3 (2026-06-15) — Stuck Detection

Added stuck: bool and idle_seconds to health endpoint. Monitor restarts on stuck: true. Added 300s restart debounce. Queue re-registers after 30min of no events.