diff --git a/contract-registry.yaml b/contract-registry.yaml index 019be2b..ea07da0 100644 --- a/contract-registry.yaml +++ b/contract-registry.yaml @@ -628,7 +628,7 @@ contracts: sensitivity: high status: active owner: abiba - version: 3.3.0 + version: 3.4.0 trigger: type: scheduled cadence: '*/15 * * * *' diff --git a/zulip-health.prose.md b/zulip-health.prose.md index bce42d6..c28fbfb 100644 --- a/zulip-health.prose.md +++ b/zulip-health.prose.md @@ -1,9 +1,9 @@ --- kind: responsibility name: zulip-health -description: Multi-platform health monitor for the Zulip messaging mesh spanning Platform A (pi/Abiba Zulip bridge), Platform B (Tanko on DSH), and Platform C (Agent Zero Docker). Verifies bot registration, DM delivery, and cross-platform connectivity. The kagentz Zulip adapter leg is retired (its code no longer exists) — Agent Zero is probed for A2A liveness only. Mumuni is no longer monitored from this host — she runs on her own container (kagentz CT 105 on minipve, .14) and is monitored on her side. +description: Multi-platform health monitor for the Zulip messaging mesh spanning Platform A (pi/Abiba Zulip bridge), Platform B (Tanko on DSH), and Platform C (Agent Zero Docker). Verifies bot registration, DM delivery, and cross-platform connectivity. The kagentz Zulip adapter leg is retired (its code no longer exists) — Agent Zero is probed for A2A liveness, A2A response, and public-path access only. Mumuni is no longer monitored from this host — she runs on her own container (kagentz CT 105 on minipve, .14) and is monitored on her side. title: Zulip Mesh Health Monitor — Multi-Platform -version: 3.3.0 +version: 3.4.0 runtime_contract: 2 agent: abiba report_only_agents: @@ -30,7 +30,7 @@ session start. - **Zulip API key** for `abiba-bot@chat.sysloggh.net` in `$ZULIP_API_KEY` - **SSH access** to amdpve (192.168.68.15) for Tanko — CT 112 reached via `pct exec` (direct SSH to .122 is not a dependency of this contract: per-worker key availability varies); and the Agent Zero Docker host (192.168.68.14) - **PM2** on localhost for pi process management -- **Network access** to `chat.sysloggh.net`, `localhost:9200` +- **Network access** to `chat.sysloggh.net`, `kagentz.sysloggh.net` (C3 public path), `localhost:9200` - **Write access** to `/root/zulip-health-monitor.log` and `/tmp/zulip-monitor-debounce` - **Relay access** via RA-H OS MCP for alert delivery @@ -129,10 +129,10 @@ grep -c "async def edit_message" ~/.hermes/plugins/*/zulip*/adapter.py ### Liveness rule (scoped) -Any HTTP response proves the service is ALIVE. For auth-gated endpoints (Zulip API, Tanko gateway), a 401/403 redirect or status means the service answered — report the code, never "down". Only a failed CONNECTION (curl status 000, timeout, refused) is a failed probe. +Any HTTP response proves the service is ALIVE. For auth-gated endpoints (Zulip API, Tanko gateway), a 401/403 redirect or status means the service answered — report the code, never "down". Only a failed CONNECTION (curl status 000, timeout, refused) is a failed probe. **Proxy-fronted exception:** a reverse proxy's `502`/`504` (upstream refused/unreachable) is not the backend answering — for proxy-fronted public endpoints such as `https://kagentz.sysloggh.net/` (C3), treat it as DOWN. **STANDING PROBE RULES (2026-09-14, from defect report 1150.msg):** -1. **Any HTTP status means ALIVE.** 200, 301, 302, 401, 403, 404 all prove the service answered — report the code, never "down". A redirect is not a failure. Only a failed CONNECTION (curl status 000, timeout, refused) is a failed probe. +1. **Any HTTP status means ALIVE.** 200, 301, 302, 401, 403, 404 all prove the service answered — report the code, never "down". A redirect is not a failure. Only a failed CONNECTION (curl status 000, timeout, refused) is a failed probe (proxy `502`/`504` excepted, above). 2. **A failed probe is never a service verdict.** Print `probe-failed: ` naming the exact URL/host/port and the failure kind (timeout, refused, no-route, dns), retry once at a longer timeout, and only then report. 3. **Say which probe produced each number.** "API: 000" is unusable; "API https://chat.sysloggh.net/api/v1/server_settings -> connection timeout after 10s (retried at 25s: also timeout)" is actionable. @@ -469,7 +469,8 @@ ssh root@192.168.68.15 "pct exec 112 -- curl -s -b /tmp/dsh-new.jar -o /dev/null > monitor issued a restart for something that could not start, posting a false > kagentz-adapter-down alert on every run. Do NOT re-add an adapter-process, > heartbeat/queue, or adapter-restart step. Agent Zero is probed for A2A -> liveness only, and a probe must never restart a platform. +> liveness/response and public-path access only, and a probe must never restart +> a platform. **C1: A2A Server Health (no credential needed)**