PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
Per defect report 1150.msg: - Add standing probe rules section (2026-09-14) - Step 1 (Zulip API): retry once at 25s on 000, print target + code - Step 2 (Platform A): retry once at 25s on 000, print target + code, note must run on Abiba host - Apply same shape as infrastructure-monitoring: any HTTP status = ALIVE; only 000/timeout = probe-failed Verified: all probes now return real HTTP codes (Zulip API 200, Platform A 200, Tanko 401, Agent Zero 401)
561 lines
26 KiB
Markdown
561 lines
26 KiB
Markdown
---
|
||
kind: responsibility
|
||
name: zulip-health
|
||
description: Multi-platform health monitor for the Zulip messaging mesh spanning Platform A (pi/Abiba Zulip bridge), Platform B (Tanko on DSH), and Platform C (Agent Zero Docker). Verifies bot registration, DM delivery, and cross-platform connectivity. The kagentz Zulip adapter leg is retired (its code no longer exists) — Agent Zero is probed for A2A liveness only. Mumuni is no longer monitored from this host — she runs on her own container (kagentz CT 105 on minipve, .14) and is monitored on her side.
|
||
title: Zulip Mesh Health Monitor — Multi-Platform
|
||
version: 3.3.0
|
||
runtime_contract: 2
|
||
agent: abiba
|
||
report_only_agents:
|
||
- koby # ⛔ KOBY IS NEVER REPAIRED (Rule 17, 2026-08-17) — detect + report, never fix on .129
|
||
---
|
||
|
||
# Zulip Mesh Health Monitor
|
||
|
||
Monitors the Zulip-connected agents under this host's operational control (pi,
|
||
DSH, Agent Zero). Runs every 15 minutes in the background. Also triggers on
|
||
session start.
|
||
|
||
> **Mumuni is NOT monitored from this host (captain ruling 2026-09-10).** She
|
||
> moved off this host onto her own container — kagentz CT 105 on minipve
|
||
> (192.168.68.14), running a dedicated `hermes` user — and is monitored on her
|
||
> side. No step in this contract, and no leg of `scripts/zulip-monitor.sh`, may
|
||
> ssh to her old deployment, read her `~/.hermes/gateway_state.json`, or alert on
|
||
> her state. The former Platform-B-for-Mumuni steps (gateway process, heartbeat,
|
||
> response delivery) are retired: they always read "unknown" against the
|
||
> decommissioned deployment and produced a false 🔴 alert on every run.
|
||
|
||
## Requires
|
||
|
||
- **Zulip API key** for `abiba-bot@chat.sysloggh.net` in `$ZULIP_API_KEY`
|
||
- **SSH access** to amdpve (192.168.68.15) for Tanko — CT 112 reached via `pct exec` (direct SSH to .122 is not a dependency of this contract: per-worker key availability varies); and the Agent Zero Docker host (192.168.68.14)
|
||
- **PM2** on localhost for pi process management
|
||
- **Network access** to `chat.sysloggh.net`, `localhost:9200`
|
||
- **Write access** to `/root/zulip-health-monitor.log` and `/tmp/zulip-monitor-debounce`
|
||
- **Relay access** via RA-H OS MCP for alert delivery
|
||
|
||
## Maintains
|
||
|
||
- zulip_server_status: "healthy" | "down"
|
||
- agents: map of per-platform agent snapshots (see schema below)
|
||
- last_check: timestamp — When the last full diagnostic ran
|
||
- overall_severity: "healthy" | "degraded" | "critical"
|
||
- restart_debounce: timestamp — Last restart action (enforces 300s minimum)
|
||
|
||
### Per-Agent Snapshot Schema
|
||
|
||
```json
|
||
{
|
||
"abiba": {
|
||
"platform": "pi",
|
||
"connected": true,
|
||
"response_pipeline": "healthy",
|
||
"queue_healthy": true,
|
||
"pm2_status": "online",
|
||
"echo_loop_bot_msgs_15min": 12,
|
||
"edit_fail_rate_pct": 3,
|
||
"severity": "healthy"
|
||
},
|
||
"tanko": {
|
||
"platform": "dsh",
|
||
"service_state": "active",
|
||
"http_status": 200,
|
||
"severity": "healthy"
|
||
}
|
||
}
|
||
```
|
||
|
||
### Postconditions
|
||
|
||
- Every platform is independently checked; one failure doesn't block others
|
||
- Restart actions respect 300s debounce window
|
||
- Critical conditions generate relay alerts to user
|
||
- All checks logged to `/root/zulip-health-monitor.log` with timestamps
|
||
|
||
## Strategies
|
||
|
||
### When Zulip server is unreachable
|
||
Skip all per-platform checks — they'll all fail downstream. Report "Zulip server down" and alert.
|
||
|
||
### When all agents are silent
|
||
Differentiate: if Zulip server returns 200, likely a shared infrastructure issue (Netbird, DNS). If server is down, it's upstream — wait.
|
||
|
||
### When restart is indicated but debounce window hasn't passed
|
||
Log the condition as "pending restart" with the timestamp. If the condition persists after debounce window, apply restart. Never bypass debounce for non-critical conditions.
|
||
|
||
### When a platform agent is unreachable via SSH
|
||
Log as "unreachable" — don't treat as critical unless it persists for 3+ consecutive checks.
|
||
|
||
### Severity escalation
|
||
|
||
## Streaming Support (2026-07-05)
|
||
|
||
Zulip agents now support progressive message editing during agent generation.
|
||
When a Zulip agent under this monitor's scope (Tanko on DSH) processes a message, the response is
|
||
streamed in real-time via Zulip's `PATCH /api/v1/messages/{id}` API:
|
||
|
||
- Adapter implements `edit_message()` using `_api_patch()` helper
|
||
- Gateway stream consumer progressively edits the Zulip message
|
||
- User sees real-time agent thinking instead of waiting for full response
|
||
- Verified: Tanko (CT 112) has streaming active; Mumuni's (kagentz CT 105) is verified on her own host, not from here
|
||
|
||
### Verification
|
||
```bash
|
||
# Check if agent has streaming:
|
||
grep -c "async def edit_message" ~/.hermes/plugins/*/zulip*/adapter.py
|
||
# Must return 1 — streaming is active
|
||
```
|
||
- Single agent warning → log only
|
||
- Single agent critical → relay message to user
|
||
- Two or more agents critical → immediate relay + attempt auto-recovery
|
||
- All agents critical + server up → Netbird/DNS likely down
|
||
|
||
## Invariants
|
||
|
||
- **Never restart more than once per 300s** per agent
|
||
- **Never restart if PM2 crashes > 10/h** — alert user instead
|
||
- **Never send duplicate alerts** — check last alert timestamp before relaying
|
||
- **Health checks are read-only** — diagnostics don't mutate state except for logged restarts
|
||
- **Respect Zulip API rate limits** — no more than 200 requests in rapid succession
|
||
|
||
## Continuity
|
||
|
||
- **On session start**: Run full diagnostic pass
|
||
- **Every 15 minutes**: Scheduled background check while Abiba is running
|
||
- **On `zulip-status` command**: Run on-demand and report to user
|
||
- **On critical alert**: Escalate to relay message immediately, don't wait for schedule
|
||
|
||
## Execution
|
||
|
||
### Liveness rule (scoped)
|
||
|
||
Any HTTP response proves the service is ALIVE. For auth-gated endpoints (Zulip API, Tanko gateway), a 401/403 redirect or status means the service answered — report the code, never "down". Only a failed CONNECTION (curl status 000, timeout, refused) is a failed probe.
|
||
|
||
**STANDING PROBE RULES (2026-09-14, from defect report 1150.msg):**
|
||
1. **Any HTTP status means ALIVE.** 200, 301, 302, 401, 403, 404 all prove the service answered — report the code, never "down". A redirect is not a failure. Only a failed CONNECTION (curl status 000, timeout, refused) is a failed probe.
|
||
2. **A failed probe is never a service verdict.** Print `probe-failed: <target> <kind>` naming the exact URL/host/port and the failure kind (timeout, refused, no-route, dns), retry once at a longer timeout, and only then report.
|
||
3. **Say which probe produced each number.** "API: 000" is unusable; "API https://chat.sysloggh.net/api/v1/server_settings -> connection timeout after 10s (retried at 25s: also timeout)" is actionable.
|
||
|
||
### Step 1: Zulip Server Liveness
|
||
|
||
```bash
|
||
# Probe the Zulip API (authenticated, any HTTP status = ALIVE)
|
||
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 https://chat.sysloggh.net/api/v1/server_settings -u 'abiba-bot@chat.sysloggh.net:$ZULIP_API_KEY')
|
||
if [ "$code" == "000" ]; then
|
||
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 25 https://chat.sysloggh.net/api/v1/server_settings -u 'abiba-bot@chat.sysloggh.net:$ZULIP_API_KEY')
|
||
echo "Zulip API https://chat.sysloggh.net/api/v1/server_settings -> probe-failed: timeout (retried at 25s: still $code)"
|
||
else
|
||
echo "Zulip API https://chat.sysloggh.net/api/v1/server_settings -> $code"
|
||
fi
|
||
```
|
||
|
||
Expected: `200` (authenticated). Any HTTP status = ALIVE; only 000/timeout = probe-failed. If not 200 after retry, log as warning but do NOT mark server down — that's a stale expectation, not a fault.
|
||
|
||
### Step 2: Platform A — pi (Abiba, localhost)
|
||
|
||
**A1: Health Endpoint**
|
||
|
||
```bash
|
||
# Probe the Abiba extension health endpoint (loopback, any HTTP status = ALIVE)
|
||
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 http://127.0.0.1:9200/health)
|
||
if [ "$code" == "000" ]; then
|
||
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 25 http://127.0.0.1:9200/health)
|
||
echo "Abiba extension http://127.0.0.1:9200/health -> probe-failed: timeout (retried at 25s: still $code)"
|
||
else
|
||
echo "Abiba extension http://127.0.0.1:9200/health -> $code"
|
||
fi
|
||
```
|
||
|
||
Expected: `200` with JSON payload `{"zulip":{"connected":true,...}}`. Any HTTP status = ALIVE; only 000/timeout = probe-failed. **NOTE: This probe MUST run on the Abiba host (CT 100) where 127.0.0.1:9200 is the extension. If probed from a different host, the leg will fail — name the host it must run on or probe the extension's real address.**
|
||
|
||
Check the JSON payload:
|
||
|
||
| Field | Healthy | Critical |
|
||
|-------|---------|----------|
|
||
| `connected` | `true` | `false` |
|
||
| `response_pipeline` | `healthy` | `blocked` |
|
||
| `queue_healthy` | `true` | `false` |
|
||
| `stuck` | `false` | `true` |
|
||
| `idle_seconds` | < 1800 | ≥ 1800 |
|
||
| `pending_count` | 0 | > 0 with `response_pipeline: blocked` |
|
||
| `agent_busy_duration_seconds` | < 300 | ≥ 600 |
|
||
| `last_error` | `null` | non-null string |
|
||
| `retry_count` | 0–2 | 3+ |
|
||
| `queue_id` | non-null string | `null` |
|
||
|
||
**A2: PM2 Process**
|
||
|
||
```bash
|
||
pm2 show abiba-zulip --no-color 2>/dev/null
|
||
```
|
||
|
||
Check: `status=online`, `restarts < 10/h`, `uptime > 60s`.
|
||
|
||
**A3: Echo Loop Detection**
|
||
|
||
```bash
|
||
grep -a "Skipped.*bot msgs" /root/.pm2/logs/abiba-zulip-out.log | tail -5
|
||
```
|
||
|
||
> 100 skipped in 15min → info only (echo loop prevention working).
|
||
|
||
**A4: Response Delivery**
|
||
|
||
```bash
|
||
grep -a "Finalized\|Failed to finalize" /root/.pm2/logs/abiba-zulip-out.log | tail -20
|
||
```
|
||
|
||
> 50% fail rate → critical — check editMessage API.
|
||
|
||
**Platform A Actions**
|
||
|
||
| Condition | Action |
|
||
|-----------|--------|
|
||
| `response_pipeline: blocked` | `pm2 restart abiba-zulip` |
|
||
| `connected: false` | `pm2 restart abiba-zulip` |
|
||
| `queue_healthy: false` | `pm2 restart abiba-zulip` |
|
||
| `stuck: true` | `pm2 restart abiba-zulip` |
|
||
| `retry_count >= 3` | `pm2 restart abiba-zulip` |
|
||
| `response_pipeline: degraded` | Monitor — no action, watchdog handles |
|
||
| `last_error` set | Log and monitor |
|
||
| Crash loop >10/h | Alert user |
|
||
|
||
|
||
### Step 3: Platform B — Tanko (DSH on amdpve CT 112)
|
||
|
||
Mumuni is out of scope for this host (see the note above): she runs on her own
|
||
container and is monitored on her side.
|
||
|
||
Tanko runs on DSH (DeepSeek Harness) — it no longer runs a Hermes gateway, so
|
||
there is no `~/.hermes/gateway_state.json` on CT 112. Tanko's Zulip gateway runs
|
||
as the `dsh-web` systemd unit inside **CT 112**, which resides on the **amdpve**
|
||
PVE host (**192.168.68.15**). Direct SSH to 192.168.68.122 is not a dependency
|
||
of this contract — per-worker key availability varies — so CT 112 probes run
|
||
from the amdpve vantage via `pct exec`:
|
||
|
||
```bash
|
||
ssh root@192.168.68.15 "pct exec 112 -- <command>"
|
||
```
|
||
|
||
> **By design (verified 2026-09-08):** the `dsh-web` gateway binds
|
||
> `127.0.0.1:3080` **loopback-only**. A remote probe against
|
||
> `192.168.68.122:3080` gets connection-refused — that is EXPECTED, NOT a fault,
|
||
> and must never be raised as Tanko down. Only loopback probes from inside
|
||
> CT 112 (or the public-URL fallback below) are valid health signals.
|
||
|
||
**B1: Gateway Service State (Tanko)**
|
||
|
||
```bash
|
||
ssh root@192.168.68.15 "pct exec 112 -- systemctl is-active dsh-web"
|
||
```
|
||
|
||
Expected: `active`. Anything else → gateway service down → apply the Tanko heal
|
||
(restart via DSH service, Platform B Actions table below).
|
||
|
||
**B2: Gateway HTTP Liveness (Tanko — loopback-only :3080)**
|
||
|
||
```bash
|
||
ssh root@192.168.68.15 "pct exec 112 -- curl -s --connect-timeout 5 --max-time 10 -o /dev/null -w '%{http_code}' http://127.0.0.1:3080/"
|
||
```
|
||
|
||
Alive = **ANY** HTTP status response from the endpoint — the expected set is
|
||
`200`/`301`/`302`/`307`/`308`/`401`/`403` (the gateway UI is token-gated and
|
||
legitimately answers with redirects/auth-challenges, so never require a bare
|
||
`200`), and any other status, including `404`/`5xx`, also counts alive: a
|
||
process answering `503` is running and self-heal must NOT restart-loop it.
|
||
Down = connection refused (`000`) or timeout only. Statuses outside the
|
||
expected set are logged/reported as a warning — reported, never healed on.
|
||
|
||
**B3: Public-URL Fallback Probe (Tanko — for nodes without pct/ssh access to amdpve)**
|
||
|
||
```bash
|
||
curl -s --connect-timeout 10 --max-time 15 -o /dev/null -w '%{http_code}' https://tankodhs.sysloggh.net/
|
||
```
|
||
|
||
Fallback only — used when the monitoring node has no pct/SSH path to amdpve.
|
||
Alive = **ANY** HTTP status response from the endpoint — healthy signals are
|
||
`302` (authentik proxy-auth redirect) and `401` (auth-gated), and any other
|
||
status, including `404`/`5xx`, also counts alive: the endpoint is up and
|
||
answering and must NOT be restart-looped. Down = connection refused (`000`) or
|
||
timeout only. Never expect a bare `200` — the public URL terminates in the
|
||
token-gated authentik chain. Statuses outside the healthy set are
|
||
logged/reported as a warning — reported, never healed on.
|
||
|
||
**Platform B Actions**
|
||
|
||
| Condition | Action |
|
||
|-----------|--------|
|
||
| `dsh-web` service not `active` | Restart Tanko via DSH service |
|
||
| HTTP `:3080` connection refused/timeout (`000`) | Same as above |
|
||
| HTTP status outside the expected set | Log/report as a warning — reported, never healed on |
|
||
|
||
**B4: dsh-web Authentication (Tanko — restart-persistent login)**
|
||
|
||
The dsh-web UI is token-gated. On every start the process prints a random
|
||
launch token to the journal:
|
||
|
||
```
|
||
dsh web: http://127.0.0.1:3080/?token=<TOKEN>
|
||
```
|
||
|
||
The token only bootstraps an authority-bound, HMAC-signed browser cookie with a
|
||
30-day lifetime. The signing secret is durable in
|
||
`/root/.dsh/.credentials.yaml` (key `client-connection/browser-session`), so a
|
||
cookie minted once keeps working across `dsh-web` restarts; the launch token
|
||
itself rotates on every restart.
|
||
|
||
**Login endpoint (public, Authentik-gated):**
|
||
`https://tankodhs.sysloggh.net/dsh-web-login`
|
||
|
||
It lives inside the Authentik-gated `:80` server block
|
||
(`/etc/nginx/sites-available/dsh`, symlinked from
|
||
`/etc/nginx/sites-enabled/dsh`) as `location = /dsh-web-login`, guarded by
|
||
`auth_request /outpost.goauthentik.io/auth/nginx`. It proxies to dsh-web with
|
||
`Host: tankodhs.sysloggh.net`, so the minted cookie is bound to the public
|
||
authority — never to `127.0.0.1:3080`. The token-dependent line is isolated in
|
||
the generated include `/etc/dsh-web/nginx-login.conf`:
|
||
|
||
```
|
||
proxy_pass http://127.0.0.1:3080/?token=<TOKEN>;
|
||
```
|
||
|
||
**Token refresh (non-disruptive):**
|
||
`/opt/deepseek-harness/capture-dsh-token.sh` (source:
|
||
`scripts/capture-dsh-token.sh`) reads candidate launch tokens from the journal
|
||
**scoped to the service's current systemd invocation**
|
||
(`systemctl show -p InvocationID` + `_SYSTEMD_INVOCATION_ID=`), re-sampling the
|
||
invocation on every pass so a restart that lands during the wait switches to the
|
||
new invocation; a restarted process's stale token is never considered while its
|
||
new startup banner is still pending and there is no whole-journal or
|
||
cross-invocation fallback. Each candidate
|
||
is then functionally verified against dsh-web with `Host: tankodhs.sysloggh.net`,
|
||
using the first the running process accepts with `303`. It waits up to 120s for
|
||
a restarted process to accept a token and re-probes every current-invocation
|
||
candidate on each pass, so a token that briefly returns `000` while the service
|
||
is still starting is not disqualified. If none is accepted it leaves the include
|
||
untouched and exits so the timer retries (exiting non-zero when a pending reload
|
||
is still outstanding). It writes
|
||
`/etc/dsh-web/launch-token` and regenerates `/etc/dsh-web/nginx-login.conf`,
|
||
reloading nginx only when the on-disk include differs from the generated one or
|
||
the applied-state stamp does not match the token (`nginx -t` guards the reload,
|
||
and the stamp is written only after a successful `nginx -s reload`, so a failed
|
||
or interrupted reload is retried on the next run). Any failed reload records a
|
||
pending-reload marker under `/etc/dsh-web/`; the next run attempts the reload
|
||
before the token wait, independent of token state, and clears the marker only
|
||
once the reload succeeds, so a disabled legacy `:8081` file can never leave the
|
||
running nginx unreloaded. The generated include is recreated before any
|
||
`nginx -t` if it is missing, so a failed run cannot wedge recovery.
|
||
Runs are serialized with `flock` on `/run/capture-dsh-token.lock`. It **never
|
||
stops or starts `dsh-web`**.
|
||
It is triggered by the `dsh-web.service` drop-in
|
||
`/etc/systemd/system/dsh-web.service.d/20-token-refresh.conf`
|
||
(`ExecStartPost=/bin/systemctl --no-block start dsh-web-token.service`) and by
|
||
`dsh-web-token.timer` every 2 minutes for reconciliation.
|
||
|
||
<details><summary>Installed systemd wiring (CT 112)</summary>
|
||
|
||
```ini
|
||
# /etc/systemd/system/dsh-web-token.service
|
||
[Unit]
|
||
Description=Refresh the dsh-web launch token for the nginx login endpoint
|
||
After=dsh-web.service
|
||
[Service]
|
||
Type=oneshot
|
||
TimeoutStartSec=180
|
||
ExecStart=/opt/deepseek-harness/capture-dsh-token.sh
|
||
|
||
# /etc/systemd/system/dsh-web-token.timer
|
||
[Unit]
|
||
Description=Periodically refresh the dsh-web login token
|
||
[Timer]
|
||
OnBootSec=90s
|
||
OnUnitActiveSec=120s
|
||
AccuracySec=10s
|
||
Persistent=true
|
||
[Install]
|
||
WantedBy=timers.target
|
||
|
||
# /etc/systemd/system/dsh-web.service.d/20-token-refresh.conf
|
||
[Service]
|
||
ExecStartPost=/bin/systemctl --no-block start dsh-web-token.service
|
||
```
|
||
|
||
</details>
|
||
|
||
> **Do NOT reintroduce the `:8081` endpoint.** It listened on `0.0.0.0:8081`
|
||
> with no `auth_request` and was a full Authentik bypass for anyone on the LAN.
|
||
> The script now removes `/etc/nginx/sites-enabled/dsh.token` automatically if
|
||
> it ever reappears.
|
||
|
||
**Authentication flow:**
|
||
1. `GET https://tankodhs.sysloggh.net/dsh-web-login`
|
||
2. Unauthenticated → Authentik sign-in; once authenticated the request reaches
|
||
dsh-web with `Host: tankodhs.sysloggh.net`.
|
||
3. dsh-web accepts the launch token on `GET /`, writes the
|
||
`dsh-auth-<authority-hash>` cookie (30 days, `HttpOnly`, `SameSite=Strict`)
|
||
and returns `303` to `/`.
|
||
4. Every later request through `/` presents that cookie; the token is not needed
|
||
again until the cookie expires or a new browser is used.
|
||
|
||
**Verification** (amdpve vantage):
|
||
```bash
|
||
# 1. Login endpoint is Authentik-gated: unauthenticated -> 302 (not 200/303).
|
||
ssh root@192.168.68.15 "pct exec 112 -- curl -s -o /dev/null -w '%{http_code}\n' \
|
||
-H 'Host: tankodhs.sysloggh.net' http://127.0.0.1/dsh-web-login"
|
||
# Expected: 302
|
||
|
||
# 2. Legacy :8081 endpoint is gone (connection refused -> 000).
|
||
ssh root@192.168.68.15 "pct exec 112 -- curl -s --max-time 3 -o /dev/null \
|
||
-w '%{http_code}\n' http://192.168.68.122:8081/"
|
||
# Expected: 000
|
||
|
||
# 3. Backend cookie mint + reuse (exactly what /dsh-web-login proxies to).
|
||
TOKEN=$(ssh root@192.168.68.15 "pct exec 112 -- cat /etc/dsh-web/launch-token")
|
||
ssh root@192.168.68.15 "pct exec 112 -- curl -s -c /tmp/dsh.jar -o /dev/null \
|
||
-H 'Host: tankodhs.sysloggh.net' 'http://127.0.0.1:3080/?token=$TOKEN'"
|
||
ssh root@192.168.68.15 "pct exec 112 -- curl -s -b /tmp/dsh.jar -o /dev/null \
|
||
-w '%{http_code}\n' -H 'Host: tankodhs.sysloggh.net' http://127.0.0.1:3080/"
|
||
# Expected: 200 — the minted dsh-auth-... cookie (authority
|
||
# tankodhs.sysloggh.net) is replayed on the next request and accepted.
|
||
|
||
# 4. Token refresh is non-disruptive and idempotent.
|
||
ssh root@192.168.68.15 "pct exec 112 -- /opt/deepseek-harness/capture-dsh-token.sh"
|
||
# Expected: "token unchanged; nginx not reloaded" when nothing changed
|
||
```
|
||
|
||
**Restart durability (acceptance):** after `systemctl restart dsh-web`, (a) the
|
||
cookie minted before the restart still returns `200` on `/`, and (b) the
|
||
refreshed `/etc/dsh-web/nginx-login.conf` carries the new token and mints a
|
||
fresh cookie. Both verified live 2026-09-11.
|
||
|
||
```bash
|
||
# 5. Cookie survives a dsh-web restart, and the new token mints a new cookie.
|
||
ssh root@192.168.68.15 "pct exec 112 -- systemctl restart dsh-web"
|
||
# dsh-web is Type=simple: restart returns before :3080 is listening. Bounded-poll
|
||
# until the socket answers (any status but 000) before asserting the cookie.
|
||
for i in $(seq 1 60); do
|
||
UP=$(ssh root@192.168.68.15 "pct exec 112 -- curl -s -o /dev/null -w '%{http_code}' \
|
||
-H 'Host: tankodhs.sysloggh.net' http://127.0.0.1:3080/")
|
||
[ "$UP" != "000" ] && break
|
||
sleep 2
|
||
done
|
||
ssh root@192.168.68.15 "pct exec 112 -- curl -s -b /tmp/dsh.jar -o /dev/null \
|
||
-w '%{http_code}\n' -H 'Host: tankodhs.sysloggh.net' http://127.0.0.1:3080/"
|
||
# Expected: 200 — the pre-restart cookie is still accepted.
|
||
# The restart's ExecStartPost (or the 2-minute timer) refreshes the include. A
|
||
# manual run may no-op on the flock, so poll until the include carries a token
|
||
# the running process accepts (bounded wait) before the mint+reuse check.
|
||
for i in $(seq 1 60); do
|
||
TOKEN=$(ssh root@192.168.68.15 "pct exec 112 -- sed -n 's/.*token=//p' /etc/dsh-web/nginx-login.conf | tr -d ';\n'")
|
||
CODE=$(ssh root@192.168.68.15 "pct exec 112 -- curl -s -o /dev/null -w '%{http_code}' \
|
||
-H 'Host: tankodhs.sysloggh.net' 'http://127.0.0.1:3080/?token=$TOKEN'")
|
||
[ "$CODE" = "303" ] && break
|
||
sleep 2
|
||
done
|
||
# Expected: 303 — the include now holds the token the running process accepts.
|
||
ssh root@192.168.68.15 "pct exec 112 -- curl -s -c /tmp/dsh-new.jar -o /dev/null \
|
||
-H 'Host: tankodhs.sysloggh.net' 'http://127.0.0.1:3080/?token=$TOKEN'"
|
||
ssh root@192.168.68.15 "pct exec 112 -- curl -s -b /tmp/dsh-new.jar -o /dev/null \
|
||
-w '%{http_code}\n' -H 'Host: tankodhs.sysloggh.net' http://127.0.0.1:3080/"
|
||
# Expected: 200 — the refreshed token minted a fresh cookie.
|
||
```
|
||
|
||
|
||
### Step 4: Platform C — Agent Zero (kagentz, CT 105 via Docker host .14)
|
||
|
||
> **The kagentz Zulip adapter leg is retired (2026-09-12).** Its code
|
||
> (`/a0/usr/kagentz-zulip/`) no longer exists in the agent-zero container, so
|
||
> the former adapter-process and heartbeat/queue checks always failed and the
|
||
> monitor issued a restart for something that could not start, posting a false
|
||
> kagentz-adapter-down alert on every run. Do NOT re-add an adapter-process,
|
||
> heartbeat/queue, or adapter-restart step. Agent Zero is probed for A2A
|
||
> liveness only, and a probe must never restart a platform.
|
||
|
||
**C1: A2A Server Health**
|
||
|
||
```bash
|
||
# A2A listens on :80 inside the agent-zero container (host-mapped to :50080) and
|
||
# is auth-gated: an unauthenticated probe gets 401, which means the server is up.
|
||
ssh root@192.168.68.14 "docker exec agent-zero curl -s --connect-timeout 5 -o /dev/null -w '%{http_code}' http://127.0.0.1:80/a2a/"
|
||
```
|
||
|
||
Expected: `401` (auth-gated, A2A server is up and responding) or `200` (if no auth required). Connection refused (`000`) → A2A server down. Any other status → running but unexpected: log/report it, never restart.
|
||
|
||
**C2: A2A Response Verification**
|
||
|
||
```bash
|
||
# A2A listens on :80 inside the container and is auth-gated (401 expected unauthenticated).
|
||
ssh root@192.168.68.14 "docker exec agent-zero curl -s -X POST http://127.0.0.1:80/a2a \
|
||
-H 'Content-Type: application/json' \
|
||
-H 'Authorization: Bearer $LITELLM_KEY' \
|
||
-d '{\"jsonrpc\":\"2.0\",\"method\":\"tasks/send\",\"params\":{\"message\":{\"role\":\"user\",\"parts\":[{\"text\":\"ping\"}]}},\"id\":1}'"
|
||
```
|
||
|
||
Expected: task ID with "working" status. Poll for completion with `tasks/get`. If 401, check LITELLM_KEY is set.
|
||
|
||
**Platform C Actions**
|
||
|
||
| Condition | Action |
|
||
|-----------|--------|
|
||
| A2A returns `000` (connection refused/timeout) | Alert only — never restart the platform; investigate the agent-zero container |
|
||
| A2A returns a status other than `200`/`401` | Log/report as a warning — reported, never healed on |
|
||
| LiteLLM 401 | Check API key in a2a_agent.py `LITELLM_KEY` |
|
||
|
||
### Step 5: Global Checks
|
||
|
||
**Cross-Agent Echo Loop Detection**
|
||
|
||
Check each agent's log for excessive bot-to-bot chatter:
|
||
- Abiba: `Skipped.*bot msgs` count
|
||
- Tanko: Repeated DM exchanges between bots
|
||
|
||
If any bot processes >50 bot-originated messages in 15min → warning.
|
||
|
||
### Step 6: Compile and Report
|
||
|
||
1. Compile all platform checks and severity
|
||
2. Determine `overall_severity` from worst per-agent severity
|
||
3. If restart action needed, check `/tmp/zulip-monitor-debounce` — apply only if >300s since last restart
|
||
4. Log full diagnostic to `/root/zulip-health-monitor.log` with timestamp
|
||
5. If any agent critical or >2 degraded: send relay message to user
|
||
6. Update `last_check` timestamp in `### Maintains` snapshot
|
||
|
||
### Restart Debounce
|
||
|
||
All restart actions MUST debounce: minimum 300s between restarts.
|
||
Track via `/tmp/zulip-monitor-debounce` (unix timestamp of last restart).
|
||
|
||
---
|
||
|
||
## History
|
||
|
||
### Gen 5 (2026-07-02) — Rate Limit Death Spiral Fix
|
||
|
||
**Root Cause**: Proactive Queue Rotation at 25 min triggered queue re-registration every cycle. Each re-registration + retry loop (3 attempts) + monitor restart = 8-12 API calls per cycle. Combined with monitor's own API calls (server check, stream alerts), `abiba-bot` hit Zulip's rate limit (429 RATE_LIMIT_HIT). Each restart reset the cycle, creating a death spiral: 111 restarts in 24 hours.
|
||
|
||
**Fixes — Extension (`index.js`)**:
|
||
1. **Rotation extended to 55 min** (from 25) with **±90s jitter** — avoids aligning with cron/monitor cycles
|
||
2. **Rate-limit-aware retry** — if connection fails with 429, skip the retry loop entirely, wait 120s, try once
|
||
|
||
**Fixes — Monitor (`zulip-monitor.sh`)**:
|
||
3. **Smart Triage** instead of instant restart:
|
||
- **Rate-limit detection**: If error log shows recent 429s, wait — don't add more API load
|
||
- **Self-healing detection**: If `retry_count` is 1-2, extension is already retrying — don't interrupt
|
||
- **Rotation window awareness**: If `queue_age` is 25-35min, disconnection is likely transient rotation — wait
|
||
- **Persistent failure threshold**: Only restart after 3 consecutive failed checks (45 min) AND no self-healing in progress
|
||
- **Debounce gate**: Skip all checks entirely if recently restarted
|
||
|
||
### Gen 4 (2026-06-29) — Response Pipeline Fix
|
||
|
||
**Root Cause**: LLM hangs due to GPU saturation (503 QUEUE_TIMEOUT) → pi's `agent_end` never fires → pending Zulip replies accumulate with no timeout. Health endpoint showed `connected: true, stuck: false` while user experienced complete silence.
|
||
|
||
**Four-Layer Defense**:
|
||
1. **Response Watchdog** — 30s deadline timer; 5min placeholder edit; 10min error message + dequeue
|
||
2. **Queue Health Ping** — Every 2min `/api/v1/events?dont_block=true` to detect silent expiry
|
||
3. **Proactive Queue Rotation** — New queue every 25min + 600 empty poll reconnection threshold
|
||
4. **Health Endpoint v2** — Added `response_pipeline`, `pending_count`, `queue_healthy`, `agent_busy_duration_seconds`
|
||
|
||
### Gen 3 (2026-06-15) — Stuck Detection
|
||
|
||
Added `stuck: bool` and `idle_seconds` to health endpoint. Monitor restarts on `stuck: true`. Added 300s restart debounce. Queue re-registers after 30min of no events.
|