What this fixes (live, reproduced before the change).scripts/zulip-monitor.sh's Abiba leg read the extension health payload at the top level (d.get('connected',False)) while the state is nested at zulip.connected. The read could therefore never return true, and its own rule - "not connected -> pm2 restart abiba-zulip" - fired on every pass. Consequence measured today: four Abiba verdicts (04:23, 05:35, 06:55, 09:35 UTC), allDisconnected - restarted, and zeroConnected lines, while the chat server answered 200 in the same runs and the bot's own log showed a clean connect with continuing heartbeats. The bot was healthy; the watchdog was killing it. This also corrects two earlier fleet readings - "reconnect defect" (wrong: the bot is fine) and "restarts are the watchdog working" (wrong: they are self-inflicted).
Change.
Connection state is read from the real shape: zulip.connected / zulip.last_error. The retry-counter branch was dropped, not migrated: the endpoint emits no such field.
New fail-safe contract: a fetch error, non-2xx status, empty or unparseable body, missing zulip object, or non-boolean zulip.connected is classified probe-failed, which alerts and never restarts. pm2 restart runs only on affirmative zulip.connected=false.
A healthy-but-erroring leg reports degraded; the verdict is logged as one line.
Test coverage (this is the point of the change, not an add-on).tests/zulip-monitor-abiba.sh verbatim-extracts the leg from the script and drives it against two fixtures (zulip-health-connected.json, zulip-health-disconnected.json), asserting: connected -> no restart; disconnected -> restart; garbage/missing-key/empty body -> probe-failed alert with no restart. That pins the producer/consumer shape so this class cannot return silently. Regression suite 36/36 plus a live smoke pass; full pipeline green (intent/rebase/review/test/document/lint).
Scope discipline. Script + tests only. infrastructure-control.prose.md Sec 7 still calls this script "Disabled/Replaced" while it is the live monitor - deliberately not rewritten here: the truth of that statement depends on a deployment audit (which host runs which copy) that this diff cannot supply, and it is tracked with the measured evidence in the reconciliation item prose-zulip-monitor-executor-reconcile-20260909. Two sibling defects found on the way - daily-infra-report.py repeats the same nested-vs-top-level mismatch, and zulip-health.prose.md lists payload fields the producer never emits - are filed as monitor-consumer-shape-siblings-20260909 rather than folded in here.
Head: f57923b4fac5e4cc1932e75647d332b32c9c044f on fm/zulip-monitor-false-selfheal-20260909 (pipeline reported the PR step provider-unsupported for Gitea, so firstmate opened this PR).
**What this fixes (live, reproduced before the change).** `scripts/zulip-monitor.sh`'s Abiba leg read the extension health payload at the top level (`d.get('connected',False)`) while the state is nested at `zulip.connected`. The read could therefore never return true, and its own rule - "not connected -> `pm2 restart abiba-zulip`" - fired on every pass. Consequence measured today: four Abiba verdicts (04:23, 05:35, 06:55, 09:35 UTC), **all** `Disconnected - restarted`, and **zero** `Connected` lines, while the chat server answered 200 in the same runs and the bot's own log showed a clean connect with continuing heartbeats. The bot was healthy; the watchdog was killing it. This also corrects two earlier fleet readings - "reconnect defect" (wrong: the bot is fine) and "restarts are the watchdog working" (wrong: they are self-inflicted).
**Change.**
- Connection state is read from the real shape: `zulip.connected` / `zulip.last_error`. The retry-counter branch was **dropped**, not migrated: the endpoint emits no such field.
- New fail-safe contract: a fetch error, non-2xx status, empty or unparseable body, missing `zulip` object, or non-boolean `zulip.connected` is classified `probe-failed`, which **alerts and never restarts**. `pm2 restart` runs only on affirmative `zulip.connected=false`.
- A healthy-but-erroring leg reports `degraded`; the verdict is logged as one line.
**Test coverage (this is the point of the change, not an add-on).** `tests/zulip-monitor-abiba.sh` verbatim-extracts the leg from the script and drives it against two fixtures (`zulip-health-connected.json`, `zulip-health-disconnected.json`), asserting: connected -> no restart; disconnected -> restart; garbage/missing-key/empty body -> `probe-failed` alert with **no** restart. That pins the producer/consumer shape so this class cannot return silently. Regression suite 36/36 plus a live smoke pass; full pipeline green (intent/rebase/review/test/document/lint).
**Scope discipline.** Script + tests only. `infrastructure-control.prose.md` Sec 7 still calls this script "Disabled/Replaced" while it is the live monitor - deliberately **not** rewritten here: the truth of that statement depends on a deployment audit (which host runs which copy) that this diff cannot supply, and it is tracked with the measured evidence in the reconciliation item `prose-zulip-monitor-executor-reconcile-20260909`. Two sibling defects found on the way - `daily-infra-report.py` repeats the same nested-vs-top-level mismatch, and `zulip-health.prose.md` lists payload fields the producer never emits - are filed as `monitor-consumer-shape-siblings-20260909` rather than folded in here.
Head: `f57923b4fac5e4cc1932e75647d332b32c9c044f` on `fm/zulip-monitor-false-selfheal-20260909` (pipeline reported the PR step provider-unsupported for Gitea, so firstmate opened this PR).
The monitor's Abiba leg parsed the :9200/health payload at the top level
(d.get('connected',False)) while the pi Zulip extension serves connection
state NESTED at zulip.connected / zulip.last_error (verified against the
extension's startHealthServer handler and the live endpoint). PI_CONNECTED
was therefore always False and every run took the DISCONNECTED path,
restarting a healthy bot: pm2 restarts=8 with the process created
2026-09-09T09:35:09Z, four ❌ Abiba verdicts today (04:23/05:35/06:55/09:35
UTC) and zero ✅, while the Zulip server answered HTTP 200 and the bot kept
heartbeating. The watchdog was the fault, not the connection.
Changes (Abiba leg only; every other leg byte-identical):
- Read the real nested shape: zulip.connected, zulip.last_error and
zulip.messages_processed. The old retry_count branch is DROPPED — the
payload exposes no retry counter (the extension keeps retryCount internal
and never serialises it), so the branch is fabricated and cannot stay.
- Fail-safe restart decision: a fetch error (HTTP 000), non-2xx response,
empty/unparseable body, or payload missing a boolean zulip.connected is a
clearly-labelled PROBE FAILURE (🟠 alert + ⚠️ log line naming reason and
HTTP code) and NEVER calls pm2 restart. pm2 restart runs only on
affirmative zulip.connected=false (🔴/❌ path unchanged in wording).
connected=true with last_error keeps the degraded 🟡 warn-no-restart path.
Regression test (new tests/zulip-monitor-abiba.sh, 36 checks): extracts the
real Abiba leg from between the # -- abiba-leg-start/-end markers in the
shipped script and executes it verbatim with stubbed curl/notify/pm2 against
fixtures of the real payload shape — asserts connected=true → no restart,
connected=false → restart, and empty/garbage/missing-key/non-boolean/HTTP
000/HTTP 500 bodies → probe-failure alert with zero restarts. The suite
fails loudly on the pre-fix base and on a marker-intact top-level-parse
variant, so this class of bug cannot return silently.
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
What this fixes (live, reproduced before the change).
scripts/zulip-monitor.sh's Abiba leg read the extension health payload at the top level (d.get('connected',False)) while the state is nested atzulip.connected. The read could therefore never return true, and its own rule - "not connected ->pm2 restart abiba-zulip" - fired on every pass. Consequence measured today: four Abiba verdicts (04:23, 05:35, 06:55, 09:35 UTC), allDisconnected - restarted, and zeroConnectedlines, while the chat server answered 200 in the same runs and the bot's own log showed a clean connect with continuing heartbeats. The bot was healthy; the watchdog was killing it. This also corrects two earlier fleet readings - "reconnect defect" (wrong: the bot is fine) and "restarts are the watchdog working" (wrong: they are self-inflicted).Change.
zulip.connected/zulip.last_error. The retry-counter branch was dropped, not migrated: the endpoint emits no such field.zulipobject, or non-booleanzulip.connectedis classifiedprobe-failed, which alerts and never restarts.pm2 restartruns only on affirmativezulip.connected=false.degraded; the verdict is logged as one line.Test coverage (this is the point of the change, not an add-on).
tests/zulip-monitor-abiba.shverbatim-extracts the leg from the script and drives it against two fixtures (zulip-health-connected.json,zulip-health-disconnected.json), asserting: connected -> no restart; disconnected -> restart; garbage/missing-key/empty body ->probe-failedalert with no restart. That pins the producer/consumer shape so this class cannot return silently. Regression suite 36/36 plus a live smoke pass; full pipeline green (intent/rebase/review/test/document/lint).Scope discipline. Script + tests only.
infrastructure-control.prose.mdSec 7 still calls this script "Disabled/Replaced" while it is the live monitor - deliberately not rewritten here: the truth of that statement depends on a deployment audit (which host runs which copy) that this diff cannot supply, and it is tracked with the measured evidence in the reconciliation itemprose-zulip-monitor-executor-reconcile-20260909. Two sibling defects found on the way -daily-infra-report.pyrepeats the same nested-vs-top-level mismatch, andzulip-health.prose.mdlists payload fields the producer never emits - are filed asmonitor-consumer-shape-siblings-20260909rather than folded in here.Head:
f57923b4fac5e4cc1932e75647d332b32c9c044fonfm/zulip-monitor-false-selfheal-20260909(pipeline reported the PR step provider-unsupported for Gitea, so firstmate opened this PR).The monitor's Abiba leg parsed the :9200/health payload at the top level (d.get('connected',False)) while the pi Zulip extension serves connection state NESTED at zulip.connected / zulip.last_error (verified against the extension's startHealthServer handler and the live endpoint). PI_CONNECTED was therefore always False and every run took the DISCONNECTED path, restarting a healthy bot: pm2 restarts=8 with the process created 2026-09-09T09:35:09Z, four ❌ Abiba verdicts today (04:23/05:35/06:55/09:35 UTC) and zero ✅, while the Zulip server answered HTTP 200 and the bot kept heartbeating. The watchdog was the fault, not the connection. Changes (Abiba leg only; every other leg byte-identical): - Read the real nested shape: zulip.connected, zulip.last_error and zulip.messages_processed. The old retry_count branch is DROPPED — the payload exposes no retry counter (the extension keeps retryCount internal and never serialises it), so the branch is fabricated and cannot stay. - Fail-safe restart decision: a fetch error (HTTP 000), non-2xx response, empty/unparseable body, or payload missing a boolean zulip.connected is a clearly-labelled PROBE FAILURE (🟠 alert + ⚠️ log line naming reason and HTTP code) and NEVER calls pm2 restart. pm2 restart runs only on affirmative zulip.connected=false (🔴/❌ path unchanged in wording). connected=true with last_error keeps the degraded 🟡 warn-no-restart path. Regression test (new tests/zulip-monitor-abiba.sh, 36 checks): extracts the real Abiba leg from between the # -- abiba-leg-start/-end markers in the shipped script and executes it verbatim with stubbed curl/notify/pm2 against fixtures of the real payload shape — asserts connected=true → no restart, connected=false → restart, and empty/garbage/missing-key/non-boolean/HTTP 000/HTTP 500 bodies → probe-failure alert with zero restarts. The suite fails loudly on the pre-fix base and on a marker-intact top-level-parse variant, so this class of bug cannot return silently.