On 2026-09-19 the Agent Zero web/A2A server inside the kagentz container (CT 105) exited during its own /api/restart call at 01:12:34Z. The container stayed Up, supervisorctl kept reporting run_ui RUNNING (it supervises the wrapper that survived), and nothing restarted the HTTP server. The public URL https://kagentz.sysloggh.net/ returned the NetBird 502 page for about 20 hours.
Our own monitoring detected it: zulip-health.prose.md step C1 already says a refused A2A connection is an incident, and scripts/zulip-monitor.sh logged Result: 🔴 1 issue(s) found and fired the red notification. The outage still sat unreported because the monitoring lane summarised that run as OK and dismissed the A2A failure as "expected (no credentials configured)" - conflating C2 (response verification, which needs LITELLM_KEY) with C1 (liveness, which needs no credential).
What this changes
Run verdict is non-optimistic. The monitor's Result: line is now 🔴 INCIDENT — N issue(s) found when N > 0 and ✅ 0 issues (all healthy) otherwise, so a run with findings can no longer be summarised as healthy.
C1 / C2 / C3 are separated and documented. C1 = A2A liveness, no credential needed, 000 is an incident. C2 = A2A response verification, requires LITELLM_KEY, and a 401 there is a credential issue, not a server-down incident. C3 = public access path, new.
New C3 public access path leg. Probes https://kagentz.sysloggh.net/ - 200/302/401 alive, 502 (proxy's "upstream refused" page) or a failed connection an incident, anything else a reported warning. This is the leg that would have caught the outage from the captain's point of view.
Nothing is auto-healed. The contract's existing rule that a probe must never restart the platform is preserved and restated for C3.
The Platform C action table now carries a row per leg (C1/C2/C3).
Verification
Behavioural tests in tests/test_zulip_kagentz_legs.py run the shipped monitor in the repo's sandbox (stub ssh/curl, log rewritten) and assert on the run's own verdict: C3 502 -> incident, C3 000 -> incident, healthy control C1 401 + C3 302 -> 0 issues, C1 000 -> incident.
The pre-existing behavioural suite tests/test_mumuni_monitor_removal.py was updated to carry the C3 stub and the new leg labels, so the legs stay covered by execution rather than by source-text assertions.
Bite proof (independent re-run by firstmate outside the repo): on this head, 13 passed in the two suites; with origin/master's scripts/zulip-monitor.sh restored, the same suites give 8 failed, 5 passed - the new cases fail pre-fix.
Two test_daily_report_* cases error in both runs because scripts/daily-infra-report.py requires ZULIP_API_KEY and refuses to run without it; that script is untouched by this branch and the behaviour is by design.
Scope
Contract, monitor script, tests, registry version bump, and a one-line doc reconcile. No changes to the agent-zero container, CT 105, or the NetBird proxy. The repair itself (restarting the UI) was done separately and verified: the public URL serves the Agent Zero login page and LiteLLM POST /v1/a2a/discover returns the agent card.
## Why
On 2026-09-19 the Agent Zero web/A2A server inside the kagentz container (CT 105) exited during its own `/api/restart` call at 01:12:34Z. The container stayed `Up`, `supervisorctl` kept reporting `run_ui RUNNING` (it supervises the wrapper that survived), and nothing restarted the HTTP server. The public URL `https://kagentz.sysloggh.net/` returned the NetBird 502 page for about 20 hours.
Our own monitoring detected it: `zulip-health.prose.md` step C1 already says a refused A2A connection is an incident, and `scripts/zulip-monitor.sh` logged `Result: 🔴 1 issue(s) found` and fired the red notification. The outage still sat unreported because the monitoring lane summarised that run as `OK` and dismissed the A2A failure as "expected (no credentials configured)" - conflating C2 (response verification, which needs `LITELLM_KEY`) with C1 (liveness, which needs no credential).
## What this changes
- **Run verdict is non-optimistic.** The monitor's `Result:` line is now `🔴 INCIDENT — N issue(s) found` when `N > 0` and `✅ 0 issues (all healthy)` otherwise, so a run with findings can no longer be summarised as healthy.
- **C1 / C2 / C3 are separated and documented.** C1 = A2A liveness, no credential needed, `000` is an incident. C2 = A2A response verification, requires `LITELLM_KEY`, and a `401` there is a credential issue, not a server-down incident. C3 = public access path, new.
- **New C3 public access path leg.** Probes `https://kagentz.sysloggh.net/` - `200`/`302`/`401` alive, `502` (proxy's "upstream refused" page) or a failed connection an incident, anything else a reported warning. This is the leg that would have caught the outage from the captain's point of view.
- **Nothing is auto-healed.** The contract's existing rule that a probe must never restart the platform is preserved and restated for C3.
- The Platform C action table now carries a row per leg (C1/C2/C3).
## Verification
- Behavioural tests in `tests/test_zulip_kagentz_legs.py` run the shipped monitor in the repo's sandbox (stub `ssh`/`curl`, log rewritten) and assert on the run's own verdict: C3 502 -> incident, C3 000 -> incident, healthy control C1 401 + C3 302 -> 0 issues, C1 000 -> incident.
- The pre-existing behavioural suite `tests/test_mumuni_monitor_removal.py` was updated to carry the C3 stub and the new leg labels, so the legs stay covered by execution rather than by source-text assertions.
- **Bite proof** (independent re-run by firstmate outside the repo): on this head, `13 passed` in the two suites; with `origin/master`'s `scripts/zulip-monitor.sh` restored, the same suites give `8 failed, 5 passed` - the new cases fail pre-fix.
- Two `test_daily_report_*` cases error in both runs because `scripts/daily-infra-report.py` requires `ZULIP_API_KEY` and refuses to run without it; that script is untouched by this branch and the behaviour is by design.
## Scope
Contract, monitor script, tests, registry version bump, and a one-line doc reconcile. No changes to the agent-zero container, CT 105, or the NetBird proxy. The repair itself (restarting the UI) was done separately and verified: the public URL serves the Agent Zero login page and LiteLLM `POST /v1/a2a/discover` returns the agent card.
- (b) Changed Result line to say 'INCIDENT' when ISSUES > 0, '0 issues (all healthy)' when ISSUES = 0
- (c) Documented C1 (no credential needed), C2 (requires LITELLM_KEY) distinction
- (d) Added C3 public access path leg for https://kagentz.sysloggh.net/
- C3 treats 200/302/401 as alive, 502/000 as incident
- Added tests/test_zulip_kagentz_legs.py to verify all changes
- C3 502 → INCIDENT
- C3 000 → INCIDENT
- C1 401 + C3 302 → 0 issues, all healthy
- C1 000 → INCIDENT
Each test asserts from the run's own log/verdict, not from file text.
Prose assertions kept as secondary.
Proven to bite: run against pre-fix script (origin/master) shows all 4
behavioural cases fail because C3 leg doesn't exist and Result line
doesn't say 'INCIDENT'.
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
Why
On 2026-09-19 the Agent Zero web/A2A server inside the kagentz container (CT 105) exited during its own
/api/restartcall at 01:12:34Z. The container stayedUp,supervisorctlkept reportingrun_ui RUNNING(it supervises the wrapper that survived), and nothing restarted the HTTP server. The public URLhttps://kagentz.sysloggh.net/returned the NetBird 502 page for about 20 hours.Our own monitoring detected it:
zulip-health.prose.mdstep C1 already says a refused A2A connection is an incident, andscripts/zulip-monitor.shloggedResult: 🔴 1 issue(s) foundand fired the red notification. The outage still sat unreported because the monitoring lane summarised that run asOKand dismissed the A2A failure as "expected (no credentials configured)" - conflating C2 (response verification, which needsLITELLM_KEY) with C1 (liveness, which needs no credential).What this changes
Result:line is now🔴 INCIDENT — N issue(s) foundwhenN > 0and✅ 0 issues (all healthy)otherwise, so a run with findings can no longer be summarised as healthy.000is an incident. C2 = A2A response verification, requiresLITELLM_KEY, and a401there is a credential issue, not a server-down incident. C3 = public access path, new.https://kagentz.sysloggh.net/-200/302/401alive,502(proxy's "upstream refused" page) or a failed connection an incident, anything else a reported warning. This is the leg that would have caught the outage from the captain's point of view.Verification
tests/test_zulip_kagentz_legs.pyrun the shipped monitor in the repo's sandbox (stubssh/curl, log rewritten) and assert on the run's own verdict: C3 502 -> incident, C3 000 -> incident, healthy control C1 401 + C3 302 -> 0 issues, C1 000 -> incident.tests/test_mumuni_monitor_removal.pywas updated to carry the C3 stub and the new leg labels, so the legs stay covered by execution rather than by source-text assertions.13 passedin the two suites; withorigin/master'sscripts/zulip-monitor.shrestored, the same suites give8 failed, 5 passed- the new cases fail pre-fix.test_daily_report_*cases error in both runs becausescripts/daily-infra-report.pyrequiresZULIP_API_KEYand refuses to run without it; that script is untouched by this branch and the behaviour is by design.Scope
Contract, monitor script, tests, registry version bump, and a one-line doc reconcile. No changes to the agent-zero container, CT 105, or the NetBird proxy. The repair itself (restarting the UI) was done separately and verified: the public URL serves the Agent Zero login page and LiteLLM
POST /v1/a2a/discoverreturns the agent card.