fix(zulip-monitor): add C3 public access path, make run verdict non-optimistic, separate C1/C2/C3 #124

Merged
abiba-bot merged 6 commits from fm/kagentz-a2a-outage-masked-as-expected-20260919 into master 2026-09-19 22:50:06 +00:00
Owner

Why

On 2026-09-19 the Agent Zero web/A2A server inside the kagentz container (CT 105) exited during its own /api/restart call at 01:12:34Z. The container stayed Up, supervisorctl kept reporting run_ui RUNNING (it supervises the wrapper that survived), and nothing restarted the HTTP server. The public URL https://kagentz.sysloggh.net/ returned the NetBird 502 page for about 20 hours.

Our own monitoring detected it: zulip-health.prose.md step C1 already says a refused A2A connection is an incident, and scripts/zulip-monitor.sh logged Result: 🔴 1 issue(s) found and fired the red notification. The outage still sat unreported because the monitoring lane summarised that run as OK and dismissed the A2A failure as "expected (no credentials configured)" - conflating C2 (response verification, which needs LITELLM_KEY) with C1 (liveness, which needs no credential).

What this changes

  • Run verdict is non-optimistic. The monitor's Result: line is now 🔴 INCIDENT — N issue(s) found when N > 0 and ✅ 0 issues (all healthy) otherwise, so a run with findings can no longer be summarised as healthy.
  • C1 / C2 / C3 are separated and documented. C1 = A2A liveness, no credential needed, 000 is an incident. C2 = A2A response verification, requires LITELLM_KEY, and a 401 there is a credential issue, not a server-down incident. C3 = public access path, new.
  • New C3 public access path leg. Probes https://kagentz.sysloggh.net/ - 200/302/401 alive, 502 (proxy's "upstream refused" page) or a failed connection an incident, anything else a reported warning. This is the leg that would have caught the outage from the captain's point of view.
  • Nothing is auto-healed. The contract's existing rule that a probe must never restart the platform is preserved and restated for C3.
  • The Platform C action table now carries a row per leg (C1/C2/C3).

Verification

  • Behavioural tests in tests/test_zulip_kagentz_legs.py run the shipped monitor in the repo's sandbox (stub ssh/curl, log rewritten) and assert on the run's own verdict: C3 502 -> incident, C3 000 -> incident, healthy control C1 401 + C3 302 -> 0 issues, C1 000 -> incident.
  • The pre-existing behavioural suite tests/test_mumuni_monitor_removal.py was updated to carry the C3 stub and the new leg labels, so the legs stay covered by execution rather than by source-text assertions.
  • Bite proof (independent re-run by firstmate outside the repo): on this head, 13 passed in the two suites; with origin/master's scripts/zulip-monitor.sh restored, the same suites give 8 failed, 5 passed - the new cases fail pre-fix.
  • Two test_daily_report_* cases error in both runs because scripts/daily-infra-report.py requires ZULIP_API_KEY and refuses to run without it; that script is untouched by this branch and the behaviour is by design.

Scope

Contract, monitor script, tests, registry version bump, and a one-line doc reconcile. No changes to the agent-zero container, CT 105, or the NetBird proxy. The repair itself (restarting the UI) was done separately and verified: the public URL serves the Agent Zero login page and LiteLLM POST /v1/a2a/discover returns the agent card.

## Why On 2026-09-19 the Agent Zero web/A2A server inside the kagentz container (CT 105) exited during its own `/api/restart` call at 01:12:34Z. The container stayed `Up`, `supervisorctl` kept reporting `run_ui RUNNING` (it supervises the wrapper that survived), and nothing restarted the HTTP server. The public URL `https://kagentz.sysloggh.net/` returned the NetBird 502 page for about 20 hours. Our own monitoring detected it: `zulip-health.prose.md` step C1 already says a refused A2A connection is an incident, and `scripts/zulip-monitor.sh` logged `Result: 🔴 1 issue(s) found` and fired the red notification. The outage still sat unreported because the monitoring lane summarised that run as `OK` and dismissed the A2A failure as "expected (no credentials configured)" - conflating C2 (response verification, which needs `LITELLM_KEY`) with C1 (liveness, which needs no credential). ## What this changes - **Run verdict is non-optimistic.** The monitor's `Result:` line is now `🔴 INCIDENT — N issue(s) found` when `N > 0` and `✅ 0 issues (all healthy)` otherwise, so a run with findings can no longer be summarised as healthy. - **C1 / C2 / C3 are separated and documented.** C1 = A2A liveness, no credential needed, `000` is an incident. C2 = A2A response verification, requires `LITELLM_KEY`, and a `401` there is a credential issue, not a server-down incident. C3 = public access path, new. - **New C3 public access path leg.** Probes `https://kagentz.sysloggh.net/` - `200`/`302`/`401` alive, `502` (proxy's "upstream refused" page) or a failed connection an incident, anything else a reported warning. This is the leg that would have caught the outage from the captain's point of view. - **Nothing is auto-healed.** The contract's existing rule that a probe must never restart the platform is preserved and restated for C3. - The Platform C action table now carries a row per leg (C1/C2/C3). ## Verification - Behavioural tests in `tests/test_zulip_kagentz_legs.py` run the shipped monitor in the repo's sandbox (stub `ssh`/`curl`, log rewritten) and assert on the run's own verdict: C3 502 -> incident, C3 000 -> incident, healthy control C1 401 + C3 302 -> 0 issues, C1 000 -> incident. - The pre-existing behavioural suite `tests/test_mumuni_monitor_removal.py` was updated to carry the C3 stub and the new leg labels, so the legs stay covered by execution rather than by source-text assertions. - **Bite proof** (independent re-run by firstmate outside the repo): on this head, `13 passed` in the two suites; with `origin/master`'s `scripts/zulip-monitor.sh` restored, the same suites give `8 failed, 5 passed` - the new cases fail pre-fix. - Two `test_daily_report_*` cases error in both runs because `scripts/daily-infra-report.py` requires `ZULIP_API_KEY` and refuses to run without it; that script is untouched by this branch and the behaviour is by design. ## Scope Contract, monitor script, tests, registry version bump, and a one-line doc reconcile. No changes to the agent-zero container, CT 105, or the NetBird proxy. The repair itself (restarting the UI) was done separately and verified: the public URL serves the Agent Zero login page and LiteLLM `POST /v1/a2a/discover` returns the agent card.
abiba-bot added 6 commits 2026-09-19 22:44:37 +00:00
- (b) Changed Result line to say 'INCIDENT' when ISSUES > 0, '0 issues (all healthy)' when ISSUES = 0
- (c) Documented C1 (no credential needed), C2 (requires LITELLM_KEY) distinction
- (d) Added C3 public access path leg for https://kagentz.sysloggh.net/
- C3 treats 200/302/401 as alive, 502/000 as incident
- Added tests/test_zulip_kagentz_legs.py to verify all changes
- C3 502 → INCIDENT
- C3 000 → INCIDENT
- C1 401 + C3 302 → 0 issues, all healthy
- C1 000 → INCIDENT

Each test asserts from the run's own log/verdict, not from file text.
Prose assertions kept as secondary.

Proven to bite: run against pre-fix script (origin/master) shows all 4
behavioural cases fail because C3 leg doesn't exist and Result line
doesn't say 'INCIDENT'.
no-mistakes(document): docs: reconcile zulip-monitor status in infrastructure-control
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
c72436b406
abiba-bot merged commit 6fa0a255df into master 2026-09-19 22:50:06 +00:00
Sign in to join this conversation.
No Reviewers
No labels
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: SyslogSolution/prose-contracts#124