fix(monitoring): probe precision - any HTTP status means alive, failed probes never become service verdicts #92

Merged
abiba-bot merged 2 commits from fix/probe-precision-20260914 into master 2026-09-14 15:07:36 +00:00
Owner

Stops two monitoring contracts from announcing outages that are not happening.

The reported failures (all disproved at the authority, same host, minutes later)

From the 14:16-14:17 UTC runs:

contract reported measured
Grafana: 000 (down) http://192.168.68.116:3001/ -> 302
Zulip API: 000 (public URL failing) https://chat.sysloggh.net/api/v1/server_settings -> 200
Public URL: 000 https://chat.sysloggh.net/ -> 302
Abiba extension: 000 http://127.0.0.1:9200/health -> 200 {"zulip":{"connected":true,...}}
GPU exporters: 404 (all) http://192.168.68.8:9400/metrics and .110:9400/metrics -> 200
Prometheus: 302 (reported as a fault) :9090/-/healthy -> 200

The lane's own proxmox-monitor run 3.5 hours earlier had measured Grafana 200 on the same endpoint, so the contracts also disagreed with each other. A line like "public URL failing" is exactly the kind of claim that gets escalated to the captain; it was false.

Changes

  • infrastructure-monitoring.prose.md (127+/103-): three standing probe rules - any HTTP status means ALIVE; a failed connection is probe-failed: <target> <kind> with the target and kind named; every number names the probe that produced it. Each leg (GPU exporters, LiteLLM health, PVE API, Prometheus, Grafana, Docker metrics, Zulip post) now captures its own status code, retries once at a longer timeout on 000, and reports the code. The GPU exporters are probed per host on /metrics (the scrape path) via a loop over .8, .110, .15 - coverage is unchanged, no leg was deleted to make the output clean.
  • zulip-health.prose.md (32+/4-): the same three rules plus retry/probe-failed shape for the Zulip API leg and the Abiba extension leg, with an explicit note that the extension probe MUST run on the Abiba host (CT 100) where 127.0.0.1:9200 is the extension - probed from anywhere else that leg can never succeed.

Verification performed by firstmate before this PR

  • Every URL above measured directly, with the returned status recorded (table).
  • Re-read the whole three-dot diff for coverage: the GPU exporter loop still covers all three hosts; no monitoring leg was removed. Expect exactly two files, +159/-107.
  • Branch base is 5ed6f817 (the merged #91 branch head); the three-dot diff against master (6a1c496) is the two files listed, i.e. #91's content is not re-included.

Reviewers: run the probe commands for the Zulip API and extension legs by hand and confirm they return real codes; confirm an auth-gated 401 and a redirect are reported as ALIVE rather than as faults; confirm the extension leg names the host it must run on; and confirm no monitored target was dropped in the rewrite.

Stops two monitoring contracts from announcing outages that are not happening. ## The reported failures (all disproved at the authority, same host, minutes later) From the 14:16-14:17 UTC runs: | contract reported | measured | |---|---| | `Grafana: 000 (down)` | `http://192.168.68.116:3001/` -> **302** | | `Zulip API: 000 (public URL failing)` | `https://chat.sysloggh.net/api/v1/server_settings` -> **200** | | `Public URL: 000` | `https://chat.sysloggh.net/` -> **302** | | `Abiba extension: 000` | `http://127.0.0.1:9200/health` -> **200** `{"zulip":{"connected":true,...}}` | | `GPU exporters: 404 (all)` | `http://192.168.68.8:9400/metrics` and `.110:9400/metrics` -> **200** | | `Prometheus: 302` (reported as a fault) | `:9090/-/healthy` -> **200** | The lane's own `proxmox-monitor` run 3.5 hours earlier had measured Grafana 200 on the same endpoint, so the contracts also disagreed with each other. A line like "public URL failing" is exactly the kind of claim that gets escalated to the captain; it was false. ## Changes - **`infrastructure-monitoring.prose.md`** (127+/103-): three standing probe rules - any HTTP status means ALIVE; a failed connection is `probe-failed: <target> <kind>` with the target and kind named; every number names the probe that produced it. Each leg (GPU exporters, LiteLLM health, PVE API, Prometheus, Grafana, Docker metrics, Zulip post) now captures its own status code, retries once at a longer timeout on 000, and reports the code. The GPU exporters are probed per host on `/metrics` (the scrape path) via a loop over `.8`, `.110`, `.15` - **coverage is unchanged, no leg was deleted to make the output clean**. - **`zulip-health.prose.md`** (32+/4-): the same three rules plus retry/probe-failed shape for the Zulip API leg and the Abiba extension leg, with an explicit note that the extension probe MUST run on the Abiba host (CT 100) where `127.0.0.1:9200` is the extension - probed from anywhere else that leg can never succeed. ## Verification performed by firstmate before this PR - Every URL above measured directly, with the returned status recorded (table). - Re-read the whole three-dot diff for coverage: the GPU exporter loop still covers all three hosts; no monitoring leg was removed. Expect exactly two files, +159/-107. - Branch base is `5ed6f817` (the merged #91 branch head); the three-dot diff against `master` (`6a1c496`) is the two files listed, i.e. #91's content is not re-included. Reviewers: run the probe commands for the Zulip API and extension legs by hand and confirm they return real codes; confirm an auth-gated 401 and a redirect are reported as ALIVE rather than as faults; confirm the extension leg names the host it must run on; and confirm no monitored target was dropped in the rewrite.
abiba-bot added 2 commits 2026-09-14 14:56:00 +00:00
Per defect report 1150.msg:
- GPU exporters: probe /metrics (Prometheus scrape target), not bare /
- Grafana: correct port 3001 (not 3000)
- All probes: print target name + full URL + HTTP code
- All probes: retry once at 25s on 000/timeout
- Apply standing rules: any HTTP status = ALIVE; only 000/timeout/refused = probe-failed

Verified: all probes now return real HTTP codes (GPU 200, Grafana 200, Router 200, LiteLLM 301)
fix: probe precision — add retry + probe-failed reporting to zulip-health
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
86d2987ad8
Per defect report 1150.msg:
- Add standing probe rules section (2026-09-14)
- Step 1 (Zulip API): retry once at 25s on 000, print target + code
- Step 2 (Platform A): retry once at 25s on 000, print target + code, note must run on Abiba host
- Apply same shape as infrastructure-monitoring: any HTTP status = ALIVE; only 000/timeout = probe-failed

Verified: all probes now return real HTTP codes (Zulip API 200, Platform A 200, Tanko 401, Agent Zero 401)
abiba-bot merged commit 4fe4f3621d into master 2026-09-14 15:07:36 +00:00
Sign in to join this conversation.
No Reviewers
No labels
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: SyslogSolution/prose-contracts#92