From 19ed186d0ad83838d9ac2572bd2305950b98755e Mon Sep 17 00:00:00 2001 From: abiba-bot Date: Thu, 10 Sep 2026 02:04:07 +0000 Subject: [PATCH] no-mistakes(document): Scope gpu-monitor liveness rule to bare-200 probes --- gpu-monitor.prose.md | 38 +++++++++++++++++++++----------------- 1 file changed, 21 insertions(+), 17 deletions(-) diff --git a/gpu-monitor.prose.md b/gpu-monitor.prose.md index 05098e1..881bac4 100644 --- a/gpu-monitor.prose.md +++ b/gpu-monitor.prose.md @@ -72,16 +72,21 @@ This replaces the previous DM-only delivery. All agents on the mesh can see and ## Alert Thresholds -### Liveness rule (any-HTTP-response) +### Liveness rule (scoped) -A probe is **ALIVE** if the endpoint returns **ANY** HTTP status — including -redirects and auth challenges. A bare `200` is not required. **DOWN = connection -refused (`000`) or timeout only.** This is the same rule zulip-health adopted for -Tanko (loopback `:3080` + public URL). Applied here: the router's -`/health/unified` answers `301 Moved Permanently` → `/gpu/gpu-data` (the same -payload), so `301` is healthy and a bare-200 expectation would false-alarm. -Statuses outside the expected set on an otherwise-alive endpoint are reported as -a warning, never as DOWN. +The any-HTTP-response rule applies ONLY to redirect/auth-gated liveness +endpoints, where any HTTP answer proves a listener is up. Applied here: the +router's `/health/unified` answers `301 Moved Permanently` → `/gpu/gpu-data` +(the same payload) and LiteLLM's `/litellm/health` answers `301` → +`/litellm/health/liveliness`. For those endpoints a probe is **ALIVE** on +**ANY** HTTP status — `3xx` redirects and `401`/`403` auth challenges included — +and **DOWN = connection refused (`000`) or timeout only**. Same scoped rule as +zulip-health (Tanko) and infrastructure-monitoring. + +Probes whose success condition is specifically a bare `200` are NOT covered by +the any-HTTP rule. On those — the GPU `:8080/health` endpoints, the router +`/health`, and the dashboard — an unexpected status (`401`/`403`, `5xx`, or +anything other than the expected `200`) is an **ALERT**, not "alive". | Metric | Warning | Critical | |--------|---------|----------| @@ -133,9 +138,9 @@ pwd -P # GPU host health — DIRECT on :8080. NEVER probe bare port 80 on a GPU host: # http://192.168.68.8/health has no listener and returns 000 → false DEGRADED. curl -s -o /dev/null -w '%{http_code}' http://192.168.68.8:8080/health -# Expected: 200 (alive = any HTTP status; 000/timeout = DOWN) +# Expected: 200 (bare-200 probe: any other status is an alert; 000/timeout = DOWN) curl -s -o /dev/null -w '%{http_code}' http://192.168.68.110:8080/health -# Expected: 200 (alive = any HTTP status; 000/timeout = DOWN) +# Expected: 200 (bare-200 probe: any other status is an alert; 000/timeout = DOWN) # Router unified health (source of truth; 301 → /gpu/gpu-data is HEALTHY) curl -s -o /dev/null -w '%{http_code}' http://192.168.68.116/health/unified @@ -143,7 +148,7 @@ curl -s -o /dev/null -w '%{http_code}' http://192.168.68.116/health/unified # Router basic health (via nginx on port 80 — router .116 only, never a GPU host) curl -s -o /dev/null -w '%{http_code}' http://192.168.68.116/health -# Expected: 200 +# Expected: 200 (bare-200 probe: any other status is an alert; 000/timeout = DOWN) # LiteLLM health (via nginx on port 80) curl -s -o /dev/null -w '%{http_code}' http://192.168.68.116/litellm/health @@ -151,16 +156,15 @@ curl -s -o /dev/null -w '%{http_code}' http://192.168.68.116/litellm/health # Dashboard curl -s -o /dev/null -w '%{http_code}' http://192.168.68.116/dashboard/ -# Expected: 200 +# Expected: 200 (bare-200 probe: any other status is an alert; 000/timeout = DOWN) ``` **Report format**: Begin every report with the **absolute path the probe executed from** (`pwd -P`, or the monitor script's absolute path) so a stale-consumer report is distinguishable from a real fault at read time. Summarize actual -results from each probe. Apply the any-HTTP-response liveness rule above: only -connection-refused (`000`) or timeout is DOWN. Flag an alert only when a probe is -DOWN, or when an alive endpoint returns an unexpected status. Never probe a GPU -host on bare port 80. +results from each probe. Apply the scoped liveness rule above: on auth-gated +endpoints only connection-refused (`000`) or timeout is DOWN; on bare-200 probes +any other status is an alert. Never probe a GPU host on bare port 80. ### view-dashboard Open `http://localhost:9100/` in browser — Live HTML dashboard