no-mistakes(document): Scope gpu-monitor liveness rule to bare-200 probes
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s
This commit is contained in:
+21
-17
@@ -72,16 +72,21 @@ This replaces the previous DM-only delivery. All agents on the mesh can see and
|
||||
|
||||
## Alert Thresholds
|
||||
|
||||
### Liveness rule (any-HTTP-response)
|
||||
### Liveness rule (scoped)
|
||||
|
||||
A probe is **ALIVE** if the endpoint returns **ANY** HTTP status — including
|
||||
redirects and auth challenges. A bare `200` is not required. **DOWN = connection
|
||||
refused (`000`) or timeout only.** This is the same rule zulip-health adopted for
|
||||
Tanko (loopback `:3080` + public URL). Applied here: the router's
|
||||
`/health/unified` answers `301 Moved Permanently` → `/gpu/gpu-data` (the same
|
||||
payload), so `301` is healthy and a bare-200 expectation would false-alarm.
|
||||
Statuses outside the expected set on an otherwise-alive endpoint are reported as
|
||||
a warning, never as DOWN.
|
||||
The any-HTTP-response rule applies ONLY to redirect/auth-gated liveness
|
||||
endpoints, where any HTTP answer proves a listener is up. Applied here: the
|
||||
router's `/health/unified` answers `301 Moved Permanently` → `/gpu/gpu-data`
|
||||
(the same payload) and LiteLLM's `/litellm/health` answers `301` →
|
||||
`/litellm/health/liveliness`. For those endpoints a probe is **ALIVE** on
|
||||
**ANY** HTTP status — `3xx` redirects and `401`/`403` auth challenges included —
|
||||
and **DOWN = connection refused (`000`) or timeout only**. Same scoped rule as
|
||||
zulip-health (Tanko) and infrastructure-monitoring.
|
||||
|
||||
Probes whose success condition is specifically a bare `200` are NOT covered by
|
||||
the any-HTTP rule. On those — the GPU `:8080/health` endpoints, the router
|
||||
`/health`, and the dashboard — an unexpected status (`401`/`403`, `5xx`, or
|
||||
anything other than the expected `200`) is an **ALERT**, not "alive".
|
||||
|
||||
| Metric | Warning | Critical |
|
||||
|--------|---------|----------|
|
||||
@@ -133,9 +138,9 @@ pwd -P
|
||||
# GPU host health — DIRECT on :8080. NEVER probe bare port 80 on a GPU host:
|
||||
# http://192.168.68.8/health has no listener and returns 000 → false DEGRADED.
|
||||
curl -s -o /dev/null -w '%{http_code}' http://192.168.68.8:8080/health
|
||||
# Expected: 200 (alive = any HTTP status; 000/timeout = DOWN)
|
||||
# Expected: 200 (bare-200 probe: any other status is an alert; 000/timeout = DOWN)
|
||||
curl -s -o /dev/null -w '%{http_code}' http://192.168.68.110:8080/health
|
||||
# Expected: 200 (alive = any HTTP status; 000/timeout = DOWN)
|
||||
# Expected: 200 (bare-200 probe: any other status is an alert; 000/timeout = DOWN)
|
||||
|
||||
# Router unified health (source of truth; 301 → /gpu/gpu-data is HEALTHY)
|
||||
curl -s -o /dev/null -w '%{http_code}' http://192.168.68.116/health/unified
|
||||
@@ -143,7 +148,7 @@ curl -s -o /dev/null -w '%{http_code}' http://192.168.68.116/health/unified
|
||||
|
||||
# Router basic health (via nginx on port 80 — router .116 only, never a GPU host)
|
||||
curl -s -o /dev/null -w '%{http_code}' http://192.168.68.116/health
|
||||
# Expected: 200
|
||||
# Expected: 200 (bare-200 probe: any other status is an alert; 000/timeout = DOWN)
|
||||
|
||||
# LiteLLM health (via nginx on port 80)
|
||||
curl -s -o /dev/null -w '%{http_code}' http://192.168.68.116/litellm/health
|
||||
@@ -151,16 +156,15 @@ curl -s -o /dev/null -w '%{http_code}' http://192.168.68.116/litellm/health
|
||||
|
||||
# Dashboard
|
||||
curl -s -o /dev/null -w '%{http_code}' http://192.168.68.116/dashboard/
|
||||
# Expected: 200
|
||||
# Expected: 200 (bare-200 probe: any other status is an alert; 000/timeout = DOWN)
|
||||
```
|
||||
|
||||
**Report format**: Begin every report with the **absolute path the probe executed
|
||||
from** (`pwd -P`, or the monitor script's absolute path) so a stale-consumer
|
||||
report is distinguishable from a real fault at read time. Summarize actual
|
||||
results from each probe. Apply the any-HTTP-response liveness rule above: only
|
||||
connection-refused (`000`) or timeout is DOWN. Flag an alert only when a probe is
|
||||
DOWN, or when an alive endpoint returns an unexpected status. Never probe a GPU
|
||||
host on bare port 80.
|
||||
results from each probe. Apply the scoped liveness rule above: on auth-gated
|
||||
endpoints only connection-refused (`000`) or timeout is DOWN; on bare-200 probes
|
||||
any other status is an alert. Never probe a GPU host on bare port 80.
|
||||
|
||||
### view-dashboard
|
||||
Open `http://localhost:9100/` in browser — Live HTML dashboard
|
||||
|
||||
Reference in New Issue
Block a user