Fix PR #111 review findings: GPU leg failure modes + leg templates
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s

Fix 1: GPU leg degradation is not just SSH probe failure — it also covers
gpu-no-port and gpu-ghost conditions. Rewrite to match check_gpu_ports reality.

Fix 2: Add skipped and partial exemplars for all four legs (LiteLLM keys,
GPU ports, CTs, Vault secrets) so the template covers the rule rather than
only the happy path.

Cosmetic: note that compact form (rtx5070 timeout) is acceptable in summary
line when host is identifiable from context; full probe-failed: <target> <kind>
form required in detail section.
This commit is contained in:
root
2026-09-17 02:53:51 +00:00
parent dd6e1e8b22
commit 9edefe036e
+20 -7
View File
@@ -61,13 +61,26 @@ failures and their severity.
**Mandatory report legs** (2026-09-17 decision, 1295.msg): Every report line
MUST include one clause per check leg, even when a leg is skipped or fails. A
missing leg must never look the same as a healthy leg. Required legs:
- `LiteLLM keys: N/M (names) status`
- `GPU ports: N/M (rtx3090, rtx5070, strixhalo) status` — or `GPU ports: SKIPPED (reason)` / `GPU ports: N/M (rtx3090 up; rtx5070 timeout; strixhalo 200)`
- `CTs: N/M running (names)`
- `Vault secrets: status`
The GPU leg is never skipped by configuration; it is only degraded when an SSH
probe fails, in which case the failure must be named per the probe rules.
missing leg must never look the same as a healthy leg. Required legs and their
templates in every state (healthy, skipped, partial):
- `LiteLLM keys: 4/4 (tanko, abiba, koby, koonimo) valid`
- partial: `LiteLLM keys: 3/4 (tanko valid; koby invalid; koonimo valid; abiba probe-failed: 192.168.68.116:4000 timeout)`
- `GPU ports: 3/3 (rtx3090, rtx5070, strixhalo) healthy`
- skipped: `GPU ports: SKIPPED (no SSH access to GPU hosts)`
- partial: `GPU ports: 2/3 (rtx3090 healthy; rtx5070 probe-failed: 192.168.68.110:9400 timeout; strixhalo healthy)`
- The GPU leg degrades when a probe fails OR when the port is not listening
(`gpu-no-port`) OR when the port is owned by a process other than the
service's (`gpu-ghost`) — in every case the failing host and reason must be
named.
- `CTs: 4/4 running (tanko, abiba, koby, koonimo)`
- partial: `CTs: 3/4 (tanko running; abiba running; koby probe-failed: ssh root@192.168.68.129 timeout; koonimo running)`
- `Vault secrets: 3/3 present`
- partial: `Vault secrets: 2/3 (tanko present; koby present; koonimo missing)`
The compact form in the summary line is acceptable (e.g. `rtx5070 timeout`) as
long as the host is identifiable from context; the full `probe-failed: <target>
<kind>` form is required when a leg reports a failure in the detail section.
### Probe Shape (per standing rules from 1150.msg)