The busy line was rendering as:
'busy (completion timed out after retry; host healthy host healthy (200))'
because host_detail already contains 'host healthy (200)' and the prefix
also said 'host healthy'. Fixed to:
'busy (completion timed out after retry; host healthy (200))'
F1 cosmetic fix from PR #123 verify.
Three states:
- healthy: passed, exit 0 (unchanged)
- busy (completion timed out after retry AND host /health answered):
⚠️ DEGRADED line, does NOT fail the run, exit 0
- host unreachable or real fault: ❌, exit 1 (unchanged)
Summary now reports degraded count:
- All pass, no degraded: '✅ All checks passed'
- All pass, 1+ degraded: '✅ All checks passed (1 degraded: gpu-dense)'
- Some failed: '❌ Some checks failed' or '❌ Some checks failed (1 degraded: ...)'
Host health mapping verified:
- gpu-dense -> 192.168.68.8:8080/health
- gpu-vision -> 192.168.68.110:8080/health
- strix-moe -> 192.168.68.15:8080/health
1. TIMEOUT KIND FIX: run_command returns (1, '', 'TIMEOUT') when its own
timeout fires. probe_http now checks for this before falling through to
'curl exit <rc>', so a 30s timeout reports 'timeout after 30s' not
'curl exit 1'.
2. BUSY/DEGRADED DETECTION: After both model probes fail, check the
model's host health endpoint (e.g. 192.168.68.8:8080/health for
gpu-dense). If the host answers 200, report 'busy (completion timed
out after retry; host healthy 200)' — do NOT fail the run on that
alone. If the host does not answer, that's a real FAIL.
3. RETRY TIMEOUT INCREASED: Single-host retry timeout raised from 45s to
90s. Worst-case prefill on a single-slot .8 host is ~76s (observed
83K-token prompt at 1078 tok/s), so 90s covers it.
New line shapes:
- Busy: 'gpu-dense: busy (completion timed out after retry; host healthy 200)'
- Real failure: 'probe-failed: gpu-dense timeout after 30s then timeout after 90s (2 attempts)'
The failure line now preserves both attempts' failure kinds instead of
hardcoding 'timeout after retry, 45s'. If both attempts fail, the report
shows: 'probe-failed: <model> <first kind> then <retry kind> (2 attempts)'.
This fixes the self-contradictory output when the first attempt timed out
but the retry failed with connection refused, and prevents the duration
from appearing twice when both attempts were timeouts.
Example outputs:
- timeout then timeout: 'probe-failed: gpu-dense timeout after 30s then timeout after 45s (2 attempts)'
- timeout then refused: 'probe-failed: gpu-dense timeout after 30s then connection refused (2 attempts)'
- refused then refused: 'probe-failed: gpu-dense connection refused then connection refused (2 attempts)'
Single-host models (gpu-dense, gpu-vision, strix-moe) now retry once at
45s on initial 30s timeout failure before declaring probe-failed. This
prevents a single transient timeout (cold prefill ~13s or concurrent
generation hold) from failing the entire health digest.
Evidence: 2026-09-19 ~06:55Z digest failed gpu-dense at 30s; 06:56Z
direct probe 200 in 1.04s.
The failed-probe-fails-the-run property is preserved: if both attempts
fail, the script still exits non-zero with the target and duration named.
Closes: daily-health-digest false negative on single transient timeout
The key-count line in litellm-health-check now reports self-describing
output: '18 total (10 on page 1)' instead of bare '18' or '4'. Uses
total_count from the paginated API response and names what was
counted. Previous bare numbers could not reconcile changes between
runs; now a reader sees both the total and the page 1 sample.
- get_response_body() returns first 200 chars of response body (single line)
- On 401/403 model probe: report code + body + key_alias
- Monitor key alias: monitor-20260813 (from /etc/litellm-monitor.env on CT 116)
- Failed connections stay probe-failed, 200 stays plain 200
- Do not turn other statuses into credential faults
Signed-off-by: Abiba
- probe_http now returns (code, failure_kind) tuple
- Model probes report 'probe-failed: <model> <kind> (Ns timeout)' on 000
- Do not assert a service verdict from a failed probe
- 30s timeout for single-host aliases (RTX 3090 needs long warmup/prefill)
- 60s timeout for syslog-auto pool alias with retry on 000
Signed-off-by: Abiba
1. Timeout fix for pool alias (syslog-auto):
- Single-host aliases (gpu-dense, gpu-vision, strix-moe): 30s timeout
- Pool alias (syslog-auto): 60s timeout, retry once on 000 before failing
- Cold first request to pool alias can take ~13s; 10s was too short
2. Remove DEBUG prints from output:
- Removed 'DEBUG: keylen=...' and 'DEBUG: response=...' lines
- These leaked key inventory to status logs
- Success output now shows only counts (e.g., 'Admin Key List: 10 keys')
Implements all 11 checks from litellm-health contract:
- Liveliness, Containers, Prometheus, Grafana health probes
- Model probes for gpu-dense, gpu-vision, strix-moe, syslog-auto
- Admin API key list (10 keys), GitHub status, Docker Stats metrics
Fixed quoting for SSH commands and response parsing (dict with 'keys' field).
Backend edge uses internal IP 192.168.68.116, not public URL.
Docker Stats fetched from CT 116 host itself (127.0.0.1:9324/metrics).
All 11 checks passing consistently.