Closes the recurring failure that produced the same wrong verdict every four hours: Liveliness: 401 (auth required), Containers: 401 (auth required) on a healthy gateway.
Why
litellm-health.prose.md was the last monitoring contract with no executor script. Prose alone meant the lane improvised its commands each run and substituted the public URL, which legitimately requires authentication for those paths — so a healthy gateway was reported as needing auth, four times in one day. The contract's own steps already specified the backend edge; nothing enforced it.
agent-health-check had the same class of problem and stopped producing false verdicts once it gained scripts/agent-health-check.py and the lane was told to run the script and paste its output. This PR does the same for litellm-health.
Changes
scripts/litellm-health-check.py — implements the contract's checks with the documented endpoints and prints one line per check (name, target probed, result). Specifically: backend edge 192.168.68.116 for the health/model probes; monitor key read from /etc/litellm-monitor.env on CT 116; master key from the harness-litellm container; the /key/list admin call run on the CT 116 host (the container has no curl/wget, PR #84); one model probe per GPU host plus the syslog-auto pool; the docker-stats exporter read from 127.0.0.1:9324 on CT 116 because it binds to localhost; credential-missing and admin-call-failed as labelled outcomes rather than a bare 401 or "0 keys"; non-zero exit on real failure.
litellm-health.prose.md — an explicit executor section: run the script from the clone and paste its output; hand-rolled probes are not a substitute; backend-edge checks must not use the public URL.
Verification
The lane ran the script twice, both 11/11: liveliness 200, 12 containers, Prometheus 200, Grafana 200, four model probes 200, GitHub 301, key list 10 keys, docker-stats 200. Firstmate independently verified the two hardest endpoints before the fix: /key/list returns 10 keys via the host-side curl, and 127.0.0.1:9324/metrics returns 200 on the CT 116 host while being unreachable remotely (which is why the earlier attempts read as failures).
Reviewers should run the script from the clone and paste the per-check output, and confirm no credential is hardcoded (keys are fetched at runtime).
Closes the recurring failure that produced the same wrong verdict every four hours: `Liveliness: 401 (auth required), Containers: 401 (auth required)` on a healthy gateway.
## Why
`litellm-health.prose.md` was the last monitoring contract with **no executor script**. Prose alone meant the lane improvised its commands each run and substituted the public URL, which legitimately requires authentication for those paths — so a healthy gateway was reported as needing auth, four times in one day. The contract's own steps already specified the backend edge; nothing enforced it.
`agent-health-check` had the same class of problem and stopped producing false verdicts once it gained `scripts/agent-health-check.py` and the lane was told to run the script and paste its output. This PR does the same for litellm-health.
## Changes
- **`scripts/litellm-health-check.py`** — implements the contract's checks with the documented endpoints and prints one line per check (name, target probed, result). Specifically: backend edge `192.168.68.116` for the health/model probes; monitor key read from `/etc/litellm-monitor.env` on CT 116; master key from the `harness-litellm` container; the `/key/list` admin call run on the CT 116 **host** (the container has no curl/wget, PR #84); one model probe per GPU host plus the `syslog-auto` pool; the docker-stats exporter read from `127.0.0.1:9324` on CT 116 because it binds to localhost; `credential-missing` and `admin-call-failed` as labelled outcomes rather than a bare 401 or "0 keys"; non-zero exit on real failure.
- **`litellm-health.prose.md`** — an explicit executor section: run the script from the clone and paste its output; hand-rolled probes are not a substitute; backend-edge checks must not use the public URL.
## Verification
The lane ran the script twice, both **11/11**: liveliness 200, 12 containers, Prometheus 200, Grafana 200, four model probes 200, GitHub 301, key list 10 keys, docker-stats 200. Firstmate independently verified the two hardest endpoints before the fix: `/key/list` returns 10 keys via the host-side curl, and `127.0.0.1:9324/metrics` returns 200 on the CT 116 host while being unreachable remotely (which is why the earlier attempts read as failures).
Reviewers should run the script from the clone and paste the per-check output, and confirm no credential is hardcoded (keys are fetched at runtime).
Implements all 11 checks from litellm-health contract:
- Liveliness, Containers, Prometheus, Grafana health probes
- Model probes for gpu-dense, gpu-vision, strix-moe, syslog-auto
- Admin API key list (10 keys), GitHub status, Docker Stats metrics
Fixed quoting for SSH commands and response parsing (dict with 'keys' field).
Backend edge uses internal IP 192.168.68.116, not public URL.
Docker Stats fetched from CT 116 host itself (127.0.0.1:9324/metrics).
All 11 checks passing consistently.
Add 'Executor Script' section:
- Run scripts/litellm-health-check.py from the clone
- Hand-rolled probes not acceptable substitute
- Backend-edge checks use internal IP 192.168.68.116, not public URL
- Docker Stats fetched from CT 116 host (127.0.0.1:9324/metrics)
- Admin Key List requires proper quoting for SSH commands
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
Closes the recurring failure that produced the same wrong verdict every four hours:
Liveliness: 401 (auth required), Containers: 401 (auth required)on a healthy gateway.Why
litellm-health.prose.mdwas the last monitoring contract with no executor script. Prose alone meant the lane improvised its commands each run and substituted the public URL, which legitimately requires authentication for those paths — so a healthy gateway was reported as needing auth, four times in one day. The contract's own steps already specified the backend edge; nothing enforced it.agent-health-checkhad the same class of problem and stopped producing false verdicts once it gainedscripts/agent-health-check.pyand the lane was told to run the script and paste its output. This PR does the same for litellm-health.Changes
scripts/litellm-health-check.py— implements the contract's checks with the documented endpoints and prints one line per check (name, target probed, result). Specifically: backend edge192.168.68.116for the health/model probes; monitor key read from/etc/litellm-monitor.envon CT 116; master key from theharness-litellmcontainer; the/key/listadmin call run on the CT 116 host (the container has no curl/wget, PR #84); one model probe per GPU host plus thesyslog-autopool; the docker-stats exporter read from127.0.0.1:9324on CT 116 because it binds to localhost;credential-missingandadmin-call-failedas labelled outcomes rather than a bare 401 or "0 keys"; non-zero exit on real failure.litellm-health.prose.md— an explicit executor section: run the script from the clone and paste its output; hand-rolled probes are not a substitute; backend-edge checks must not use the public URL.Verification
The lane ran the script twice, both 11/11: liveliness 200, 12 containers, Prometheus 200, Grafana 200, four model probes 200, GitHub 301, key list 10 keys, docker-stats 200. Firstmate independently verified the two hardest endpoints before the fix:
/key/listreturns 10 keys via the host-side curl, and127.0.0.1:9324/metricsreturns 200 on the CT 116 host while being unreachable remotely (which is why the earlier attempts read as failures).Reviewers should run the script from the clone and paste the per-check output, and confirm no credential is hardcoded (keys are fetched at runtime).