Follow-up to PR #87 (merged). Firstmate ran the merged script from a clean checkout right after it landed and got 10/11 while both the author and the reviewer reported 11/11 - two real robustness defects.
Defect 1 - false FAIL on the pool alias
probe_http() defaults to a 10s timeout. syslog-auto is the POOL alias, so a cold first request pays full prefill on whichever host answers - measured at ~13s on this fleet. The merged script therefore printed ❌ syslog-auto: 0 while gpu-dense, gpu-vision and strix-moe all returned 200 in the same run. Re-running that exact request with a longer allowance returned 200 in 0.13s, confirming a timeout rather than a fault.
Fix: single-host aliases get 30s; the syslog-auto pool alias gets 60s and one retry on 000. Cheap GET checks keep short timeouts so a genuinely unreachable endpoint still fails promptly.
Defect 2 - the script dumped the key list into its output
The merged version printed DEBUG: keylen=43 and DEBUG: response={"keys":["bf3d58bb...", ...]}, putting the gateway's virtual-key inventory into the contract's status line and the lane's archived status log. Hashes rather than raw keys, but unnecessary exposure.
Fix: both DEBUG prints removed; the count is still reported (Admin Key List: 10 keys) and nothing else.
Verification (firstmate, from this branch, independent of the author)
scripts/litellm-health-check.py from this branch run twice: 11/11 both times, including syslog-auto: 200.
grep -c DEBUG over a full run: 0.
Reviewers: run the branch script and paste the output; confirm 11/11 and that no DEBUG line appears.
Follow-up to PR #87 (merged). Firstmate ran the merged script from a clean checkout right after it landed and got **10/11** while both the author and the reviewer reported 11/11 - two real robustness defects.
## Defect 1 - false FAIL on the pool alias
`probe_http()` defaults to a 10s timeout. `syslog-auto` is the POOL alias, so a cold first request pays full prefill on whichever host answers - measured at ~13s on this fleet. The merged script therefore printed `❌ syslog-auto: 0` while `gpu-dense`, `gpu-vision` and `strix-moe` all returned 200 in the same run. Re-running that exact request with a longer allowance returned **200 in 0.13s**, confirming a timeout rather than a fault.
**Fix:** single-host aliases get 30s; the `syslog-auto` pool alias gets 60s **and one retry on 000**. Cheap GET checks keep short timeouts so a genuinely unreachable endpoint still fails promptly.
## Defect 2 - the script dumped the key list into its output
The merged version printed `DEBUG: keylen=43` and `DEBUG: response={"keys":["bf3d58bb...", ...]}`, putting the gateway's virtual-key inventory into the contract's status line and the lane's archived status log. Hashes rather than raw keys, but unnecessary exposure.
**Fix:** both DEBUG prints removed; the count is still reported (`Admin Key List: 10 keys`) and nothing else.
## Verification (firstmate, from this branch, independent of the author)
- `scripts/litellm-health-check.py` from this branch run twice: **11/11 both times**, including `syslog-auto: 200`.
- `grep -c DEBUG` over a full run: **0**.
Reviewers: run the branch script and paste the output; confirm 11/11 and that no DEBUG line appears.
1. Timeout fix for pool alias (syslog-auto):
- Single-host aliases (gpu-dense, gpu-vision, strix-moe): 30s timeout
- Pool alias (syslog-auto): 60s timeout, retry once on 000 before failing
- Cold first request to pool alias can take ~13s; 10s was too short
2. Remove DEBUG prints from output:
- Removed 'DEBUG: keylen=...' and 'DEBUG: response=...' lines
- These leaked key inventory to status logs
- Success output now shows only counts (e.g., 'Admin Key List: 10 keys')
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
Follow-up to PR #87 (merged). Firstmate ran the merged script from a clean checkout right after it landed and got 10/11 while both the author and the reviewer reported 11/11 - two real robustness defects.
Defect 1 - false FAIL on the pool alias
probe_http()defaults to a 10s timeout.syslog-autois the POOL alias, so a cold first request pays full prefill on whichever host answers - measured at ~13s on this fleet. The merged script therefore printed❌ syslog-auto: 0whilegpu-dense,gpu-visionandstrix-moeall returned 200 in the same run. Re-running that exact request with a longer allowance returned 200 in 0.13s, confirming a timeout rather than a fault.Fix: single-host aliases get 30s; the
syslog-autopool alias gets 60s and one retry on 000. Cheap GET checks keep short timeouts so a genuinely unreachable endpoint still fails promptly.Defect 2 - the script dumped the key list into its output
The merged version printed
DEBUG: keylen=43andDEBUG: response={"keys":["bf3d58bb...", ...]}, putting the gateway's virtual-key inventory into the contract's status line and the lane's archived status log. Hashes rather than raw keys, but unnecessary exposure.Fix: both DEBUG prints removed; the count is still reported (
Admin Key List: 10 keys) and nothing else.Verification (firstmate, from this branch, independent of the author)
scripts/litellm-health-check.pyfrom this branch run twice: 11/11 both times, includingsyslog-auto: 200.grep -c DEBUGover a full run: 0.Reviewers: run the branch script and paste the output; confirm 11/11 and that no DEBUG line appears.