fix: credential-sourcing fix for litellm-health and infrastructure-monitoring
- litellm-health.prose.md: Make monitor key retrieval explicit (ssh from CT 100 to CT 116) with executable commands; add syslog-auto alias; document credential-missing failure condition (not bare 401 or 0 keys) - infrastructure-monitoring.prose.md: Fix Zulip POST probe to retrieve keys from CT 116 via ssh instead of sourcing local env file that doesn't exist on executor host - Verify model inference probes return 200 for gpu-dense, gpu-vision, strix-moe, syslog-auto with corrected credential retrieval
This commit is contained in:
@@ -129,9 +129,11 @@ the any-HTTP rule. On those — the authenticated Zulip POST and the router
|
||||
pwd -P
|
||||
|
||||
# Zulip API health (POST ping)
|
||||
source /etc/litellm-monitor.env
|
||||
# NOTE: /etc/litellm-monitor.env exists only on CT 116, retrieve keys from CT 116 via:
|
||||
monitor_key=$(ssh root@192.168.68.116 "grep LITELLM_MONITOR_KEY /etc/litellm-monitor.env | cut -d= -f2")
|
||||
zulip_key=$(ssh root@192.168.68.116 "grep ZULIP_BOT_KEY /etc/zulip-bot.env | cut -d= -f2" 2>/dev/null || echo "not-found")
|
||||
ZULIP_USER="abiba-bot@chat.sysloggh.net"
|
||||
curl -s -o /dev/null -w '%{http_code}' -X POST https://chat.sysloggh.net/api/v1/messages -u "${ZULIP_USER}:${ZULIP_BOT_KEY}"
|
||||
curl -s -o /dev/null -w '%{http_code}' -X POST https://chat.sysloggh.net/api/v1/messages -u "${ZULIP_USER}:${zulip_key}"
|
||||
# Expected: 200 (HTTP 000 = unreachable/cache)
|
||||
|
||||
# PM2 process health
|
||||
|
||||
Reference in New Issue
Block a user