From d9368467ffd1d87c35e58f9096c7d8d8ca7fc572 Mon Sep 17 00:00:00 2001 From: root Date: Sat, 12 Sep 2026 23:13:24 +0000 Subject: [PATCH 1/2] fix: credential-sourcing fix for litellm-health and infrastructure-monitoring - litellm-health.prose.md: Make monitor key retrieval explicit (ssh from CT 100 to CT 116) with executable commands; add syslog-auto alias; document credential-missing failure condition (not bare 401 or 0 keys) - infrastructure-monitoring.prose.md: Fix Zulip POST probe to retrieve keys from CT 116 via ssh instead of sourcing local env file that doesn't exist on executor host - Verify model inference probes return 200 for gpu-dense, gpu-vision, strix-moe, syslog-auto with corrected credential retrieval --- infrastructure-monitoring.prose.md | 6 ++++-- litellm-health.prose.md | 18 ++++++++++++++---- 2 files changed, 18 insertions(+), 6 deletions(-) diff --git a/infrastructure-monitoring.prose.md b/infrastructure-monitoring.prose.md index c87c35a..66f9ed5 100644 --- a/infrastructure-monitoring.prose.md +++ b/infrastructure-monitoring.prose.md @@ -129,9 +129,11 @@ the any-HTTP rule. On those — the authenticated Zulip POST and the router pwd -P # Zulip API health (POST ping) -source /etc/litellm-monitor.env +# NOTE: /etc/litellm-monitor.env exists only on CT 116, retrieve keys from CT 116 via: +monitor_key=$(ssh root@192.168.68.116 "grep LITELLM_MONITOR_KEY /etc/litellm-monitor.env | cut -d= -f2") +zulip_key=$(ssh root@192.168.68.116 "grep ZULIP_BOT_KEY /etc/zulip-bot.env | cut -d= -f2" 2>/dev/null || echo "not-found") ZULIP_USER="abiba-bot@chat.sysloggh.net" -curl -s -o /dev/null -w '%{http_code}' -X POST https://chat.sysloggh.net/api/v1/messages -u "${ZULIP_USER}:${ZULIP_BOT_KEY}" +curl -s -o /dev/null -w '%{http_code}' -X POST https://chat.sysloggh.net/api/v1/messages -u "${ZULIP_USER}:${zulip_key}" # Expected: 200 (HTTP 000 = unreachable/cache) # PM2 process health diff --git a/litellm-health.prose.md b/litellm-health.prose.md index 5e47053..bd11d32 100644 --- a/litellm-health.prose.md +++ b/litellm-health.prose.md @@ -70,7 +70,15 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2) - LiteLLM master key for admin endpoints only (`/key/list`, `/key/generate`, `/key/info`) - The dedicated `monitor` agent key on CT 116 at `/etc/litellm-monitor.env` (root-only 0600) for model inference checks, scoped for every alias step 7 probes (`gpu-dense`, `gpu-vision`, - `strix-moe`) — the master key must never be used for inference + `strix-moe`, `syslog-auto`). Retrieve from the executor's host via: + + ``` + monitor key: ssh root@192.168.68.116 "grep LITELLM_MONITOR_KEY /etc/litellm-monitor.env | cut -d= -f2" + master key: ssh root@192.168.68.116 "docker exec harness-litellm printenv LITELLM_MASTER_KEY" + ``` + + If credentials are missing or unreadable, the probe must report `credential-missing` (not bare 401 or "0 keys"). + The master key must never be used for inference. ## GPU Fleet Topology @@ -148,9 +156,11 @@ contracts — read them there. Do not re-add retired names (`gemma-4-12b`, `gpu- - POST http://{{backend_host}}/litellm/v1/chat/completions model=gpu-dense → expect 200 (RTX 3090, .8) - POST http://{{backend_host}}/litellm/v1/chat/completions model=gpu-vision → expect 200 (RTX 5070, .110) - POST http://{{backend_host}}/litellm/v1/chat/completions model=strix-moe → expect 200 (Strix Halo, .15) - - Auth uses the dedicated `monitor` agent key, read on CT 116 from - `/etc/litellm-monitor.env` (root-only 0600). Do NOT use the master key for inference — - the master key is for admin endpoints only (`/key/list`, `/key/generate`, `/key/info`). + - Auth uses the dedicated `monitor` agent key. Retrieve via: + `ssh root@192.168.68.116 "grep LITELLM_MONITOR_KEY /etc/litellm-monitor.env | cut -d= -f2"` + Do NOT use the master key for inference — the master key is for admin endpoints only + (`/key/list`, `/key/generate`, `/key/info`). Retrieve master key via: + `ssh root@192.168.68.116 "docker exec harness-litellm printenv LITELLM_MASTER_KEY"` - KEY SCOPE: the `monitor` key MUST be scoped for all three probed aliases (`gpu-dense`, `gpu-vision`, `strix-moe`), otherwise the probe returns 403 and the host is not covered. If a probe returns 403, widen the monitor key's model list on CT 116 (add the missing -- 2.54.0 From 6f40a3be601214d8baa3da0ef3bbab25597c933b Mon Sep 17 00:00:00 2001 From: root Date: Sat, 12 Sep 2026 23:17:27 +0000 Subject: [PATCH 2/2] fix: address credential-sourcing review findings - infrastructure-monitoring: Correct Zulip key path to /etc/litellm-monitor.env (not /etc/zulip-bot.env which doesn't exist); add credential-missing check; remove unused monitor_key variable - litellm-health: Clarify that syslog-auto is a fallback pool alias, not a step 7 probe; monitor key must be scoped for all four aliases --- infrastructure-monitoring.prose.md | 13 ++++++++----- litellm-health.prose.md | 5 +++-- 2 files changed, 11 insertions(+), 7 deletions(-) diff --git a/infrastructure-monitoring.prose.md b/infrastructure-monitoring.prose.md index 66f9ed5..68544a7 100644 --- a/infrastructure-monitoring.prose.md +++ b/infrastructure-monitoring.prose.md @@ -130,11 +130,14 @@ pwd -P # Zulip API health (POST ping) # NOTE: /etc/litellm-monitor.env exists only on CT 116, retrieve keys from CT 116 via: -monitor_key=$(ssh root@192.168.68.116 "grep LITELLM_MONITOR_KEY /etc/litellm-monitor.env | cut -d= -f2") -zulip_key=$(ssh root@192.168.68.116 "grep ZULIP_BOT_KEY /etc/zulip-bot.env | cut -d= -f2" 2>/dev/null || echo "not-found") -ZULIP_USER="abiba-bot@chat.sysloggh.net" -curl -s -o /dev/null -w '%{http_code}' -X POST https://chat.sysloggh.net/api/v1/messages -u "${ZULIP_USER}:${zulip_key}" -# Expected: 200 (HTTP 000 = unreachable/cache) +zulip_key=$(ssh root@192.168.68.116 "grep ZULIP_BOT_KEY /etc/litellm-monitor.env | cut -d= -f2") +if [ -z "$zulip_key" ]; then + echo "credential-missing: ZULIP_BOT_KEY not found in /etc/litellm-monitor.env" +else + ZULIP_USER="abiba-bot@chat.sysloggh.net" + curl -s -o /dev/null -w '%{http_code}' -X POST https://chat.sysloggh.net/api/v1/messages -u "${ZULIP_USER}:${zulip_key}" + # Expected: 200 (HTTP 000 = unreachable/cache) +fi # PM2 process health pm2 jlist diff --git a/litellm-health.prose.md b/litellm-health.prose.md index bd11d32..5ac5853 100644 --- a/litellm-health.prose.md +++ b/litellm-health.prose.md @@ -161,8 +161,9 @@ contracts — read them there. Do not re-add retired names (`gemma-4-12b`, `gpu- Do NOT use the master key for inference — the master key is for admin endpoints only (`/key/list`, `/key/generate`, `/key/info`). Retrieve master key via: `ssh root@192.168.68.116 "docker exec harness-litellm printenv LITELLM_MASTER_KEY"` - - KEY SCOPE: the `monitor` key MUST be scoped for all three probed aliases (`gpu-dense`, - `gpu-vision`, `strix-moe`), otherwise the probe returns 403 and the host is not covered. + - KEY SCOPE: the `monitor` key MUST be scoped for the three probed aliases (`gpu-dense`, + `gpu-vision`, `strix-moe`) plus the `syslog-auto` fallback pool, otherwise the probe + returns 403 and the host is not covered. If a probe returns 403, widen the monitor key's model list on CT 116 (add the missing alias) and re-run — never drop the host from the probe to make the check pass. - `gemma-4-12b` was retired and returns 400 `Invalid model name` — do not re-add it to -- 2.54.0