diff --git a/infrastructure-monitoring.prose.md b/infrastructure-monitoring.prose.md index c87c35a..66f9ed5 100644 --- a/infrastructure-monitoring.prose.md +++ b/infrastructure-monitoring.prose.md @@ -129,9 +129,11 @@ the any-HTTP rule. On those — the authenticated Zulip POST and the router pwd -P # Zulip API health (POST ping) -source /etc/litellm-monitor.env +# NOTE: /etc/litellm-monitor.env exists only on CT 116, retrieve keys from CT 116 via: +monitor_key=$(ssh root@192.168.68.116 "grep LITELLM_MONITOR_KEY /etc/litellm-monitor.env | cut -d= -f2") +zulip_key=$(ssh root@192.168.68.116 "grep ZULIP_BOT_KEY /etc/zulip-bot.env | cut -d= -f2" 2>/dev/null || echo "not-found") ZULIP_USER="abiba-bot@chat.sysloggh.net" -curl -s -o /dev/null -w '%{http_code}' -X POST https://chat.sysloggh.net/api/v1/messages -u "${ZULIP_USER}:${ZULIP_BOT_KEY}" +curl -s -o /dev/null -w '%{http_code}' -X POST https://chat.sysloggh.net/api/v1/messages -u "${ZULIP_USER}:${zulip_key}" # Expected: 200 (HTTP 000 = unreachable/cache) # PM2 process health diff --git a/litellm-health.prose.md b/litellm-health.prose.md index 5e47053..bd11d32 100644 --- a/litellm-health.prose.md +++ b/litellm-health.prose.md @@ -70,7 +70,15 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2) - LiteLLM master key for admin endpoints only (`/key/list`, `/key/generate`, `/key/info`) - The dedicated `monitor` agent key on CT 116 at `/etc/litellm-monitor.env` (root-only 0600) for model inference checks, scoped for every alias step 7 probes (`gpu-dense`, `gpu-vision`, - `strix-moe`) — the master key must never be used for inference + `strix-moe`, `syslog-auto`). Retrieve from the executor's host via: + + ``` + monitor key: ssh root@192.168.68.116 "grep LITELLM_MONITOR_KEY /etc/litellm-monitor.env | cut -d= -f2" + master key: ssh root@192.168.68.116 "docker exec harness-litellm printenv LITELLM_MASTER_KEY" + ``` + + If credentials are missing or unreadable, the probe must report `credential-missing` (not bare 401 or "0 keys"). + The master key must never be used for inference. ## GPU Fleet Topology @@ -148,9 +156,11 @@ contracts — read them there. Do not re-add retired names (`gemma-4-12b`, `gpu- - POST http://{{backend_host}}/litellm/v1/chat/completions model=gpu-dense → expect 200 (RTX 3090, .8) - POST http://{{backend_host}}/litellm/v1/chat/completions model=gpu-vision → expect 200 (RTX 5070, .110) - POST http://{{backend_host}}/litellm/v1/chat/completions model=strix-moe → expect 200 (Strix Halo, .15) - - Auth uses the dedicated `monitor` agent key, read on CT 116 from - `/etc/litellm-monitor.env` (root-only 0600). Do NOT use the master key for inference — - the master key is for admin endpoints only (`/key/list`, `/key/generate`, `/key/info`). + - Auth uses the dedicated `monitor` agent key. Retrieve via: + `ssh root@192.168.68.116 "grep LITELLM_MONITOR_KEY /etc/litellm-monitor.env | cut -d= -f2"` + Do NOT use the master key for inference — the master key is for admin endpoints only + (`/key/list`, `/key/generate`, `/key/info`). Retrieve master key via: + `ssh root@192.168.68.116 "docker exec harness-litellm printenv LITELLM_MASTER_KEY"` - KEY SCOPE: the `monitor` key MUST be scoped for all three probed aliases (`gpu-dense`, `gpu-vision`, `strix-moe`), otherwise the probe returns 403 and the host is not covered. If a probe returns 403, widen the monitor key's model list on CT 116 (add the missing