diff --git a/infrastructure-monitoring.prose.md b/infrastructure-monitoring.prose.md index c87c35a..68544a7 100644 --- a/infrastructure-monitoring.prose.md +++ b/infrastructure-monitoring.prose.md @@ -129,10 +129,15 @@ the any-HTTP rule. On those — the authenticated Zulip POST and the router pwd -P # Zulip API health (POST ping) -source /etc/litellm-monitor.env -ZULIP_USER="abiba-bot@chat.sysloggh.net" -curl -s -o /dev/null -w '%{http_code}' -X POST https://chat.sysloggh.net/api/v1/messages -u "${ZULIP_USER}:${ZULIP_BOT_KEY}" -# Expected: 200 (HTTP 000 = unreachable/cache) +# NOTE: /etc/litellm-monitor.env exists only on CT 116, retrieve keys from CT 116 via: +zulip_key=$(ssh root@192.168.68.116 "grep ZULIP_BOT_KEY /etc/litellm-monitor.env | cut -d= -f2") +if [ -z "$zulip_key" ]; then + echo "credential-missing: ZULIP_BOT_KEY not found in /etc/litellm-monitor.env" +else + ZULIP_USER="abiba-bot@chat.sysloggh.net" + curl -s -o /dev/null -w '%{http_code}' -X POST https://chat.sysloggh.net/api/v1/messages -u "${ZULIP_USER}:${zulip_key}" + # Expected: 200 (HTTP 000 = unreachable/cache) +fi # PM2 process health pm2 jlist diff --git a/litellm-health.prose.md b/litellm-health.prose.md index 5e47053..5ac5853 100644 --- a/litellm-health.prose.md +++ b/litellm-health.prose.md @@ -70,7 +70,15 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2) - LiteLLM master key for admin endpoints only (`/key/list`, `/key/generate`, `/key/info`) - The dedicated `monitor` agent key on CT 116 at `/etc/litellm-monitor.env` (root-only 0600) for model inference checks, scoped for every alias step 7 probes (`gpu-dense`, `gpu-vision`, - `strix-moe`) — the master key must never be used for inference + `strix-moe`, `syslog-auto`). Retrieve from the executor's host via: + + ``` + monitor key: ssh root@192.168.68.116 "grep LITELLM_MONITOR_KEY /etc/litellm-monitor.env | cut -d= -f2" + master key: ssh root@192.168.68.116 "docker exec harness-litellm printenv LITELLM_MASTER_KEY" + ``` + + If credentials are missing or unreadable, the probe must report `credential-missing` (not bare 401 or "0 keys"). + The master key must never be used for inference. ## GPU Fleet Topology @@ -148,11 +156,14 @@ contracts — read them there. Do not re-add retired names (`gemma-4-12b`, `gpu- - POST http://{{backend_host}}/litellm/v1/chat/completions model=gpu-dense → expect 200 (RTX 3090, .8) - POST http://{{backend_host}}/litellm/v1/chat/completions model=gpu-vision → expect 200 (RTX 5070, .110) - POST http://{{backend_host}}/litellm/v1/chat/completions model=strix-moe → expect 200 (Strix Halo, .15) - - Auth uses the dedicated `monitor` agent key, read on CT 116 from - `/etc/litellm-monitor.env` (root-only 0600). Do NOT use the master key for inference — - the master key is for admin endpoints only (`/key/list`, `/key/generate`, `/key/info`). - - KEY SCOPE: the `monitor` key MUST be scoped for all three probed aliases (`gpu-dense`, - `gpu-vision`, `strix-moe`), otherwise the probe returns 403 and the host is not covered. + - Auth uses the dedicated `monitor` agent key. Retrieve via: + `ssh root@192.168.68.116 "grep LITELLM_MONITOR_KEY /etc/litellm-monitor.env | cut -d= -f2"` + Do NOT use the master key for inference — the master key is for admin endpoints only + (`/key/list`, `/key/generate`, `/key/info`). Retrieve master key via: + `ssh root@192.168.68.116 "docker exec harness-litellm printenv LITELLM_MASTER_KEY"` + - KEY SCOPE: the `monitor` key MUST be scoped for the three probed aliases (`gpu-dense`, + `gpu-vision`, `strix-moe`) plus the `syslog-auto` fallback pool, otherwise the probe + returns 403 and the host is not covered. If a probe returns 403, widen the monitor key's model list on CT 116 (add the missing alias) and re-run — never drop the host from the probe to make the check pass. - `gemma-4-12b` was retired and returns 400 `Invalid model name` — do not re-add it to