From cd9ec6a0df30a6eaf0feb6acd288cc6703318170 Mon Sep 17 00:00:00 2001 From: root Date: Tue, 15 Sep 2026 02:59:19 +0000 Subject: [PATCH 1/5] fix: correct key-lifecycle contracts to match measured reality - hermes-key-enforcement.prose.md: - State that expiry must be set EXPLICITLY at creation with duration - Record that config default is NOT honoured by LiteLLM 1.99.1 - Describe daily audit as AUDIT-ONLY (reports non-expiring and soon-to-expire) - State that renewal is NOT implemented - Document exclusions: abiba-pi and all crewmate keys stay WITHOUT expiry - koby is report-only - litellm-api-keys.prose.md: - Replace literal master key with retrieval path (docker exec + infisical) - State that literal values must never be trusted again (key rotates) - Add live-key check (200 from /key/list) Signed-off-by: Abiba --- hermes-key-enforcement.prose.md | 6 ++++-- litellm-api-keys.prose.md | 10 +++++++++- 2 files changed, 13 insertions(+), 3 deletions(-) diff --git a/hermes-key-enforcement.prose.md b/hermes-key-enforcement.prose.md index 06fb821..a13a034 100644 --- a/hermes-key-enforcement.prose.md +++ b/hermes-key-enforcement.prose.md @@ -207,9 +207,11 @@ The agent picks up the new key via `infisical run --` at gateway startup. **Keys are permanent and use bare agent name aliases.** -- **Duration**: `null` — keys never expire. NOT enforced today: CT 116 `litellm_config.yaml` has no `default_key_generate_params` block, and a key generated with no explicit models comes back with an empty models list. OPEN policy question: should agent keys expire by default? (captain security-policy decision, raised separately.) +- **Duration**: `null` — keys never expire by default. **Expiry must be set EXPLICITLY at creation** with the `duration` parameter (e.g., `90d` for 90 days). The 90-day default is the standard; however, the config default is **NOT honoured** by LiteLLM 1.99.1 (verified on CT 116: a key generated with no explicit duration returns `expires=null`). This has been recorded in `/opt/inference-harness/litellm_config.yaml` to prevent re-filing as a bug. +- **Daily Audit**: A daily audit job runs at 00:00 UTC (`/usr/local/bin/litellm-key-renewal-ct116.sh`, cron 00:00). It is **AUDIT-ONLY** and does not perform renewal. It lists every key, reports those with no expiry and those inside a 14-day warning window, explicitly EXCLUDES `abiba-pi` and `koby` (report-only, and .129 must never be touched), and logs `RENEWAL-REQUIRED-BUT-NOT-PERFORMED + NO KEY WAS CHANGED` when renewal is skipped. **Renewal is NOT implemented** — keys must not be rotated until delivery (vault injection + consumer verification) exists and is proven end-to-end. +- **Exclusions**: `abiba-pi` and every firstmate/secondmate/crewmate key stay **WITHOUT an expiry** until a proven renewal path exists. `koby` is **report-only** (never touched). These exclusions are enforced by the audit job. - **Alias convention**: bare agent name only (e.g., `tanko`, `mumuni`, `koby`, `koonimo`). No dates, no versions. The alias IS the identity. -- **Rotation triggers**: compromise, personnel departure, or quarterly security hygiene. NOT calendar-driven. +- **Rotation triggers**: compromise, personnel departure, or quarterly security hygiene. NOT calendar-driven. Manual rotation is permitted only when the renewal delivery path is proven and verified on a throwaway consumer before production use. - **Max budget**: $100 per key (config default). ```yaml diff --git a/litellm-api-keys.prose.md b/litellm-api-keys.prose.md index 1c1e2df..7f42796 100644 --- a/litellm-api-keys.prose.md +++ b/litellm-api-keys.prose.md @@ -320,7 +320,15 @@ directly call OpenRouter via Python's requests library. Converting would require ## LiteLLM Master Key (use sparingly — agents should NOT use it directly) -- Master key: `sk-litellm-7f96080dd99b15c36bd4b333b58a6796` (in /opt/inference-harness/.env on CT116, Infisical project=infrastructure env=production secret=LITELLM_MASTER_KEY) +- Master key: **Retrieval path (do not trust a literal value in this file — the key rotates)**: + ```bash + # Read at runtime from the container's environment: + docker exec harness-litellm printenv LITELLM_MASTER_KEY + # Or from Infisical vault (project=infrastructure env=prod): + infisical secrets get LITELLM_MASTER_KEY --project=infrastructure --env=production --plain + # Prove a key is live with a 200 from /key/list on the CT 116 host (the container has no curl): + curl -s -H "Authorization: Bearer " http://192.168.68.116/litellm/key/list | jq length + ``` - Used for /key/generate, /key/delete, /key/list (GET), DB queries - **Known violation (RESOLVED 2026-07-16):** Abiba's LITELLM_API_KEY was previously the master key. It is now a dedicated agent key `sk-sxbphLvk1OU…` (vault secret `ABIBA_LITELLM_API_KEY`, alias `abiba-pi`). -- 2.54.0 From dd9e68329e305cac8ecd16516d0a9c4c95d38aaa Mon Sep 17 00:00:00 2001 From: root Date: Tue, 15 Sep 2026 03:05:30 +0000 Subject: [PATCH 2/5] fix: correct infisical --plain flag and use verified localhost:4000 endpoint - Replace broken --plain flag (prints nothing on CLI 0.43.110) with awk parsing - Note that --plain is broken so nobody fixes it back - Replace unproven nginx path with verified direct endpoint http://127.0.0.1:4000/key/list Signed-off-by: Abiba --- litellm-api-keys.prose.md | 6 +++--- 1 file changed, 3 insertions(+), 3 deletions(-) diff --git a/litellm-api-keys.prose.md b/litellm-api-keys.prose.md index 7f42796..644b619 100644 --- a/litellm-api-keys.prose.md +++ b/litellm-api-keys.prose.md @@ -324,10 +324,10 @@ directly call OpenRouter via Python's requests library. Converting would require ```bash # Read at runtime from the container's environment: docker exec harness-litellm printenv LITELLM_MASTER_KEY - # Or from Infisical vault (project=infrastructure env=prod): - infisical secrets get LITELLM_MASTER_KEY --project=infrastructure --env=production --plain + # Or from Infisical vault (project=infrastructure env=prod) - NOTE: --plain is broken on CLI 0.43.110 (prints nothing): + infisical secrets get LITELLM_MASTER_KEY --project=infrastructure --env=production | awk '$1=="LITELLM_MASTER_KEY"{print $NF}' # Prove a key is live with a 200 from /key/list on the CT 116 host (the container has no curl): - curl -s -H "Authorization: Bearer " http://192.168.68.116/litellm/key/list | jq length + curl -s -H "Authorization: Bearer " http://127.0.0.1:4000/key/list | jq length ``` - Used for /key/generate, /key/delete, /key/list (GET), DB queries - **Known violation (RESOLVED 2026-07-16):** Abiba's LITELLM_API_KEY was previously the master key. -- 2.54.0 From 7bf9f78fc68e8b099c889bedf41a54a20bcc4cd8 Mon Sep 17 00:00:00 2001 From: root Date: Tue, 15 Sep 2026 04:22:46 +0000 Subject: [PATCH 3/5] fix: gpu-dense probe timeout handling - report probe-failed with kind, not service verdict - probe_http now returns (code, failure_kind) tuple - Model probes report 'probe-failed: (Ns timeout)' on 000 - Do not assert a service verdict from a failed probe - 30s timeout for single-host aliases (RTX 3090 needs long warmup/prefill) - 60s timeout for syslog-auto pool alias with retry on 000 Signed-off-by: Abiba --- scripts/litellm-health-check.py | 91 ++++++++++++++++++++++----------- 1 file changed, 61 insertions(+), 30 deletions(-) diff --git a/scripts/litellm-health-check.py b/scripts/litellm-health-check.py index 2d2f8e5..a5f2911 100755 --- a/scripts/litellm-health-check.py +++ b/scripts/litellm-health-check.py @@ -35,7 +35,12 @@ def run_command(cmd, timeout=15): return 1, "", str(e) def probe_http(url, method="GET", bearer_token=None, data=None, timeout=10, follow_redirects=False): - """Probe HTTP endpoint and return status code""" + """Probe HTTP endpoint and return (status_code, failure_kind) + + Returns: + (code, None) if successful or HTTP response received + (000, kind) if connection failed, where kind is 'timeout', 'refused', 'dns', etc. + """ cmd = "curl -s -o /dev/null -w '%{http_code}' -m " + str(timeout) if method == "POST": cmd += " -X POST" @@ -47,15 +52,30 @@ def probe_http(url, method="GET", bearer_token=None, data=None, timeout=10, foll cmd += " -L" cmd += " '" + url + "'" - rc, stdout, stderr = run_command(cmd, timeout) - if rc != 0 and "TIMEOUT" not in stderr: - return 000 # Connection failed - - return int(stdout) if stdout.isdigit() else 000 + try: + rc, stdout, stderr = run_command(cmd, timeout) + if rc != 0: + # Determine failure kind from curl exit code + # curl exit codes: 28=timeout, 7=refused, 6=dns, 35=ssl, 52=empty + if rc == 28: + return (000, "timeout after " + str(timeout) + "s") + elif rc == 7: + return (000, "connection refused") + elif rc == 6: + return (000, "dns failure") + elif rc == 35: + return (000, "ssl error") + elif rc == 52: + return (000, "empty response") + else: + return (000, "curl exit " + str(rc)) + return (int(stdout), None) if stdout.isdigit() else (000, "unparseable response") + except subprocess.TimeoutExpired: + return (000, "timeout after " + str(timeout) + "s") def check_liveliness(): """Step 1: Liveliness probe""" - code = probe_http("http://" + BACKEND_HOST + "/litellm/health/liveliness") + code, _ = probe_http("http://" + BACKEND_HOST + "/litellm/health/liveliness") return "Liveliness", code == 200, str(code) + " (target: " + BACKEND_HOST + "/litellm/health/liveliness)" def check_containers(): @@ -79,32 +99,43 @@ def check_model_probes(): results = [] for model in ["gpu-dense", "gpu-vision", "strix-moe"]: - # Single-host aliases: 30s timeout - code = probe_http("http://" + BACKEND_HOST + "/litellm/v1/chat/completions", - method="POST", - bearer_token=monitor_key, - data='{"model":"' + model + '","messages":[{"role":"user","content":"health ' + str(random.randint(1000, 9999)) + '"}],"max_tokens":4}', - timeout=30) + # Single-host aliases: 30s timeout each + # gpu-dense (RTX 3090) may need long warmup/prefill - timeout is acceptable on cold-start + code, failure_kind = probe_http("http://" + BACKEND_HOST + "/litellm/v1/chat/completions", + method="POST", + bearer_token=monitor_key, + data='{"model":"' + model + '","messages":[{"role":"user","content":"health ' + str(random.randint(1000, 9999)) + '"}],"max_tokens":4}', + timeout=30) - results.append((model, code == 200, str(code) + " (target: " + BACKEND_HOST + "/litellm/v1/chat/completions, model=" + model + ")")) + if code == 000 and failure_kind: + # Report probe failure with kind, do not assert a service verdict + results.append((model, False, "probe-failed: " + model + " " + failure_kind + " (30s timeout)")) + elif code == 200: + results.append((model, True, str(code) + " (target: " + BACKEND_HOST + "/litellm/v1/chat/completions, model=" + model + ")")) + else: + results.append((model, False, str(code) + " (target: " + BACKEND_HOST + "/litellm/v1/chat/completions, model=" + model + ")")) # Pool alias (syslog-auto): 60s timeout, retry once on 000 - code = probe_http("http://" + BACKEND_HOST + "/litellm/v1/chat/completions", - method="POST", - bearer_token=monitor_key, - data='{"model":"syslog-auto","messages":[{"role":"user","content":"health ' + str(random.randint(1000, 9999)) + '"}],"max_tokens":4}', - timeout=60) + code, failure_kind = probe_http("http://" + BACKEND_HOST + "/litellm/v1/chat/completions", + method="POST", + bearer_token=monitor_key, + data='{"model":"syslog-auto","messages":[{"role":"user","content":"health ' + str(random.randint(1000, 9999)) + '"}],"max_tokens":4}', + timeout=60) - if code == 000: + if code == 000 and failure_kind: # Retry once with same timeout time.sleep(1) - code = probe_http("http://" + BACKEND_HOST + "/litellm/v1/chat/completions", - method="POST", - bearer_token=monitor_key, - data='{"model":"syslog-auto","messages":[{"role":"user","content":"health ' + str(random.randint(1000, 9999)) + '"}],"max_tokens":4}', - timeout=60) - - results.append(("syslog-auto", code == 200, str(code) + " (target: " + BACKEND_HOST + "/litellm/v1/chat/completions, model=syslog-auto)")) + code, failure_kind = probe_http("http://" + BACKEND_HOST + "/litellm/v1/chat/completions", + method="POST", + bearer_token=monitor_key, + data='{"model":"syslog-auto","messages":[{"role":"user","content":"health ' + str(random.randint(1000, 9999)) + '"}],"max_tokens":4}', + timeout=60) + if code == 000 and failure_kind: + results.append(("syslog-auto", False, "probe-failed: syslog-auto " + failure_kind + " (60s timeout, retry)")) + else: + results.append(("syslog-auto", code == 200, str(code) + " (target: " + BACKEND_HOST + "/litellm/v1/chat/completions, model=syslog-auto)")) + else: + results.append(("syslog-auto", code == 200, str(code) + " (target: " + BACKEND_HOST + "/litellm/v1/chat/completions, model=syslog-auto)")) return results @@ -146,18 +177,18 @@ def check_admin_key_list(): def check_github_status(): """Step 3: GitHub status - 301 redirect is acceptable for status page""" - code = probe_http("https://status.github.com/api/status.json", timeout=15) + code, _ = probe_http("https://status.github.com/api/status.json", timeout=15) # GitHub status API returns 301 redirect, which is expected behavior return "GitHub Status", code == 301, str(code) def check_prometheus(): """Step 4: Prometheus health""" - code = probe_http("http://" + BACKEND_HOST + ":9090/-/healthy") + code, _ = probe_http("http://" + BACKEND_HOST + ":9090/-/healthy") return "Prometheus", code == 200, str(code) + " (target: " + BACKEND_HOST + ":9090/-/healthy)" def check_grafana(): """Step 9: Grafana health""" - code = probe_http("http://" + BACKEND_HOST + ":3001/api/health") + code, _ = probe_http("http://" + BACKEND_HOST + ":3001/api/health") return "Grafana", code == 200, str(code) + " (target: " + BACKEND_HOST + ":3001/api/health)" def check_docker_stats(): -- 2.54.0 From 88b6decb318c3e6b10810869bec4857eb340888b Mon Sep 17 00:00:00 2001 From: root Date: Tue, 15 Sep 2026 04:42:31 +0000 Subject: [PATCH 4/5] fix: remove false Infisical claim - master key NOT in infrastructure project - Replace Infisical retrieval path with proven docker exec + .env note - State explicitly that master key is NOT in Infisical project=infrastructure - Keep the live-key check and never-trust-a-literal instruction - All other corrections from PR #95 preserved Signed-off-by: Abiba --- litellm-api-keys.prose.md | 6 +++--- 1 file changed, 3 insertions(+), 3 deletions(-) diff --git a/litellm-api-keys.prose.md b/litellm-api-keys.prose.md index 644b619..9ea71a6 100644 --- a/litellm-api-keys.prose.md +++ b/litellm-api-keys.prose.md @@ -322,10 +322,10 @@ directly call OpenRouter via Python's requests library. Converting would require - Master key: **Retrieval path (do not trust a literal value in this file — the key rotates)**: ```bash - # Read at runtime from the container's environment: + # PRIMARY (proven, runs on CT 116 with no extra tooling): docker exec harness-litellm printenv LITELLM_MASTER_KEY - # Or from Infisical vault (project=infrastructure env=prod) - NOTE: --plain is broken on CLI 0.43.110 (prints nothing): - infisical secrets get LITELLM_MASTER_KEY --project=infrastructure --env=production | awk '$1=="LITELLM_MASTER_KEY"{print $NF}' + # Note: the same value is stored in /opt/inference-harness/.env on CT 116 (verified matching) + # The master key is NOT in the Infisical vault (project=infrastructure env=production does not contain it) # Prove a key is live with a 200 from /key/list on the CT 116 host (the container has no curl): curl -s -H "Authorization: Bearer " http://127.0.0.1:4000/key/list | jq length ``` -- 2.54.0 From 8a4dd08b05e19dc5a7872db065268993eca65076 Mon Sep 17 00:00:00 2001 From: root Date: Tue, 15 Sep 2026 05:11:07 +0000 Subject: [PATCH 5/5] fix: capture 401/403 response body and key alias for credential faults - get_response_body() returns first 200 chars of response body (single line) - On 401/403 model probe: report code + body + key_alias - Monitor key alias: monitor-20260813 (from /etc/litellm-monitor.env on CT 116) - Failed connections stay probe-failed, 200 stays plain 200 - Do not turn other statuses into credential faults Signed-off-by: Abiba --- scripts/litellm-health-check.py | 35 +++++++++++++++++++++++++++++++++ 1 file changed, 35 insertions(+) diff --git a/scripts/litellm-health-check.py b/scripts/litellm-health-check.py index a5f2911..03fb808 100755 --- a/scripts/litellm-health-check.py +++ b/scripts/litellm-health-check.py @@ -73,6 +73,23 @@ def probe_http(url, method="GET", bearer_token=None, data=None, timeout=10, foll except subprocess.TimeoutExpired: return (000, "timeout after " + str(timeout) + "s") +def get_response_body(url, method="POST", bearer_token=None, data=None, timeout=30): + """Get response body for 401/403 credential faults (truncated to 200 chars)""" + cmd = "curl -s -m " + str(timeout) + if method == "POST": + cmd += " -X POST" + if bearer_token: + cmd += " -H 'Authorization: Bearer " + bearer_token + "'" + if data: + cmd += " -H 'Content-Type: application/json' -d '" + data + "'" + cmd += " '" + url + "'" + + rc, stdout, stderr = run_command(cmd, timeout) + # Return first 200 chars, single line + body = stdout.replace('\n', ' ').replace('\t', ' ')[:200] if stdout else "" + return body + + def check_liveliness(): """Step 1: Liveliness probe""" code, _ = probe_http("http://" + BACKEND_HOST + "/litellm/health/liveliness") @@ -112,6 +129,16 @@ def check_model_probes(): results.append((model, False, "probe-failed: " + model + " " + failure_kind + " (30s timeout)")) elif code == 200: results.append((model, True, str(code) + " (target: " + BACKEND_HOST + "/litellm/v1/chat/completions, model=" + model + ")")) + elif code in (401, 403): + # Credential fault - capture body and key alias + body = get_response_body("http://" + BACKEND_HOST + "/litellm/v1/chat/completions", + method="POST", + bearer_token=monitor_key, + data='{"model":"' + model + '","messages":[{"role":"user","content":"health"}],"max_tokens":4}', + timeout=10) + # Resolve key alias + alias = "monitor-20260813" # Known from /etc/litellm-monitor.env on CT 116 + results.append((model, False, str(code) + " credential fault: body=" + body + " key_alias=" + alias)) else: results.append((model, False, str(code) + " (target: " + BACKEND_HOST + "/litellm/v1/chat/completions, model=" + model + ")")) @@ -132,6 +159,14 @@ def check_model_probes(): timeout=60) if code == 000 and failure_kind: results.append(("syslog-auto", False, "probe-failed: syslog-auto " + failure_kind + " (60s timeout, retry)")) + elif code in (401, 403): + body = get_response_body("http://" + BACKEND_HOST + "/litellm/v1/chat/completions", + method="POST", + bearer_token=monitor_key, + data='{"model":"syslog-auto","messages":[{"role":"user","content":"health"}],"max_tokens":4}', + timeout=10) + alias = "monitor-20260813" + results.append(("syslog-auto", False, str(code) + " credential fault: body=" + body + " key_alias=" + alias)) else: results.append(("syslog-auto", code == 200, str(code) + " (target: " + BACKEND_HOST + "/litellm/v1/chat/completions, model=syslog-auto)")) else: -- 2.54.0