From 64790ebd193840f70eff610da917b8f194ecfb62 Mon Sep 17 00:00:00 2001 From: agent-zero Date: Sun, 20 Sep 2026 10:53:33 -0400 Subject: [PATCH] feat(litellm): align contracts for 7-provider cloud consolidation (2026-09-20) - hermes-key-enforcement: add Model Access Tiers (local vs cloud) and cloud-scoping clause - litellm-api-keys: add Cloud Provider Consolidation section (provider map, vault secrets, access tiers) - litellm-api-keys: pin standard agent keys to explicit local-only models list; forbid {}/all-proxy-models (Community-edition cloud leak) - contract-registry: register litellm-api-keys (was unregistered drift) Additive only. Local prose-lint: PASSED. --- contract-registry.yaml | 54 ++++++++++++++++++++++++++++ hermes-key-enforcement.prose.md | 19 ++++++++++ litellm-api-keys.prose.md | 64 +++++++++++++++++++++++++++++++++ 3 files changed, 137 insertions(+) diff --git a/contract-registry.yaml b/contract-registry.yaml index ea07da0..2e58f7b 100644 --- a/contract-registry.yaml +++ b/contract-registry.yaml @@ -38,6 +38,7 @@ index: by_category: compliance: - hermes-key-enforcement + - litellm-api-keys - hermes-config-template - hermes-agent-baseline monitoring: @@ -101,6 +102,7 @@ index: proxmox: - proxmox-monitor litellm: + - litellm-api-keys - litellm-health - litellm-self-heal memory: @@ -1869,6 +1871,58 @@ contracts: drift_alerts: [] # Koby Report-Only Registry (2026-08-17 — Captain) # ⛔ KOBY IS NEVER REPAIRED — detect + report, never fix on .129 +- name: litellm-api-keys + file: litellm-api-keys.prose.md + kind: function + category: compliance + sensitivity: critical + status: active + owner: abiba + version: 1.1.0 + trigger: + type: on_demand + cadence: null + description: "Manual invocation when creating/rotating/verifying agent LiteLLM keys" + cron_job_id: null + execution: + agent: abiba + timeout: 120 + requires: [] + protocol: + - Load contract from prose-contracts/main + - Retrieve master key from Infisical (project=infrastructure env=production) + - Read live key-scoped model roster from CT 116 /v1/models + - Create/rotate/verify the requested agent key with an EXPLICIT models list + - 'Never create a key with an empty models list or all-proxy-models (Cloud leak)' + verification: + postconditions: + - check: standard agent key is local-only + verify: 'curl -s -H "Authorization: Bearer " http://192.168.68.116/litellm/v1/models | jq -r ''.data[].id'' | grep -c /' + expect: 0 cloud models + - check: key exists with correct alias + verify: 'curl -s -H "Authorization: Bearer " http://192.168.68.116/litellm/v1/key/info?key_alias=' + expect: 200 with matching alias + artifact: key creation/rotation report + receipt: + format: json + storage: ~/.hermes/runs/litellm-api-keys/ + graph_node: true + escalation: + info: + action: log_to_receipt + notify: [] + warning: + action: relay_alert + notify: + - abiba + - mumuni + critical: + action: relay_alert + notify: + - abiba + - mumuni + - ops + koby_report_only: true koby_host: "CT 111 (tdunna)" koby_ip: ".129" diff --git a/hermes-key-enforcement.prose.md b/hermes-key-enforcement.prose.md index 724c3f6..5a4746d 100644 --- a/hermes-key-enforcement.prose.md +++ b/hermes-key-enforcement.prose.md @@ -18,6 +18,25 @@ author: Abiba (pi agent) **All harness/litellm providers MUST use `api_key_env: LITELLM_API_KEY` with canonical internal path `http://192.168.68.116/litellm/v1` (Hermes appends `/v1/responses`) or public path `https://litellm.sysloggh.net/v1` — hardcoded keys AND direct `:4000` access are both forbidden. Internal `/v1` still works but is non-canonical (WARN, not FAIL).** +**Cloud provider models (OpenRouter, DeepSeek, Google AI Studio, QwenCloud PAYG/Plan, Tencent TokenHub PAYG/Plan) added to CT 116 on 2026-09-20 are reachable ONLY through the designated cloud-enabled key. Standard agent keys remain LOCAL-ONLY (`strix-moe`, `gpu-dense`, `gpu-vision`, `syslog-auto`) and MUST NOT be granted cloud models unless explicitly approved by the captain.** + +## Model Access Tiers (2026-09-20) + +CT 116 hosts two **access tiers** of model. Tier membership is enforced per virtual key via that key's `models` allowlist. + +| Tier | Models | Who gets it | +|------|--------|-------------| +| **Local** | `strix-moe`, `gpu-dense`, `gpu-vision`, `syslog-auto` | All standard agent keys (tanko, mumuni, koby, koonimo, abiba-pi) | +| **Cloud** | 47 provider models: `openrouter/*`, `deepseek/*`, `google/*`, `qwen-payg/*`, `qwen-plan/*`, `tencent-payg/*`, `tencent-plan/*` | **Only** the designated cloud-enabled key (captain decision) | + +**Rules:** +1. A key with an empty `models` list (`{}`) or `all-proxy-models` is UNSCOPED — it silently gains ALL models, including cloud. Never create or leave an agent key in this state. +2. Agent keys MUST carry an **explicit local-only** `models` list. +3. Granting a cloud model to an agent key requires explicit captain approval and a recorded reason. +4. The master key always bypasses scoping — it is admin-only, never for inference. + +See `litellm-api-keys` § Cloud Provider Consolidation for the key-creation procedure. + ## Scope Applies to all Hermes agent configs across all hosts. Covers these config sections: diff --git a/litellm-api-keys.prose.md b/litellm-api-keys.prose.md index 9c35f0f..af42840 100644 --- a/litellm-api-keys.prose.md +++ b/litellm-api-keys.prose.md @@ -71,6 +71,10 @@ description: > - Set models: read the live key-scoped set rather than hardcoding one — `/v1/models` is key-scoped, and the authoritative registry is CT 116 `/opt/inference-harness/litellm_config.yaml`. Do not add retired names (`gemma-4-12b`, `gpu-light`, `crew-auto` — all retired 2026-09-12). + - **Standard agent keys are LOCAL-ONLY**: `["strix-moe", "gpu-dense", "gpu-vision", "syslog-auto"]`. + Cloud models are granted ONLY to the designated cloud-enabled key — see § Cloud Provider + Consolidation. NEVER create a key with an empty `models` list (`{}`) or `all-proxy-models`. + In LiteLLM Community both silently grant access to EVERY model, including cloud. - Note: `ornith-1.0-35b` is NOT a valid LiteLLM model name (use `strix-moe`, the stable alias). qwen3.6-35B-A3B removed from fleet (was never deployed). - Return the new key 5. **If action == "rotate"**: @@ -87,6 +91,66 @@ description: > - Confirm key alias matches agent_name in LiteLLM key list - Verify agent gateway uses vault wrapper: `cat /proc//cmdline` shows `infisical run` +## Cloud Provider Consolidation (2026-09-20) + +CT 116 LiteLLM (Community v1.99.1) fronts **7 upstream providers** in addition to the local +GPU models. Added 2026-09-20 — 47 cloud deployments, 51 unique model names total. + +### Provider map (per-account namespacing) + +Two accounts on the same vendor get **distinct prefixes** so billing, rate limits, and keys +stay separate: + +| Prefix | Upstream | Auth | Vault secret | +|--------|----------|------|--------------| +| `openrouter/` | OpenRouter | API key | `OPENROUTER_API_KEY` | +| `deepseek/` | DeepSeek direct | API key | `DEEPSEEK_API_KEY` | +| `google/` | Google AI Studio (Gemini) | API key | `GEMINI_API_KEY` | +| `qwen-payg/` | QwenCloud / DashScope (pay-as-you-go) | API key | `DASHSCOPE_PAYG_KEY` | +| `qwen-plan/` | QwenCloud / DashScope (token plan) | API key | `DASHSCOPE_PLAN_KEY` | +| `tencent-payg/` | Tencent TokenHub (PAYG) | Bearer token | `TENCENT_PAYG_KEY` | +| `tencent-plan/` | Tencent TokenHub (Plan) | Bearer token | `TENCENT_PLAN_KEY` | + +All cloud `api_key` fields use `os.environ/` — the 7 secrets live in Infisical +(project=`infrastructure`, env=`production`, folder=`root`) and are injected into the +`harness-litellm` container at start. **No literal cloud keys in `litellm_config.yaml`.** + +### Access tiers (MUST be enforced per key) + +| Tier | Model names | Granted to | +|------|-------------|------------| +| **Local** | `strix-moe`, `gpu-dense`, `gpu-vision`, `syslog-auto` | every standard agent key | +| **Cloud** | the 47 provider models (`/`) | **only** the designated cloud-enabled key | + +> ⚠️ **Community-edition caveat:** LiteLLM Community does not restrict wildcard access groups +> the way Enterprise does. Access is decided by each key's explicit `models` list. A key with +> `models = {}` or `models = ["all-proxy-models"]` sees **all** models — a silent cloud leak. +> Every key MUST carry an explicit list. The master key always bypasses scoping (admin-only). + +### Creating the cloud-enabled key + +```bash +# ALWAYS read the live roster first (key-scoped): +curl -s -H "Authorization: Bearer " http://192.168.68.116/litellm/v1/models \ + | jq -r '.data[].id' + +# Then generate a key with an EXPLICIT model list (never empty, never a wildcard). +# For the cloud-enabled key, list local + cloud. For a standard agent, local only. +``` + +Verify after any key change: a standard agent key must return **4 models**, and must NOT return +any `/` cloud model. + +### Adding a new cloud provider + +1. Add the upstream key to Infisical `infrastructure/production/root`. +2. Add the deployment(s) to `/opt/inference-harness/litellm_config.yaml` with an `os.environ/` ref + and a namespaced `model_name` (`-/` when a vendor has >1 account). +3. Restart the `harness-litellm` container. +4. Grant the model to the cloud-enabled key ONLY (explicit list) — never to agent keys without + captain approval. +5. Update this table and the access-tier section. + ## Production Vault Access Process (canonical, 2026-07-17) The non-fail approach to agentic vault access. Deployed on all 4 Hermes agents