feat(litellm): align contracts for 7-provider cloud consolidation (2026-09-20)
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s
- hermes-key-enforcement: add Model Access Tiers (local vs cloud) and cloud-scoping clause
- litellm-api-keys: add Cloud Provider Consolidation section (provider map, vault secrets, access tiers)
- litellm-api-keys: pin standard agent keys to explicit local-only models list; forbid {}/all-proxy-models (Community-edition cloud leak)
- contract-registry: register litellm-api-keys (was unregistered drift)
Additive only. Local prose-lint: PASSED.
This commit is contained in:
@@ -38,6 +38,7 @@ index:
|
|||||||
by_category:
|
by_category:
|
||||||
compliance:
|
compliance:
|
||||||
- hermes-key-enforcement
|
- hermes-key-enforcement
|
||||||
|
- litellm-api-keys
|
||||||
- hermes-config-template
|
- hermes-config-template
|
||||||
- hermes-agent-baseline
|
- hermes-agent-baseline
|
||||||
monitoring:
|
monitoring:
|
||||||
@@ -101,6 +102,7 @@ index:
|
|||||||
proxmox:
|
proxmox:
|
||||||
- proxmox-monitor
|
- proxmox-monitor
|
||||||
litellm:
|
litellm:
|
||||||
|
- litellm-api-keys
|
||||||
- litellm-health
|
- litellm-health
|
||||||
- litellm-self-heal
|
- litellm-self-heal
|
||||||
memory:
|
memory:
|
||||||
@@ -1869,6 +1871,58 @@ contracts:
|
|||||||
drift_alerts: []
|
drift_alerts: []
|
||||||
# Koby Report-Only Registry (2026-08-17 — Captain)
|
# Koby Report-Only Registry (2026-08-17 — Captain)
|
||||||
# ⛔ KOBY IS NEVER REPAIRED — detect + report, never fix on .129
|
# ⛔ KOBY IS NEVER REPAIRED — detect + report, never fix on .129
|
||||||
|
- name: litellm-api-keys
|
||||||
|
file: litellm-api-keys.prose.md
|
||||||
|
kind: function
|
||||||
|
category: compliance
|
||||||
|
sensitivity: critical
|
||||||
|
status: active
|
||||||
|
owner: abiba
|
||||||
|
version: 1.1.0
|
||||||
|
trigger:
|
||||||
|
type: on_demand
|
||||||
|
cadence: null
|
||||||
|
description: "Manual invocation when creating/rotating/verifying agent LiteLLM keys"
|
||||||
|
cron_job_id: null
|
||||||
|
execution:
|
||||||
|
agent: abiba
|
||||||
|
timeout: 120
|
||||||
|
requires: []
|
||||||
|
protocol:
|
||||||
|
- Load contract from prose-contracts/main
|
||||||
|
- Retrieve master key from Infisical (project=infrastructure env=production)
|
||||||
|
- Read live key-scoped model roster from CT 116 /v1/models
|
||||||
|
- Create/rotate/verify the requested agent key with an EXPLICIT models list
|
||||||
|
- 'Never create a key with an empty models list or all-proxy-models (Cloud leak)'
|
||||||
|
verification:
|
||||||
|
postconditions:
|
||||||
|
- check: standard agent key is local-only
|
||||||
|
verify: 'curl -s -H "Authorization: Bearer <KEY>" http://192.168.68.116/litellm/v1/models | jq -r ''.data[].id'' | grep -c /'
|
||||||
|
expect: 0 cloud models
|
||||||
|
- check: key exists with correct alias
|
||||||
|
verify: 'curl -s -H "Authorization: Bearer <MASTER>" http://192.168.68.116/litellm/v1/key/info?key_alias=<AGENT>'
|
||||||
|
expect: 200 with matching alias
|
||||||
|
artifact: key creation/rotation report
|
||||||
|
receipt:
|
||||||
|
format: json
|
||||||
|
storage: ~/.hermes/runs/litellm-api-keys/
|
||||||
|
graph_node: true
|
||||||
|
escalation:
|
||||||
|
info:
|
||||||
|
action: log_to_receipt
|
||||||
|
notify: []
|
||||||
|
warning:
|
||||||
|
action: relay_alert
|
||||||
|
notify:
|
||||||
|
- abiba
|
||||||
|
- mumuni
|
||||||
|
critical:
|
||||||
|
action: relay_alert
|
||||||
|
notify:
|
||||||
|
- abiba
|
||||||
|
- mumuni
|
||||||
|
- ops
|
||||||
|
|
||||||
koby_report_only: true
|
koby_report_only: true
|
||||||
koby_host: "CT 111 (tdunna)"
|
koby_host: "CT 111 (tdunna)"
|
||||||
koby_ip: ".129"
|
koby_ip: ".129"
|
||||||
|
|||||||
@@ -18,6 +18,25 @@ author: Abiba (pi agent)
|
|||||||
|
|
||||||
**All harness/litellm providers MUST use `api_key_env: LITELLM_API_KEY` with canonical internal path `http://192.168.68.116/litellm/v1` (Hermes appends `/v1/responses`) or public path `https://litellm.sysloggh.net/v1` — hardcoded keys AND direct `:4000` access are both forbidden. Internal `/v1` still works but is non-canonical (WARN, not FAIL).**
|
**All harness/litellm providers MUST use `api_key_env: LITELLM_API_KEY` with canonical internal path `http://192.168.68.116/litellm/v1` (Hermes appends `/v1/responses`) or public path `https://litellm.sysloggh.net/v1` — hardcoded keys AND direct `:4000` access are both forbidden. Internal `/v1` still works but is non-canonical (WARN, not FAIL).**
|
||||||
|
|
||||||
|
**Cloud provider models (OpenRouter, DeepSeek, Google AI Studio, QwenCloud PAYG/Plan, Tencent TokenHub PAYG/Plan) added to CT 116 on 2026-09-20 are reachable ONLY through the designated cloud-enabled key. Standard agent keys remain LOCAL-ONLY (`strix-moe`, `gpu-dense`, `gpu-vision`, `syslog-auto`) and MUST NOT be granted cloud models unless explicitly approved by the captain.**
|
||||||
|
|
||||||
|
## Model Access Tiers (2026-09-20)
|
||||||
|
|
||||||
|
CT 116 hosts two **access tiers** of model. Tier membership is enforced per virtual key via that key's `models` allowlist.
|
||||||
|
|
||||||
|
| Tier | Models | Who gets it |
|
||||||
|
|------|--------|-------------|
|
||||||
|
| **Local** | `strix-moe`, `gpu-dense`, `gpu-vision`, `syslog-auto` | All standard agent keys (tanko, mumuni, koby, koonimo, abiba-pi) |
|
||||||
|
| **Cloud** | 47 provider models: `openrouter/*`, `deepseek/*`, `google/*`, `qwen-payg/*`, `qwen-plan/*`, `tencent-payg/*`, `tencent-plan/*` | **Only** the designated cloud-enabled key (captain decision) |
|
||||||
|
|
||||||
|
**Rules:**
|
||||||
|
1. A key with an empty `models` list (`{}`) or `all-proxy-models` is UNSCOPED — it silently gains ALL models, including cloud. Never create or leave an agent key in this state.
|
||||||
|
2. Agent keys MUST carry an **explicit local-only** `models` list.
|
||||||
|
3. Granting a cloud model to an agent key requires explicit captain approval and a recorded reason.
|
||||||
|
4. The master key always bypasses scoping — it is admin-only, never for inference.
|
||||||
|
|
||||||
|
See `litellm-api-keys` § Cloud Provider Consolidation for the key-creation procedure.
|
||||||
|
|
||||||
## Scope
|
## Scope
|
||||||
|
|
||||||
Applies to all Hermes agent configs across all hosts. Covers these config sections:
|
Applies to all Hermes agent configs across all hosts. Covers these config sections:
|
||||||
|
|||||||
@@ -71,6 +71,10 @@ description: >
|
|||||||
- Set models: read the live key-scoped set rather than hardcoding one — `/v1/models` is key-scoped,
|
- Set models: read the live key-scoped set rather than hardcoding one — `/v1/models` is key-scoped,
|
||||||
and the authoritative registry is CT 116 `/opt/inference-harness/litellm_config.yaml`. Do not add
|
and the authoritative registry is CT 116 `/opt/inference-harness/litellm_config.yaml`. Do not add
|
||||||
retired names (`gemma-4-12b`, `gpu-light`, `crew-auto` — all retired 2026-09-12).
|
retired names (`gemma-4-12b`, `gpu-light`, `crew-auto` — all retired 2026-09-12).
|
||||||
|
- **Standard agent keys are LOCAL-ONLY**: `["strix-moe", "gpu-dense", "gpu-vision", "syslog-auto"]`.
|
||||||
|
Cloud models are granted ONLY to the designated cloud-enabled key — see § Cloud Provider
|
||||||
|
Consolidation. NEVER create a key with an empty `models` list (`{}`) or `all-proxy-models`.
|
||||||
|
In LiteLLM Community both silently grant access to EVERY model, including cloud.
|
||||||
- Note: `ornith-1.0-35b` is NOT a valid LiteLLM model name (use `strix-moe`, the stable alias). qwen3.6-35B-A3B removed from fleet (was never deployed).
|
- Note: `ornith-1.0-35b` is NOT a valid LiteLLM model name (use `strix-moe`, the stable alias). qwen3.6-35B-A3B removed from fleet (was never deployed).
|
||||||
- Return the new key
|
- Return the new key
|
||||||
5. **If action == "rotate"**:
|
5. **If action == "rotate"**:
|
||||||
@@ -87,6 +91,66 @@ description: >
|
|||||||
- Confirm key alias matches agent_name in LiteLLM key list
|
- Confirm key alias matches agent_name in LiteLLM key list
|
||||||
- Verify agent gateway uses vault wrapper: `cat /proc/<pid>/cmdline` shows `infisical run`
|
- Verify agent gateway uses vault wrapper: `cat /proc/<pid>/cmdline` shows `infisical run`
|
||||||
|
|
||||||
|
## Cloud Provider Consolidation (2026-09-20)
|
||||||
|
|
||||||
|
CT 116 LiteLLM (Community v1.99.1) fronts **7 upstream providers** in addition to the local
|
||||||
|
GPU models. Added 2026-09-20 — 47 cloud deployments, 51 unique model names total.
|
||||||
|
|
||||||
|
### Provider map (per-account namespacing)
|
||||||
|
|
||||||
|
Two accounts on the same vendor get **distinct prefixes** so billing, rate limits, and keys
|
||||||
|
stay separate:
|
||||||
|
|
||||||
|
| Prefix | Upstream | Auth | Vault secret |
|
||||||
|
|--------|----------|------|--------------|
|
||||||
|
| `openrouter/` | OpenRouter | API key | `OPENROUTER_API_KEY` |
|
||||||
|
| `deepseek/` | DeepSeek direct | API key | `DEEPSEEK_API_KEY` |
|
||||||
|
| `google/` | Google AI Studio (Gemini) | API key | `GEMINI_API_KEY` |
|
||||||
|
| `qwen-payg/` | QwenCloud / DashScope (pay-as-you-go) | API key | `DASHSCOPE_PAYG_KEY` |
|
||||||
|
| `qwen-plan/` | QwenCloud / DashScope (token plan) | API key | `DASHSCOPE_PLAN_KEY` |
|
||||||
|
| `tencent-payg/` | Tencent TokenHub (PAYG) | Bearer token | `TENCENT_PAYG_KEY` |
|
||||||
|
| `tencent-plan/` | Tencent TokenHub (Plan) | Bearer token | `TENCENT_PLAN_KEY` |
|
||||||
|
|
||||||
|
All cloud `api_key` fields use `os.environ/<NAME>` — the 7 secrets live in Infisical
|
||||||
|
(project=`infrastructure`, env=`production`, folder=`root`) and are injected into the
|
||||||
|
`harness-litellm` container at start. **No literal cloud keys in `litellm_config.yaml`.**
|
||||||
|
|
||||||
|
### Access tiers (MUST be enforced per key)
|
||||||
|
|
||||||
|
| Tier | Model names | Granted to |
|
||||||
|
|------|-------------|------------|
|
||||||
|
| **Local** | `strix-moe`, `gpu-dense`, `gpu-vision`, `syslog-auto` | every standard agent key |
|
||||||
|
| **Cloud** | the 47 provider models (`<prefix>/<model>`) | **only** the designated cloud-enabled key |
|
||||||
|
|
||||||
|
> ⚠️ **Community-edition caveat:** LiteLLM Community does not restrict wildcard access groups
|
||||||
|
> the way Enterprise does. Access is decided by each key's explicit `models` list. A key with
|
||||||
|
> `models = {}` or `models = ["all-proxy-models"]` sees **all** models — a silent cloud leak.
|
||||||
|
> Every key MUST carry an explicit list. The master key always bypasses scoping (admin-only).
|
||||||
|
|
||||||
|
### Creating the cloud-enabled key
|
||||||
|
|
||||||
|
```bash
|
||||||
|
# ALWAYS read the live roster first (key-scoped):
|
||||||
|
curl -s -H "Authorization: Bearer <AGENT_KEY>" http://192.168.68.116/litellm/v1/models \
|
||||||
|
| jq -r '.data[].id'
|
||||||
|
|
||||||
|
# Then generate a key with an EXPLICIT model list (never empty, never a wildcard).
|
||||||
|
# For the cloud-enabled key, list local + cloud. For a standard agent, local only.
|
||||||
|
```
|
||||||
|
|
||||||
|
Verify after any key change: a standard agent key must return **4 models**, and must NOT return
|
||||||
|
any `<prefix>/` cloud model.
|
||||||
|
|
||||||
|
### Adding a new cloud provider
|
||||||
|
|
||||||
|
1. Add the upstream key to Infisical `infrastructure/production/root`.
|
||||||
|
2. Add the deployment(s) to `/opt/inference-harness/litellm_config.yaml` with an `os.environ/` ref
|
||||||
|
and a namespaced `model_name` (`<provider>-<account>/<model>` when a vendor has >1 account).
|
||||||
|
3. Restart the `harness-litellm` container.
|
||||||
|
4. Grant the model to the cloud-enabled key ONLY (explicit list) — never to agent keys without
|
||||||
|
captain approval.
|
||||||
|
5. Update this table and the access-tier section.
|
||||||
|
|
||||||
## Production Vault Access Process (canonical, 2026-07-17)
|
## Production Vault Access Process (canonical, 2026-07-17)
|
||||||
|
|
||||||
The non-fail approach to agentic vault access. Deployed on all 4 Hermes agents
|
The non-fail approach to agentic vault access. Deployed on all 4 Hermes agents
|
||||||
|
|||||||
Reference in New Issue
Block a user