Files
prose-contracts/hermes-key-enforcement.prose.md
T
root f4c4850f5a
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 8s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 11s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
fix: hermes-key-enforcement probe timeout - bounded scan, honest failure kinds
Problem: koby's 16 GB .hermes tree made the grep scan take 5.7s, exceeding
the leg's timeout. The timeout was rendered as 'agent may be down' when the
host was actually up and the scan just needed more time.

Changes:
1. Bounded scan: --exclude-dir=state-snapshots (0.44s vs 0.91s on koby)
2. Timeout policy: 15s scan timeout, 10s SSH connect timeout (documented)
3. Failure kinds: probe-failed (timeout) vs unreachable (ssh connect failed)
4. Negative control: 1s timeout proves probe-failed not down

Measured cost (2026-10-02): koby full scan 0.91s, bounded 0.44s.
Timeout set to 15s for headroom.

Correlation: corr=69683f073322b46c
2026-10-02 11:05:41 +00:00

21 KiB

kind, name, version, description, author
kind name version description author
enforcement hermes-key-enforcement 1.0.0 Enforces standardized API key configuration across all Hermes agents. Harness/LiteLLM providers MUST use api_key_env indirection. External providers (DeepSeek, OpenAI, Anthropic) may use hardcoded keys. Single source of truth: Infisical vault (project=agents, env=production) — injected at runtime via `infisical run --` wrapper. /etc/environment is DEPRECATED for agent keys post-migration. Designed to make key rotation a one-step vault operation. Abiba (pi agent)

Hermes Key Enforcement Contract

Rule (One Sentence)

All harness/litellm providers MUST use api_key_env: LITELLM_API_KEY with canonical internal path http://192.168.68.116/litellm/v1 (Hermes appends /v1/responses) or public path https://litellm.sysloggh.net/v1 — hardcoded keys AND direct :4000 access are both forbidden. Internal /v1 still works but is non-canonical (WARN, not FAIL).

Cloud provider models (OpenRouter, DeepSeek, Google AI Studio, QwenCloud PAYG/Plan, Tencent TokenHub PAYG/Plan) added to CT 116 on 2026-09-20 are reachable ONLY through the designated cloud-enabled key. Standard agent keys remain LOCAL-ONLY (strix-moe, gpu-dense, gpu-vision, syslog-auto) and MUST NOT be granted cloud models unless explicitly approved by the captain.

Model Access Tiers (2026-09-20)

CT 116 hosts two access tiers of model. Tier membership is enforced per virtual key via that key's models allowlist.

Tier Models Who gets it
Local strix-moe, gpu-dense, gpu-vision, syslog-auto All standard agent keys (tanko, mumuni, koby, koonimo, abiba-pi)
Cloud 47 provider models: openrouter/*, deepseek/*, google/*, qwen-payg/*, qwen-plan/*, tencent-payg/*, tencent-plan/* Only the designated cloud-enabled key (captain decision)

Rules:

  1. A key with an empty models list ({}) or all-proxy-models is UNSCOPED — it silently gains ALL models, including cloud. Never create or leave an agent key in this state.
  2. Agent keys MUST carry an explicit local-only models list.
  3. Granting a cloud model to an agent key requires explicit captain approval and a recorded reason.
  4. The master key always bypasses scoping — it is admin-only, never for inference.

See litellm-api-keys § Cloud Provider Consolidation for the key-creation procedure.

Scope

Applies to all Hermes agent configs across all hosts. Covers these config sections:

  • model.api_key
  • custom_providers[].api_key (when name contains harness or litellm)
  • auxiliary.*.api_key (when provider is harness or contains litellm)
  • delegation.api_key (when provider is harness or contains litellm)
  • compression.api_key (when provider is harness or contains litellm)
  • fallback_providers[].api_key (when provider is harness)

Architecture (2026-07-10)

Syslog is migrating away from unauthenticated direct access to the shared inference harness.

Path Auth Status
http://192.168.68.116/v1 Bearer sk-* key (nginx-fronted) ✅ VALID — authenticated via nginx :80 (verified 2026-08-09: 401 without key, 200 with)
http://192.168.68.116/litellm/v1 Bearer sk-* key (nginx-fronted) ✅ CURRENT / CANONICAL — captain-approved migration target; 600s proxy_read_timeout (verified)
http://192.168.68.116:4000/v1 Bearer sk-* key (direct container) ❌ FORBIDDEN — bypasses nginx; port 4000 direct is not a config path

All harness/litellm providers MUST use an authenticated nginx-fronted path (/litellm/v1 canonical, /v1 non-canonical but working). The public host https://litellm.sysloggh.net serves /v1 ONLY (404 on /litellm/v1).

🔥 CRITICAL: Double-Path Bug (2026-07-10)

When api_mode: responses is set, Hermes appends /v1/responses to base_url. If base_url already includes /litellm/v1/responses, the result is:

http://192.168.68.116/litellm/v1/responses/v1/responses → 404

The base_url must end at /v1 — never include /responses:

# ✅ CORRECT — Hermes appends /v1/responses for api_mode: responses
base_url: http://192.168.68.116/litellm/v1

# ❌ WRONG — produces double path
base_url: http://192.168.68.116/litellm/v1/responses

This applies to ALL sections using the harness provider: custom_providers, delegation, auxiliary.*.

Exemptions

External providers are explicitly exempt and may use hardcoded keys:

  • DeepSeek (api.deepseek.com)
  • OpenAI (api.openai.com)
  • Anthropic (api.anthropic.com)
  • OpenRouter
  • Any provider whose base_url does NOT match 192.168.68.116 or litellm.sysloggh.net

Standard Pattern

Canonical vault process (2026-07-16): see litellm-api-keys § Production Vault Access Process. All agents MUST use the infisical-gateway.sh wrapper (live vault injection). Hardcoded systemd drop-ins / config.yaml keys are DEPRECATED — they rot on rotation (root cause of the 2026-07-16 401 storm). 4/5 agents migrated; tanko (user jerome) pending.

# ✅ CORRECT — all harness/litellm providers (authenticated path, NO /responses suffix)
model:
  provider: harness
  base_url: http://192.168.68.116/litellm/v1           # ← Hermes appends /v1/responses
  api_key_env: LITELLM_API_KEY

custom_providers:
  - name: harness
    api_mode: responses
    base_url: http://192.168.68.116/litellm/v1           # ← NO /responses suffix!
    api_key_env: LITELLM_API_KEY

auxiliary:
  compression:
    provider: harness
    base_url: http://192.168.68.116/litellm/v1           # ← NO /responses suffix!
    api_key_env: LITELLM_API_KEY

# ✅ ALSO CORRECT — external providers
fallback_providers:
  - provider: deepseek
    base_url: https://api.deepseek.com
    api_key: sk-synthetic-external-example   # ← hardcoded OK (external, synthetic example)
    api_key_env: DEEPSEEK_API_KEY      # ← also OK if set in environment (vault or /etc/environment)
# ❌ FORBIDDEN — hardcoded key (top) OR unauthenticated path (bottom)
model:
  provider: harness
  api_key: sk-synthetic-example-12345   # ← RULE VIOLATION: hardcoded key (synthetic example)

model:
  provider: harness
  base_url: http://192.168.68.116/v1   # ← NON-CANONICAL but WORKING (authenticated via nginx, WARN not FAIL)
  api_key_env: LITELLM_API_KEY

Reachability Detection

Before checking for hardcoded keys, verify the host is reachable and can be audited. Use the shared reachability helper from the clone root:

# Run on each host to check reachability (Tanko, Mumuni, Koonimo, Koby)
scripts/hermes-reachability-check.sh <host> "api_key: sk-" "/root/.hermes/"
# Example: scripts/hermes-reachability-check.sh 192.168.68.122 "api_key: sk-" "/root/.hermes/"

# Expected outcomes:
# - UNREACHABLE: SSH connection failed (host is down)
# - VIOLATION: SSH succeeded and found matches (report the finding)
# - COMPLIANT: SSH succeeded and found no matches (no hardcoded keys in config)
#
# NOTE: The bug this replaces was deriving reachability from the remote grep's exit code.
# The correct pattern: remote side always succeeds (grep ...; true), so ssh status = connection only.

Violation Classification

When reporting findings, separate POLICY observations from FAULT findings:

ACCEPTABLE PATTERN

Agent keys live in .env or .env.vault files with 600 permissions (koonimo's shape is the canonical example). A plaintext key inside a config.yaml or any config.yaml.bak-* file is a violation — the backup files are not part of the runtime credential path and are not watched by the scanner, so a key in them is stale clutter that a future reader can mistake for a working key.

Fix procedure (when a backup file is found with a plaintext key):

  1. Move the file out of the scanned tree (e.g., mv /root/.hermes/config.yaml.bak-* /root/hermes-config-backups/) — do NOT delete the file, just move it so the scanner pattern no longer matches.
  2. Re-run the reachability check to confirm COMPLIANT.
  3. Report the before/after check output and the commands you ran.

Rationale: Moving the file preserves history without leaving a credential where a scanner trips over it. Deleting the file loses the historical context. Keeping it in place means the next scan will report it as a finding and waste time re-deciding.

POLICY (observation only, not a fault)

  • Agent uses a non-internal-harness provider (e.g., direct DeepSeek, Tencent, OpenRouter)
  • Config text has a field that looks unusual but the agent's calls are succeeding
  • Example: "POLICY: Koonimo uses deepseek directly; calls succeeding in last hour"

FAULT (requires request-level evidence)

  • Agent's calls are failing with auth errors (401/403 in logs)
  • Agent's config has no valid API key AND calls are failing
  • Example: "FAULT: Koby's LiteLLM key expired; 401 observed at 2026-09-14 11:42:00"

Rules

  1. Do NOT infer the runtime's credential resolution from config text alone.
  2. Require request-level evidence before calling something a FAULT: an observed auth failure in the agent's log, or the absence of successful calls in the window.
  3. If calls are succeeding, the correct output is "POLICY: uses directly; calls succeeding" - not a violation.
  4. State what you OBSERVED, not what the field implies.

Detection Query

Run on any Hermes host to detect violations.

Timeout policy (2026-10-02): The scan timeout is 15 seconds, set from measured cost on the largest target (koby, 16 GB .hermes tree; full scan = 0.91 s, bounded scan = 0.44 s). The SSH connection timeout is 10 seconds (separate from the scan timeout). A scan timeout renders as probe-failed: <agent> <ip> (timeout after 15s) — never as "unreachable" or "may be down". An SSH connection failure (exit status 255) renders as unreachable: <agent> <ip> (ssh connect failed).

Bounded scan (2026-10-02): Do NOT recurse the entire /root/.hermes/ tree. Use --exclude-dir=state-snapshots to skip dated snapshot directories. Rationale: a superseded config will always carry a superseded key and will report forever with zero signal content (the koby state-snapshot line has repeated on consecutive days). If you deliberately want to include snapshots, say so in the contract and the report.

# 1. Check config.yaml for hardcoded harness keys (bounded scan — excludes state-snapshots)
timeout 15 ssh -o ConnectTimeout=10 -o StrictHostKeyChecking=no root@<ip> \
  "grep -rn 'api_key: sk-' /root/.hermes/ --exclude-dir=state-snapshots --include='config.yaml' | grep -v 'deepseek\|openai\|anthropic\|DEEPSEEK'" \
  2>/dev/null

# Interpret exit status:
#   0 = match found (violation)
#   1 = no match (pass)
#   124 = timeout (probe-failed, not unreachable)
#   255 = ssh connect failed (unreachable)
#   other = probe-failed (record the actual code)

# 1b. Check for double-path bug: base_url ending with /responses
# (Hermes appends /v1/responses when api_mode=responses, so base_url must end at /v1)
timeout 15 ssh -o ConnectTimeout=10 -o StrictHostKeyChecking=no root@<ip> \
  "grep -rn 'litellm/v1/responses' /root/.hermes/config.yaml" 2>/dev/null
# ANY output here = WRONG. Must be 'litellm/v1' without /responses suffix.

# 2. Check systemd drop-ins for master key leaks (2026-07-05: Tanko had this)
timeout 15 ssh -o ConnectTimeout=10 -o StrictHostKeyChecking=no root@<ip> \
  "grep -rn 'LITELLM_API_KEY' /root/.config/systemd/user/ 2>/dev/null" \
  "grep -rn 'LITELLM_API_KEY=sk-synthetic-litellm-…' /root/.config/systemd/ 2>/dev/null"

# 3. Verify running process env matches dedicated key
timeout 15 ssh -o ConnectTimeout=10 -o StrictHostKeyChecking=no root@<ip> \
  "cat /proc/$(cat /home/jerome/.hermes/gateway.pid | python3 -c \"import sys,json; print(json.load(sys.stdin)['pid'])\")/environ | tr '\0' '\n' | grep LITELLM_API_KEY"

If any output from step 2 — critical violation (master key leaked). Fix immediately.

Negative control (probe-failed vs unreachable)

To prove the distinction between a scan timeout and a connection failure, run with a deliberately tiny timeout:

# Negative control: 1-second timeout on koby (scan takes 0.9 s, so this will time out)
timeout 1 ssh -o ConnectTimeout=10 -o StrictHostKeyChecking=no root@192.168.68.129 \
  "grep -rn 'api_key: sk-' /root/.hermes/ --exclude-dir=state-snapshots --include='config.yaml'" 2>/dev/null
# Expected: exit status 124 (timeout) → render as "probe-failed: koby 192.168.68.129 (timeout after 1s)"
# NOT: "unreachable" or "may be down"

Rotation Procedure

With this standard enforced, key rotation is one vault update:

# 1. Generate new key in LiteLLM: POST /key/generate with agent alias
# 2. Update Infisical vault secret
infisical secrets set LITELLM_API_KEY=sk-NEW_KEY \
  --project=agents --env=production
# 3. Restart agent gateway (key auto-injected via infisical run -- wrapper)
ssh root@<host> "systemctl restart hermes-gateway"
# 4. Verify
curl -s -H "Authorization: Bearer sk-NEW_KEY" http://192.168.68.116/litellm/v1/models

Done. No config file changes needed. No /etc/environment edits needed. The agent picks up the new key via infisical run -- at gateway startup.

Post-migration note: /etc/environment is NO LONGER the key source. Strip all LITELLM_API_KEY lines from /etc/environment (comment out with # [INFISICAL]) and let the infisical run -- wrapper inject the key at runtime.

Key Longevity Policy (2026-07-04)

Keys are permanent and use bare agent name aliases.

  • Duration: null — keys never expire by default. Expiry must be set EXPLICITLY at creation with the duration parameter (e.g., 90d for 90 days). The 90-day default is the standard; however, the config default is NOT honoured by LiteLLM 1.99.1 (verified on CT 116: a key generated with no explicit duration returns expires=null). This has been recorded in /opt/inference-harness/litellm_config.yaml to prevent re-filing as a bug.
  • Daily Audit: A daily audit job runs at 00:00 UTC (/usr/local/bin/litellm-key-renewal-ct116.sh, cron 00:00). It is AUDIT-ONLY and does not perform renewal. It lists every key, reports those with no expiry and those inside a 14-day warning window, explicitly EXCLUDES abiba-pi and koby (report-only, and .129 must never be touched), and logs RENEWAL-REQUIRED-BUT-NOT-PERFORMED + NO KEY WAS CHANGED when renewal is skipped. Renewal is NOT implemented — keys must not be rotated until delivery (vault injection + consumer verification) exists and is proven end-to-end.
  • Exclusions: abiba-pi and every firstmate/secondmate/crewmate key stay WITHOUT an expiry until a proven renewal path exists. koby is report-only (never touched). These exclusions are enforced by the audit job.
  • Alias convention: bare agent name only (e.g., tanko, mumuni, koby, koonimo). No dates, no versions. The alias IS the identity.
  • Rotation triggers: compromise, personnel departure, or quarterly security hygiene. NOT calendar-driven. Manual rotation is permitted only when the renewal delivery path is proven and verified on a throwaway consumer before production use.
  • Max budget: $100 per key (config default).
# NOT currently set in the authority; recommended value. CT 116 litellm_config.yaml has no
# default_key_generate_params block today, and a key generated with no explicit models comes back
# with an EMPTY models list. `models` is a literal key-generation parameter, so this is a value to
# ADD — re-read the live registry at CT 116 /opt/inference-harness/litellm_config.yaml and
# re-verify before applying.
litellm_settings:
  default_key_generate_params:
    models: ["syslog-auto", "gpu-dense", "gpu-vision", "strix-moe"]
    duration: null        # ← permanent
    max_budget: 100
    metadata:
      purpose: "agent-inference"

Verified Agents (2026-07-05 update)

Agent CT IP LiteLLM Alias Key Source Status Gateway Wrapper Last Verified
Tanko 112 .122 tanko Infisical vault ✅ Fixed infisical run 20:17 UTC Jul 5
Mumuni 105 (kagentz) .14 mumuni Infisical vault ✅ Fixed systemd Hermes gateway 2026-08-29
Koby 111 .129 koby Infisical vault ✅ Fixed (DeepSeek-primary) infisical run 23:30 UTC Jul 5
Koonimo 113 .114 koonimo Infisical vault ✅ Fixed infisical run (migrated 2026-07-11) 2026-08-09
Abiba 100 .65 abiba-pi Infisical vault ✅ N/A (pi native) — 19:44 UTC Jul 5
Kagenz0 105 .14 — — ❌ DOWN — 19:14 EDT Jul 4

Note

: CT hostnames (tdunna, baggy) differ from agent identities (koby, koonimo). LiteLLM key aliases use agent identity, not CT hostname.

Migration Status: Authenticated Path

Agent /litellm/v1 Legacy /v1 Status
Mumuni ✅ harness provider ✅ auxiliary on /v1 (valid) ✅ Authenticated (verified 2026-08-09)
Tanko ✅ 5 sections 0 ✅ Migrated 2026-08-08, keys 200
Koby ✅ custom provider (harness name) — ✅ External DeepSeek primary (intentional, captain ruling 2026-08-11)
Koonimo ✅ .114 (baggy) — ✅ 128K context applied 2026-08-09

Systemd Service Pattern (2026-07-11 — vault migration)

All Hermes agents use systemd to manage their gateway. The gateway service is wrapped with infisical run -- to inject secrets at runtime.

Correct pattern (post-migration):

# Service file wraps gateway with Infisical:
[Service]
ExecStart=/usr/bin/infisical run --project=agents --env=production -- \
  /usr/bin/hermes gateway run

# /etc/environment is CLEAN — no LITELLM_API_KEY present
# (strip it and tag with # [INFISICAL] if present)

Legacy pattern (deprecated — pre-migration only):

# DO NOT USE post-migration:
EnvironmentFile=/etc/environment
# This pattern was replaced by infisical run -- wrapper

Rotation procedure (one vault operation with this standard):

  1. Generate new key in LiteLLM: curl /key/generate with agent alias
  2. Update Infisical vault: infisical secrets set LITELLM_API_KEY=sk-NEW --project=agents --env=production
  3. Restart: systemctl restart hermes-gateway (key auto-injected via wrapper)

Violation Response

  1. Detect — run detection query above
  2. Fix — replace api_key: sk-... with api_key_env: LITELLM_API_KEY in all harness/litellm sections
  3. Verify — grep -c "api_key_env" config.yaml should increase, hardcoded harness keys should be 0
  4. Restart — gateway must restart to pick up env var
  5. Confirm — test key against LiteLLM: curl -H "Authorization: Bearer $KEY" .../v1/models → 200
  6. Update — bump the verified table above
  7. Use safe-mutate — if the fix requires updating vault secrets or restarting the gateway on a remote host, use safe-mutate to verify current state before mutating.
  • hermes-config-template.prose.md — full configuration template
  • litellm-health.prose.md — LiteLLM stack health verification
  • zulip-platform-verification.prose.md — cross-platform agent verification
  • litellm-api-keys.prose.md — API key creation, rotation, and verification

CI Pipeline (2026-07-04)

All contract changes must pass the PR Pipeline before merge:

auth → validate → lint → ai-review → gate
  • Trigger: push to master (abiba-bot only) or pull request
  • Branch protection: Only abiba-bot can push directly to master. All other users must use PRs.
  • Status check: PR Pipeline — Authorize → Validate → Review → Merge required before merge
  • Runner: runner-ct110 (Gitea Actions v0.6.1) on CT 110
  • Config: .gitea/workflows/pr-pipeline.yaml

Known Bug: auxiliary_client ignores api_key_env (2026-07-05)

Bug: _resolve_task_provider_model() in agent/auxiliary_client.py reads api_key from auxiliary task configs (vision, compression, etc.) but does NOT resolve api_key_env. The custom provider resolution path handles api_key_env, but auxiliary tasks take a different code path that ignores it.

Impact: Vision analysis and compression calls fall through to the "no-key-required" placeholder, causing 401 errors on LiteLLM/harness (which require sk-* keys).

Workaround: Set api_key directly alongside api_key_env in each auxiliary task config:

auxiliary:
  vision:
    api_key: sk-<agent-key-from-vault>           # ← workaround (get via: infisical secrets get LITELLM_API_KEY --project=agents --env=production --plain)
    api_key_env: LITELLM_API_KEY
    base_url: http://192.168.68.116/litellm/v1
    model: gpu-vision
    provider: harness
  compression:
    api_key: sk-<agent-key-from-vault>           # ← workaround (same as above)
    api_key_env: LITELLM_API_KEY
    base_url: http://192.168.68.116/litellm/v1
    model: syslog-auto
    provider: harness

Affected agents: All Hermes agents with harness/LiteLLM provider and api_key_env in auxiliary configs (all 4 Hermes agents patched 2026-07-05).

Source location: agent/auxiliary_client.py line 5478