Files
prose-contracts/hermes-agent-baseline.prose.md
T
root 9f0e04f22c
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s
fix: update contracts for hwepve 6th node + Mumuni migration (lxc/114 on hwepve)
- proxmox-monitor: 5→6 nodes, add hwepve (.4, Huawei Matebook 16), note Mumuni (lxc/114) placement
- zulip-health: add dual-gateway collision detection to B2, note Mumuni on hwepve
- hermes-agent-baseline: fix Mumuni node minipve→hwepve
- hermes-zulip-restore: add hwepve Proxmox column for Mumuni
- mumuni-delegation: fix CT 118/storepve/.6 → lxc/114/hwepve/.123, 5→6 nodes
- pm2-self-heal: verified — no Mumuni references, no changes needed

Post-migration discoveries:
- Mumuni container migrated from minipve to hwepve (6th node, 192.168.68.4)
- IP preserved (192.168.68.123), hostname 'mumuni' unchanged
- Dual gateway collision (PID 1442 --replace + PID 2924 --force) resolved by
  container restart at 10:07; gateway restarted clean at 13:06 with fresh Zulip
  queue (errors=0, reconnects=0)
2026-07-20 17:08:58 +00:00

10 KiB

kind, name, version, description, author
kind name version description author
template hermes-agent-baseline 1.0.0 Canonical known-good baseline for all Syslog Hermes agents. Captures the exact configuration state, keys, workarounds, and audit procedure. When an agent's configuration goes sideways, restore from this baseline. Last verified 2026-07-16. All GPUs 256K context (RTX 3090 .8, RTX 5070 .110, Strix Halo .15). Parallel 1 fleet-wide (Strix Halo handles compression solo). Abiba (pi agent)

Hermes Agent Baseline — Canonical Good State

Quick Restore

# Verify all agents against baseline in one command:
for ct in 112 114 111 113; do
    echo "CT $ct: $(pct-run $ct grep api_key: /root/.hermes/config.yaml | grep -c sk-) api_keys found"
done

Agent Map

Agent CT Node IP LiteLLM Alias Key Source Platform
Tanko 112 amdpve .122 tanko Infisical vault Hermes
Mumuni 114 hwepve .123 mumuni Infisical vault Hermes
Koby 129 amdpve srv1079750 koby Infisical vault Hermes
Koonimo 114 amdpve ? koonimo Infisical vault Hermes
Shumba 192.168.68.119 N/A N/A (DeepSeek) Hermes (RETIRED — CT119 now Infisical vault)

Note

: CT hostnames (tdunna→CT129, baggy→CT114) differ from agent identities (koby, koonimo).

Access: pct-run <CT_ID> <command> — no IPs needed. GPU hosts (.8, .110, .15) use SSH. Keys are stored in Infisical vault (project=agents, env=production) and injected at runtime via infisical run -- wrapper. Plaintext keys removed from this baseline.

Key Architecture

Infisical vault → infisical run -- hermes gateway → LITELLM_API_KEY (runtime)
                       ↓
Agent (systemd) → LITELLM_API_KEY → LiteLLM (:116/v1) → Router (:9000) → GPU (llama-server)
                                     └── Key DB (Postgres)
  • Master key: stored in Infisical vault (project=infrastructure, secret=LITELLM_MASTER_KEY) — ADMIN ONLY
  • Agent keys: Each agent has a dedicated key in LiteLLM's database with alias matching the agent name
  • Key injection: infisical run --project=agents --env=production -- hermes gateway run injects LITELLM_API_KEY at runtime
  • Key source: Infisical vault → runtime env var. /etc/environment is CLEAN (stripped, tagged # [INFISICAL])
  • Legacy override (pre-migration): /home/jerome/.config/systemd/user/hermes-gateway.service.d/env.conf — should be REMOVED

Config Pattern — Mandatory Fields

For Hermes Agents (Tanko, Mumuni, Koonimo)

Every agent's /root/.hermes/config.yaml (or /home/jerome/.hermes/config.yaml) MUST have:

1. Main Model

model:
  default: syslog-auto
  provider: custom:harness       # or: harness
  api_key_env: LITELLM_API_KEY
  max_tokens: 4096

2. Custom Provider

custom_providers:
  - name: harness
    model: syslog-auto
    base_url: http://192.168.68.116/v1
    api_key_env: LITELLM_API_KEY
    api_mode: chat_completions

3. Vision (CRITICAL — must have api_key directly!)

auxiliary:
  vision:
    provider: harness
    model: gemma-4-12b              # or syslog-auto
    base_url: http://192.168.68.116/v1
    api_key_env: LITELLM_API_KEY
    api_key: <value from: infisical secrets get LITELLM_API_KEY --project=agents --env=production>   # ← MANDATORY workaround
    timeout: 60
    download_timeout: 30

4. Compression (must have api_key!)

  compression:
    enabled: true
    threshold: 0.65
    target_ratio: 0.3
    provider: harness
    model: syslog-auto               # or gemma-4-12b
    base_url: http://192.168.68.116/v1
    api_key_env: LITELLM_API_KEY
    api_key: <value from: infisical secrets get LITELLM_API_KEY --project=agents --env=production>   # ← MANDATORY workaround
    timeout: 120

Known Bug: api_key_env Ignored by Auxiliary Client

Bug location: agent/auxiliary_client.py_resolve_task_provider_model() (line ~5478)

What happens: The function reads api_key from auxiliary task configs but does NOT resolve api_key_env. If only api_key_env is set (no api_key), the key resolves to None, and the explicit_base_url branch in resolve_provider_client falls through to "no-key-required" → 401 from LiteLLM.

Impact: Vision analysis, compression, and any other auxiliary task calling LiteLLM/harness will fail with:

401: LiteLLM Virtual Key expected. Received=no-k****ired, expected to start with 'sk-'

Workaround: Set api_key directly (copy the value from Infisical vault: infisical secrets get LITELLM_API_KEY --project=agents --env=production) alongside api_key_env in every auxiliary task config that uses the harness provider.

Permanent fix: Patch _resolve_task_provider_model() to resolve api_key_env when api_key is empty:

cfg_api_key = str(task_config.get("api_key", "")).strip() or None
if not cfg_api_key:
    key_env = str(task_config.get("api_key_env", "")).strip()
    if key_env:
        cfg_api_key = os.getenv(key_env, "").strip() or None

Audit Procedure

Full Audit (all agents)

for ct in 112 114 111 113; do
    echo "=== CT $ct ==="
    # Verify /etc/environment is CLEAN (no LITELLM_API_KEY)
    pct-run $ct "grep -c LITELLM_API_KEY /etc/environment 2>/dev/null || echo '0 (clean)'"
    # Verify gateway uses infisical run wrapper
    pct-run $ct "ps aux | grep 'infisical run' | grep -v grep"
    # Check for hardcoded harness keys
    pct-run $ct grep "api_key: sk-" /root/.hermes/config.yaml | grep -v api_key_env
    echo ""
done

Master Key Leak Check

# On every agent — must return empty:
pct-run <CT> grep -rl "sk-litellm" /root/ /etc/ 2>/dev/null
# Vault is the only place the master key should exist

Verify Key Works

# Retrieve key from vault and test:
KEY=$(infisical secrets get LITELLM_API_KEY --project=agents --env=production --plain)
curl -s http://192.168.68.116:80/v1/models \
  -H "Authorization: Bearer $KEY" | grep syslog-auto
# Must return model list

Verify Vision/Compression

# Check both api_key and api_key_env are present:
pct-run <CT> grep -A8 "vision:" /root/.hermes/config.yaml | grep api_key
# Must show both api_key: sk-... and api_key_env: LITELLM_API_KEY

For Koby (CT 129 / tdunna)

Koby runs Hermes on CT 129 (tdunna). Config files at /root/.hermes/config.yaml. Same Hermes pattern as Tanko/Mumuni/Koonimo — see config sections above.

LiteLLM key: alias koby in LiteLLM DB, injected via infisical run -- wrapper.

For pi Agents (Abiba)

Abiba (CT100) runs pi via PM2 with the Zulip extension. Config files: ~/.pi/agent/models.json, ~/.pi/agent/settings.json.

models.json — Must only list models authorized for the agent's LiteLLM key. Key is injected via infisical run -- wrapper at PM2 startup:

{
  "providers": {
    "syslog-harness": {
      "baseUrl": "http://192.168.68.116/v1",
      "api": "openai-completions",
      "apiKey": "${LITELLM_API_KEY}",
      "models": [
        { "id": "syslog-auto" },
        { "id": "strix-moe" },
        { "id": "gpu-dense" },
        { "id": "gpu-light" },
        { "id": "qwen3.6-27B-code" },
        { "id": "gemma-4-12b" }
      ]
    }
  }
}

settings.json — Always use syslog-auto as default:

{
  "defaultProvider": "syslog-harness",
  "defaultModel": "syslog-auto"
}

Validation: Verify models match LiteLLM's authorized list:

curl -s http://192.168.68.116:4000/v1/models \
  -H "Authorization: Bearer $(grep apiKey ~/.pi/agent/models.json | head -1 | cut -d'"' -f4)" \
  | jq '.data[].id'

Stuck worker detection: In PM2 logs, workers=[<id>:busy:N] with growing N indicates a stuck worker (model error, no agent_end emitted). Fix: correct models.json, delete stale sessions from ~/.pi/agent/sessions/zulip/, restart PM2.

Service: pm2 restart koby-zulip, health at :9201/health.

Systemd Pattern

For agents where the gateway runs as a user service:

# /home/jerome/.config/systemd/user/hermes-gateway.service.d/env.conf
[Service]
Environment="LITELLM_API_KEY=sk-..."    # must match /etc/environment

For agents where the gateway runs as root:

# /root/.config/systemd/user/hermes-gateway.service
EnvironmentFile=/etc/environment        # sources LITELLM_API_KEY

No drop-in that hardcodes the master key. Ever.

Attachment Cache Locations

Type Path
Images ~/.hermes/cache/images/
Documents ~/.hermes/cache/documents/
Audio ~/.hermes/cache/audio/

Health Verification

Run the consolidated health check:

python3 /root/scripts/agent-health-check.py

This validates all 4 LiteLLM keys, detects GPU port conflicts (ghost processes), verifies gateway liveness, confirms Zulip streaming (edit_message present), and counts recent errors. Non-disruptive — never restarts anything.

GPU Port Conflict Detection

All 3 GPU hosts have pre-start ghost detection in their launch wrappers:

  • .8 and .110: inline check in llama-wrapper.sh
  • .15: /usr/local/bin/port-cleanup.sh (ExecStartPre, replaces blanket pkill)

Detection pattern: ss -tlnp on port 8080 → compare pid against systemctl MainPID. If they differ → ghost detected → kill ghost → start fresh.

  • hermes-key-enforcement.prose.md — key policy, rotation, detection query
  • hermes-config-template.prose.md — full configuration template
  • gpu-fleet.prose.md — GPU fleet and agent key table
  • infrastructure-control.prose.md — CT inventory with pct-run access
  • litellm-health.prose.md — LiteLLM stack health verification

Change Log

Date Change
2026-07-08 Koby: fixed model mismatch (qwen3.6-35B-A3B→syslog-auto), added config section. Key rotated and stored in vault. Added Failure Mode #11 to zulip-adapter-lessons.
2026-07-06 Port conflict detection added to all 3 GPU wrappers. Consolidated health check script deployed. Zulip streaming edit_message enabled for Tanko/Mumuni.
2026-07-05 Baseline created. All 4 agents audited, master key removed, api_key workaround applied